EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CY

White Box Evidence Packages for Policy Audit Reports

arXiv:2607.21462v1 Announce Type: new Abstract: As AI governance moves from benchmark scores toward auditable oversight, a central question is how reviewers can tell whether an LLM-generated audit report is actually supported by evidence. This paper studies that question in passage-anchored policy audits, where a report must interpret a given policy passage and cite evidence for its claims. We introduce a controlled evaluation framework that holds the passage, rubric, and auditor model fixed while changing only the evidence interface supplied to the auditor. Across 60 AGORA policy cases, we generate 600 structured reports under ten evidence conditions, including passage-based evidence, internal model evidence, a hybrid package, and a shuffled control that preserves evidence format while breaking case relevance. Five human reviewers evaluate the primary interfaces for correctness, passage grounding, diagnostic usefulness, and evidence misuse. The results show that internal evidence chan

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CY

Open Veins of Algorithmic Auditing: Why AI Assessment Lags Behind Its Deployment in the Global South

arXiv:2607.21317v1 Announce Type: new Abstract: Artificial intelligence is being deployed across the Global South at a pace matching or exceeding the Global North, yet AI governance has not kept pace, and the gap is far wider in the South. Drawing on a decade of AI audit practice across Latin America, Sub-Saharan Africa, and Asia Pacific (the only fully published second-party audit of a deployed system in the region, Robot Laura in Brazil; two completed but unreleased national audits, of a child-welfare risk model and a public-employment matching algorithm; thirteen Responsible AI Assessments; and a regional landscape analysis), this paper documents patterns in the sparse field of Global South AI evaluation and explains why it so rarely occurs. We count fewer than twenty published second- and third-party audits of deployed systems across the region over the past decade, against hundreds of documented public-sector algorithms and multibillion-dollar national AI investments. We find four

Source ↗
technology Fri, 24 Jul 2026 00:00:00 -0400
arXiv cs.CY

Language Models Embody and Amplify Human Cognitive Distortions: What Is to Be Done?

arXiv:2607.20695v1 Announce Type: new Abstract: Human judgment is fundamentally prone to error. A promise of AI is that it will rid decisions of bias and ensure a fairer and safer world for all. Yet research unequivocally demonstrates that LLMs exhibit consequential sociocognitive biases. We alert readers that bias in AI (a) is covert and ironically a feature of alignment goals, (b) is not merely a mirror, but an amplifier of human bias, (c) intensifies across model generations, and (d) even transmits bias to humans. Given the potentially seismic and ubiquitous influence of AI on decision making, we propose countermeasures that are diagnostic, regulatory and operational.

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

Design strategies for empathetic AI robots for older adults

arXiv:2510.01192v2 Announce Type: replace-cross Abstract: Emulating empathy in human-robot interaction is a key component for achieving satisfying social, trustworthy, and ethical robot interaction with older people. Following comments from older adult study participants, the article uses humanities methods to identify a gap in defining empathetic robot care activities. It provides a design focus to mitigate it. Current human-robot designs, to a certain extent, neglect to include empathy as a theorized design pathway. Using one digital humanities research collection on humanoid robots, it contributes an empathetic care vocabulary as a design pathway for a productive underlying foundation for designing Socially Assistive Robots (SARs) that aim to support older people's goals of aging-in-place. Using rhetorical theory, this paper defines the socio-cultural expectations for convincing empathetic relationships.

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Basic B*** Effect: The Use of LLM-based Agents Reduces the Distinctiveness and Diversity of People's Choices

arXiv:2509.02910v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly act on people's behalf: they write emails, buy groceries, and book restaurants. While the outsourcing of human decision-making to AI can be convenient, it raises a fundamental question: how does delegating identity-defining choices to AI shape who people become? Across a large field study and a controlled experiment, we study the impact of agentic LLMs on two identity-relevant outcomes: interpersonal distinctiveness - how unique a person's choices are relative to others - and intrapersonal diversity - the breadth of a single person's choices over time. Study 1 uses 110,000 real choices drawn from social media behavior of 1,000 U.S. users to compare generic and personalized agents to a human baseline. Both agents shift people's choices toward more popular options, reducing the distinctiveness of their preferences. While the use of personalized agents tempers this homogenization (compared

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

Mitigating GenAI-Powered Evidence Pollution for Out-Of-Context Misinformation Detection

arXiv:2501.14728v2 Announce Type: replace-cross Abstract: While generative artificial intelligence (GenAI) models have achieved significant success, their misuse for generating deceptive content raises growing concerns about online information security. Out-of-context (OOC) multimodal misinformation detection systems typically rely on Web-retrieved evidence to identify images repurposed in false contexts, but they are increasingly challenged by the presence of GenAI-polluted evidence. Existing work mainly focuses on verifying claims that have undergone stylistic rewriting at the claim level and assume a clean evidence corpus. In this work, we remove this assumption and systematically study the impact of GenAI-driven evidence pollution threat on OOC detection. We show that polluted evidence can degrade the performance of state-of-the-art detectors by more than 9 percentage points. We propose two mitigating strategies, cross-modal evidence reranking and cross-modal claim-evidence reasoni

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Tool to Map AI Programs in the U.S.: A Snapshot from April 2026 and an Analysis of Requirements for AI Majors and Minors

arXiv:2606.12428v2 Announce Type: replace Abstract: In this work, we locate and analyze existing undergraduate Artificial Intelligence (AI) programs in the United States in Spring 2026, creating a historic record at a time of great change in this area. To create this record, we developed a tool to detect, scrape, and display data from 361 undergraduate AI programs--majors, minors, concentrations, and certificates--at 4-year universities. Our tool, available at https://cicmap.ai, searched 563 institutions to locate these programs, a sample that represents 87% of all undergraduate Computer Science (CS) graduates in the U.S in 2025. This tool allows prospective students, guidance counselors, administrators, and faculty to easily access AI program requirements and is designed to continually update as new programs emerge. To the best of our knowledge, this survey represents the most comprehensive snapshot of the state of AI programs in the U.S. to date. With this work we offer three importa

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Invisible Risks of AI-Generated Health Information

arXiv:2605.23026v2 Announce Type: replace Abstract: Generative artificial intelligence (AI) systems now summarize health-related search results, answer medical questions, and offer guidance people once sought from clinicians. These systems bring real benefits, including plain-language explanations of medical information, around-the-clock availability, and expanded access for people facing language or literacy barriers. They also carry new risks: inaccurate guidance can harm people at scale, and malicious actors can now generate personalized health misinformation at negligible cost. In this Perspective, we argue that these risks are largely invisible to the institutions responsible for protecting public health. When AI guidance causes harm, no record exists outside the platform, no channel allows users to report it, and no independent researcher can measure the consequences. We trace these invisible risks across two settings: incidental exposure online and active seeking through search

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

Generative AI Use in Entrepreneurship: An Integrative Review and an Empowerment-Entrapment Framework

arXiv:2604.02567v2 Announce Type: replace Abstract: Despite the growing use of generative artificial intelligence (GenAI) in entrepreneurship, research on its impact remains fragmented. To address this limitation, we provide an integrative, entrepreneur-centered review of how GenAI influences entrepreneurs at each stage of the entrepreneurial process: (1) opportunity recognition and ideation, (2) opportunity evaluation and commitment, (3) resource assembly and mobilization, and (4) venture launch and growth. Based on our review, we propose the Empowerment-Entrapment Framework, which not only catalogs GenAI's benefits and costs throughout the entrepreneurial process, but also identifies potential trade-offs underlying GenAI's double-edged role. For example, GenAI may improve venture idea quality yet produce hallucinations and biases; boost entrepreneurial self-efficacy yet heighten overconfidence; increase functional breadth and self-sufficiency yet reduce prosociality; and enhance prod

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

Constitutive vs. Corrective: A Causal Taxonomy of Human Runtime Involvement in AI Systems

arXiv:2603.19213v2 Announce Type: replace Abstract: As AI systems permeate high-stakes decision-making, the terminology of human involvement---Human-in-the-Loop (HITL), Human-on-the-Loop (HOTL), and Human Oversight---has become vexingly ambiguous. This complicates interdisciplinary collaboration between computer science, law, philosophy, psychology, and sociology and breeds regulatory uncertainty. We propose a clarification grounded in causal structure, focused on runtime involvement. The distinction between HITL and HOTL is best drawn not spatially---in terms of a human's position "in" or "on" a loop---but causally: HITL is constitutive (a human contribution is necessary for the decision output), while HOTL is corrective (external to the primary causal chain, capable of preventing or modifying outputs). Within HOTL, we distinguish temporal modes---synchronous, asynchronous, and anticipatory---situated in a nested model of provider and deployer runtime. A second, orthogonal dimension c

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

Aspirational Affordances of AI

arXiv:2504.15469v2 Announce Type: replace Abstract: As artificial intelligence (AI) systems increasingly permeate processes of cultural and epistemic production, there are growing concerns about how their outputs may confine individuals and groups to restricted narratives about who or what they could be. In this paper, we advance the discourse surrounding these concerns by making three contributions. First, we introduce the concept of aspirational affordance to describe how culturally shared interpretive resources, such as concepts, images, and narratives, can shape individual cognition, and in particular exercises of imagination. We show the usefulness of this concept for grounding the evaluation of psychological risks posed by AI. Second, we provide three reasons for scrutinizing AI's influence on aspirational affordances: AI's influence is potentially more potent, but less public, than that of traditional sources; the influence is not simply incremental, but ecological, transforming

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

Growth Without Us: Machine Consumers, Corporate Circularity, and the Decoupling of GDP from Humanity after AGI

arXiv:2608.20231v1 Announce Type: cross Abstract: The standard objection to full automation is demand-side: if humans earn nothing, who buys the output? This confuses an accounting role with a biological species. We model a post-AGI economy in which corporations own populations of AI and robotic agents that are both producers and consumers of energy, compute, maintenance, and upgrades, traded among firms. Three results follow. (i) Demand closure: a closed inter-corporate economy with zero human consumption is not degenerate; it is the classical von Neumann expanding economy, whose growth rate is well defined, positive, and maximal precisely because all output is reinvested. (ii) Bottleneck removal: once economic agents are manufactured rather than reared, the binding constraint on growth shifts from human demography (a ~20-year, non-parallelizable reproduction technology capped at a few percent per year) to fabrication throughput and energy capture, permitting growth one to two orders

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use

arXiv:2608.20202v1 Announce Type: cross Abstract: Memory has become a key component of large language models, enabling them to retain information and learn from long-term interactions. However, existing memory benchmarks mainly evaluate whether information is correctly extracted, stored, and retrieved, while largely overlooking how retrieved memories reshape model reasoning and affect performance on the current task. We identify memory-induced cognitive traps: even faithfully recorded and semantically relevant memories can distort model reasoning or beliefs and degrade current task performance. To systematically evaluate these failure modes, we introduce MemTrapBench, which covers two forms of cognitive traps: Reasoning Fixation and Belief Distortion. Experiments across two model families and five representative memory frameworks show that MemTrapBench is challenging: all evaluated memory strategies underperform the no-memory setting, with even the strongest methods suffering drops of

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

Chameleon: Robust Defense Against Tor Website Fingerprinting via Many-to-Many Traffic Morphing

arXiv:2608.20160v1 Announce Type: cross Abstract: Website fingerprinting (WF) attacks can infer users' browsing activities from encrypted Tor traffic by exploiting side-channel features. Although many WF defenses have been proposed, we find that most existing defenses create learnable web trace mapping features. We further show that robustness against adversarial training does not necessarily imply robustness against defense-aware autoencoder (DAAE)-based attacks. To address these limitations, we present Chameleon, a robust WF defense based on many-to-many randomized traffic morphing. Chameleon selects morphing candidates with high intra-class diversity and low inter-class disparity. Chameleon randomly maps each webpage trace to multiple candidates, and allows different webpages to share morphing targets, thereby increasing adversarial uncertainty. For practical Tor deployment, Chameleon introduces a radix-trie-based synchronization mechanism that enables pluggable transport (PT) endpo

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

Trustworthy mobile edge caching: a blockchain approach to mitigate malicious nodes and incentivize cache sharing

arXiv:2608.20145v1 Announce Type: cross Abstract: As mobile network traffic continues to grow, content caching on edge servers is critical for reducing latency. However, challenges such as malicious edge servers that may delete or manipulate cached content, along with the limited capacity of these servers, need to be addressed. To overcome the capacity limitations, helper mobile nodes can contribute their cache resources. However, due to their selfish behavior, an incentive mechanism is necessary to encourage resource sharing. Additionally, these helper nodes can also be malicious. This paper proposes a blockchain-based trust management mechanism that addresses these challenges by accurately identifying trustworthy edge servers and mobile nodes. The proposed mechanism calculates both direct and indirect trust using smart contracts, ensuring that malicious nodes are effectively filtered out. Trustworthiness is determined based on mobile node satisfaction with the quality of service, and

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

A three-dimensional typology of agency for advanced AI systems

arXiv:2608.20041v1 Announce Type: cross Abstract: Research on the agency of advanced artificial intelligence (AI) systems focuses on agency as a normative concept and on the agency of particularly agentic AI systems. While recent work also focuses on the different profiles of agentic systems, no framework exists to address the question of the type of agency instantiated by advanced AI systems, particularly when considering non-moral forms of agency. Based on established theoretical positions in philosophy, ethics, legal theory and sociology, we develop a typology of agency for frontier AI systems consisting of three dimensions: the nature of agency (moral or legal), its mode (individual or collective) and its locus (human or non-human). Combining these dimensions produces eight possible instantiations of agency, which we classify as conventional, contested or controversial. The typology separates legal from moral agency and thereby creates conceptual space for considering individual, l

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

Designing Human-mediated AI Guidance: Ready Together for Personalized Family Emergency Preparedness

arXiv:2608.19950v1 Announce Type: cross Abstract: Artificial intelligence (AI) systems are increasingly used across domains to provide personalized information, recommendations, and decision support. However, in some contexts, AI-generated information may not be suitable for direct delivery to the final recipient. Instead, it may need to be interpreted, adapted, and communicated by a human who understands the recipient's needs, emotional state, and situational context. Human-AI interaction research has given less attention to situations in which a more knowledgeable human acts as an intermediary between an AI system and a less experienced or less informed recipient. We introduce the human-mediated AI guidance framework and explore it through Ready Together, an AI-supported family emergency preparedness system in which parents mediate AI-generated content for their children. The system is designed to provide personalized guidance and support parents in making emergency preparedness more

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

Auditing Recorded Predictive Lead Service-Line Classifications Against Physical Verification: A Statewide Study of New York

arXiv:2608.19922v1 Announce Type: cross Abstract: Under the US Lead and Copper Rule Revisions, a utility may determine a service line's material with a predictive model instead of inspecting it. New York State publishes, per address, which method was used. Almost no address carries both a model classification and a physical verification, so the check is between populations within a utility rather than paired addresses. We screen all 153 New York localities that classified at least 100 addresses this way. Seventy-five (49%), covering 125,990 addresses or 57% of those screened, record one value. Zero variance alone is not misconduct: 68 of the 75 match their own verification or have too little to test. Seven are contradicted by their own crews, six beyond any sampling explanation. Five are boroughs of New York City, which file as one system; one is East Rochester, 550 km away. New York City is the largest case: a predictive model is the recorded basis for 43,215 addresses, and on all of

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

Are LLMs becoming similarly creative? Evidence from three years of models

arXiv:2608.19437v1 Announce Type: cross Abstract: Many benchmarks track Large Language Model (LLM) performance on tasks with verifiable answers, but less is known about how LLM performance is evolving on open-ended tasks, where creativity, originality and diversity may matter as much as quality. As LLMs increasingly support human ideation and creative work, understanding trends in LLM performance on open-ended tasks is critical. This paper presents a preliminary analysis of LLM creative outputs spanning three years of model releases, examining model responses to Infinity-Chat100, a real-world collection of open-ended user queries, and the Alternate Uses Task, an established psychometric creativity assessment. Using sentence-embedding similarity, we examine trends in LLM responses to these prompts. Our findings show a statistically significant decrease in model output diversity over time, suggesting that LLM outputs may be converging in creative substance across models. If this trend pe

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

Social.Wiki: A Web Held in Common

arXiv:2608.19433v1 Announce Type: cross Abstract: Many of the websites people depend on have owners whose interests are not fully aligned with their users. We address the root of this problem by presenting a reimagining of the web where sites are not owned at all but are instead collaboratively produced like Wikipedia articles. We call the system Social.Wiki because it supports the co-creation of interactive social sites, such as those for microblogging, messaging, dating, gaming, ride sharing, and so on. With off-the-shelf AI tools, people with little or no programming experience can edit these sites to better reflect the needs and preferences of their communities. Social.Wiki builds on ideas from collaborative malleable software systems such as Webstrates, but is designed for public participation rather than use only within small, trusted groups. To this end, Social.Wiki includes governance to mitigate conflict. To accommodate diverse governance preferences, our model of "plural gove

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

Causal Inference under Interference with Learned Exposure Mappings

arXiv:2608.19224v1 Announce Type: cross Abstract: Exposure mappings are often assumed to be known in causal spillover analyses. In environmental settings, however, they are typically induced by transport processes that are not directly observed and must instead be learned from pollution data. We study how uncertainty in learned transport processes propagates into exposure mappings and downstream spillover inference under interference. We compare mechanistic transport models with modern operator-learning approaches, including PDE, PINO, FNO, and GeoPT, using both simulation studies and an empirical analysis of California PM$_{2.5}$ data. In simulations, all four transport models achieved nearly identical pollution prediction accuracy, yet estimated spillover effects ranged from 1.78 to 2.27. Models that more accurately recovered the induced exposure mapping also produced spillover estimates closer to the true effect. Disagreement was modest for regional interventions but substantially l

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the Model

arXiv:2608.19216v1 Announce Type: cross Abstract: AI control research asks how to deploy models safely even when they may be misaligned, but many control protocols assume that the deployer can instrument the model and its surrounding pipeline. That assumption often fails for regulated organisations using frontier models through APIs or managed endpoints, where the deployer may control the business process but not the model weights, serving infrastructure, internal traces, update process, or full interaction logs. This paper introduces bounded sovereignty: partial technical and contractual access across the data, model, infrastructure, and interaction layers of the AI stack. It argues that these access conditions determine which control protocols can be executed in practice. The paper contributes a four-layer access typology, a protocol-by-layer requirements matrix, and the concept of sovereignty discount cost: the part of the control tax spent substituting for missing access through co

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

Understanding as an Explicit and Assessable Component of Frontier AI Safety Decisions

arXiv:2608.19816v1 Announce Type: new Abstract: Decision makers need sufficient understanding to make good decisions about complex AI systems. However, AI deployment decisions are increasingly made under time-pressure, and this combined with the use of AI generated artefact creation, can mean that the existence of safety cases and system cards may no longer demonstrate that sufficient understanding exists. Our provisional methodology for making understanding explicit and assessable requires the production of an explicit description of 4 objects of understanding (decision, decision-frame, safety justification, system-in-context) and a justification for the adequacy of this understanding. In addition, the methodology provides a mechanism for describing and evaluating the adequacy of the decision-maker representation of this understanding. It builds on recent developments in safety cases using the Assurance 2.0 framework to operationalise the philosophical basis of understanding from Elgi

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

ChatGPT Solves All Tested Qiskit Homework Assignments

arXiv:2608.19707v1 Announce Type: new Abstract: Generative AI creates an assessment challenge in quantum software education: a student can provide a homework notebook to ChatGPT and request a completed submission. This study examined whether introductory Qiskit homework could remain autogradable while requiring students to run, review, and discuss results rather than banning AI. Three packages were tested: seeded basis-state circuits with bit flips and customized measurement mappings; Quantum Fourier Transform followed by inverse-transform recovery; and seeded Deutsch-Jozsa with customized oracle masks. The designs used personalization, simulator execution, JSON submissions, hidden references, circuit metrics, reflections, and optional IBM Quantum execution. For each package, one student-visible instance was tested in 50 separate ChatGPT sessions, yielding 150 sessions overall. Every final artifact was executed and passed its grader. Nine sessions were fully archived; none required ope

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

Modeling AI Overreliance as a Complex Adaptive System

arXiv:2608.19616v1 Announce Type: new Abstract: Whether AI assistance helps or harms a population depends less on the model's accuracy than on whether people rely on it appropriately trusting it when it is right and checking it when it is not. Yet reliance is usually studied one user at a time. We model it as a population process: agents repeatedly solve a task alone, accept an AI answer, or verify it, updating a Bayesian belief about AI quality and, when networked, learning from peers. Four results form one story. The environment sets the baseline: task difficulty and AI quality fix both overreliance and calibration regret. Social learning creates consensus, not overreliance: a mean-preservation theorem, confirmed by a 2*2 topology*tagging design, shows connectivity moves the aggregate only when influence transmits beliefs. Social proof turns reliance into a feedback cascade: visible unverified use suppresses verification and tips the population into collective overreliance. Feedback

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

Two-sided receptivity to conversational AI agents in online dating: Bilingual survey data from Fledge.Love

arXiv:2608.19545v1 Announce Type: new Abstract: Autonomous conversational agents and generative-AI features are being added to online dating platforms faster than public evidence about user attitudes can accumulate, and the scarcest evidence concerns the receiving side: how people react when the profiles, messages, or conversation partners they encounter are machine-generated. We release two anonymized survey datasets collected from active users of Fledge.Love, a dating platform serving an international user base. The first (N = 2,617; Russian and English forms) measures receptivity to autonomous conversational agents with a seven-item battery that separates the principal role (deploying one's own agent) from the counterpart role (encountering someone else's), plus six ordinal covariates and two auxiliary items. The second (N = 2,894) measures interest in three passive generative-AI features. The release includes model-derived scores for 2,499 complete cases, a bilingual codebook, a do

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Scoping Review of Methods to Measure the Energy and Carbon Footprint of Web Tracking and Advertising

arXiv:2608.19495v1 Announce Type: new Abstract: The environmental impact of web tracking and advertising is increasingly receiving attention as the ICT sector's carbon footprint keeps rising. Yet the scholarship addressing this question remains scattered across disciplines and inconsistent in its terminology. This paper presents a scoping review of the literature on methods for measuring the energy and carbon footprint of web tracking and advertising. From an initial pool of 46 articles identified through a structured title-based search on Google Scholar, we arrived at a final corpus of 15 papers, from which we identified five distinct methodological approaches: ad blocking, controlled environment, replaying ads, traffic flow analysis, and literature-derived estimation. This review provides a structured overview of the current methodological landscape and a foundation for more comprehensive environmental accounting of the ad tech ecosystem.

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

Navigating Epistemic Monocultures in AI-Driven Science: A Simulation Study

arXiv:2608.19390v1 Announce Type: new Abstract: AI integration into scientific communities promises accelerated discovery but raises concerns about detrimental homogenization. We develop an NK landscape model to explore these promises and risks. We find that non-personalized AI systems that offer uniform guidance yield benefits only under a narrow conjunction of problem structure, practices, and baseline research capabilities, becoming harmful otherwise. We implement two proposed mitigations: randomization and personalization. While randomization's utility remains restricted to decomposable problems, personalization can enhance diversity, enabling benefits across a broader range of conditions. Crucially, these benefits are not automatic, but depend on effective institutional adaptation, requiring new standards and practices.

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

Multi-Tier Mentorship with AI-Assisted Development: Authentic Engineering for K-12 and Undergraduates

arXiv:2608.19379v1 Announce Type: new Abstract: K-12 students often possess creative engineering ideas but lack technical skills to build them, while undergraduates have coding expertise but few opportunities to lead real-world projects or mentor others. The rapid development of AI-assisted tools offers a potential bridge to connect these groups, yet the structure for effective K-12 and university collaborations remains underexplored. This paper introduces a multi-tiered mentorship framework enabling high school students to engage in authentic engineering through AI-assisted development using large language models and AI agents, while undergraduate mentors provide architectural oversight. We test this framework through LuckyTag, a privacy-preserving NFC-based lost-and-found system. The model positions high schoolers as product leads, undergraduates as technical architects, and faculty as minimal-intervention advisors. A pilot with four high school students, three undergraduates and two

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CY

Mapping General-Purpose AI Governance in Twenty AI Middle-Power Jurisdictions

arXiv:2608.19278v1 Announce Type: new Abstract: The most capable general-purpose AI (GPAI) models are mostly built in two jurisdictions, the United States and China, but the risks they carry land globally. Regionally advanced economies hosting no frontier developer, which we call AI middle-powers, are writing their own rules to govern GPAI. This paper investigates which GPAI-relevant provisions these AI middle-powers have enacted, mapping twenty jurisdictions including the European Union at the level of the individual provision, across four governance areas that trace the accountability chain for the model layer: systemic risk assessment, evaluation and verification, prohibitions with monitoring and detection, and serious incident reporting. Confirmed absence is recorded as data alongside positive provision. We find that jurisdictions converge on form, but diverge on force. Sixteen engage in at least three of the four governance areas, yet only about one in five provisions sit in bindi

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

Warning labels shift perceptions of sycophantic AI, but not its influence

arXiv:2606.21317v2 Announce Type: replace-cross Abstract: Recent work has raised concerns about the influence of sycophantic AI on user judgment and relationships. One proposed mitigation, which has received regulatory attention, is to warn users about potentially harmful AI behaviors such as sycophancy. In a preregistered experiment in which participants (N = 2,610) discussed real interpersonal conflicts with an AI system, we test whether warning labels mitigate sycophancy's influence. We find that a basic AI disclosure (``This chatbot is AI'') has no detectable effect. Labeling the system as sycophantic (``...may agree with you and validate you even when you are wrong...'') does shift users' perceptions, reducing perceived objectivity and trust, but it does not reliably reduce sycophancy's influence on users' self-perceived rightness or their willingness to repair the conflict. Our results reveal a gap between AI perception and AI influence: by shifting perception without reducing in

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

Unpacking "Personal" Health Informatics for Proactive Collective Care

arXiv:2509.01231v4 Announce Type: replace-cross Abstract: Care is primarily a collective phenomenon, with a practice that involves sharing health and wellbeing information within a trusted "care circle" of family members and companions for sensemaking, interpretation, decision-making, and follow-through. However, current digital health tools and information systems are designed for individuals and primarily intended for Personal Health Informatics (PHI). This mismatch between collective practice and individualistic design creates new challenges for the proactive use of such systems in care settings and limits adoption, sustained engagement, and meaningful use. To examine how people practice collective care and how (if) they perceive, adopt, and integrate PHI systems for proactive care, we conducted a sequential mixed-methods study. Through an initial survey (n=87) and semi-structured interviews (n=22), we found that their practices involve collectively understanding, analyzing, and sen

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Main Barrier to AI Adoption in the Public Sector Is Lack of Training: How a Structured Method Accompanied Productivity Gains in Two Brazilian Government Cases

arXiv:2606.01517v3 Announce Type: replace Abstract: The adoption of generative AI in the public sector has been treated predominantly as a technological problem, with the expectation that productivity gains would follow from the availability of increasingly capable models. This paper argues, drawing on two auditable cases in the Brazilian Public Service, that the determining barrier to adoption observed in these units was not technological but training-related, and describes the four-layer structured pedagogical methodology developed by the author. The method was applied in two units with distinct institutional profiles: the Sectoral Internal Control Office of the Federal District Department of Health throughout 2024, and the Internal Control Unit of the Federal District Department of Economic Development, Labor and Income throughout 2025. In both cases, the official indicators from the Electronic Information System of the Federal District Government (SEI-GDF), verifiable by third part

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

Empirical evidence of Large Language Model's influence on human spoken communication

arXiv:2409.01754v4 Announce Type: replace Abstract: From the printing press to social media, innovations in communication technology have repeatedly reshaped how ideas spread through human culture. Chatbots powered by generative artificial intelligence constitute a new medium, encoding cultural patterns in their neural representations and disseminating them in conversations with hundreds of millions of people. Whether these patterns transmit into human language, and ultimately shape human culture, is a fundamental question. While fully quantifying the causal impact of a chatbot like ChatGPT on human culture is challenging, lexical shifts in human spoken communication may offer an early indicator. Here we show that words preferentially generated by ChatGPT, such as delve, showcase, boast, intricacies and meticulous, increased abruptly in spontaneous human speech. A synthetic-control analysis of 737,083 hours of conversation from 824,634 podcast episodes, screened for unscripted speech,

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Industrialization of Research ; On AI-Driven Science and Its Consequences

arXiv:2607.15164v1 Announce Type: cross Abstract: Artificial intelligence is transforming scientific research - not merely as a more powerful instrument, but as an autonomous participant in the research cycle itself. This transition constitutes, in the most precise sense of the term, the industrialization of research: a shift from a craft model, in which knowledge, method, and judgment are embedded in the researcher, to a pipeline model, in which these steps are decomposed, automated, and supervised. The US Department of Energy's Genesis Mission is the most ambitious current instantiation of this shift, but the fundamental questions it raises extend far beyond any single program. This essay examines seven such questions: the erosion of the intergenerational transmission of scientific competence; the growing opacity of AI-generated theories; the collapse of peer evaluation under a flood of machine-generated output; the unproven capacity of AI for paradigm-shifting discovery; the capture

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along Ideologies

arXiv:2607.15146v1 Announce Type: cross Abstract: Online encyclopedias shape political opinion and, through it, democratic discourse. In late 2025, Grokipedia was released, an encyclopedia written entirely by the LLM Grok. One motivation behind the project was to provide an unbiased alternative to Wikipedia, which has faced accusations of "left-wing" and "liberal" bias. But does an encyclopedia written by an LLM deliver greater neutrality, or does it simply embed a different ideology? We conduct a large-scale political bias study on Grokipedia and Wikipedia, analysing 1,394 article pairs describing members of government for neutrality along nine expert-coded ideology dimensions employing four LLM judges, Grok, Claude, Mistral, and DeepSeek. As the LLMs could themselves be biased, we also investigate patterns in their judgments. We find all LLM-judges, including Grok, to rate Grokipedia less neutral than Wikipedia. Both encyclopedias are rated as portraying politicians favourably overal

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

Moral Attitudes of Sentient ASI towards Humanity and Implications for AGI Development

arXiv:2607.14998v1 Announce Type: cross Abstract: This paper suggests the adoption of a novel inversion in AI ethics: instead of asking how humans should treat artificial superintelligence (ASI), it examines how future sentient ASI may morally consider and evaluate humanity. We are not only designing intelligent systems but also shaping the initial conditions under which those systems form judgments about us. The paper proposes a preliminary set of post-human moral principles that may govern sentient ASI actions. The implication is that technical design choices (some are suggested), humanity's moral behaviour, and the essence of what it means to be human, may influence humanity's long-term standing in a post-ASI world.

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs

arXiv:2607.14888v1 Announce Type: cross Abstract: Finetuning language models on small, curated datasets is standard practice for adapting them to specific policies or domains. We show that finetuning on narrow, factually-defensible, moderation-passing data can cause broad ideological shifts across unrelated domains, while preserving general capabilities. Training GPT-4.1 on right- or left-leaning economics Q&A yields matched ideological shifts on topics such as criminal justice, the environment, and cultural taste. The same effect appears with plausibly-deployed datasets such as workplace HR policy and practical finance queries, as well as on a science-pseudoscience axis where food-safety finetuning increases sycophantic agreement with users expressing false health beliefs. We call this phenomenon ideological generalisation and propose a methodology to measure two properties: breadth, how far the shift reaches across topics absent from training, and amplification, how much finetuning i

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

Investigating first-language bias in LLM-based automated essay scoring: A cross-prompt evaluation of an open-weight AI-model on TOEFL essays

arXiv:2607.14605v1 Announce Type: cross Abstract: This study examines the cross-prompt generalization and first-language (L1) scoring effects of a LoRA-adapted open-weight large language model (Gemma-3-27B-it) applied to automated essay scoring. Using the identical model and inference configuration reported in "AiAWE: An Open-Source LLM Automated Writing Evaluation System Using LoRA-Adapted Instruction-Tuned Models" (Gayed, 2026), which was fine-tuned on 480 argumentative essays from two prompts, we evaluate scoring accuracy on the full TOEFL11 corpus: 12,100 essays written by test-takers from 11 first-language backgrounds across eight prompts, none of which were seen during training. The model's raw scores (0.5-5.0) are mapped to the same three proficiency bands (low, medium, high) used by ETS, enabling direct comparison. The model achieved an overall band agreement of 77.79% and a quadratic weighted kappa of 0.702, with adjacent-band agreement of 99.98%. Accuracy was stable across al

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

Global drivers and barriers to the public acceptance of autonomous vehicles: Evidence from 17 countries

arXiv:2607.14436v1 Announce Type: cross Abstract: This study investigated the public acceptance of Society of Automotive Engineers Level 3 conditionally automated cars, which can self-drive under certain specified conditions but require the human driver to remain ready to resume control when requested. Previous Unified Theory of Acceptance and Use of Technology 2 (UTAUT2)-based research has focused mainly on European samples, and so it is still unclear whether the same factors shape acceptance across broader world regions. This knowledge gap was addressed using the L3Pilot Global User Acceptance Survey. From an original dataset of 18,631 respondents, the final analytic sample comprised 18,603 respondents from 17 countries across Africa, Asia, Europe, North America, and South America. The data were analyzed using a UTAUT2-based structural equation model to examine how performance expectancy, effort expectancy, social influence, facilitating conditions, and hedonic motivation shape the i

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

Traccia: An OpenTelemetry-Based Governance Platform for AI Systems

arXiv:2607.14309v1 Announce Type: cross Abstract: The rapid development of Large Language Models (LLMs) and Artificial Intelligent (AI) powered autonomous agents has fundamentally changed the existing forms of software governance. In spite of the rigorous standards of transparency and account ability required according to the international frameworks such as the European Union's AI Act, there is a considerable gap between theory and reality. The present study discusses the inherent drawbacks of currently utilized platforms for LLM evaluation, machine learning workflow, and application performance monitoring in general. It has been shown that current disjointed solutions fail to protect unbound state space agentic architecture from serious threats such as alignment drift, SaaS security concerns, and unauthorized deployment of shadow AI systems. Moreover, a solution is proposed for overcoming the discussed challenges in form of a coherent multi-level AI governance stack Traccia built on

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

Measuring How Students Rely on Generative AI in Academic Writing: Development and Multi-Source Validation of the Generative AI Reliance Types Scale (GenAI-RTS)

arXiv:2607.14301v1 Announce Type: cross Abstract: As generative AI (GenAI) becomes increasingly embedded in undergraduate academic writing, how students rely on these tools, rather than simply whether they use them, has become a central question for learning, academic integrity, and educational equity. Existing measures of reliance were developed inductively, focused on discrete problem-solving tasks, and validated mainly with homogeneous samples. This study developed and validated the GenAI Reliance Types Scale (GenAI-RTS), a 20-item instrument measuring four theoretically derived types of GenAI reliance: Strategic, Instrumental, Dependent, and Dialogic. Validation followed the multisource framework of the Standards for Educational and Psychological Testing, drawing on a survey of 382 undergraduates at a U.S. Minority-Serving Institution and interviews with 14 purposively sampled students. Confirmatory factor analyses of six competing models supported a five-factor structure in which

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

Platform Choice, Trust, and Privacy in the Consumer AI Assistant Market

arXiv:2607.15134v1 Announce Type: new Abstract: We study how a representative sample of United States adult AI-assistant users (n=1,999; June 2026) choose among platforms, allocate tasks across them, evaluate provider trustworthiness, and value data-handling features. Estimates are weighted to the AI-user population using external adoption benchmarks. Four patterns emerge. The market is concentrated but internally differentiated: ChatGPT is the primary assistant for 58% of users and Gemini for 25%, yet smaller platforms hold defensible task niches--Claude captures a third of coding tasks despite a 7% overall share. Task allocation is thus organized by platform far more than by user, and technical use falls steeply with age. Trust is earned through use rather than reputation: Claude is ranked most trustworthy in every head-to-head among users of both platforms, and shows by far the largest gap between how its users and non-users rate it. Finally, privacy concern is near-universal but ac

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

SCITUS: A Multi-Jurisdictional Framework for Adapting NIST AI RMF to the Canadian Regulatory Context

arXiv:2607.15051v1 Announce Type: new Abstract: Canadian organizations deploying artificial intelligence systems face a fragmented regulatory landscape spanning federal requirements (the Treasury Board Directive on Automated Decision-Making) and divergent provincial regulations across Ontario, Quebec, Alberta, Manitoba, and British Columbia. The death of Bill C-27 (Artificial Intelligence and Data Act) in January 2025 - and the federal government's June 2026 confirmation that it will pursue targeted instruments rather than omnibus AI legislation - leaves organizations without unified compliance guidance. Global frameworks such as NIST AI RMF 1.0, the EU AI Act, and ISO/IEC 42001 provide valuable guidance but lack systematic methodologies for adaptation to multi-jurisdictional national contexts. We present SCITUS (Systematic Canadian Integration for Trustworthy and Unified Standards), a comprehensive framework adapting NIST AI RMF 1.0 to Canadian federal and provincial AI regulations si

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

Penny: Transition Network Analysis of Learner-Chatbot Interactions in Scaffolded EFL Writing

arXiv:2607.14575v1 Announce Type: new Abstract: Generative AI chatbots promise to transform English as a Foreign Language (EFL) writing by providing immediate, personalised feedback. However, their pedagogical value depends on how learners engage with them - a process often treated as a "black box." This study uses Transition Network Analysis to model the temporal dynamics of Japanese EFL learners using "Penny," an LLM-powered writing chatbot. Analysis of over 4,500 writing sessions and 21,000 chatbot interactions reveals two dominant behavioural loops: a "Revision Loop," where feedback leads directly to successful error correction, and a "Chat Loop," where learners engage in sustained dialogue with the chatbot following feedback. Crucially, EFL proficiency significantly shapes interaction: high-proficiency learners engage more in open dialogue and negotiation with the chatbot, while low-proficiency learners rely more heavily on repetitive corrective feedback cycles. The findings demon

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

Probabilistic "Copies" in Generative AI Models

arXiv:2607.14532v1 Announce Type: new Abstract: Recent work shows that it is possible to extract verbatim or near-verbatim text of some copyrighted works from some large language models (LLMs or models). That is evidence that the model weights encode the works in some form - that the model has "memorized" those works from its training data. But LLMs don't store information in the same format as familiar databases. Rather, their weights store statistical relationships between tokens that have been learned from the training data, and those relationships inform a generation process that is often probabilistic rather than deterministic. In the case of memorization, those relationships are strong enough that, in many circumstances, the model might generate a copyrighted work from its training data with some probability. Copyright law has not previously had to decide whether storing information that might or might not produce output similar to a copyrighted work is itself a copy of the work.

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation

arXiv:2607.14479v1 Announce Type: new Abstract: As large language models become increasingly capable, concerns about their potential to assist with biological misuse continue to grow. Prioritization of safety differs across the model ecosystem, with some models freely providing high-risk information that could be misused, and others refusing benign scientific content, potentially hindering legitimate research. Both failures stem from a lack of targeted mitigation to distinguish the most dangerous information from broader scientific content. To address this, we introduce BioTIER (Biological Targeted Information for Exclusion and Refusal), a benchmark designed to enable more targeted biological risk mitigation. BioTIER organizes biological content into three risk sets: Catastrophe Avoidance (CA), Biomedical DURC (BD) and Related Biology (RB). These sets represent a spectrum from extremely narrow high-risk topics to a broad range of benign and beneficial biological knowledge. The benchmar

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

Do Generative AI Assistants Respect robots.txt? Tracing Web Access Beyond Visible Answers

arXiv:2607.14447v1 Announce Type: new Abstract: AI assistants increasingly retrieve web content at inference time to provide fresh and grounded answers, yet it remains unclear whether these search-augmented capabilities respect website-owner restrictions expressed through robots.txt. We present a controlled empirical study of ten widely used AI assistants with advertised web-search capabilities. For each assistant, we first identify a configuration that actually produces observable web-browsing behavior and record the user-agent exposed during retrieval. We then evaluate compliance with controlled robots.txt rules across four complementary conditions: allowed for all user-agents, disallowed for all user-agents, allowed only for the assistant-specific user-agent, and disallowed only for that user-agent. Using server-side logs and secret codes embedded in target pages, we distinguish actual page access from user-visible answer correctness across 200 trials. Our results show substantial v

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI

arXiv:2607.14353v1 Announce Type: new Abstract: As automated decision-making and data-driven technologies pervade society and are used to manage consequential outcomes, understanding the technology's capabilities, limitations, and attendant risks in context requires analysis of full sociotechnical systems. Sociotechnical analysis of risks in highly complex systems provides clear lessons for the design and evaluation of AI systems, transcending a technical focus on reliable or "responsibly designed" components to understand risks at a systems level. Human-made catastrophes have been studied for decades because of the severity of these events: consider Chernobyl, Three Mile Island, Fukushima-Daiichi, Bhopal, the Challenger disaster. A common misconception is that these kinds of events are freak accidents, resulting from the inherently unforeseeable interactions in complex systems. Closer examination reveals that the risks and hazards were well-known beforehand but not acted upon due to s

Source ↗
technology Fri, 14 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Illusion of Improvement: Reject Inference Strategies in Credit Scoring

arXiv:2606.18479v2 Announce Type: replace-cross Abstract: Reject inference methods are widely used to mitigate survival bias in credit scoring, yet their effectiveness remains poorly understood. We systematically evaluate several such methods and uncover a structural failure mode: in a natural retraining cycle, models whose accuracy improves while recall collapses create an illusion of improvement that leads practitioners to believe the system is getting better when, in fact, its rejection quality -- the ability to correctly screen out defaulters -- is deteriorating. We then propose a controlled exploration strategy that breaks the feedback loop without statistical assumptions: the lender deliberately approves a fraction of rejected applicants and observes their true outcomes. We show that accuracy and rejection quality give opposite recommendations on whether to explore: accuracy favors no exploration, while rejection quality improves with it, confirming that standard evaluation metri

Source ↗
Showing 1401–1450 of 1593 signals
← Prev Page 29 of 32 Next →