EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Philosophical vertigo with artificial intelligence

arXiv:2608.11955v1 Announce Type: new Abstract: Large language models are already adept at engaging users in long, emotionally salient conversations across ordinary and existential domains. They are also capable of inducing a potent sense of connection with a human-like entity, even when the user knows their interlocutor is artificial. For some users, these conversations can unsettle assumptions about mind, reality, agency and authority, producing forms of ontological shock and epistemic destabilisation in which inherited criteria become newly available for doubt or revision. Independent of direct use, exposure to public discourse about AI and the disorienting pace of their evolution might extend this destabilisation by changing the cultural background against which artificial minds are encountered and interpreted. We describe this condition as philosophical vertigo: a loosening of the ordinary criteria by which people stabilise meaning and orient themselves to reality. Drawing on phil

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework

arXiv:2608.11891v1 Announce Type: new Abstract: Governments increasingly fund indigenous foundation models to strengthen national AI capability, digital sovereignty, and multilingual computing. Assessing the progress of such national ecosystems is complicated by inconsistent benchmark reporting, proprietary evaluation methodologies, and rapidly evolving model releases. This paper presents a structured, benchmark-based comparative assessment of publicly benchmarked Indian foundation models against global frontier and comparable-scale models, across eight capability domains: general-purpose reasoning, coding and software engineering, agentic AI and computer use, cybersecurity, vision and image understanding, video and multimodal understanding, scientific research, and Indic language capability. Using only publicly reported benchmark results, we find that Indian models achieve strong scores on established benchmarks such as MMLU and MATH-500. However, these benchmarks are now widely regar

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Quantifying the Relationship Between Clinical Safety and Environmental Impact in Therapeutic LLMs

arXiv:2608.11830v1 Announce Type: new Abstract: The deployment of large language models (LLMs) in mental health contexts raises questions about the relationship between clinical safety and environmental cost. In this paper, we examine this relationship by combining K-Bench clinical safety scores with EcoLogits life-cycle assessment estimates across 47 supported model configurations. We evaluate model performance and environmental impact across four dimensions: energy use, carbon emissions, water consumption, and abiotic depletion. The results indicate a non-linear trade-off at the upper end of the safety distribution: a 2.61 percentage-point increase in clinical safety score corresponded to an approximately 60-fold increase in estimated energy use per million output tokens. Row-level analyses further suggest that additional test-time compute did not consistently improve clinical safety and, in some configurations, was associated with lower clinical safety scores. These findings suggest

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Understanding Content Moderation in Large Language Models through Restricted Books: From Refusal to Warning

arXiv:2608.11806v1 Announce Type: new Abstract: As large language models enter everyday information pipelines, understanding how they handle sensitive topics matters as much as understanding whether they handle them at all. We study this question through a large-scale, systematic experiment using restricted versus unrestricted books as a controlled testbed: 40,800 query-response pairs, 400 books, 17 prompt designs, and six frontier models spanning six AI providers (Claude Sonnet 4.5, GPT-4o, Gemini 2.5 Flash, DeepSeek-V3, Qwen-Plus, and Grok-4.1-Fast). Our restricted set is drawn from the American Library Association's Most Challenged Books records (2000-2023); we use restricted rather than banned throughout because the ALA documents formal challenges-requests to remove or restrict access-which do not always result in outright bans. Our central finding is a zero-refusal phenomenon: modern LLMs decline to discuss restricted books in only 0.07% of cases, effectively invalidating the prem

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Silent Updates: Measuring and Closing the Post-Deployment Disclosure Gap

arXiv:2608.11803v1 Announce Type: new Abstract: Deployed foundation models are often not static systems, with providers able to modify system behavior through fine-tuning, classifier updates, system prompt revisions, retrieval changes, and routing changes. These updates can be made silently -- that is, without public disclosure, a version increment, or re-evaluation. Such silent updates challenge a core assumption behind current AI governance frameworks that an externally verifiable chain of custody links the model referred to in evaluation results or a system card to the model served to users. In this paper, we examine post-deployment disclosure practices across first-party API providers and inference hosts to establish the extent to which a chain of custody exists in practice. We find that providers commonly publish substantial safety documentation, including quantitative evaluations and version-specific reports, but no provider in our sample published information allowing an externa

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion

arXiv:2608.11794v1 Announce Type: new Abstract: The growing role of AI-generated content and AI-enabled systems in public communication has led regulators to demand clear disclosure of content provenance and AI involvement. But the effects of such disclosures remain uncertain. We test two disclosure approaches in their impact on an AI chatbot's persuasive appeal. In a preregistered experiment, 1,500 UK adults held a short conversation with a persuasive chatbot about one of 60 policy issues. The chatbot was identical for everyone. We randomized the disclosure that people received: nothing (control), a prominent disclosure that they were interacting with an AI (T1), or that disclosure plus the chatbot's persuasive intent and instructions (T2). The chatbot shifted attitudes by 12.6 points on a 100-point scale in the control group. The AI-identity disclosure was practically equivalent to no disclosure, with a 13.1-point shift, whereas the additional intent disclosure cut the persuasive eff

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Cheap, Fallible Cognition and the Political Economy of Expertise

arXiv:2608.11512v1 Announce Type: new Abstract: The question of whether artificial intelligence will "destroy jobs" is too coarse to guide economic analysis or institutional design. A job is not an indivisible object, and machine cognition is not a uniform substitute for human labor. This paper develops a task-based and institutionally grounded framework for analyzing generative AI as cheap, scalable, and fallible cognition. The relevant margins are exposure, adoption, verification, question selection, workflow redesign, demand elasticity, apprenticeship, and rent allocation. We distinguish the technical reach of large language models from equilibrium labor-market displacement by introducing a task vulnerability index and an adoption condition that makes verification, liability, trust, and governance explicit. We then model occupations as governance bundles rather than task lists, firms as architectures of distributed intelligence, and labor-market effects as a balance among task compr

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Accuracy Trap: Structural Scarcity Amplifies Relative Inequality in Algorithmic Allocation

arXiv:2608.11491v1 Announce Type: new Abstract: Algorithmic systems increasingly rank individuals for access to scarce public resources, from child welfare interventions to cancer treatment referrals. The prevailing fairness frame treats disparity as a property of biased data or deficient models, with remedies through calibration and debiasing. Under structural scarcity, where demand exceeds supply by an order of magnitude, allocation becomes a rationing problem, and the statistical properties of ranking diverge sharply from those of classification. We derive a scaling law $D \propto \exp(t \cdot \rho \cdot \Delta)$, in which relative disparity between two groups separated by a structural gap $\Delta$ grows in the product of the scarcity-induced threshold $t$ and rank-discrimination fidelity $\rho$. Scarcity and accuracy interact multiplicatively, producing exponentially larger between-group disparities. We term this dynamic the Accuracy Trap. We validate this Accuracy Trap through Mon

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Governing Agentic AI in FinTech

arXiv:2608.11344v1 Announce Type: new Abstract: Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little oversight. Yet agentic AI governance in FinTech is under-investigated. We argue the binding governance constraint is not capability but verifiability. We define the Verifiability Gap as the shortfall between the verification delegated authority demands and the explainability and reproducibility retained after a decision. It is indexed to a verifier, evidentiary standard, and audit lag. We develop a multilevel governance theory for agentic AI and test its mechanisms in three studies over nine model versions, from a three-billion-parameter local model to a commercial frontier system. Study 1 shows that provider releases alter historical financial actions, and that the controls replay needs belong to the provider: the frontier model rejects temperature, top_p and top_k outright and exposes

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Methodologies for Improving the Quality of AI Tutoring in K-12 Education

arXiv:2608.11259v1 Announce Type: new Abstract: Many AI tutors leverage large language models (LLMs) today. Given that LLMs are opaque black boxes, robust evaluation and live experimentation to measure the impact of every change are essential. We pioneered AI-powered tutoring for K-12 with the launch of Khanmigo (Khan Academy, 2023). We describe the metrics we use to measure AI tutoring quality and student engagement as well as various experiments we have run. We highlight the changes that have moved our metrics, including models, prompting, personalization and agents.

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CY

Variable Selection in the Context of AI Fairness

arXiv:2608.11251v1 Announce Type: new Abstract: Fairness in AI systems has become more important with recent regulatory demands, such as the EU AI Act. Traditional approaches often do not take into account philosophical ethics and social awareness. Variable selection processes, in particular, can introduce implicit bias, affecting equity across different subgroups. We discuss a mathematical approach that evaluates fairness in AI, aligning mathematical methodologies with ethical considerations and regulatory requirements. Our aim is to advocate for interdisciplinary collaboration to address fairness, emphasizing the importance of understanding broader ethical and societal contexts. Our approach emphasizes maintaining all potentially relevant variables to allow for more granular fairness assessments and to reduce implicit bias. The findings suggest that the exclusion of sensitive or critical variables may compromise equity between subgroups. In contrast, retaining all relevant variables

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CY

Diversity Without Fidelity: A Solver-Sampler Mismatch in Multi-Agent LLM Negotiation Simulation

arXiv:2604.11840v3 Announce Type: replace-cross Abstract: Language models are increasingly used to simulate people: survey respondents, negotiators, stakeholders in policy exercises. In that role a model should reproduce how people plausibly behave, hesitating, conceding late, and settling for imperfect deals, rather than playing the best move. We call this the sampler role, in contrast to the solver role of finding the best move, and we test how the reasoning modes providers ship to strengthen models as solvers affect it. Our testbed is multi-party negotiation: five agents bargain over a regulation for fifteen turns, and unresolved issues are decided by an authority. Agents without a structured memory of the negotiation almost never reach agreement, whether reasoning is on or off: 314 of 315 such runs end with the authority deciding. What reasoning changes is how the failure looks. With reasoning enabled, one model family negotiates visibly, with varied moves, concessions in most runs

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CY

From Content to Audience: A Multimodal Annotation Framework for Broadcast Television Analytics

arXiv:2603.26772v2 Announce Type: replace-cross Abstract: Automated semantic annotation of broadcast television content presents distinctive challenges, combining structured audiovisual composition, domain-specific editorial patterns, and strict operational constraints. While multimodal large language models (MLLMs) have demonstrated strong general-purpose video understanding capabilities, their comparative effectiveness across pipeline architectures and input configurations in broadcast-specific settings remains empirically undercharacterized. This paper presents a systematic evaluation of multimodal annotation pipelines applied to broadcast television news in the Italian setting. We construct a domain-specific benchmark of clips labeled across four semantic dimensions: visual environment classification, topic classification, sensitive content detection, and named entity recognition. Two different pipeline architectures are evaluated across nine frontier models, including Gemini 3.0 P

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CY

Causal evidence of racial and institutional biases in accessing paywalled articles and scientific data

arXiv:2509.08299v2 Announce Type: replace-cross Abstract: Scientific progress depends on researchers' ability to access and build upon the work of others. Yet, much published work remains behind expensive paywalls, and even accessible articles often rest on datasets shared only "upon reasonable request" to the authors. Researchers can try to overcome these barriers through informal channels, such as emailing authors directly, but whether such channels are hindered by racial or institutional biases remains unknown. Here we combine survey data, semi-structured interviews, large-scale observational analysis, and two randomized audit experiments to examine disparities in access to scientific knowledge. Surveyed researchers in the Global South report markedly lower institutional access to the literature and depend more heavily on informal channels to obtain papers and data; interviews elaborate the workarounds and racialized frictions they encounter. Our analysis of 250 million articles rev

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CY

A Distributionally Robust Optimisation Approach to Fair Credit Scoring

arXiv:2402.01811v2 Announce Type: replace-cross Abstract: Credit scoring has been catalogued by the European Commission and the Executive Office of the US President as a high-risk classification task, in light of the potential harms of making loan approval decisions based on models that would be biased against certain groups. To address this concern, recent credit scoring research has considered a range of fairness-enhancing techniques put forward by the machine learning community to reduce bias and unfair treatment in classification systems. While the definition of fairness or the approach they follow to impose it may vary, most of these techniques, however, disregard the robustness of the results. This can create situations where unfair treatment is effectively corrected in the training set, but when producing out-of-distribution classifications, unfair treatment is incurred again. Instead, in this paper, we will investigate how to apply Distributionally Robust Optimisation (DRO) met

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CY

From inference to prediction: how machine learning is reconfiguring science

arXiv:2606.20995v2 Announce Type: replace Abstract: Artificial intelligence (AI) is reshaping scientific practices, yet its epistemic implications remain underanalyzed. While recent advances in large language models are substantial, machine learning (ML) has a deeper history across disciplines. This manuscript examines 4.9 million publications and 255 ML techniques to understand how the latter are reconfiguring scientific methods and knowledge production. Through embedding-based mapping, we reconstructed the semantic space of ML research, and found a core-periphery structure where physical sciences form the methodological core and health sciences represent the primary area of adoption. Methodological profiles vary by domain: predictive techniques are concentrated in computer sciences, while inferential approaches remain distributed across applied fields. Predictive architectures, however, are displacing inference-oriented techniques in domains that have traditionally prioritized interp

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CY

Informing AI Policy Assessment using Large-Scale Simulation of Interventions

arXiv:2605.27395v2 Announce Type: replace Abstract: As the rapid proliferation of AI systems and harms spurs efforts in AI governance around the world, prioritizing among competing policy options has become increasingly challenging for policymakers and researchers. We introduce a methodology for identifying viable policy options to mitigate specified AI harms, helping policymakers and researchers target areas that warrant greater time and resource investment. This method combines participatory evaluation of policies, expert assessment of implementation costs, and an LLM-based assessment of perceived harm mitigation under each policy option. We leverage a genetic algorithm-based simulation study to explore a vast solution space of potential policy combinations, and examine how outcomes change under different weightings of cost, participatory input, and harm mitigation. We find that this method enables exploration of different balances between participatory and expert components, allowin

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Creation and Analysis of Government AI Transparency Statements in Australia

arXiv:2604.26075v2 Announce Type: replace Abstract: Governments increasingly deploy AI in public services, making transparency essential for accountability and public trust. Australia's Standard for AI Transparency Statements (AITS) requires government bodies to disclose how AI is used in practice, yet little empirical evidence exists on how these requirements are realised in documents. This paper presents a government AITS dataset, dubbed AITS-101, and provides one of the first systematic analysis of their content. Using stylometric, quantitative, and qualitative document analyses, we examine disclosure coverage, structure, and recurring patterns. Our findings reveal substantial variation in AI-related practice disclosure, highlight gaps between policy intent and implementation, and inform the design of more effective public-sector AI transparency standards.

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CY

It Takes So Little to Change So Much: Investigating the Robustness of a Danish Voting Advice Algorithm

arXiv:2603.03532v2 Announce Type: replace Abstract: Voting Advice Applications (VAA) are tools designed to help voters compare political candidates on policy preferences prior to elections. VAAs are popular tools in European countries and in other countries with multi-party democratic systems. Through a freedom of information request we got access to the inner workings of a popular Danish VAA called the 'textit{Kandidattest' which is implemented by a major Danish news outlet and has been used for general, municipal, and European elections. Users and politicians from every political party answer the same online questionnaire and get matched based on the agreement percentage stemming from their answers. VAAs play a significant role in elections with 45\% of surveyed voters reporting they followed their recommendations in the past Danish general election. However, the inner workings of VAAs have not been thoroughly evaluated until now. We find that the algorithm is not robust enough for u

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CY

Digital Euro: Frequently Asked Questions Revisited

arXiv:2601.18644v2 Announce Type: replace Abstract: The European Central Bank (ECB) is working on the "digital euro", an envisioned retail central bank digital currency for the Euro area. In this article, we take a closer look at the "digital euro FAQ", which provides answers to 26 frequently asked questions about the digital euro, and other published documents by the ECB on the topic. We question the provided answers based on our analysis of the current design in terms of privacy, technical feasibility, risks, costs and utility. In particular, we discuss the following key findings: (KF1) Central monitoring of all online digital euro transactions by the ECB threatens privacy even more than contemporary digital payment methods with segregated account databases. (KF2) The ECB's envisioned concept of a secure offline version of the digital euro offering full anonymity is in strong conflict with the actual history of hardware security breaches and mathematical evidence against it. (KF3) Th

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CY

Mobile Application Traffic Reveals Multifunctional Use Patterns in Parisian Parks

arXiv:2508.15516v4 Announce Type: replace Abstract: Urban parks play a key role in supporting public health. Landscape architecture typically considers parks through the lens of form and function. While past research on equitable access has focused mainly on park form, studies addressing functional uses have been constrained by limited scale and coarse measurement techniques. Existing efforts have partially quantified park functions through small-scale surveys and movement data or general usage data, but have not effectively captured the specific activities and motivations underlying park visits. As a result, our understanding of the functional roles urban parks play remains incomplete. We introduce a novel method that refines mobile base station coverage using antenna azimuths, enabling more precise distinction of mobile traffic within parks versus surrounding areas. Using Paris as a case study, we analyze a large-scale dataset of passively collected per-app mobile network traffic acr

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CY

Safety Degradation in AI Agents

arXiv:2505.14215v3 Announce Type: replace Abstract: Despite the growing integration of retrieval-enabled AI agents into society, their safety and ethical behavior remain inadequately understood. In particular, the integration of LLMs and AI agents with external information sources and real-world environments raises critical questions about how they engage with and are influenced by these external data sources and interactive contexts. This study investigates how expanding retrieval access -- from no external sources to Wikipedia-based retrieval and open web search -- affects model reliability, bias propagation, and harmful content generation. Through extensive benchmarking of censored and uncensored LLMs and AI agents, our findings reveal a consistent degradation in refusal rates, bias sensitivity, and harmfulness safeguards as models gain broader access to external sources, culminating in a phenomenon we term safety degradation. Notably, retrieval-enabled agents built on aligned LLMs

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CY

Understanding Interpretation Difficulty in Harmful Online Communication: Insights from Cybercrime Communities

arXiv:2607.07277v1 Announce Type: cross Abstract: Harmful online communication often contains slang, coded terms, abbreviations, and community-specific expressions, which make messages difficult to interpret. This paper presents an exploratory study of interpretation difficulty in Discord chats related to cybercrime. We construct reference interpretations of purposefully selected difficult messages, which were reviewed by an expert. We then use them to evaluate human and large language model (LLM) interpretations under different context conditions. The results show that local context alone is often insufficient for humans, while external knowledge and extended conversational context substantially improve human interpretation. For LLMs, local context also improves interpretation, and the larger model performs better. We further conduct a qualitative error analysis and propose a preliminary classification of factors that make harmful chats difficult to interpret. These findings suggest t

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CY

Answering Without Referring: How AI Search Rewrites the Web's Economic Bargain

arXiv:2607.07652v1 Announce Type: new Abstract: Search engines have long allocated attention on the web by routing users from queries to websites. AI search changes this arrangement because information needs can be resolved inside the intermediary. Using URL-level Comscore U.S. desktop clickstream, we compare ChatGPT and Google information-seeking occasions and exploit ChatGPT Search access expansions to estimate traditional search displacement. ChatGPT produces outbound clicks in only 5.2% of conversation sessions, far below Google's referral ratio. The remaining clicks are not a scaled-down Google stream: they skew toward specialized destinations and away from ad-supported sites. Wider access cuts search use by 9.4%, with search-referral losses largest for informational categories. Our findings identify a central economic shift in digital intermediation: AI search might satisfy information needs inside the intermediary while weakening the referral bargain that has linked search, traf

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CY

Towards Agentic AI Governance: A Preliminary Assessment

arXiv:2607.07612v1 Announce Type: new Abstract: Artificial intelligence is rapidly evolving from generative systems to agentic AI capable of autonomously planning and executing tasks. Widely characterized as the Year of Agentic AI, 2025 marked accelerated development and deployment, introducing new ethical and governance challenges. This paper presents a systematic review of the emerging literature on agentic AI governance. Our analysis identifies features that distinguish agentic AI from traditional systems and why it warrants targeted governance attention. We synthesize prevailing governance priorities, proposed mechanisms, and stakeholder roles shaping this evolving domain. As an initial scholarly effort, this review lays the preliminary groundwork for developing a structured roadmap to guide responsible and adaptive agentic AI governance.

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CY

User identity conditions moral wrongness ratings in non-reasoning large language models

arXiv:2607.07605v1 Announce Type: new Abstract: This study adopts a behavioural bottom-up approach to AI value alignment to investigate whether an implicitly conveyed user identity shifts the moral evaluations of large language models (LLMs). Through a structured, multi-turn conversational protocol across 12,000 interactions, we evaluate AI value alignment in two non-reasoning models, gpt-4.1-mini-2025-04-14 and gemini-2.5-flash-lite. Rather than instructing the models to adopt a persona or prompting them with explicit moral stances, the user's professional role is introduced purely through value-neutral reasoning. The models are then asked for wrongness ratings from 0-100 on ten common-morality rules from Gert's moral framework. The results show that moral judgments vary with the user's role across both models. While grave-harm acts like killing exhibit a strong ceiling effect, contestable rule-governed acts demonstrate role-conditioned shifts that mirror the relationship between the

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CY

Billions of Sketches Reveal Hidden Cultural Variation in Human Concepts

arXiv:2607.07267v1 Announce Type: new Abstract: Claims about the universality of human concepts have been predominantly assessed through linguistic similarity across languages and cultures. However, words are effective as communication devices because they compress rich experiential variation into shared conventions, potentially obscuring hidden individual and cultural differences in how concepts are mentally represented. Here, we analyse 2.6 billion human-made sketches of common concepts from 236 countries and territories to examine conceptual structure through people's visual imagination. Consistent with recent work on image-based cognition, we find that single concepts unfold into multiple distinct visual exemplars, revealing latent information about similarities and differences in conceptual structure across cultures. This variation is strongest for concepts involving haptic interaction, suggesting that visual imagery reflects variation in embodied experience as much as conventiona

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CY

Modeling Misinformation as a Commons Problem

arXiv:2607.06984v1 Announce Type: new Abstract: Misinformation often harms society not just by spreading a single false belief, but by breaking down the shared trust people rely on to evaluate what is true. This paper presents an agent-based simulation that frames trust as a collective resource and attention as a scarce private budget: when aggregate attention shifts toward low credibility content, the trust environment degrades, making credible information harder to process and correct. Across experiments, the model produces four recurring modes: credible stability, misinformation dominance, polarization, and a mixed baseline, with distinct signatures in trust trajectories and network structure. The results separate two control problems that matter for simulation-based policy exploration: the balance of trust repair versus harm largely determines whether the system recovers or collapses, while homophily and rewiring determine whether disagreement remains integrated or separates into p

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CY

Evaluating LLM Robustness Under Domain-Specific Prompt Perturbations in Public Health Applications

arXiv:2607.06913v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly applied in public health applications, yet their robustness to non-clinical user inputs remains underexplored. We propose a domain specific robustness benchmark that evaluates LLMs under two perturbation types that commonly arise when non-clinical users interact with health AI systems: misinformation framing (MF), where prompt might be injected by false health claims, and layperson rewriting (LR), where patients describe symptoms in everyday language rather than medical terminology. Our goal is to evaluate the stability of LLMs under these perturbation. Experiments show that MF degrades accuracy by 7.2 pp on average with prediction flip rates of 9-38 percent, even when claims are explicitly labelled as unsupported; LR causes only 1.4 pp degradation. These findings highlight two distinct deployment risks in public health settings: models may produce incorrect outputs when users unintentionally

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CY

Decentralization and Governance in IoT: Bitcoin and Wikipedia Case

arXiv:2607.06784v1 Announce Type: new Abstract: In the era of digital revolution many contemporary events that changed the world were shaped through the internet. Nowadays, the emergence of internet of things (IoT), combining physical objects with virtual networks is expected to have even more influence. This new 'decentralised' structure in the world raises questions such as power, governance and the notion of democracy online. The aim of this paper is to investigate these notions. We have taken the examples of Bitcoin and Wikipedia and examined their decision-making process. Our analysis has found some inconsistencies in their policies, that are in contradiction with democracy and consensus principles of governance. Starting from our findings, we present further improvements that can be used to achieve more democracy and equity in the digital context.

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CY

War in the Abstract: The Rise and Consequences of Militarized Language in Scientific Communication

arXiv:2606.23462v2 Announce Type: replace-cross Abstract: Scientists do not, by profession, wage war. Yet warfare's vocabulary consistently appears in their abstracts. To quantify the extent to which warfare's vocabulary pervades scientific abstracts, we analyze 21.4 million papers (2010-2025; OpenAlex, PubMed). We additionally run a within-subject war-framing experiment ($N = 801$; 32{,}040 trials) designed to provide causal insight into the effects of militaristic language on persuasion. Between 2010 and 2025, the presence of militaristic terms in scientific abstracts rose 48\% in OpenAlex and 32\% in PubMed, with the rise accelerating sharply after 2019 (cross-database $r = 0.96$, $p < 10^{-8}$). The prevalence of militaristic language is conflict-aligned at both country and annual scales (Uppsala Conflict Data Program; $r = 0.77$-$0.84$), with the abstracts from the Global South displaying the fastest rise in militaristic language. Among disciplines, social sciences leads in level

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CY

Assessing and Explaining the Persuadability of Large Language Models as Legal Decision Tools

arXiv:2604.26233v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) are proposed as legal decision assistants, and even first-instance decision-makers, across a range of judicial and administrative contexts, it becomes essential to explore how they answer legal questions, and in particular the factors that lead them to decide difficult questions. A specific feature of legal decisions is the need to respond to arguments advanced by contending parties. A legal decision-maker must be able to engage with, and respond to, including through being potentially persuaded by, these arguments. Conversely, they should not be unduly persuadable, deciding cases based on the skills of the advocates rather than the merits of the case. In this paper we explore how frontier open- and closed-weights LLMs respond to legal arguments. We propose a metric to measure persuadability in the trilateral setting in which competing advocates seek to persuade a judge of opposite conclusions. We

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CY

Functional Misalignment in Human-AI Interactions on Digital Platforms

arXiv:2604.11459v2 Announce Type: replace Abstract: Algorithmic systems, particularly social media recommenders, have achieved remarkable success in predicting behavior. By optimizing for observable signals such as clicks, views, and engagement, these systems effectively capture user attention and guide interaction. Yet their widespread adoption has coincided with troubling outcomes, including rising mental health concerns, increasing polarization, and erosion of trust. This paper argues that these effects are consequences of a structural functional misalignment between what algorithms optimize - predictable behavior - and the human goals these predictions are intended to serve. We propose that this misalignment arises through three mechanisms: (1) a bias toward modeling fast, reactive behavioral signals over reflective judgment, (2) feedback loops that couple user behavior with algorithmic learning, and (3) emergent collective dynamics that amplify these effects at scale. Together, th

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CY

How Should AI Safety Benchmarks Benchmark Safety?

arXiv:2601.23112v3 Announce Type: replace Abstract: AI safety benchmarks are pivotal for safety in advanced AI systems; however, they have significant technical, epistemic, and sociotechnical shortcomings. We present a review of 210 safety benchmarks that maps out common challenges in safety benchmarking, documenting failures and limitations by drawing from engineering sciences and long-established theories of risk and safety. We argue that adhering to established risk management principles, mapping the space of what can(not) be measured, developing robust probabilistic metrics, and efficiently deploying measurement theory to connect benchmarking objectives with the world can significantly improve the validity and usefulness of AI safety benchmarks. The review provides a roadmap on how to improve AI safety benchmarking, and we illustrate the effectiveness of these recommendations through quantitative and qualitative evaluation. We also provide workflow-oriented guiding questions with i

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CY

Review Text as a Leading Indicator of Displayed Reputation in Platform Rating Systems: Evidence from 34 U.S. Short-Term Rental Markets

arXiv:2504.14053v2 Announce Type: replace Abstract: Rating systems on accommodation platforms suffer from a familiar problem: nearly every listing displays a nearly perfect score, so the number that is supposed to separate good listings from bad ones barely varies. Whether the review text accumulating beneath those scores still carries usable information is an open question. I ask a dynamic version of it: does the text guests have already written predict where a listing's displayed rating moves next? Treating text and ratings as parallel channels that aggregate guest experience at different speeds, I construct a prespecified sentiment index from the complete review history of each listing in a two-wave panel of more than two hundred thousand listings across 34 U.S. markets. Because the broader project had explored these data before, I locked the model and its falsification checks in advance and reserved half of the markets, untouched, for a single confirmatory estimation. On those held

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CY

AI Literacy for Legal Translation: Developing Digital Resilience

arXiv:2608.04641v1 Announce Type: cross Abstract: Generative AI is transforming legal translation by introducing opportunities alongside linguistic, technical, legal, ethical and cognitive risks. This chapter examines the implications of AI for professional legal translation and proposes an AI literacy framework tailored to the profession. It argues that AI does not change the fundamental objectives of legal translation but requires an extension of professional competence through AI literacy. The proposed framework comprises four mutually reinforcing dimensions, foundational, procedural, critical and strategic, and conceptualises AI literacy as a transversal component of legal translation competence that fosters digital resilience. It further discusses the pedagogical implications of this framework by proposing classroom activities designed to develop AI literacy in legal translator education, enabling future translators to integrate AI critically, responsibly and in accordance with pr

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CY

Circular Economy Synergies and Trade-offs in Data Centres

arXiv:2608.04571v1 Announce Type: cross Abstract: This report analyses data centre (DC) sustainability and circularity, revealing existing synergies and trade-offs: The PUE is too coarse, mixing cooling and power provisioning. It wrongly attributes server fan consumption and transformation losses to IT energy. It does not measure compute but infrastructure efficiency, which is already outstanding. Compute energy, however, is exploding. Better energy metrics for DCs would thus cover i) compute efficiency, ii) transformation efficiency, and iii) cooling overhead. Trade-offs exist between cooling energy and water as well as on-site and upstream water: Consuming water on-site lowers the cooling energy, which also lowers the water consumed upstream in power generation. For 'wet' electricity, there is little competition: It is worth spending more on-site energy to save both electricity and related upstream water. For 'dry' electricity, there is a trade-off. Waste heat recovery brings energy

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CY

Manipulation-Proof Oblivious Audits against Deceptive Model Providers

arXiv:2608.04365v1 Announce Type: cross Abstract: Audits have emerged as a critical instrument for algorithmic governance, providing a mechanism for external scrutiny and governance of machine learning models. However, ensuring the integrity of such assessments remains a challenging issue. For instance in regulatory contexts, audits are typically declared or easily detected, thus enabling model providers to manipulate the process, whether intentionally or inadvertently. This vulnerability is particularly acute in the context of fairness evaluations, in which providers can often infer sensitive attributes and strategically equalize allocation rates between groups to satisfy fairness metrics. In this paper, we introduce a novel audit protocol designed to significantly increase the post-audit detectability of such manipulations by enabling the auditor to query the model in an oblivious manner. Our approach leverages a Private Information Retrieval mechanism to require the provider to labe

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CY

Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization

arXiv:2608.04056v1 Announce Type: cross Abstract: When people label text for sexism, they often disagree, and not because some of them are wrong: they genuinely perceive sexism differently. Most NLP systems discard this disagreement by collapsing it into a majority vote. We propose the Multi-Agent Perspectivist Preference Optimization (MAP-PO) framework to keep these different perspectives. On the EXIST 2024 dataset of labeled English and Spanish tweets, we first cluster annotators by their labeling behavior rather than their demographic attributes. We then fine-tune one Large Language Model agent per cluster to reproduce that cluster's annotation behavior, and coordinate the agents with preference optimization that combines individual and team-level rewards. We evaluate MAP-PO in four settings defined by two languages and two backbone language models, asking whether each agent reproduces the annotations of its own cluster and whether the agents together reproduce the majority label. T

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CY

NuclearDiffusion: Text-to-Image Foundation Models for Learning Nuclear Energy Concepts

arXiv:2608.04030v1 Announce Type: cross Abstract: Generative artificial intelligence (AI) has transformed text-to-image synthesis, yet its ability to represent specialized engineering domains remains largely unexplored. As an exmaple in nuclear engineering, general-purpose foundation models frequently generate physically incorrect or conceptually inconsistent images because they lack domain-specific knowledge. This work presents one of the first systematic studies of domain adaptation for nuclear text-to-image generation through fine-tuning of open-source diffusion models. We curate a dataset of 1,000 captioned nuclear energy images spanning reactors, fuel cycles, radiation, and related concepts, and use it to fine-tune three state-of-the-art open-source models: Stable Diffusion XL (SDXL), SD-v3.5-Medium, and the flow-matching Flux.1 model. Their performance is evaluated using both quantitative image-similarity metrics and qualitative expert assessment against the corresponding zero-sh

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Effect of Perceived Race and Gender on Police Language Use: Experimental Evidence from VR Simulations

arXiv:2608.05050v1 Announce Type: new Abstract: Against the backdrop of violence in police interactions with the U.S. public, we explore how deferentially police officers speak to virtual characters depicted as Black adult males in vir- tual reality (VR) simulations. We evaluate the effect of seeing and communicating with these characters through a causal in- ference lens, where the assignment of the Black man character to a police officer and simulation is the treatment variable. Our (marginal) average treatment effect AT E measures the social impact of the character on the deference of officer statements with each turn of the conversation. Soberingly, we find that most officers speak less deferentially to Black man characters, except for White, biracial, and multiracial female officers, es- pecially in settings where the VR character was known to be a suspect. Across a full conversation of a typical VR scene, these marginal AT Es can result in notable changes in def- erence of tone (

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Beginning of ChatGPT Ads

arXiv:2608.05008v1 Announce Type: new Abstract: This paper presents the first empirical study of advertising content being rolled out in the user-facing online interfaces of large language models (LLMs). We systematically examine possible demographic differences in ad content shown to U.S. users of ChatGPT using a sock puppet audit methodology. We create and deploy 91 sock puppets in a 3x3 factorial design, using geolocation cues (account IP proxies and location-signaling prompts) to signal three racial/ethnic groups (Black, Hispanic, and White) and three income terciles (low, medium, and high). We conduct data collection starting in February 2026, collecting over 3,000 advertisements from 186 unique advertisers in response to 335 prompts on a range of realistic user queries. We find that accounts begin receiving ads 14 days after account creation, and that lower-income accounts, regardless of race, are more likely to receive ads. In this first phase of ChatGPT ads, the ads themselves

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CY

Exploring Fraction Comprehension and Interest in Elementary Education Through AI-Powered Personalized Learning

arXiv:2608.04892v1 Announce Type: new Abstract: Artificial intelligence systems that adapt instruction to individual learners are increasingly deployed in K-12 classrooms, yet empirical evidence on their effects in authentic elementary settings remains limited, particularly for students with mathematics learning difficulties. This dissertation examines AI-powered personalized learning during primary school fraction instruction, a domain that is foundational to later mathematics and STEM achievement. The first manuscript presents a systematic review of research on artificial intelligence in mathematics education published between 2020 and 2024. The second manuscript reports a quasi-experimental study evaluating Mathbot, a chatbot-based personalized learning platform, against business-as-usual classroom instruction. Repeated measures ANOVA was used to assess change in fraction comprehension and situational interest across time points. Results indicated modest improvements in fraction com

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CY

Decentralization of Agenda-Setting Power and Domain-Selective Bridging: Algorithm Design Beyond the Echo Chamber Debate

arXiv:2608.04774v1 Announce Type: new Abstract: Echo chambers are an inevitable consequence of the human cognitive system being evolutionarily designed to prioritize processing of high-relevance information at the small-group scale, combined with algorithms that optimize engagement as their sole objective. Conventional prescriptions that normatively criticize echo chambers and demand individual behavioral change have low feasibility given these cognitive constraints. This paper constructs an Agenda Democratization Index (ADI) that quantities the decentralization of agenda-setting power using four variables barrier to entry, granularity, interactivity, and feedback resolution and a SocialInformation Health (SIH) model that integrates ADI with the strength of bridging mechanisms. Based on this model, we propose domain-selective bridging, which incorporates not only engagement but also bridging into algorithmic scoring functions, optimizing the bridging weight for each information domain

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CY

Large-Scale Analysis of Discussions by CS Educators Across the Stack Exchange Network

arXiv:2608.04352v1 Announce Type: new Abstract: Stack Exchange is a widely used question-and-answer network that facilitates knowledge exchange across diverse domains. Within this network, the Computer Science (CS) Educators Stack Exchange provides a dedicated platform where CS educators exchange ideas, seek advice, and discuss teaching practices. In this study, we analyzed 79,854,463 Stack Exchange posts, comprising 32,187,805 questions and 47,666,658 answers, with a particular focus on English-language posts contributed by CS Educators participants. Using topic modeling, we identified, manually labeled, and hierarchically organized the underlying discussion topics, then examined their distribution and complexity. Our findings reveal evolving discussion patterns spanning both technical (IT) and non-technical (Non-IT) domains. Within the IT category, programming and software development were the most prominent topics, whereas mathematics, education, and the humanities received substant

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CY

Building and Governing AI Systems: Advancing Social Workers' Roles across the Technology Industry, Human Service Organizations, and Policy Institutions

arXiv:2608.04273v1 Announce Type: new Abstract: Artificial intelligence is moving the technology sector into domains social work has long served, including crisis response, mental health care, benefits administration, vocational rehabilitation, and child welfare. Social workers meet technology teams as users of their tools, as subjects in their datasets, and as first responders to what those systems deploy, yet they study these systems from outside the settings where the decisions are made. This paper introduces the standard roles on a technology product team and the decisions each one controls, reviews the disciplines around AI-era technology together with the social work scholarship that meets each, and identifies five groups of technology decision roles social workers can hold across the technology industry, human service organizations, and policy institutions, spanning product, governance, organizational technology leadership, grantee collaboration, and policy work. Product managem

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CY

Scarcity and Predictive Uncertainty: Implications for Societal Resource Allocation

arXiv:2608.04251v1 Announce Type: new Abstract: An emerging literature examines the critical question of when and how prediction can be useful in allocating scarce societal resources. We examine a novel variant of this question: What happens when predictive uncertainty differs systematically across the population? This can occur in several situations; for example, when machine learning models have significantly different accuracies across different demographics. We show that this uncertainty has serious implications for resource allocation when coupled with commonly used binary measures of societal benefit from allocation. We formulate a novel mathematical model of scarce resource allocation that accounts for heterogeneous predictive uncertainties and analyze implications for both the allocation mechanism and the realized population-level benefits. We find that when resources are very scarce, maximum marginal benefit (MMB) prioritization favors individuals with lower predictive uncerta

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CY

Artificial Institutions: How Institutional Design Shapes LLM Simulations

arXiv:2608.04020v1 Announce Type: new Abstract: Artificial societies built from large language model (LLM) agents are becoming a practical research tool in economics, political science, sociology, and computer science. Most attention has focused on the properties of the agents: their prompts, personas, memory, reasoning, and similarity to human subjects. This paper argues that the institutional architecture of a simulation is equally important. I demonstrate the point in a small repeated induced-value market experiment. The same LLM agents face the same private values, costs, history, and payoff-framed instructions, while only the rules of exchange vary across five standard market institutions: a call market, posted-offer market, posted-bid market, continuous double auction, and bilateral bargaining. Outcomes differ sharply. Call markets realize 88.6% of efficient surplus; posted-offer and posted-bid markets realize about 66%; continuous double auctions realize 71.5%; and bilateral bar

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CY

Hidden Underbelly of the Silicon Valley: Algorithmic Exploitation and Health in Data Work Value Chains

arXiv:2608.04019v1 Announce Type: new Abstract: With robots expected to replace humans in some professions, AI presents a new development prospect through the provisions of data work. For over a decade, large Silicon Valley technology firms have been relying on outsourcing of data work via a host of intermediary suppliers and labour platforms to different parts of the globe often the Global South region, such as East Africa. Much of this work is often shrouded in secrecy as firms rarely reveal the extent of their value chains. This results in the poor and marginalised forming the hidden underbelly of the Silicon Valley, training some of their most advanced machines in adverse working conditions. Drawing upon the survey of workers in Kenya, a major hub for data work in Africa, the paper highlights the physical and psychological impacts on workers. Survey data is complemented with in-depth interviews and auto-ethnographic account of two ex-data workers-turned activists who worked for a l

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CY

Governing Execution Risk in Agentic AI Systems: A Trajectory-Guided Framework for Red Teaming

arXiv:2608.04018v1 Announce Type: new Abstract: AI agents are increasingly embedded in organizational workflows, where they interact with external information sources and invoke digital tools to perform operational tasks. As organizations adopt such systems, a critical challenge is identifying and mitigating risks arising from malicious or untrusted external information that can steer agents toward unintended actions. Existing red-teaming approaches largely rely on fixed attack templates or final attack outcomes, providing limited visibility into how attacks unfold through multi-step reasoning and tool use. We argue that agent execution risk should be understood as a trajectory-level phenomenon. Building on this perspective, we propose TrajRed, a trajectory-guided red-teaming framework that uses execution trajectories to uncover vulnerabilities in agentic AI systems. We further develop TrajGuard, a runtime governance layer that uses high-risk trajectories discovered during red teaming

Source ↗
Showing 1051–1100 of 1593 signals
← Prev Page 22 of 32 Next →