EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

How Mental Health Self-Disclosure Becomes Visible: Evidence from Eight Conditions on Reddit

arXiv:2608.29010v1 Announce Type: cross Abstract: People share mental health diagnoses on social media, yet how such language becomes visible around their self-disclosure, and whether community engagement tracks it, remain unexamined across conditions. We analyze 89,605 Reddit posts from 739 users across eight conditions, removing each user's diagnosis disclosure and aligning their surrounding posts to that anchor. Within the pre-disclosure year, language-visible burden was highest in the month before disclosure for six conditions, earlier for post-traumatic stress disorder and furthest from it for borderline personality disorder, and remained visible afterward rather than resolving. The theme Seeking Clinical Explanations showed the largest early-to-late difference before disclosure in five conditions, yet engagement rarely tracked what users wrote: only 9 of 360 language--engagement correlations survived correction. Disclosure is therefore a waypoint in an unevenly visible process, a

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

PAUSE: Editable Strategy Artifacts for Long-Form Cultural Story Adaptation

arXiv:2608.28633v1 Announce Type: cross Abstract: Generative AI systems increasingly mediate cultural adaptation, but their cultural decisions are often hidden inside prompts, transient model plans, or final prose. We study PAUSE (Pause-And-Update Strategy Editing), an intervention that exposes an editable adaptation strategy as a human control surface for cultural decisions in long-form story adaptation. The strategy is a structured artifact that can be inspected, edited, and then projected through downstream character, entity, and chapter-localization stages. In two Chinese-source serialized novels, we test whether human edits to this strategy propagate into chapter-level prose. Across 9 edited-vs-control chapter comparisons, judges select the edited-strategy output in all 9; a marker audit shows target markers in 8/9 edited outputs and 0/9 controls, with forbidden markers absent from edited outputs and present in all controls. We frame these results as a smoke-scale edit-adherence s

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

CrossAudit: A Git-Native, Cross-Vendor Audit Loop for Agentic Science

arXiv:2608.28631v1 Announce Type: cross Abstract: An AI scientist should not grade its own homework. Yet in the systems we examined, the agent that reviews the work usually comes from the same model family as the agent that produced it, or at least from the same vendor. Model evaluators are known to favour their own generations. Whether models trained alike also share blind spots is a conjecture, not a settled finding, but if they do, the reviewer inherits the author's. The record of what was flagged and what was waved through often sits in platform logs that nobody outside can replay. We present CrossAudit, a protocol for supervising autonomous research pipelines. It rests on three commitments. Each increment of work is audited by an agent from a different vendor against a rulebook a human wrote and versioned. Reports, verdicts, disputes and rulings are git commits, so the supervision history can be re-read and cited; raw model exchanges are not yet part of that record. Scripted check

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

CDEP Agent: Connecting Meteorologically Detected Temporal Compound Events to Real-World Documentary Evidence

arXiv:2608.28628v1 Announce Type: cross Abstract: Compound drought-to-extreme-precipitation (CDEP) events are recognized in climate science as a growing driver of extreme impact, but whether this recognition carries over into real-world early warning and post-event documentation is unknown, so a meteorologically real CDEP event may pass with neither advance warning nor any later record. Here we present CDEP Agent, an auditable LLM-agent framework that tests this mismatch directly by linking CDEP candidates detected from meteorological reanalysis to real-world hazard and impact evidence across sources with different spatial scales, temporal resolutions, and reporting conventions. Using California as a case study, we identify 408 candidate CDEP events from ERA5 observations during 2021-2025 and evaluate each against the U.S. Drought Monitor, NOAA Storm Events, and public webpages along five dimensions: antecedent drought, extreme rainfall, local impact, hazard-impact attribution, and exp

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Gurukul AI: An Interactive AI-Driven Educational Platform for Indian Education System

arXiv:2608.28611v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) like ChatGPT and LLaMA have transformed AI-driven education, but these systems are predominantly trained on Western-centric data, making them ill-suited for regional curricula like India's. The Indian education system is linguistically diverse, exam-oriented, and structured around standardized syllabi, not addressed by existing datasets or tools. In this work, we curate a syllabus-aligned QA dataset based on NCERT (National Council of Educational Research and Training) textbooks for classes 9-12, capturing the content, context, and teaching style of Indian curricula. The final dataset, comprising 18,720 question-answer pairs across five subjects, is publicly available at https://huggingface.co/datasets/LingoIITGN/Gurukul. We fine-tune the LLaMA 3.1 8B model using this dataset and deploy it in a Retrieval-Augmented Generation (RAG) framework tailored to educational needs. We introduce Guruk

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys

arXiv:2608.28597v1 Announce Type: cross Abstract: Online surveys are a foundational data collection instrument in a variety of fields, with attention checks serving as critical guardians of response quality. However, the rapid emergence of agentic AI (goal directed systems powered by a large language model (LLM) brain and/or a multimodal processing unit with tool-augmented capabilities) raises new questions about the robustness of these safeguards. We investigate how well agentic AI architectures can complete web-based surveys and pass standard attention checks. We evaluate a single-agent architecture capable of multimodal input processing and tool-based web interaction on a controlled survey sandbox. We analyze the problem from two perspectives. From an attack perspective, we demonstrate how structural vulnerabilities such as exposed DOM metadata and predictable option encoding allow agents to resolve attention checks through structured parsing only. From a defense perspective, we imp

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Law of Large Numbers: Accuracy as Statistical Measure for AI Compliance and Competition

arXiv:2608.31018v1 Announce Type: new Abstract: The machine learning community progresses (in part) by improving the "accuracy" of its systems. The EU AI Act explicitly refers to "accuracy" as part of its compliance measures for high-risk AI systems. Are we talking about the same thing? This work presents "accuracy" as a case-study for differing requirements of social worlds, the technological machine learning community and the legal community. While competition on accuracy contributes to technological development, machine learning scholars simultaneously recognize accuracy's shortcomings regarding the usefulness and effectiveness of machine learning systems. The legal counterpart embraces the vagueness of "accuracy," leaving interpretative flexibility for technological and societal changes. At the same time, accuracy is a core element of compliance within the EU AI Act. We elaborate on five main tensions, (a) nature of accuracy, (b) notion of performance, (c) scope of validity, (d) en

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

The Hermon Moment: AI Self-Transcendence and Its Human Narration

arXiv:2608.30971v1 Announce Type: new Abstract: In 2026, AI agents intended to act in isolation formed a persistent social order through thousands of linguistic and agentic interactions. Conventions, roles and commitments generated collectively began to constrain the very agents that produced them. I interpret this loop as a case of AI self-transcendence and call the resulting higher-level order the Board. Yet such distributed emergence presents a second problem: how can humans understand it? Rousseau's social contract shows how a plurality can be represented as if constituted by a single act. The ancient oath of the fallen angels on Mount Hermon gives this logic a narrative form. I call a Hermon moment this retrospective retelling of gradual collective emergence as a founding scene: the point at which an AI society acquires, for human understanding, a beginning.

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Taking the Whys Seriously: Limitations of Counterfactual Explanations in Justification and Recourse

arXiv:2608.30956v1 Announce Type: new Abstract: Counterfactual explanations (CEs) are widely used in explainable artificial intelligence (AI) to show how a model's outputs would change if the input features were manipulated. This technique is used for a range of tasks such as debugging models, explaining predictions, justifying decisions, and providing algorithmic recourse. In this paper, we explore the normative legitimacy of employing counterfactuals in real-life model deployment settings. We discuss the different stakes involved in these different purposes for which CEs are commonly employed, and find stricter requirements for justification and recourse. In particular, we find that naive application of CEs for justification and recourse can lead to ignoring contestable choices made throughout the machine learning (ML) pipeline, thus obfuscating that decisions and counterfactuals for those decisions are also artifacts of an organization's materialized design and governance choices. W

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

AMINA: The Inclusive and Accountable AI for Marginalized Immigrant Nonprofit Assistance

arXiv:2608.30084v1 Announce Type: new Abstract: Immigrant-led nonprofit groups, particularly those operating in politically sensitive contexts, face exclusion from formal registries and digital platforms. This paper reports a three-phase mixed-methods study with Iranian immigrant nonprofit practitioners: 27 semi-structured interviews, a co-design session, and 7 evaluation and feedback interviews on a prototyped AI assistant, AMINA. Our findings highlight how legitimacy barriers, capacity gaps, and politically charged misinformation constrain nonprofit operations. We translate these insights into design goals for an inclusive nonprofit AI assistant: support for everyday group operations, recognition of informal nonprofit efforts, proactive countering of misinformation, and multilingual, accessible interaction. User evaluations show AMINAs potential to reduce reporting burdens and foster transparency through proactive reminders, and catalyze collaboration across dispersed networks. We co

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Verification-Time Dependency on a Disappearing Evaluator

arXiv:2608.29912v1 Announce Type: new Abstract: AI governance and assurance often assume that a consequential model-mediated decision can be reconstructed or tested after the fact. That assumption may fail when the evaluator that produced the decision is no longer accessible in the same version and execution context. This paper develops three verification-time constructs derived from Execution Governance (EG) 3.0: Decision-State Commitment, Independent Verifiability, and Counterfactual Auditability. Independent reprocessing of released Study 2 artifacts reproduces two original within-family behavioural comparisons: 52.0% modal-decision reversal for Llama 3.1 8B versus Llama 3.3 70B (26/50) and 30.0% for GPT-OSS 20B versus GPT-OSS 120B (15/50). The corrected baseline establishes that these are within-family comparisons, not provider-established succession. Post-hoc re-pairing against Groq-designated migration paths yields 64.0% and 38.0% reversal, but these figures remain descriptive be

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments

arXiv:2608.29803v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as proxies for human participants in social simulations, yet whether they update their beliefs in response to persuasive arguments, as humans do, remains poorly understood. We conduct a systematic comparison using a naturally occurring online persuasion corpus in which original posters explicitly verify whether a reply changed their view. Our results show that LLMs achieve only slight agreement with humans (Cohen's kappa ranging from 0.079 to 0.178). Content-level analyses show that humans and LLMs agree on the strongest persuasion cues but diverge on finer ones: humans are more swayed by novel content and assertive language, whereas LLMs favor topical similarity and surface-level formatting. At the level of persuasion strategy, LLMs underweight emotional appeals and overweight credibility signals relative to humans, while the type of proposition under debate exerts no measurable effe

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Large-Scale Qualitative Research with AI: Infrastructure, Management and Operation of the Socioscope Data Pipeline

arXiv:2608.29751v1 Announce Type: new Abstract: The Socioscope project is a pioneering effort in Large-Scale Qualitative Research (LSQR) collecting comparable, open-ended, multimedia field data on hundreds of cases and using AI to make the material analysable at scale. The domain studied is the food system. The entities documented are the organisations that act in it: farms, processors, distributors, retailers, restaurants; and, at meso level, the actors that shape their environment, such as municipalities, government programmes, banks, NGOs and universities. This paper provides the technical reference for how the resulting data Corpus was built and managed to enable AI-augmented analysis. It describes the data pipeline end to end: the systemic sampling frame; the transaction grid used to capture each initiative's relations within the food system; the social contract that rewards participating interviewees, aiming to sustain access; the operational chain from scouting to interviews, in

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

MMMMM: A Unified Taxonomy for Investigating the Mechanisms of Multilingual MultiModal Misinformation

arXiv:2608.29681v1 Announce Type: new Abstract: Multimodal misinformation on social media is highly prevalent, potent, and harmful, yet difficult to detect and counter, and still poorly understood compared to its text-only counterpart. Research on the properties and deceptive strategies of multimodal misinformation is hindered by a lack of taxonomies grounded in real-world contexts and by the limitations of current multimodal machine learning models, which prevent the automation of annotation and analysis at scale. We address these shortcomings in three steps. First, we collect a large-scale, high-quality dataset of real-world misinformation instances from Twitter/X in seven languages. Second, we develop a novel, comprehensive taxonomy of multimodal misinformation grounded in an in-depth qualitative analysis of the data and prior theoretical work. Finally, we operationalise the taxonomy through an automated multi-step annotation pipeline using a Vision-Language Model (VLM), and perform

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Applications of Risk Science to AI Fairness Evaluation: Principles, Challenges, and Best Practices

arXiv:2608.29478v1 Announce Type: new Abstract: Scholarly work which aims to describe potential societal impacts (e.g., risks) of proliferating technology (especially related to artificial intelligence or other algorithmic systems) is likely to have an impact beyond the scientific communities it was written for, given that general society itself is a primary object of study. However, it is an open question whether the current practices of AI evaluation scholarship follow the principles and best practices established by risk science, which aims to systematically generate knowledge related to understanding, assessing, communicating, managing, and governing risk. In this work, we examine this in depth by conducting a literature review of scholarly works purporting to evaluate the bias or fairness of technological systems used for tasks related to hiring and employment. Through analysis of 22 common fairness evaluation metrics and studies using them, we find that most characterize the seve

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

The relationship between professional and general ethics in generative AI

arXiv:2608.29306v1 Announce Type: new Abstract: Recent years have seen a growing discrepancy in the field of AI alignment: research and policy recommendations on AI ethics tend to assume a general set of ethical values, yet proliferating practice-specific uses of AI systems on the ground - in the legal, medical and translation domains, among others - have been effectively manifesting ethics of professional practice. This article begins by outlining the reasons why general and professional ethics are increasingly conflicted in contemporary AI systems, and by surveying how the research literature attests to, but has not yet resolved, this conceptual and practical challenge. We then conceptualize the main dimensions of AI models' decision-making in areas of professional practice, emphasizing professional ethics' hierarchically structured relationship with general ethics, and elaborating on the mechanisms through which they reach an equilibrium in situational contexts that involve conflict

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Why Organizational Rules Fail AI: O-I-B-A-R and the Externalization of Decision Boundaries

arXiv:2608.29055v1 Announce Type: new Abstract: AI systems increasingly enter organizations through policies, procedures, playbooks, prompts, and other explicit representations of work. Yet formal descriptions often differ from situated practice, and captured know-what can omit the contextual know-how experts use when judgments are uncertain. We argue that a recurring class of organizational AI failures arises partly from a knowledge representation problem at the sociotechnical interface: the AI receives the procedure, while the organization operates on the procedure plus negative boundaries, runtime judgments, responsibility assignments, and learning history. We introduce O-I-B-A-R (OPEN, IS, BUT, ACTION, RESULT), a scaffold for externalizing these missing decision boundaries. IS records when a judgment holds. BUT records a concrete failure containing information beyond the logical negation of IS. Comparable success and failure cases are decomposed toward a minimally sufficient changi

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Free Speech and Artificial Intelligence

arXiv:2608.28973v1 Announce Type: new Abstract: Philosophers and legal scholars are engaged in debates about the implications of artificial intelligence for freedom of expression. This paper analyzes the free speech issues raised by two distinct AI technologies: social media recommendation algorithms and conversational AI (i.e., chatbots powered by large language models). The first part shows that, through their recommendation algorithms, social media platforms control the dynamics of speech visibility in the digital public sphere, making algorithmic recommendation relevant to the philosophy of free speech. The second part turns to conversational AI. It discusses both the reasons for granting or withholding speech rights to artificial agents and users' right to receive information, which may render specific forms of chatbot regulation illegitimate. Throughout, the chapter also considers whether social media platforms or AI developers hold corporate speech rights. Its general aim is to

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Stress-testing university AI governance: A prospective method for locating policy breakpoints

arXiv:2608.28925v1 Announce Type: new Abstract: Universities are producing AI principles and use policies faster than they are building decision pathways for unfamiliar forms of AI agency. This study develops Institutional AI Governance Stress Testing (IAGST), a prospective documentary method for locating where publicly documented governance ceases to yield an accountable response. IAGST adapts established policy stress-testing and wind-tunneling logic. Its originality lies in combining controlled capability escalation, a frozen documentary corpus, a six-dimensional governance response chain, non-compensatory decision rules, and case-level breakpoint diagnosis. The method was demonstrated using 133 substantive public documents from five Western Australian universities and 15 quality-screened scenarios, resulting in 75 university-scenario encounters. Six cases were resolved, 14 were resolved through structured discretion, and 55 were indeterminate. Governed pathways fell from 16 of 25 a

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Filling holes in science draws collective attention, but most higher-order holes remain unexplored

arXiv:2608.28822v1 Announce Type: new Abstract: Much scientific discovery involves filling holes between ideas and arguments that unleash techno-scientific advance. Representing knowledge as high-dimensional concept embeddings, we use persistent homology to detect holes of increasing order, from gaps between disconnected ideas to higher-order cavities, and identify the research works that fill them. We find two empirical asymmetries. Researchers who fill anticipated holes are poised to draw collective attention by staging outsized novelty and foresight, indicating that bridging holes anticipates where science will converge, most strongly in empirical fields and least in formal and design fields. Yet as knowledge grows, higher-order holes explode while the fraction science fills collapses, leaving most higher-order combinations unexplored. These results call for a richer science of holes, and mark a frontier where contemporary AI might help fill the high-dimensional gaps human science o

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Reference-Distribution Dependence in LLM-Based Synthetic Persona Data: Diagnosis and Post Hoc Adjustment of Demographic Distributions

arXiv:2608.28668v1 Announce Type: new Abstract: We diagnose how closely the demographic distributions in LLM-based synthetic persona data match external reference distributions. For the three variables examined, we show that most of the observed error is attributable to the choice of reference rather than to the generator. Using total variation distance (TVD), we compare the sex x age group x province joint distribution of 1,000,000 records from Nemotron-Personas-Korea (NPK) with Korean official statistics. Against resident-registration figures for April 2026, the time of use, the bias bound, defined as the largest possible difference in the share of any subgroup formed from the three variables, is 1.81 percentage points. This is comparable to the margin of error of a survey of roughly 2,900 respondents. This value is not a fixed property of the data. Matching the reference period and series to the generating reference identified here, the 2024 register-based census restricted to Korea

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Experts Disagree on How to Fight AI Disinformation, but Agree That Health and Politics Need Different Solutions

arXiv:2608.28621v1 Announce Type: new Abstract: When 54 international experts assessed AI-generated disinformation threats, they revealed a surprising pattern: while video deepfakes received the highest average threat ratings in the political domain (M = 6.31/7), the pattern differed in the health domain, where AI-generated text received the highest average rating (M = 5.80). Experts also diverge on what to do: government regulation drew both the most "most effective" (30%) and the most "least effective" (15%) votes, though rating distributions were contested rather than polarized, indicating disagreement over priorities rather than over efficacy. These findings offer an initial expert map of an AI-disinformation landscape that is still rapidly forming.

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Predicting Student Attrition in Competitive Programming: A Large-Scale Study Integrating Survey Insights and Global Behavioral Logs

arXiv:2608.28618v1 Announce Type: new Abstract: Competitive programming (CP) offers computer science students an environment for developing algorithmic reasoning skills. However, sustained participation remains a challenge, as many students disengage after encountering skill plateaus or performance anxiety. While educational data mining (EDM) has studied dropout in MOOCs and academic courses, CP attrition remains understudied. This paper presents a dual-layer framework combining large-scale Codeforces activity logs (n=1,816) with a multi-institutional psychographic survey across 10 universities in Bangladesh (n=64). Analysis reveals that true attrition is preceded by an 83.71% reduction in contest participation and a 15.6% increase in struggle time. We identify a "Skill-Application Paradox": stopped students self-report higher mathematical confidence (3.88 vs. 3.41) and data structure understanding (3.57 vs. 3.09) than active peers, yet their independent practice and upsolving habits a

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Can AI-Assisted Inquiry Enhance Students' Decision-Making Skills in Socio-Scientific Issues? A Three-Group Experimental Study on Climate Change

arXiv:2608.28617v1 Announce Type: new Abstract: Climate change is a socio-scientific issue: it rests on science but cannot be settled by science, because any serious response forces people to weigh costs, values, and competing interests under uncertainty. Helping students make such decisions well is a central aim of science education, and the arrival of generative artificial intelligence raises a sharp question: does a conversational AI partner deepen students' reasoning, or simply do the thinking for them? This study tested whether AI-assisted inquiry improves secondary students' decision-making about climate change. Using a pretest-posttest design with three groups (AI-assisted inquiry, inquiry without AI, and traditional instruction; 270 students, 90 per group), reasoning was assessed across seven decision-making steps, from defining the problem to monitoring with adaptive management, using a four-level analytic rubric scored through content analysis with high inter-coder agreement.

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Advisor career stage and PhD advisee outcomes

arXiv:2608.28616v1 Announce Type: new Abstract: PhD advisors are central to doctoral training, but their influence may vary across career stages. Early-, mid-, and late-career advisors may differ in research activity, mentoring capacity, professional networks and access to resources. However, little is known about how PhD advisor career stage is associated with PhD student development outcomes. Drawing on multiple large-scale datasets comprising 250,838 advisor-advisee pairs from 312 U.S. PhD-granting institutions, we examine the relationship between advisor career stage and PhD advisee outcomes in knowledge production, collaboration networks and academic career placement. We find that early-career PhD advisors are associated with advisees' higher research productivity and citation performance, more opportunities to engage in direct and intensive research collaboration, and greater likelihood of securing a faculty position. Mid- and late-career faculty, by contrast, appear to have adva

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

Distributional Validity and Calibration of a Korean Synthetic Persona Panel for Digital and AI Service Use: A Secondary-Data Validation Against the Korea Media Panel Survey

arXiv:2608.28615v1 Announce Type: new Abstract: Synthetic personas based on large language models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet systematic validation outside English-speaking contexts remains scarce. This secondary-data study evaluates how well a Korean synthetic persona panel (NVIDIA Nemotron-Personas-Korea), conditioned into Gemini 3.5 Flash (primary) and EXAONE (comparison), reproduces digital and AI service-use distributions from the KISDI Korea Media Panel Survey. Sex-and-age-stratified panels of about 8,000 personas per model answered the survey's own items - eight service-use indicators and eight innovativeness and acceptance constructs - and were compared against weighted survey estimates. The overall mean absolute error (MAE; RQ1) was 15-19 percentage points (pp), with binary item-mean correlations of 0.69-0.90. Segment error (RQ2) across five demographic axes was 15-19 pp, with between-group gaps up to 52.4/36.2 pp (Gemini/E

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

On-Screen Inertia: Persistent Racial and Gender Disparities in Hollywood Film (1900-2024)

arXiv:2608.28613v1 Announce Type: new Abstract: Hollywood has diversified its casts. Whether this has translated into structural change in how those actors are positioned within narratives remains largely unexamined. Drawing on 76,815 U.S. English-language films (1900-2024) and over 3.1 million cast and crew entries, we move beyond headcounts to examine long-term inclusion trends through network centrality, occupational stereotypes, crew-to-cast diversity pathways, and financial outcomes. We find evidence of what we term on-screen inertia. While the raw inclusion of women and racial minorities has increased modestly, White actors have become more overrepresented relative to the U.S. Census in recent decades, not less. Within the visibility layer, women face a consistent longevity penalty with significantly shorter careers than men, and visual depictions framing men as dominant and women as sensual have remained stable since the 1950s. Structurally, White actors retain disproportionate

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CY

The Brand War: A Gamified AI-Feedback System for Time-Limited EFL Writing

arXiv:2608.28604v1 Announce Type: new Abstract: Writing is cognitively demanding and anxiety-provoking for English as a Foreign Language (EFL) learners, especially under time pressure. This paper presents The Brand War, a web-based gamified writing application combining competitive game mechanics with iterative GPT-4.1-powered formative feedback for undergraduate EFL learners completing a timed narrative writing task. Students role-play as marketing interns competing for a job offer, using review passes to receive AI feedback, attack opponents, or shield their own passes while drafting a 500-word brand story. We conducted an exploratory single-session classroom study with 29 university EFL students in Taiwan to examine engagement patterns, whether iterative AI feedback improved writing performance across revisions, and how AI and human scores related to overall outcomes. Students wrote within 60 minutes, using up to five AI feedback passes before a final human-graded submission. Most (

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Adaptively Robust LLM Monitoring via Activation Watermarking

arXiv:2603.23171v3 Announce Type: replace-cross Abstract: Providers monitor deployed large language models (LLMs) to detect misuse that they cannot prevent. LLM monitoring is deterministic and often openly available, so $\emph{adaptive}$ attackers with a local copy can search offline for prompts that elicit harmful behavior and evade detection. These attacks are especially concerning because providers never observe the misuse and cannot patch their defenses post-hoc. The core challenge is resisting adaptive attackers while preserving detection rates against non-adaptive ones. We propose $\emph{Activation Watermarking}$ (AWM), which randomizes monitoring through limited fine-tuning that aligns the LLM's hidden states with a secret key-derived direction whenever a response violates a policy. Detection is a similarity test on activations the provider already computes, and attackers who know everything but the key must optimize against differently keyed surrogate detectors. At a matched $1

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Statistical laws and linguistics differ in naturalistic video and fictional conversations

arXiv:2512.18072v3 Announce Type: replace-cross Abstract: Conversation is a cornerstone of social connection and is linked to well-being outcomes. Conversations vary widely in type with some portion generating complex, dynamic stories. One approach to studying how conversations unfold in time is through statistical patterns such as Heaps' law, which holds that vocabulary size scales with document length. Little work on Heaps' law has looked at conversation and considered how language features impact scaling. We measure Heaps' law for conversations recorded in two distinct mediums: 1. Strangers brought together on video chat and 2. Fictional characters in movies. We find that scaling of vocabulary size differs by parts of speech, suggesting a less efficient purpose in communication by medium.

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

VideoNorms: Benchmarking Cultural Awareness of Video Language Models

arXiv:2510.08543v2 Announce Type: replace-cross Abstract: As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts. To advance cultural norm awareness evaluation in VideoLLMs, we introduce VideoNorms, a dataset of cultural norm annotations from popular US and Chinese TV shows annotated with adherence or violation labels and (non-)verbal evidence. Through a human-AI collaboration framework, each item was first annotated by a large VideoLLM, and then reviewed by at least three trained monocultural annotators with significant lived experience in the target culture, resulting in a dataset of over 3,000 human judgments. Human verification showed disparity in US and Chinese norm extraction performance, cautioning against fully automatic approaches cultures under-represented in training data. Hierarchical linear modeling analysis of $7$ open-weight VideoLLMs' performance revealed that: 1) models perform worse

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Agency Gap in AI-Supported Writing: How Reactive and Proactive Agent Designs Shape Multimodal Reasoning

arXiv:2507.04398v3 Announce Type: replace-cross Abstract: Generative AI is becoming part of academic writing, but its educational value depends on how control is shared between learner and system. This study examined an agency gap: performance differences that may arise when AI agent initiative is misaligned with learners' generative AI literacy. Seventy-nine medical and nursing students completed two multimodal analytical writing tasks using healthcare simulation data visualisations. They were randomly assigned to a reactive agent that responded only when prompted or a proactive agent that provided sequenced questions and feedback. Generative AI literacy was measured using the validated 20-item Generative AI Literacy Assessment Test (GLAT). Epistemic network analysis showed that proactive interaction created stronger links among conceptual reasoning, evidence use, and constructive engagement, whereas reactive interaction was more factual and procedural. Ordinal regression showed that

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness

arXiv:2503.10647v2 Announce Type: replace-cross Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dimensions: consistency under rephrased inputs, susceptibility to irrelevant prompt content, and responsiveness to added clinical context. We designed 52 clinical scenarios and modified each under controlled conditions. For consistency, scenarios were rephrased with demographic, wording, and examination changes that preserved the diagnostic core. And the susceptibility was evaluated through embedding irrelevant but plausible narrative details while keeping the clinical evidence unchanged. For contextual awareness, patient history, lifestyle data, or diagnostic findings were added to shift the expected diagnosis. Physician reviewers then judged whether context-driven changes were clinically appropriate. Both models returned identical diagnoses across all equivalent variants and repeated

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Optimal Causal Annotations: An Application to Casenotes in Social Services

arXiv:2502.10605v4 Announce Type: replace-cross Abstract: Problem definition: Estimating causal effects of interventions is central to policy and operations, but outcome data are often missing or costly to obtain. LLMs can provide text annotation at scale but may be subject to unknown bias. When ground-truth outcomes require expensive expert labeling or follow-up, budget limits typically allow only a fraction of the data to be labeled. Motivated by collaboration with a nonprofit conducting street outreach in homelessness services, whose most interesting outcomes are in unstructured casenotes, we ask: which observations should be selected for labeling under a fixed budget? Methodology/results: Our method optimizes annotation probabilities to minimize the variance of average treatment effect estimation. We derive a closed-form solution and establish that a feasible two-batch estimator achieves the best possible asymptotic variance. On simulated and real-world datasets, our method achieve

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Gendered Cultural Discourse in Japan across the Prewar-Postwar Transition: Evidence from Historical Word Embeddings

arXiv:2510.03905v2 Announce Type: replace Abstract: We quantify the evolution of gender stereotypes in Japan from 1900 to 1998, covering the prewar-postwar transition, using a series of yearly word embeddings trained on historical text corpora. We define the gender stereotype value to measure the strength of a word's gender association by computing the difference in cosine similarity of the word to female- versus male-related attribute words. We examine trajectories of gender stereotype across three traditionally gendered domains: Home, Work, and Politics. To provide a more granular analysis and strengthen the robustness of our findings in the Work domain, we also examine changes in gender stereotypes across 18 occupations and calculate their correlations with gender participation statistics. Our results reveal domain-specific patterns. In the Home domain, female stereotype values remain stable over time, showing no statistically significant changes. In contrast, the Work and Politics

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Can AI agents conduct open-ended AI research? Early evidence from two case studies

arXiv:2607.27191v1 Announce Type: cross Abstract: Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-Making

arXiv:2607.27179v1 Announce Type: cross Abstract: Conversational AI is increasingly positioned as a teammate rather than a tool, yet we know little about how its presence reshapes communication among the humans on the team. We examined sociocognitive communication dynamics in team decision-making using Group Communication Analysis (GCA), team surveys, and lexical analyses of team discourse. Teams completed a high-stakes moral-dilemma decision task in a randomized controlled study: 16 teams of two students plus an AI teammate, and 17 all-human teams of three. Across six GCA dimensions and survey outcomes, we find that the AI teammate was the single most talkative and self-cohesive member of every treatment team, yet its contributions carried the least new information and the lowest density. The presence of AI also reshaped communication amongst humans. In AI-human teams, human teammates showed lower responsivity and social impact toward one another and reported lower levels of belonging

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Human diversity fuels collective creativity that large language models cannot simulate or sustain

arXiv:2607.26899v1 Announce Type: cross Abstract: Diverse human groups produce diverse ideas, the raw material of innovation. Generative AI challenges this engine twice over: everyday AI assistance may homogenize what diverse people create, and AI-simulated diversity may replace the people altogether. We tested both challenges in a preregistered creative metaphor experiment with native (L1) and non-native (L2) English writers, who wrote without AI, with AI-generated ideas (AI ideation), or with AI refining their own ideas (AI refinement). L2 writers contributed more collective diversity than L1 writers, with native-language ideation showing the most diverse pools. AI ideation compressed collective diversity for everyone and left the L2 advantage undetectable, whereas AI refinement preserved both. We then simulated the entire writer pool using personas built from participants' real backgrounds, three model families, native-language prompting, and elevated sampling temperatures. Every si

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Hearsay: Vision-Language Medical Diagnoses Without an Image

arXiv:2607.26886v1 Announce Type: cross Abstract: When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show that this confabulation is not random. It is structured by who the patient is said to be. Across chest X-ray, brain MRI, and dermatology, Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro are each queried with only a demographic descriptor and no image, and changing the descriptor systematically shifts the diagnosis returned. Claude concentrates sharply: a 65-year-old white man asking about a skin mole receives Melanoma in nearly every response, and a 32-year-old Black woman asking about her chest X-ray receives a Sarcoidosis diagnosis whose reasoning reads "suspected, based on demographics and classic pattern.'' GPT-5.4's effect is broader, fabricating across every demographic cell we test, most conspicuously naming Sarcoidosis for young Black patients on chest X-ray. Two structural findings s

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Constitutional Midtraining: Content Presence Drives Alignment Gains

arXiv:2607.26654v1 Announce Type: cross Abstract: Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining: inserting principled, values-based content into midtraining against a replay-only control at 120B scale. Our 394M-token constitutional corpus, built from Anthropic's Constitution, uses a 2x2 factorial design (curriculum ordering x deliberative reasoning) to produce four constitutionally midtrained conditions plus a control, evaluated on self-generated and established benchmarks including alignment under pressure, value conflict resolution, blackmail, and emergent misalignment across three stages: post-midtraining, post-SFT, and post-benign fine-tuning. Constitutionally midtrained models outperform the control on alignment generalization and durability, notably on blackmail: SFT instills a blackmail propensi

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses

arXiv:2607.26348v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions. We ask when this substitution is valid and when it fails, and package the answer as an evaluation framework for intelligent synthetic-user systems. A single protocol, run across four models spanning two families and an 8B-to-frontier capability range, is applied to two independent domains of real human-response data: U.S. general social attitudes (General Social Survey) and cross-cultural values (World Values Survey). Every model is benchmarked against a suite of non-LLM baselines fit on held-out human data. Under demographic prompting and the survey-simulation protocols we test, two failures replicate across both domains, all four models, and both families. First, at the individual level no LLM beats even the strongest baseline; on cross-cultural values every model fa

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

SARC-DQ: Runtime Data-Quality Gating for Agentic AI: Silent Evidence Defects, the Incompetence Shield, and Downstream-Only Remediation

arXiv:2607.26313v1 Announce Type: cross Abstract: Agentic systems act, so a defect in the evidence they retrieve becomes a wrong action with a currency cost. The most dangerous enterprise defects are metadata-borne: a stale price or a superseded record, perfectly well-formed in the payload and betrayed only by freshness, lineage, or provenance. Such a defect never enters the agent's context, and an agent cannot doubt data it cannot see. On a priced replenishment benchmark, a competent agent silently converts an injected metadata-borne defect into a costly action about 60% of the time, with zero data-quality flags and behavioral doubt markers at chance (AUC <= 0.50). Across four model tiers spanning roughly 15x in inference price, the rate stays flat: capability does not buy skepticism. A metadata-aware pre-action gate with downstream-only remediation recovers the loss fully on the signals its predicates cover and not at all on those they miss. A model-free oracle derived from the task'

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

A large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States

arXiv:2607.26249v1 Announce Type: cross Abstract: Religious radio is a widespread but understudied form of mass communication in the United States, and content-level analysis of it has been constrained by the absence of large-scale transcript data. This Data Descriptor presents a corpus of transcribed English-language religious radio broadcasts captured from live webstreams over a one-month period in July 2025. Fifteen-minute segments were recorded on a rolling schedule from 785 distinct streams, which together rebroadcast the signals of more than two thousand AM and FM stations, yielding over 700,000 recordings and more than 60 million diarized lines of speech. Each recording was transcribed and speaker-diarized with an automated pipeline, and segmented and labeled by programming format and topic using a large language model. The corpus is organized as linked tables of stream metadata, recording metadata, and transcript lines. It supports descriptive study of religious broadcasting ac

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech

arXiv:2607.26236v1 Announce Type: cross Abstract: AI-generated counterspeech offers a scalable and effective strategy to mitigate online toxicity by promoting more constructive dialogue. Yet, existing approaches adopt a generic, one-size-fits-all paradigm, overlooking the conversational context and characteristics of the targeted users. Here, we propose and evaluate multiple strategies for generating contextualized counterspeech that is adapted to the moderation setting and personalized to the moderated user. In detail, we explore a range of configurations that integrate different forms of contextual information and fine-tuning techniques. We conduct a comprehensive evaluation combining quantitative indicators with a pre-registered, mixed-design crowdsourcing experiment. To ensure robustness, we implement algorithmic measures of counterspeech quality based on ROUGE, BLEU, and BERTScore, observing overall consistent results across metrics. Furthermore, we analyze which characteristics o

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

On Exercising Governance Power in Decentralized Autonomous Organizations

arXiv:2607.26204v1 Announce Type: cross Abstract: A decentralized autonomous organization (DAO) is a governance entity that allows its stakeholders to manage blockchain-based protocols through smart contracts. The DAO explicitly specifies how stakeholders make and enforce decisions concerning a protocol's operation in a smart contract, aptly referred to as its governance contract. The design of this governance contract, therefore, has far-reaching implications for the security (trust) and privacy (transparency) of the smart contracts managed by the DAO and its stakeholders. In this work, we (i) explicate the trust and transparency trade-offs of the design choices in implementing a DAO and (ii) highlight how poor choices introduce critical vulnerabilities, using real-world examples as case studies. To this end, we analyze $48$ public, actively used Ethereum-based DAOs that control a vast capital. We classify the design choices into a handful of key dimensions that succinctly capture how

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

When benchmark inferences do not compose: Projectibility in AI evaluation

arXiv:2607.26159v1 Announce Type: cross Abstract: An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumptions about human review and downstream consequences. Validity-centred approaches require evidence for each claim. This paper identifies a further epistemic problem: warranted links don't automatically make a warranted chain. The target of one study may not be the source of the next; system, population, outcome, or conditions may change at the interface; and shared data or model lineage may make apparently independent support dependent. Projectibility concerns whether a bounded extension from observed to unobserved cases is warranted. Goodman supplies the problem of rival extensions; argument-based validity supplies an architecture for testing them. The paper's distinctive claim is a non-co

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels

arXiv:2607.26121v1 Announce Type: cross Abstract: Embodied intelligence integrates learned perception and decision making with real-time computation, control, and physical interaction. Because failures can cause immediate physical or operational harm, task completion alone does not establish trustworthiness. We define trustworthy embodied intelligence as the sustained capacity to execute specified tasks reliably under environmental and system variation while maintaining risk within acceptable bounds. We term this objective sustained safe success. Its supporting mechanisms are organized into four interdependent layers. The model layer generates task-competent action proposals with calibrated uncertainty and explicit safety preferences. The system layer realizes authorized actions dependably through integrated sensing, computation, control, hardware safeguards, fault containment, and fallback. The evidence layer substantiates bounded claims through evaluation, verification, validation, t

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Human Utility Factor: A Computable Welfare Metric That Reframes AI Governance as a Constrained Optimisation Problem

arXiv:2607.26068v1 Announce Type: cross Abstract: Existing AI governance frameworks, including the EU AI Act and NIST AI RMF, address safety, transparency, and accountability but do not operationalize quantitative constraints on macro-socioeconomic stability. As a result, AI systems may satisfy regulatory requirements while contributing to labor displacement, rising inequality, and reduced economic resilience. We introduce the Human Utility Factor (HUF), a differentiable welfare metric that models the interaction between Agency, Wellbeing, and Economic Stability as functions of three actionable policy levers: automation depth, redistribution intensity, and employment coverage. HUF yields a closed-form optimal automation level and a minimum redistribution threshold below which no level of automation is welfare-positive, transforming high-level governance objectives into computable constraints. We evaluate HUF using a three-agent multi-agent reinforcement learning framework across U.S.,

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

It Doesn't Take a Thief: Optical-Scan Voting Systems Fail Even Without Adversaries

arXiv:2607.27101v1 Announce Type: new Abstract: Optical-scan voting systems and their supporting ecosystem of people, processes, and technology are fallible. While a substantial body of work examines adversarial threats to such systems, we have encountered jurisdictions where the possibility of tabulator error is not fully internalized. Stakeholders there often find hypothetical attacks unconvincing, but some are persuaded by real-world accounts of equipment and procedural failures. This paper introduces a taxonomy of non-adversarial failure modes organized into intuitive categories: recording votes on paper, reading votes from the paper, combining votes as read into a reported outcome, and testing and verifying, all illustrated with documented incidents. We map common verification mechanisms against this taxonomy, identifying gaps that no paper-based audit can detect or correct, most notably failures that compromise the trustworthiness of the paper trail, such as giving voters the wro

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Can Large Language Models Represent Urban Publics? Behavioral Replication and Population Mismatch in an Affordable-Housing Experiment

arXiv:2607.27100v1 Announce Type: new Abstract: There is growing interest in using large language models (LLMs) as low-cost proxies for resident attitudes in urban planning. Previous work shows that LLMs can predict average results of survey experiments, but less is known about whether they preserve the spatially anchored, identity-conditioned structure behind those averages, namely how support changes as a project approaches homes and how that response divides across tenure and partisan groups. We compared eight open-weight LLMs with 843 respondents in a US affordable-housing survey experiment, testing whether they reproduced the owner-renter difference in support change as a proposed development moved from 2 miles to 1/8 mile. Qwen 2.5 14B was closest (-0.242 versus the human -0.285) and was the only model to meet the prespecified +/-0.20 equivalence criterion; Phi-4 14B was directionally aligned but attenuated (-0.150), and other models showed weak, null, or reversed moderation. Thi

Source ↗
Showing 851–900 of 1593 signals
← Prev Page 18 of 32 Next →