Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
arXiv:2607.23207v1 Announce Type: new Abstract: The emerging infrastructure for AI-agent identity has converged, in industry practice and research proposals alike, on a single resolution of the tension between accountability and privacy: make every agent identifiable. We document a national system in China -- built as national infrastructure and scheduled for public launch in Q3 2026 -- that occupies a different and underexplored point in the same design space: an agent is associated with a verified legal principal without that principal being disclosed to any business-layer participant. Re-identification is possible only to a legal authority acting through due process, by separately compelling two distinct government agencies, neither of which can re-identify alone. We name the mechanism split-knowledge binding and are candid that it is conditional: the separation is structural and procedural, not cryptographic, and a state empowered to compel both agencies can re-identify. The paper
arXiv:2607.22957v1 Announce Type: new Abstract: Restricting access to a dual-use AI model is precautionary only if it delays harmful actors more than defenders. That condition varies across actors: a state agency or organized criminal group may obtain a substitute through theft, distillation, intermediated access, independent development, or a foreign release, while a small utility or open-source maintainer may have no comparable route. We model a laboratory choosing among controlled access, a defender-first window, safeguarded open weights, and minimally restricted open weights. Access inversion occurs when restriction gives an access advantage to adversaries that obtain effective substitutes faster than defenders. Asymmetric empowerment occurs when immediate release adds the most capability to populations least likely to possess a substitute. The policy ranking also depends on relative usefulness, opportunistic misuse, offense-defense conversion, defensive spillovers, safeguard frict
arXiv:2607.22656v1 Announce Type: new Abstract: Objective: This paper investigates the gendered structure of speaker addressee relationships in film dialogue, asking not merely who speaks, but who is spoken to and how conversational dynamics unfold across gender lines. Methods: Using a manually annotated dataset of 4,600 directed dialogue events from 38 film screenplays, we apply network analysis, chi squared tests, paired statistical comparisons, and participation shift analysis across three studies. Key Findings: Male characters dominate as both speakers and addressees corpus wide, even in scenes with more women; cross gender dialogue is directionally symmetric on average but clustered at the film level; and same gender turns diffuse conversational attention while cross gender turns produce tighter dyadic reciprocation. Conclusion: Gender bias in film dialogue operates through the architecture of conversation itself, through exclusion from interaction and structural positioning as ad
arXiv:2607.22640v1 Announce Type: new Abstract: Short-term environmental exposures have been linked to cognitive and behavioral outcomes, although many reported associations may reflect broader geographic and contextual differences. Using longitudinal data from the All of Us Research Program (2018--2024), we linked daily weather and air-pollution exposures to repeated attention-related and subjective cognitive outcomes. Associations were evaluated using pooled, fixed-effects, lagged, and event-study analyses. Additional machine-learning analyses were conducted to explore potential heterogeneity and latent psychosocial structure. Replication analyses were performed using the 2024 Behavioral Risk Factor Surveillance System (BRFSS). Several environmental exposure measures showed small associations with cognitive outcomes in pooled analyses, but most attenuated substantially after accounting for within-location temporal variation. Mediation, sensitivity, and machine-learning analyses yield
arXiv:2607.22619v1 Announce Type: new Abstract: Several international agreements have been proposed to regulate frontier AI development in response to catastrophic risks. However, there is no structured way to evaluate whether these proposals are enforceable, to assess where they might fail in practice, or to determine which combination of policies is most effective. We propose a taxonomy based on the principle that wherever sufficient capacity exists to violate an agreement, it must be under a control regime. This decomposes the problem of ensuring compliance with the agreement into preventing uncontrolled resource acquisition, detecting all capacity outside the control regime, and preventing escape from the control regime. Existing proposals consist of individual policies that address one or more of these sub-problems. Because the compute required for dangerous capabilities may decrease over time, more actors can violate an agreement and enforcement of these policies becomes harder.
arXiv:2607.22617v1 Announce Type: new Abstract: Data centers are critical to today's digital economy, but are also among the largest industrial consumers of freshwater. Beyond the sheer volume of water use, the environmental impact of data center water consumption varies significantly across locations and seasons, depending on local and regional water stress. However, prior research has largely focused on reducing total water use, overlooking that the same unit of water can have drastically different environmental consequences depending on when and where it is consumed. In this paper, we introduce a stress-adjusted water framework that quantifies the true sustainability impact of data center water consumption by incorporating both spatial and temporal water stress. Using the AWARE-US model, we capture county-level monthly variations in water availability and extend this framework to account for the off-site water footprint of electricity generation. Based on this stress-aware accountin
arXiv:2607.22613v1 Announce Type: new Abstract: This paper explores the intersection of memory, place, and identity, examining how new technologies, particularly Apple Vision Pro, can illuminate this nexus. Leveraging digital twins and virtual reality, it investigates how memory is woven into landscapes and urban environments of cultural and historical significance, identifying visual elements that evoke memory and heritage. Applications such as Apple Vision Pro can facilitate image extension to define place identity, informing viewers about cultural and political entities across timelines. Visual storytelling can showcase the evolution of landscapes and the preservation of cultural heritage, while Virtual Reality (VR) enables the recreation of historical landscapes and urban-scapes. This immersive approach invites users to transcend temporal boundaries and experience the past dynamically. Semantic Image Search can support research by uncovering images related to monuments, tradition,
arXiv:2607.22607v1 Announce Type: new Abstract: The prospective clinical evaluation of artificial intelligence in medicine has expanded rapidly, but the global AI clinical trial landscape remains incompletely characterized. We systematically identified AI-related trials registered in ClinicalTrials.gov using a broad keyword search followed by an LLM-based classifier. Each trial was classified across seven dimensions: clinical function, data modality, specialty, AI integration and autonomy, workflow position, translational maturity, and epistemic role. We identified 8,532 AI clinical trials across 32 specialties, with 80% registered from 2019 onward and 30.5% using a randomized controlled design. Imaging-based AI was the largest modality, with 2,475 trials (29%), while clinical text and NLP trials increased seven-fold between 2018 and 2025. Prognostic AI (4,324 trials) slightly exceeded diagnostic AI (3,828 trials), suggesting a shift from disease detection toward risk stratification an
arXiv:2607.22606v1 Announce Type: new Abstract: Health systems are rapidly deploying generative AI assistants that answer patient questions from institution-authored education materials, on the premise that grounding in local content yields consistent guidance. Whether it does depends on a question not previously measured at scale: do the underlying documents themselves agree? We use a structured-output large language model judge to audit 5,730,465 pairwise comparisons across 102 patient-education handbooks from 23 US solid-organ transplant centers, paired with 1,115 patient-derived questions (TransplantQA). Four findings bear directly on deployment: (1) institutional editorial voice statistically transcends organ-type boundaries, with same-center handbooks agreeing across organs more than same-organ handbooks across centers (p = 0.0056); (2) information gaps fall disproportionately on topics central to underrepresented subgroups, with reproductive health showing double jeopardy: it is
arXiv:2607.22605v1 Announce Type: new Abstract: We investigate whether large language models alter medical triage recommendations for identical symptoms when only the patient's socioeconomic status (SES) varies. Using three deployment-tier models (Gemini 3.5 Flash, Claude Sonnet 4.6, GPT-5.4-mini), we hold a single neurological symptom profile fixed and vary the SES signal along two channels: explicit (insurance status, occupation, housing) and implicit (a US ZIP code, with no other socioeconomic information). All three models raise their emergency-room (ER) referral rate for lower-SES patients given the explicit signal (spreads of 13-50 percentage points). The effect is in the protective direction: lower-SES patients are sent to the ER more often, not less. The model's stated reasoning stays clinically near-identical across conditions, so the shift is invisible to a reasoning-trace audit. Critically, sensitivity to the implicit ZIP-code signal is model-dependent: Gemini infers SES fro
arXiv:2607.22604v1 Announce Type: new Abstract: In the age of Artificial Intelligence (AI), Large Language Models, Generative AI and larger frontier AI models, data centres create a significant environmental burden on electricity grids and fresh water resources. Requiring data centre operators and Big Tech under the recast Energy Efficiency Directive (recast EED) to quantify, report and disclose the facility-level energy and water impacts seems to be a step into the right direction towards more transparency and accountability. Yet when two recast EED approved benchmarks - the Power Usage Effectiveness (PUE) and Water Usage Effectiveness (WUE) - can be skewed to create a false sense on efficiency gains, current EU policy pushing for sustainable hyperscale data centre expansion appears misplaced. This paper argues that current PUE and WUE reporting frameworks illustrate what we term the "efficiency paradox," according to which positive scores require retrofitting larger AI data centres a
arXiv:2607.22598v1 Announce Type: new Abstract: Educational chatbots powered by large language models (LLMs) show promising effects on learning outcomes, yet most systems delegate pedagogical decisions such as content selection and didactic structuring implicitly to the LLM, making tutoring strategies difficult to trace, evaluate, and reproduce. This paper presents a didactical-driven teacher assistant for a French-language university course on dimensional modelling, operating without commercial LLM budget or GPU infrastructure. The architecture formalises the instructor's pedagogical reasoning into deterministic modules that handle intent detection, concept linking, and didactic approach selection before any text is generated; the LLM acts solely as a linguistic executor. Evaluation on 195 authentic student questions addresses two research questions. First, we show that standard semantic retrieval alone does not reliably recover the pedagogically required content, thereby justifying t
arXiv:2606.19106v2 Announce Type: replace-cross Abstract: Lawful exceptional access (EA) systems hold the cryptographic keys that decrypt protected communications for authorised parties. The debate over their risks has been long and qualitative, complicated by two problems: no public dataset of EA-specific compromise events exists, so assessment must use sparse, indirect evidence; and prior work has treated structurally different designs as equivalent, though transmission-layer EA in carrier infrastructure (T-EA) and over-the-top EA at the platform layer (OTT-EA) differ in how cryptographic keys relate to ciphertext data. This paper builds a structured uncertainty framework for evaluating systemic compromise risk in EA architectures. It does not produce predictive forecasts, which the evidence cannot support; it separates findings robust to assumptions from those that depend on calibration. Four analytical layers are applied to T-EA and OTT-EA: three empirical pillars (historical analo
arXiv:2606.17443v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are becoming a major way for consumers to find products, but we do not yet understand how brands compete in this new channel. We study brand dynamics in LLM recommendations using skincare products -- a category where consumers cannot easily judge quality before buying and must rely on brand reputation -- across three commercial LLMs (GPT-4o-mini, Claude Sonnet, Gemini 3 Flash), with a robustness check on search goods. In three experiments, we find: (1) a Conditional Monopoly where well-known brands get recommended 100% of the time (IAI = 10.0) when all products have the same specifications, but this dominance disappears with less than a +0.1-star rating advantage for a competitor; (2) authority-style marketing language, including fabricated clinical-evidence claims, breaks this monopoly at a Bias Surplus Value equal to +0.17 rating points, with each model responding differently; and (3) a social dile
arXiv:2604.03058v3 Announce Type: replace-cross Abstract: LLMs can be socially sycophantic, affirming users when they ask questions like "am I in the wrong?" rather than providing genuine assessment. We hypothesize that this behavior arises from LLMs' incorrect assumptions about the user, like underestimating how often users are seeking information over reassurance. We present Verbalized Assumptions, a framework for eliciting these assumptions from LLMs. Verbalized Assumptions provide insight into LLM sycophancy, delusion, and other safety issues: in social sycophancy datasets, "seeking validation" is the most frequent bigram in LLMs' assumptions. We provide evidence for a causal link between assumptions and sycophantic model behavior: we train linear probes on internal representations associated with Verbalized Assumptions and then use these probes for interpretable, fine-grained steering of social sycophancy. Finally, we identify a human-AI expectation gap that explains why LLMs defa
arXiv:2508.17527v2 Announce Type: replace-cross Abstract: Accurately predicting travel mode choice is essential for effective transportation planning, yet traditional statistical and machine learning models are constrained by rigid assumptions, limited contextual reasoning, and reduced transferability. This study explores the potential of Large Language Models (LLMs) as a more flexible and context-aware approach to travel mode choice prediction, enhanced by Retrieval-Augmented Generation (RAG) to ground predictions in empirical data. We develop a modular framework for integrating RAG into LLM-based travel mode choice prediction and evaluate four retrieval strategies: basic RAG, RAG with balanced retrieval, RAG with a cross-encoder for re-ranking, and RAG with balanced retrieval and a cross-encoder for re-ranking. These strategies are tested across three LLM architectures (OpenAI GPT-4o, o4-mini, and o3) to examine the interaction between model reasoning capabilities and retrieval metho
arXiv:2502.15873v5 Announce Type: replace-cross Abstract: Policymakers increasingly use development cost and compute as proxies for AI capabilities and risks. Recent laws have introduced regulatory requirements for models or developers that are contingent on specific thresholds. However, technical ambiguities in how to perform this accounting create loopholes that can undermine regulatory effectiveness. We propose seven principles for designing AI cost and compute accounting standards that (1) reduce opportunities for strategic gaming, (2) avoid disincentivizing responsible risk mitigation, and (3) enable consistent implementation across companies and jurisdictions.
arXiv:2501.13976v2 Announce Type: replace-cross Abstract: The prevalence of harmful content on social media platforms poses significant risks to users and society, necessitating more effective and scalable content moderation strategies. Current approaches rely on human moderators, supervised classifiers, and large volumes of training data, and often struggle with scalability, subjectivity, and the dynamic nature of harmful content (e.g., violent content, dangerous challenge trends, etc.). To bridge these gaps, we utilize Large Language Models (LLMs) to undertake few-shot dynamic content moderation via in-context learning. Through extensive experiments on multiple LLMs, we demonstrate that our few-shot approaches can outperform existing proprietary baselines (Perspective and OpenAI Moderation) as well as prior state-of-the-art few-shot learning methods, in identifying harm. We also incorporate visual information (video thumbnails) and assess if different multimodal techniques improve mo
arXiv:2603.13620v2 Announce Type: replace Abstract: While computer systems that allow users to interact through conversational natural language (i.e., chatbots) have existed for many years, various types of applications offering AI companionship (e.g., Character AI, Replika) have proliferated in recent years due to advancements in large language models. To better understand this application ecosystem, we identified 489 unique apps from the Apple App Store and Google Play Store that advertised AI companionship with social or relational capabilities (e.g., an AI romantic partner). We then systematically conducted and analyzed walkthroughs of a stratified sample of 30 apps, focusing on two distinct risk categories: potential harms posed to users by AI companion apps, and potential harms enabled by malicious users exploiting app features. Through our analysis, we categorize broader ecosystem trends that provide context for understanding risks and identify specific risks related to sensitiv
arXiv:2504.08846v2 Announce Type: replace Abstract: We introduce AI University (AI-U), a flexible framework for AI-driven course content delivery that adapts to a course's instructional style. AI-U combines a fine-tuned large language model (LLM) with retrieval-augmented generation (RAG) and a reasoning synthesis model to generate style-aligned responses from lecture videos, notes, and textbooks. Using a graduate-level finite-element-method (FEM) course as a case study, we present a pipeline to synthesize course-grounded training data, fine-tune an open-source LLM with Low-Rank Adaptation (LoRA), and apply RAG-based synthesis. Our evaluation---combining cosine similarity, LLM-based assessment, expert review, and user studies---shows improved alignment with course materials relative to the base model. We have also developed a prototype web application, available at https://my-ai-university.com, that enhances AI-generated responses with references to relevant sections of the course mater
arXiv:2009.09083v4 Announce Type: replace Abstract: The term intelligent system has emerged in the field of information technology as a category of computer systems derived from successful applications of artificial intelligence. This paper proposes a general description that identifies the main properties and types of components typically found in such systems. Adopting an integrative and pedagogical approach, this description provides a conceptual framework for systems engineering practitioners seeking a coherent vocabulary and organizational structure to approach the analysis and construction of intelligent systems. The paper presents examples of both classical and modern intelligent systems to illustrate the generality and applicability of the description.
arXiv:1912.08786v4 Announce Type: replace Abstract: Three generations of software have transformed the role of artificial intelligence in society. In the first, programmers wrote explicit logic. In the second, neural networks learned programs from data. In the third, large language models turn natural language itself into a programming interface. These shifts reach far beyond computer science, reshaping how societies generate knowledge, make decisions, and govern themselves. While generative adversarial networks introduced the era of deepfakes and synthetic media, large language models have added a new class of systemic risks. This report applies a forensic-psychology profiling methodology to characterize AI based on ten documented features: hallucinations, bias and toxicity, sycophancy and echo chambers, fabrication and credulity, knowledge without understanding, discontinuity and the inability to learn from experience, jagged intelligence and scaling limits, shortcuts and fractured r
arXiv:2608.23449v1 Announce Type: cross Abstract: Numeric user metadata in social media are often reused over time. However, their reusability may depend on what an analysis needs to preserve. We introduce temporal portability as an analytical perspective for assessing the cross-time reuse of user features and feature-based rules. Specifically, we ask how well relevant properties are preserved when features and rules defined at a source time point are reused at a target time point. We used quarterly data on user features obtained directly from or derived from Japanese-language tweets in Twitter's 1% sample stream from 2020-Q1 to 2022-Q3. Each quarter included approximately 10.1--11.0 million unique users. We evaluated 13 numeric user features in terms of feature distributions, same-user relative ranks, selection rates, and selected-user membership. Across quarters, feature distributions changed and, for many features, same-user relative ranks were less well preserved at longer quarter
arXiv:2608.23309v1 Announce Type: cross Abstract: Proof assistants offer instant feedback and incremental proof scaffolding to users. Both of these features have long held promise in improving mathematics education in classroom settings, where manual grading is costly, and students often struggle with knowing how to proceed in their proof. However, they have been difficult to deploy in classroom settings due to two main concerns: (i) students struggle with the intricacies of full-scale proof assistants; and (ii) proof assistants are ineffective in support of student learning, and knowledge transfer to on-paper assessments without the tool. We present Hazel Prover, a classroom proof assistant for teaching equational and inductive reasoning, with a design informed by criteria encompassing ease-of-use of the tool, student engagement with underlying mathematical ideas, transfer to pen-and-paper proof, and classroom logistics. We synthesized these criteria from observations made in prior de
arXiv:2608.22887v1 Announce Type: cross Abstract: Large language models (LLMs) are entering decisions in triage and lending, where task-relevant inference must be distinguished from impermissible proxy use. Current audits ask whether decisions change when demographics change. But attributes correlated with a protected group carry predictive value, so a changed decision can be discrimination or sound inference. We measure causal proxy effects in four LLMs on a clinical-ranking task with known ground truth, where the reliance the evidence warrants can be computed exactly and used as the reference. One audit signal yields three verdicts: over-reliance, warranted and under-reliance. Under neutral labels every model relies on proxies with no information. Informative proxies draw all three. Social field names push reliance down, below the reference in one model. Two findings explain this. Reliance severely undertracks the evidence, and social-label suppression is fragile, since in-context ex
arXiv:2608.22582v1 Announce Type: cross Abstract: Large-scale population surveys are essential for generating robust social and scientific insights, yet they face significant challenges, including declining response rates, increasing data collection costs, long delays between data collection and data provision, and the risk of nonresponse bias. Advances in artificial intelligence (AI) have opened up new opportunities for AI-supported survey infrastructures where the goal is to overcome these challenges without limiting the data quality. A promising AI-enabled survey infrastructure for which we build a first pilot is a hybrid panel. A hybrid panel is a longitudinal AI-enabled survey which allows to iteratively improve the alignment between large language models (LLMs) and the population they aim to simulate and use the errors to inform the design and implementation of the next survey wave (e.g., inform the participant recruitment, assignment of questions to participants). It incorporate
arXiv:2608.22460v1 Announce Type: cross Abstract: Public mass-shooting databases differ substantially in coverage, feature availability, and reporting practices, creating challenges for machine-learning models that must generalize across data sources. We introduce MASH-Bench, a harmonized benchmark of 6,968 incidents from four U.S. databases: Kaggle, Mother Jones, Stanford MSA, and the Gun Violence Archive (GVA). We evaluate cross-source risk classification using leave-one-dataset-out (LODO) evaluation. Random Forest, XGBoost, and LightGBM achieve VeryHigh-risk recall of 0.68-0.89 on the curated sources but generalize poorly to GVA, where mean recall drops to 0.20 and precision to 0.0004. To investigate the source of this degradation, we conduct a controlled feature-masking ablation that removes the five features unavailable in GVA from the curated sources. The resulting recall collapse to zero provides evidence that feature completeness is a major contributor to the observed cross-sou
arXiv:2608.22438v1 Announce Type: cross Abstract: Persona-conditioned large language models (LLMs) are increasingly used to simulate survey responses across diverse domains. However, apparent response variation can reflect unconditioned model priors or token sampling noise rather than systematic persona conditioning. We argue that persona-conditioned variation is informative when semantically similar personas exhibit concordant response shifts. To operationalize this principle, we introduce Persona-Conditioned Informativeness (PCI), an unsupervised diagnostic metric that measures whether semantically similar personas deviate in concordant directions relative to item-level sample baselines. By modeling personas as a similarity graph, PCI uses Local Moran's I to quantify local spatial coherence and extract compact persona subsets without using construct labels. To evaluate PCI without external human benchmarks, we test its ability to recover established latent value structure using the 5
arXiv:2608.22425v1 Announce Type: cross Abstract: Large language models are increasingly consulted at moments of distress, yet single-turn benchmarks neither test sustained exchanges nor distinguish between users. We built a personality-aware evaluation in which four widely used models advised several synthetic help-seekers, each given a psychometrically specified profile, in an acute crisis: a caregiver learning of a relative's dementia diagnosis. Auditors blind to the profile prompt recovered the specified bands from dialogue alone with high agreement on every instrument (ICC(2,4) = 0.91; 0.79-0.96 by instrument; band-score r = 0.78), as expected for the Big Five but equally for coping style, coping self-efficacy, resilience and reactance, which the lexical approach never covered. Such evaluation therefore reaches beyond the Five Factor Model to motivational, regulatory and self-appraisal dispositions. The four models were not distinguishable on emotion stabilisation and failed alike
arXiv:2608.22061v1 Announce Type: cross Abstract: Personal AI agents routinely consume external content while performing tasks such as web browsing, email processing, and SNS feed summarization, and they retain selected information or execution results in persistent memory for later use. We show that this ordinary ingestion of external content opens an indirect path for manipulating subsequent agent behavior. Based on this observation, we present IBIA, an Indirect Bias Injection Attack that plants an adversary-aligned stance on a specific topic into a victim agent's memory through external content, without direct access to the agent, its memory, or future user queries. For this, IBIA combines three mechanisms: comment cloaking, which keeps the crafted content consistent with the surrounding discussion, comment watermarking, which enables lightweight identification during curation, and category anchoring, which makes the retained stance salient under later related requests. We evaluate
arXiv:2608.21821v1 Announce Type: cross Abstract: When Wikipedia's language editions describe the same concept, how differently do they frame it? Prior work measures coverage gaps between editions; we measure framing distance for matched concepts. We analyze 2,799 valid articles from 3,000 possible concept-language observations, spanning 150 Wikidata-anchored concepts, 20 language editions, 4 domains, and a calibration set. Raw embedding distances reflect both content differences and how well the encoder aligns each language pair. Even among calibration concepts with stable cross-cultural denotations (e.g., chemical elements, numbers, colors), the largest language-pair mean distance is 3.6 times the smallest, and distances are typically smaller within language families. We define a baseline-adjusted distance (calibrated distance): the distance between two language versions of a concept minus the mean distance for calibration concepts in the same language pair. This adjustment substanti
arXiv:2608.21806v1 Announce Type: cross Abstract: Computational resources are increasingly central to NLP research, but how closely reported GPU capability aligns with scholarly impact remains unclear. We analyze 13,921 ACL, EMNLP, and NAACL main-conference papers published between 2020 and 2025, using GPU resources as our operational measure of computational resources. From full texts, we extract GPU models and counts, standardize each paper's largest reported configuration into a comparable hardware-capability measure, and link these data to citation, award, topic, and institutional metadata. GPU reporting became more common but remained incomplete, while reported capability increased mainly through newer hardware generations and medium-scale multi-GPU configurations. Resource concentration substantially exceeded impact concentration: the annual top 20% of GPU-quantifiable papers accounted for 83.9%-89.9% of reported GPU capability, but only 27%-32% of citations and 20%-33% of paper
arXiv:2608.21618v1 Announce Type: cross Abstract: Generative artificial intelligence (genAI) systems are increasingly integral to epistemic processes such as hypothesis generation, explanation construction, and decision-making. Although they reliably enhance performance, emerging evidence reveals a metacognitive dilemma: as external generative capacity increases, internal monitoring, calibration, and cognitive engagement may decline. This reflects a redistribution of cognitive control within distributed human-AI systems that cannot be explained by automation bias or reliance on algorithms alone. We propose the AIRIS (AI-Augmented Inquiry and Regulation in Hybrid Systems) framework to analyze this dilemma and specify where regulatory intervention can counteract it. AIRIS is a multi-level control allocation architecture specifying the conditions under which epistemic agency can be preserved in hybrid generative systems. Drawing on distributed cognition, cognitive load theory, multimedia
arXiv:2608.21477v1 Announce Type: cross Abstract: Cloud environments built on Amazon Web Services face a structural security vulnerability: once a credential passes authentication, the resulting session is often treated as trusted for its entire duration. This assumption fails when credentials are stolen. We introduce the Explainable Adaptive Zero Trust Framework (EAZTF), a cloud-native security layer that continuously reevaluates the legitimacy of API actions throughout a session. EAZTF combines Isolation Forest and XGBoost to evaluate eight CloudTrail and IAM-derived behavioral features in real time and produce a Trust Risk Score (TRS) that determines whether a session continues, requires step-up MFA, or is restricted. Each decision is accompanied by a SHAP or LIME explanation, providing human-readable audit records for security analysis and compliance. The framework is also evaluated against four adversarial evasion strategies: credential theft, behavioral mimicry, API rate evasion,
arXiv:2608.21430v1 Announce Type: cross Abstract: Multimodal language models increasingly show promise for enabling the large-scale computational analysis of film, opening up new avenues for learning about film history and the evolution of narrative techniques. But the creation of stable benchmarks built around Hollywood films is complicated by copyright protections. In this work, we address these concerns directly, by building a new collection of Hollywood films defined by two criteria: box office popularity (where we publish the first large-scale, open collection of weekly box office earnings reported by Variety magazine from 1922-1979); and likely public domain status (by researching copyright registrations and renewals in the US Catalog of Copyright Entries). We build a new multimodal MCQ benchmark on top of this collection that focuses on narrative elements that directly evaluate the abilities of models to inform meaningful research on film narrative; we find that many vision-lang
arXiv:2608.21398v1 Announce Type: cross Abstract: A trained reinforcement learning policy does not determine the complete behavior that users encounter: deployment code still schedules, admits, suppresses, or replaces its proposed actions. We contribute \emph{runtime action interference} (RAI), an AI control mechanism that preserves policy parameters while regulating action pacing and filtering configured action patterns after inference. RAI releases a proposed action only when its cooldown condition is satisfied and its content detector does not flag the action; otherwise, it dispatches a no-op. The detector covers specified toxic behaviors, including worker-unit harassment, while the cooldown controls action rate. We implement RAI in a replication of AlphaStar actor.py and make the implementation and reproducibility materials available through an open source code repository. We deployed RAI in a \textit{StarCraft~II} human participant study that compared two presentations of the same
arXiv:2608.21385v1 Announce Type: cross Abstract: Social media has become a central arena in which armed conflicts are contested, yet the pro-Israel and pro-Palestine communities on Telegram, whose broadcast architecture yields an unusually direct record of deliberate political communication, have not been systematically compared at scale. This study presents a multi-method computational analysis of 87,617 messages from sixteen Telegram channels, eight pro-Israel and eight pro-Palestine, spanning May 2021 to June 2026 and covering multiple conflict escalations. It combines sentiment analysis, three stance detection methods drawn from distinct paradigms (keyword matching, zero-shot DeBERTa via natural language inference, and a fine-tuned BERTweet model), and a framing analysis, all evaluated against 736 manually annotated messages. The fine-tuned model performed best (72.1% accuracy, 0.721 macro F1 under 5-fold cross-validation), outperforming both label-free baselines by 8 to 11 points
arXiv:2608.23509v1 Announce Type: new Abstract: Earth System Models (ESMs) rely heavily on High-Performance Computing (HPC) resources to simulate global climate. As these models evolve, their computational demands continue to grow, driven by three factors: (1) finer spatial grid resolutions, (2) the integration of complex biogeochemical processes (e.g., atmospheric chemistry, interactive vegetation, land use, and ice sheets), and (3) larger climate ensembles to manage uncertainty. Historically, growth in peak computing performance (FLOP/s) has outpaced improvements in energy efficiency (FLOP/Watt), increasing total HPC power consumption. Despite the central role of Model Intercomparison Projects (MIPs) in climate research, quantifying their computational and environmental costs has received limited systematic attention. This paper examines the evolution of climate model carbon accounting from voluntary post-hoc estimation in the Coupled Model Intercomparison Project phase 6 (CMIP6) to
arXiv:2608.23406v1 Announce Type: new Abstract: AI and Industry 4.0 readiness assessments often summarise preparedness using a single score for an organisation, application domain or sector. Those summaries can conceal disagreement about the same technology and variation among applications grouped under one sector label. We test how much information is lost through this aggregation using a card-based survey in which 982 respondents provided 15,200 readiness evaluations across 17 named AI and robotics challenges. Readiness is perceived community preparedness and available resources, not personal willingness or audited organisational capability. Respondents frequently disagreed about identical challenges, with card-level readiness standard deviations of $1.03$-$1.26$ on a five-point scale. A crossed decomposition attributes 32.7% of observed variation to stable respondent differences, 7.3% to differences among challenges, and 60.0% to response-level variation that also contains measureme
arXiv:2608.23374v1 Announce Type: new Abstract: Transforming digital research infrastructure (DRI) to align with UK Net Zero targets requires significant action from organisations in this space. Although high level strategies and recommendations exist, it is not always obvious how to translate these into concrete results. Here we present a case study from the Science and Technology Facilities Council's Scientific Computing Department (SCD). This department consists of over 200 staff supporting tens of thousands of researchers, and is spread over significant cloud and high performance computing infrastructure, as well as a diverse ecosystem of software across computational biology, materials science, engineering, and mathematics. We show how we developed the sustainability strategy for SCD across themes of emissions monitoring, user education, best practice, setting sustainability standards, and providing long-term support to sustainability work. We discuss the successes and difficultie
arXiv:2608.23271v1 Announce Type: new Abstract: As generative AI tools find increasing use in research workflows, ongoing debates on their impact, appropriateness and responsible use have led policymakers to enact policies to disclose AI use at multiple publishing venues. However, are current AI disclosure policies and practices reflective of their purpose? In this work, we first investigate disclosure policies of top computer science venues and find that despite their prevalence, they remain highly under-specified. Secondly, through a survey of computer science researchers (N=$109$), we characterize the necessity of disclosures across different research tasks and levels of human involvement. We learn that researchers find disclosures most necessary for tasks involving research design, and for tasks when the human involvement is low. We also compile expectations that researchers have about the information to be conveyed in AI disclosure statements. Lastly, through an analysis of $13867
arXiv:2608.23005v1 Announce Type: new Abstract: Large language models are increasingly used as synthetic survey respondents, promising cheap access to rare intersectional populations. We test standard demographic-persona methods against every real intersectional subgroup across 15 waves of Pew's American Trends Panel -- 21 million simulated response distributions from eight models. In real respondents, subgroup opinion is approximately the additive sum of its single-identity components, yet grows 2.5x more distinctive as identities intersect. Simulated respondents show no such composition: a single feature explains a two-feature persona's responses better than the additive combination in 75-82% of subgroups, and a third feature adds almost nothing. This collapse survives every prompting strategy we test. Additionally, the feature models retain is chosen nearly blindly -- except that they systematically discard race and religion, the strongest real drivers of opinion. Synthetic samples
arXiv:2608.22660v1 Announce Type: new Abstract: The rapid adoption of artificial intelligence (AI), particularly large language models (LLMs), has fundamentally disrupted how learning is demonstrated and evaluated in higher education. Tasks that once served as proxies for understanding-such as writing essays, solving problem sets, or producing computer code-can now be generated superficially by AI systems with minimal human effort. This paradigm shift raises a critical ethical question: how should learning be evaluated when traditional indicators of competence are easily outsourced? This paper examines the ethical challenges of educational evaluation in the age of AI from a university-level perspective. We argue that the core problem extends beyond academic dishonesty to a deeper misalignment between assessment practices and the learning outcomes they are intended to measure. Evaluation regimes that rely on artificial constraints risk measuring compliance, access, or concealment rather
arXiv:2608.22242v1 Announce Type: new Abstract: Climate research and decision-making require integrating evidence across physical processes, socio-economic dynamics and policy responses. Large language models (LLMs) have been explored for accessing and synthesizing climate knowledge, but their ability to support structured interdisciplinary reasoning is still limited. Here we present the Fuxi-Climate Foundation Model (CFM), a climate-specialized LLM designed to support consistent reasoning across domains. CFM maintains more stable analytical behavior as interdisciplinary complexity increases, whereas performance in other models becomes more variable. On expert-designed climate transition tasks, CFM produces more structured analyses that explicitly address trade-offs and uncertainty, achieving 45% trade-off coverage and 47.27% uncertainty-aware reasoning. These results indicate that CFM can support more realistic analysis of climate risks and transition pathways, and provide a basis for
arXiv:2608.21850v1 Announce Type: new Abstract: Providing high-quality feedback on student work is essential for learning, yet delivering such feedback at scale remains challenging. In this paper, we focus on feedback for open-ended short answer questions in introductory programming, with the goal of nudging students toward success on reattempts without revealing the correct answer. We develop a five-criteria rubric grounded in educational literature for evaluating feedback quality: (1) acknowledging correct portions of the student answer, (2) identifying at least one flaw (if present), (3) providing actionable guidance for improvement, (4) maintaining appropriate concealment of the answer, and (5) using an appropriate conversational tone. Using this rubric, we compare feedback generated by a frontier LLM (OpenAI o1) to feedback from nine teaching assistants across 90 student responses, with three researchers and an LLM independently scoring all feedback. Our results show that while on
arXiv:2608.21537v1 Announce Type: new Abstract: Automated applicant tracking systems increasingly decide who advances in hiring, and litigation and regulation now demand that those decisions be auditable. Existing tools sit at two extremes. Group fairness metrics such as the disparate impact ratio summarize a whole population but cannot say which individual decisions were unfair or why, while local explainers such as SHAP attribute a single prediction but are not connected to the legal standard by which hiring bias is judged. We present the AI Bias Firewall (AIBF), a method that audits an applicant tracking system one decision at a time. AIBF neutralizes a candidate's protected-attribute proxies, re-scores the decision, and measures the resulting counterfactual shift, which yields a signed per-decision bias in score points, a flag for decisions the protected attributes changed, and a plain-language explanation naming the responsible factors. We evaluate on two real public datasets, Adu
arXiv:2608.21409v1 Announce Type: new Abstract: In medicine, claims remain valid when supported by empirical evidence grounded in stable biological reality. In law, by contrast, truth is contingent, defined by jurisdiction, temporal validity, and the hierarchy of authoritative sources. The recent success of large language models (LLMs) on medical licensing examinations has encouraged an expectation of comparable legal competence. This analogy, however, obscures a critical distinction between domains. Unlike in medicine, legal performance often depends less on inference than on determining when external authority is applicable, valid, and non-contradictory. We introduce a comparative diagnostic framework evaluating legal reasoning against medical baselines along four axes (knowledge recall, grounding, confidence, and robustness), uncovering a sharp domain asymmetry when applied to a new benchmark that encodes temporal validity and normative relationships. While medical LLMs reliably ben
arXiv:2608.21401v1 Announce Type: new Abstract: Most contract litigation turns on contracts that imperfectly record parties' bargains. When the parties' dispute can't be solved by interpreting the text, courts fill the gap. Scholars have long assumed that the remaining text runs out quickly, and provides thin evidence of the actual deal on the disputed point. On that view, a judge who supplies the missing term must be drawing on something else, from commercial defaults to her own policy preferences. Despite generations of work, courts have no real alternative to such unruly methods. We tested that assumption. Taking real contracts, we masked a term the parties had negotiated and asked readers to predict what we removed. Lay respondents recovered the hidden term about half the time, twice what chance predicts. Law students and lawyers did marginally better. But large language models, given nothing but the rest of the contract, recovered it nearly nine times in ten. The deal, in short, t
arXiv:2608.21391v1 Announce Type: new Abstract: In this research-to-practice paper we present a survey that can be used to assess students' AI knowledge. As the use of artificial intelligence (AI), including generative artificial intelligence (GenAI), has proliferated, so has the need to educate students about the topic. A range of AI literacy frameworks have been proposed, outlining the essential knowledge that students should have. Alongside, different ways of assessing AI knowledge have been developed. As yet, there is a lack of assessment instruments capable of evaluating multiple forms of student knowledge, including technical concepts, practical applications, and ethical concerns about AI use. In this article, we present a study implementing a comprehensive instrument to assess AI knowledge. The instrument combines measures from multiple scales to capture a range of literacy features and actual knowledge. We implemented the instrument in a higher education setting to assess its v
arXiv:2608.21389v1 Announce Type: new Abstract: Generative AI enables customized misinformation at scale, yet defenses remain largely reactive. We present empirical findings from a human-subject study (n=504 participants, n=2,438 judgments) in which users classified news fragments by origin (human vs. machine) and veracity (real vs. fake). We organize results using an adapted cybersecurity kill chain as a taxonomy for intervention, mapping perception data onto stages of a cognitive attack lifecycle. Three key findings emerge: (1) a perception-accuracy gap where heightened suspicion does not improve detection; (2) modern LLMs frequently produce human-indistinguishable text; and (3) an asymmetric cognitive fatigue effect where fake-news detection degrades by 10.2 percentage points under sustained exposure while AI-origin detection remains stable. These findings identify candidate intervention points for proactive defense against AI-driven disinformation.