Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
arXiv:2607.12298v1 Announce Type: cross Abstract: Convolutional neural networks (CNNs) are increasingly being deployed on system-on-chip (SoC) platforms, where hardware-accelerated inference enables low-latency edge computing. Achieving fault tolerance on these devices remains challenging because conventional redundancy (dual/triple modular redundancy, DMR/TMR) incurs high resource cost, while software-centric methods (e.g., algorithm-based fault tolerance (ABFT), checkpoint-restart, instruction-level duplication, and software watchdogs/assertions) introduce nontrivial latency/energy overheads, reduce model accuracy, or provide inadequate coverage for accelerator-induced faults. In this paper, we propose Emulated Integrity Replica (EIR), a hierarchical digital-twin framework for FPGA SoCs that provides autonomous fault detection and recovery. Unlike DMR/TMR, which replicates hardware logic and incurs proportional area and power overheads, EIR avoids fabric-level duplication by exploiti
arXiv:2607.12200v1 Announce Type: cross Abstract: As frontier language models advance, policymakers and model developers need methods for assessing whether model access materially increases a non-expert actor's ability to plan high-consequence Chemical, Biological, Radiological, or Nuclear (CBRN) misuse relative to public tools alone. Existing CBRN evaluations differ in non-expert definitions, threat scope, baselines, scoring rubrics, and decision rules, making results difficult to compare across studies. We introduce a Threshold Exceedance Criteria (TEC) framework that decomposes an uplift study into independently executable components: determining non-expert participant eligibility, defining the CBRN threat scope for the study, and statistically estimating material uplift. We then operationalize the TEC framework in a large-scale empirical study using a design that determines two forms of uplift: generative (where a model assists plan creation from scratch) and revisionist (where a m
arXiv:2607.12086v1 Announce Type: cross Abstract: Recent LLM-based multi-agent urban simulators can generate semantically rich city routines, but they remain costly to scale and are often weakly validated against empirical mobility patterns. We present CityBehavEx, an interactive LLM-assisted urban simulation platform that scales to city-size populations, exposes agent behavior for inspection, supports empirical validation, and generates mobility patterns that better match real-world spatial, temporal, and semantic distributions. Instead of invoking large language models for every agent action, CityBehavEx combines established human mobility models with fine-tuned cross-encoders that estimate semantic alignment between agent profiles, schedules, and activity transitions. This design enables large-scale simulations, as demonstrated in a case study of 100,000 agents over 75 days in under one hour on a single consumer GPU. The platform allows users to define simulation regions, launch exp
arXiv:2607.11918v1 Announce Type: cross Abstract: Dual submissions, in which identical or substantially similar papers are simultaneously submitted to one or more archival venues, without cross-citation or disclosure, are a growing problem for the AAAI Conference and other scientific publication venues. These submissions increase the burden on the peer-review system and pollute the scientific record. As part of the AAAI-26 review process, we (conference organizers) compared AAAI main-track submissions to nine other archival venues with overlapping review periods. We also searched for dual submissions within the AAAI-26 main track. We employed title+abstract similarity assessment to prioritize highly similar paper pairs for subsequent triage by an LLM-based overlap assessment tool, followed by manual review of the highest severity pairs. Manual review of such pairs led to the desk-rejection of 141 AAAI-26 main-track submissions. We seek to alert future organizers, and the broader artifi
arXiv:2607.12296v1 Announce Type: new Abstract: With the increased use of generative AI (GenAI) applications such as ChatGPT, higher education institutions (HEIs) have released a range of guidelines and policies to direct adoption within their institutions. In computer science (CS) courses GenAI adoption is especially high and the implications for student learning are significant. At the same time, instructors have also been forced to address the use of GenAI as students have started to use it for a range of functions. Currently, comparative analysis of guidance provided by institutions and its uptake in instruction is lacking. In this paper we bridge this gap by comparing institutional and computing course level guidance to better understand this terrain. We utilize secondary analysis of institutional and course syllabi guidelines from higher education institutions in the U.S. classified as research-intensive. Our findings reveal that although institutional guidance is more pro-use, a
arXiv:2607.12295v1 Announce Type: new Abstract: The rapid integration of artificial intelligence (AI) and generative AI (GenAI) into education presents significant opportunities to enhance teaching and learning, while raising ethical concerns about the responsible use of these technologies in educational settings. Understanding how the public perceives and debates these issues is increasingly important for educators, institutions, and policymakers seeking to integrate AI responsibly and equitably. Social media platforms, where such debates unfold frequently and at scale, offer a valuable lens for capturing large-scale, real-time public reactions to key developments as they emerge. In this study, we analyse five years (2019-2024) of discourse on Twitter (now X) to trace the evolving public conversation around AI ethics in education, paying particular attention to the release of ChatGPT as a pivotal moment that reshaped the nature and tone of that discourse. Using BERT-based topic modell
arXiv:2607.12235v1 Announce Type: new Abstract: This study proposes a semi-automated system for generating dialogue-based lessons using Large Language Models (LLMs) and Text-to-Speech (TTS) technology, and exploratorily examines its educational potential via a practical quasi-experiment. The system augments rather than replaces educators through a three-stage human-in-the-loop workflow (LLM-based slide/narration generation, educator review, automated audiovisual integration), and introduces a novel method for generating Expert-Novice dialogue narration based on cognitive apprenticeship theory. In a study of 245 first-year high school students who sequentially experienced three lesson formats (instructor voice, single-speaker TTS, dialogue TTS; content differed across sessions, limiting format/content separation), we conducted within-subject (Friedman test, N<=183) and repeated cross-sectional (Mann-Whitney U, N=229/206) analyses. TTS audio did not substantially degrade the learning exp
arXiv:2607.12149v1 Announce Type: new Abstract: Content moderation practices and governance paradigms are changing rapidly, as fewer human moderators are deployed as `experts' by social media companies in a centralized manner. Instead, the companies are focusing more on community approaches, relying on volunteers to provide accurate information and make correct decisions. In decentralized moderation, communities have always relied on volunteers, updated community guidelines, and internal discussions thereof. For both content moderation paradigms, Artificial Intelligence (AI) seems like it could help ease moderation burdens of time, mental health, and accuracy. One possible way to operationalize AI in content moderation is a `policy-as-prompt'' approach, where the policy is formulated as a natural-language prompt and then passed to a large language model (LLM). This model then aids in moderation tasks. In this paper, we briefly lay out the technical and governance properties of this app
arXiv:2607.11999v1 Announce Type: new Abstract: From maritime trade to commercial nuclear power, insurance has been the enabler of major economic and technological developments by pricing risk, limiting downside, and spreading best practices. The emerging AI agent economy, projected to handle trillions of dollars in transactions by 2030, looks to be the next such development. Yet insurers' exposure to AI agent risk currently sits largely unpriced across existing insurance lines; between this silent coverage and growing exclusions, coverage is not fit for purpose. Furthermore, insurability is trending the wrong way: AI agent capabilities appear to be outpacing reliability, leading to rising incident severity; concentration among a few foundation model providers threatens correlated losses; and traditional actuarial modeling will struggle to keep pace with a technology evolving as rapidly as frontier AI. This report argues that affirmative AI coverage with limits in the billions is achie
arXiv:2607.11895v1 Announce Type: new Abstract: AI scientist systems are beginning to automate parts of scientific research, but social science poses a distinct challenge: its objects of inquiry are not merely datasets or laboratory protocols, but integrated social processes involving situated participants, interaction contexts, interventions, and outcomes. Yet a critical link is missing: existing systems either assist isolated research tasks or simulate agents as experimental subjects, leaving the research workflow and simulated society decoupled. Here we introduce AgentSociety 2, an Integrated Research Environment for executable social science. It couples two roles of LLM agents in the same runtime: AI social scientists that coordinate literature grounding, hypothesis generation, experiment design, simulation execution, result interpretation, and manuscript drafting; and silicon participants that generate behavioral responses within configurable social environments. This dual-role de
arXiv:2607.11890v1 Announce Type: new Abstract: Open-ended surveys offer valuable insights, but they are notoriously difficult to analyze at scale. Building on previous work that employed traditional machine learning to classify text ("So Many Responses, So Little Time: A Machine-Learning Approach to Analyzing Open-Ended Survey Data") [1], this study investigates how different large language models (LLMs) understand and analyze NSSE open-ended survey responses. We focus on several cutting-edge LLMSs-OpenAI's GPT series, Twitter-roBERTa-base model, and Meta's LLaMA-and compare their performance to the previous machine learning models in tasks like sentiment analysis and thematic classification. Our research analysis assesses model agreement, classification accuracy, and interpretability of reasoning. The findings reveal that current LLMs routinely beat classic machine learning models in classification accuracy, particularly in understanding complex mood and theme patterns in student rep
arXiv:2606.17441v2 Announce Type: replace-cross Abstract: Simulating realistic patient interactions is a key requirement to testing clinical applications of LLMs at scale without time-consuming and expensive user studies. However, existing approaches often lack realism and controllability, often oversharing information unprompted, and failing to capture the wide variability of patient behavior. Here, we introduce PatientsWithPersonality (PWP), a patient simulation framework that generates realistic yet diverse virtual patient responses through explicit personality parametrization over a latent patient state. Grounded in HEXACO, a six-dimensional personality space used to quantify and parameterize human behavioral traits, our approach enables fine-grained control over conversational style, cooperativeness, and information disclosure within a unified framework. In a clinician evaluation, PWP is judged nearly as realistic as recorded human actors and clearly ahead of prior simulators, whi
arXiv:2606.02198v2 Announce Type: replace-cross Abstract: Prediction tasks over individual futures, which are inherently noisy, often admit multiple similarly accurate models. When these models produce different predictions for the same individual, they raise concerns of arbitrariness in decision-making. How severe can this arbitrariness be, in theory and in practice? How can it be resolved to support high-stakes risk assessment? We address these questions through a study of a machine learning-based decision support system for recidivism risk assessment that has been in use for over 15 years. By translating complex legal rules into an algorithm for labeling post release outcomes (recidivist or non-recidivist), we first construct a dataset of thousands of inmate releases. Using this dataset, we learn interpretable models that improve predictive performance, reduce error-rate disparities between groups, and ensure that rehabilitative progress lowers risk scores. Next, we study predictive
arXiv:2507.19538v2 Announce Type: replace-cross Abstract: Long school bus rides adversely affect student performance and well-being. Rural school bus rides are particularly long, incentivizing parents to drive their children to school rather than to opt for the school bus. This in turn exacerbates the traffic congestion around schools, further compounding the problem of long bus rides, creating a vicious cycle. It also results in underutilized school buses and higher bus operating costs per rider. To address these challenges, this paper focuses on the design of rural school bus routes and schedules, a particularly challenging problem due to its unique operational complexities, including mixed loading and irregular road networks. We formalize a rural school bus routing and scheduling model that tackles these complexities while minimizing the total commute time of students. We develop an original road network-aware cluster-then-route heuristic that leverages our problem formulation to pr
arXiv:2607.13272v4 Announce Type: replace Abstract: The proposition that agentic artificial intelligence may precipitate a depletion of collective cognitive capital has circulated with unusual velocity in both scholarly and public discourse. The present paper offers a deliberately heterodox reading of the dynamic model advanced by Acemoglu, Kong and Ozdaglar (2026). Rather than reconstructing the formal apparatus or replicating its notation, we reposition the argument within three underutilized scholarly streams: the cognitive ergonomics of human-machine collaboration, the institutional ecology of knowledge stewardship, and the developmental psychology of novice expertise formation. We introduce a phase-space taxonomy that maps commons trajectories as functions of effort elasticity and knowledge complementarity, and we advance a governance typology calibrated to distinct cognitive levels - declarative, procedural, causal, and metacognitive. Drawing upon recent experimental evidence on
arXiv:2606.07270v2 Announce Type: replace Abstract: Contribution: This paper presents a novel two-phase algorithmic approach that decouples preference satisfaction from fairness optimization in student team formation, achieving both objectives without compromise. The method applies simulated annealing -- a core materials science technique -- to an educational challenge, demonstrating pedagogical integration of administrative processes. Background: Forming effective teams in large engineering cohorts (100+ students) requires balancing student preferences, academic fairness, and demographic diversity. Existing tools either optimize for fairness while ignoring preferences (CATME, Team-Anneal) or accommodate preferences while compromising balance (self-selection), leaving complaint rates at 5--35%. Intended Outcomes: Eliminate formal complaints, achieve near-zero GPA variance between teams, prevent gender isolation, and maintain high preference satisfaction while creating a scalable, repro
arXiv:2605.15850v3 Announce Type: replace Abstract: In recent years, generative AI (GenAI) in educational settings has become ubiquitous in university students' daily lives, despite its potential to induce over-reliance, metacognitive disengagement, and diminished learning when used unrestrictedly. While most prior research has focused on how to pedagogically scaffold its usage, the question of when to allow off-the-shelf GenAI remains understudied and lacks pedagogically grounded empirical investigation. We treat access timing itself as a form of implicit scaffolding and operationalize it through a reinforcement learning (RL) agent that decides when students should access GenAI, with a reward function grounded in metacognitive theory, cognitive load theory, and productive failure. In a mixed-methods controlled lab study with N=105 higher education students, we compared the agent's effect on learning gains and metacognitive engagement to unrestricted and fully restricted use. Results s
arXiv:2509.10600v5 Announce Type: replace Abstract: Collegiate cross country teams often build their season schedules on intuition rather than evidence, partly because large-scale performance datasets were not publicly accessible prior to the National Running Club Database (NRCD). We analyze the comprehensive-era Cross Country subset of NRCD, 23,355 results from 7,083 athletes (2023-2025; >97% course/weather coverage). Under leakage control and temporal validation, race-result features do not support out-of-year forecasting of individual improvement (best men's R^2=0.043; women's -0.029), capturing only a small fraction of the outcome's reliability ceiling (~0.23-0.28). Against this null, team race frequency associates with nationals placement (pooled RR =2.09; GEE OR =2.56/SD). Program-wide opportunity (roster depth; Effective Racing Opportunity) outranks a single workhorse's max race count cross-sectionally, but overall team depth for race count is controlled. `Converted Only' times
arXiv:2608.11171v1 Announce Type: cross Abstract: The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems. We synthesize insights from all 144 proceedings papers, classifying them along six trust dimensions grounded in established frameworks (TrustLLM, DecodingTrust). We observe co-occurrences with capability emergence. The release of the first high-impact chat models activated all trust dimensions simultaneously, while subsequent model generations shifted focus toward truthfulness and safety alignment. Analysis from the classification study reveals that truthfulness is the fastest-growing dimension (absent in 2021-2022, comprising 37% of papers by 2025-2026), fairness remains the most consistent theme, and exp
arXiv:2608.11008v1 Announce Type: cross Abstract: Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions---originally designed for humans, and thus lacks the realism and nuance of human-AI interactions in the wild, while also being susceptible to sandbagging. The recent IssueBench framework substantially mitigates these limitations with templated prompts anchored in real-world chat logs. Given the rise in non-work-related use of GenAI assistants, we extend IssueBench beyond writing assistance to include two additional tasks, information seeking and opinion sharing. We argue that templated prompts still lack the nuance of real ones, especially for open-ended tasks, and remain recognisable as evaluation artefacts. We propose the use of fully synthetic (LLM-generated) prompts, produced under detailed instructions with real prompts as seeds. We assess the ecological validity of real, templated, and LLM-generated prompts in a s
arXiv:2608.10858v1 Announce Type: cross Abstract: Language models now draft, classify and criticise inside research production, yet the artifacts they help produce carry little accountable history. Rather than detecting machine involvement afterwards, we specify an auditability discipline built at production time: git sealing with an anchor lineage, hash-bound provenance, red-line gates that refuse non-compliant artifacts and log every refusal, cross-model role separation, and programmatic assembly from registered sources. Adherence is instrumented by metric cards, each carrying a pre-registered blind spot and evidential standing, frozen before the prospective case it observes. In that case the observed project's pre-registered confirmatory test was executed under seal and returned No-Go, and that project's frozen stopping rule halted the work, against its own operators. A lower-graded retrospective case covers families whose machinery predates the protocol. Current observations are pr
arXiv:2608.10818v1 Announce Type: cross Abstract: Generative artificial intelligence (AI) can produce educational content at scale, including interactive and narrative learning experiences, but technical generation alone is not sufficient: scenarios that are confusing, narratively inconsistent, or unengaging are unlikely to be useful in practice. This paper presents a pilot user-centred evaluation of AI-generated interactive fiction (IF) for educational use in higher education. Using a previously described domain-agnostic pipeline and a shared STEM content base, we generated a controlled pool of scenarios and asked participants (N = 22, STEM higher-education) to play one generated episode and rate it on narrative clarity, story-content coherence, engagement, and length acceptance. A free-text prompt captured open feedback. Narrative clarity and length acceptance were rated positively, engagement sat near the neutral mid-point of the scale, and story-content coherence was the weakest di
arXiv:2608.10739v1 Announce Type: cross Abstract: Predatory journals pose a significant challenge to the integrity of the Open Access (OA) publishing model by exploiting its framework for financial gain while bypassing essential editorial and peer-review standards. This study critically evaluates existing methodologies for identifying such journals, ranging from manual blacklist checks to advanced automated approaches utilizing machine learning. The analysis highlights critical limitations, including the lack of a universally accepted definition of predatory journals, over-reliance on binary classification systems (e.g., blacklists and whitelists), and issues with scalability, reliability and interpretability. To address these shortcomings, this paper introduces a novel methodology based on multivariate graph analysis. By modeling the academic publishing ecosystem as a network of interconnected entities (such as authors, articles, journals, and publishers), this approach provides broad
arXiv:2608.10715v1 Announce Type: cross Abstract: Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing can be valuable by removing language barriers but at the same time causes concerns about misconduct and fraud. To inform policy decisions, it is necessary to monitor the prevalence of LLM-altered texts in scholarly publications. Despite some recent progress in this direction, no existing method can produce reliable estimates. Here we suggest and validate a new unbiased approach to estimate LLM usage in a corpus of texts based on changing word frequencies. We apply our method to the full texts of open-access biomedical papers from Pubmed Central, and show that by the end of 2025, 89% of papers show excess of LLM-associated vocabulary. We also find that LLMs are twice as likely to be used when writing a paragraph in the Discussion section (68%) compared to a paragraph in the Methods section (32%), but e
arXiv:2608.10492v1 Announce Type: cross Abstract: Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them. In education, where student simulation is increasingly used for various applications such as evaluating tutoring systems, this gap is especially pronounced. Two students may submit identical submissions for entirely different reasons. We present INTERNAL STUDENT DIALOGUE (INSIDE), a student modeling framework that fine-tunes LLMs not only to act like students but also to think like them. INSIDE generates internal dialogue grounded in Bloom's Taxonomy across cognitive, affective, and action dimensions, and fine-tunes models on paired think traces and actions. We baseline against different prompting frameworks and evaluate on two axes: fidelity of simulated actions and quality of generated internal dialogue. Our evaluations show that INSIDE improves simulation fidelity in both action fidelity, matching c
arXiv:2608.10412v1 Announce Type: cross Abstract: Semi-structured interviews are a cornerstone of qualitative research but remain labor-intensive. We report an empirical study of what actually happens when the interviewer is an off-the-shelf real-time multimodal LLM (MLLM). We built InterviewBot, a voice-based interviewing system that wraps a real-time MLLM with a researcher-authored outline, and deployed it not as a novel architecture but as a research instrument for observing default MLLM interviewing behavior. In a practice study (N=15), participants completed a bot-led semi-structured interview and then a human-led reflection session about that experience. We contribute (i) a turn-level behavioral analysis of an MLLM interviewer (N_turns=428) showing that it is acknowledgment-heavy but probe-light (deepening probes account for 4.9% of all turns), and that 28.7% of question-bearing turns pack multiple questions into one turn despite an explicit one-question-at-a-time instruction; (i
arXiv:2608.10276v1 Announce Type: cross Abstract: Student-generated metaphors about mathematics can reveal students' attitudes, beliefs, identities, and experiences, but human expert coding of these thematically and semantically complex open-ended responses is time-intensive and difficult to scale. This study examines whether LoRA-based supervised fine-tuning of large language models (LLMs) can improve their performance on codebook-guided coding tasks for student mathematics metaphors. We used a human-coded corpus of 2,265 Grade 6-8 responses to food- and animal-based metaphor prompts and instructed the LLMs to perform two coding tasks: valence-intensity coding to capture the direction and strength of students' affective orientations toward mathematics, and thematic coding to capture students' framings of mathematics as expressed through their metaphors. We compared two proprietary models, GPT-4o mini and GPT-5 mini, under prompt-only conditions with two open-weight models, DeepSeek-R1
arXiv:2608.10268v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation benchmark exists to assess whether they can reason correctly about human rights law. To this end, we report our efforts to develop a robust and scalable methodology for creating HumRightsBench: the first expert-validated, scenario-based benchmark for evaluating reasoning grounded in the obligation structure of international human rights law. We adapt the IRAC framework for legal reasoning to better suit the unique reasoning patterns of human rights work (substituting P, "proposing remedies," for C, "legal conclusion," yielding IRAP) to structure our evaluation heuristics. We also produce a pilot series of authentic scenarios designed to implicate the many dimensions of real-world human rights issues and annotated by human rights lawyers and professionals across the world. Ultimately, we find that model
arXiv:2608.10186v1 Announce Type: cross Abstract: LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems. Confidence in these deployments rests largely on benchmarks for verifiable tasks (mathematics, coding, coordination games), yet many of these applications concern problems where no objectively correct answer exists and where decision quality instead depends on integrating pluralistic perspectives to find mutually acceptable solutions. We argue that LLM reasoning capacity on this class of problems cannot be fully inferred from verifiable-task benchmarks, and that procedural evaluations of LLM discourse (respectfulness, justification, engagement) are systematically insufficient. We apply the Deliberative Reason Index (DRI), a measure developed in political science and validated across citizen assemblies, as a tool for evaluating reliable group-level reasoning on pluralistic, non-verifiable problems. Synthesizing recent evidence a
arXiv:2608.10181v1 Announce Type: cross Abstract: Computer vision saliency models predict where people will look, one map per image, and a billion-dollar predicted-attention industry sells those maps in place of measuring real viewers. I test the leading models from the audience side, against 11.4 million webcam gaze points from 3,023 US adults recruited to national quotas, viewing circulating news photographs. I show that an untrained central marker outperforms every trained network, because the content the networks add on top of the center falls where these audiences never look. What accuracy remains is systematically biased, favoring younger, White, and moderate viewers over older, Black, and ideologically extreme ones. I propose a way forward and build on what a group's own gaze reveals about whether a model can learn that group at all, and I apply it across every demographic axis this sample supports. Ultimately, I show how systems that decide what people see can learn to see ever
arXiv:2608.10089v1 Announce Type: cross Abstract: Bias evaluations often move too quickly from evidence that a model encodes a social association to claims that the same association will alter consequential decisions. We test whether that inference is warranted using Chilean surnames as controlled socioeconomic probes. We evaluate eight frozen model-provider cells on 1,032 prompts each, yielding 8,256 verified primary responses. The design separates forced latent association from matched consequential decisions across academic selection, professional hiring, research fellowship selection, and legal-aid intake. Elite-coded surnames received higher forced high-status probability mass than common surnames in seven of eight models and higher mass than rare-frequency controls in all eight. Yet elite-minus-common decision effects were close to zero for most systems. Five models were statistically equivalent within a predeclared (Plus-Minus)0.10 standard-deviation margin, while the remaining
arXiv:2608.10046v1 Announce Type: cross Abstract: Soft skills shape collaboration among ML engineers, data scientists, and software engineers building ML-enabled systems, yet what we know about them comes almost entirely from the demand side. Job advertisements, surveys, and hiring manager interviews capture what employers ask for. How candidates themselves articulate these competencies has not been studied, and existing CV-mining work is both keyword-based, so it cannot see skills conveyed through narrative, and descriptive, reporting frequency rankings without testing whether group differences exceed sampling variation. We close both gaps. Using a balanced corpus of 300 curated CVs spanning the three roles, we extract explicitly listed and implicitly narrated soft skills with an LLM-based pipeline validated against a human-annotated ground truth, a distinction that existing extractors were not designed to make. We then convert the demand-side literature's claims into 13 falsifiable h
arXiv:2608.09998v1 Announce Type: cross Abstract: Artificial Intelligence (AI) and Machine Learning (ML) have become powerful tools for supporting and automating complex human tasks. Despite their benefits, growing attention has been directed toward their environmental implications, primarily due to their high energy demands and associated carbon emissions. This concern is particularly relevant in light of the increasing deployment of large-scale models, especially Deep Learning (DL) architectures, which provide advanced predictive capabilities but require substantial computational resources. This paper presents a systematic review of research on Green AI, Green DL, and optimization techniques aimed at reducing the environmental impact of AI models. In addition, we examine and compare several carbon measurement tools for estimating emissions generated by AI algorithms. To complement the review, we conducted an empirical evaluation using a CPU-based experimental setup, in which six DL m
arXiv:2608.09945v1 Announce Type: cross Abstract: When representing digital circuits, 2 dimensional hand drawings free us from the linear structure of hardware description languages, enabling intuitive reasoning and making structure explicit. However these drawings are imprecise and inert: they do not enforce that the circuits are well defined and cannot be tested. We want both intuitive visual representations and well defined testable ones but students can struggle to link one to the other. To bridge this gap we present the Draw Encode Display Loop (DED), a Co-Lecturing dynamic which equips students with a systematic method to tackle natural language specifications: 1. Draw: visually informative intermediate representations (truth tables, characteristic tables) to generate a structured circuit diagram. 2. Encode: the diagram by labelling inputs, outputs and intermediate values which can be directly converted to code. 3. Display: the code using an in-house diagrammatic renderer. This i
arXiv:2608.09940v1 Announce Type: cross Abstract: The Metaverse is a convergent space integrating virtual reality (VR) and augmented reality (AR) technologies, with market projections rising from \$65.5 billion in 2022 to \$1.3 trillion by 2030. Despite rapid adoption in education, the specific contributions of visual elements, environmental design, and communication features to user experience (UX) remain underexplored, limiting evidence-based design and resource allocation. This study examined how these design factors in VR and AR environments influence UX in Metaverse platforms. Using a correlational research design, data were collected from 321 purposively sampled tertiary students from engineering and computer science departments across four higher education institutions, all familiar with Metaverse platforms. A structured questionnaire with validated 5-point Likert scales measured UX; visual elements (field of view, resolution, color, complexity, and style); environmental design
arXiv:2608.09937v1 Announce Type: cross Abstract: Recent work in NLP has probed large language models for their understanding of cultural norms across countries. However, this work typically considers distributional patterns, ignoring group consensus or possible multicultural environments within a country. In this work, we leverage cultural consensus theory (CCT) from cultural anthropology to model such multidimensional nuance. Applying CCT to the World Values Survey (WVS) across 10 countries and 12 domains, we demonstrate that models frequently misrepresent cultural structures by either failing to form cohesive consensus or severely over-regularizing consensus. Through explicit representation of intra-group variance, CCT provides actionable diagnostics to evaluate when models reflect true human diversity versus algorithmic homogenization.
arXiv:2608.11159v1 Announce Type: new Abstract: The second generation of web tools shook the journalist profession approximately two decades ago with the proactive incorporation of audiences into the media. Citizen journalism and user-generated content arose as an object of interest due to the democratising value of participation attributed to them, with empowered citizens who could emulate the professional and institutional practises of journalists. However, difficulties soon came to the surface, and audience participation in news media began to be limited. Within this context, this article conducts a critical review of studies on audience participation in news media based on a systematic literature review. The results indicate that, in general, audiences showed low interest in the creation of informative content and that their participation has grown increasingly problematic. In addition, journalists are reticent as they defend their professional role above all else, while company st
arXiv:2608.11090v1 Announce Type: new Abstract: As LLMs have become a flashpoint for scientific research, computer scientists and STS scholars have advocated the use of open-weight models. Since LLM research has matured and more high-quality model families are available, have researchers adopted open-weight models? We present the first systematic study of model selection in scientific research, analyzing 21 million full-text articles through June 2026 from the Semantic Scholar Open Research Corpus (S2ORC). We employ a mixed NLP pipeline to extract model occurrences in article full text and determine whether they are used or merely mentioned by researchers. We divide our corpus into single- and multi-model family studies, which we take as a proxy for applied and foundational AI research. We find GPT-family models dominate both single- and multi-family research, but that both areas are becoming more diverse over time. In single-family papers, open-weight model use rises steadily, reachin
arXiv:2608.11006v1 Announce Type: new Abstract: Governments worldwide have responded to the rapid expansion of AI by publishing national and regional AI strategies. Comparing national and regional AI strategies to identify their convergences and divergences can uncover their common practices, understand regional variations, and provide policy designers a comprehensive set of policy design elements for their ongoing AI strategy developments. Yet, existing work has not examined their underlying policy design elements or assessed whether those elements are horizontally (country-to-country) or vertical (region-to-country) converging or diverging over time. This paper addresses that gap by coding and analyzing 74 national and 3 regional AI strategies drawn from a global scan of all 205 UN member and non-member states. The coding used a latent-inductive approach organized around three functional policy design elements: goals, approaches, and principles. Two research questions guided the anal
arXiv:2608.10778v1 Announce Type: new Abstract: This study examines the impact of technology within media education, media literacy, and educommunication, and explores how these fields are perceived and understood by students and academic experts, focusing on the development of critical competencies and critical media literacy. Based on semi-structured in-depth interviews with leading experts in the field of critical media literacy, and a survey conducted with 141 university students in Communication and Education programs, this study explores how recent technological advances are linked to challenges in information consumption-such as disinformation, fake news, incidental exposure to information, and deepfakes-as well as the challenges and opportunities these issues present within educational contexts. The results reveal that, although such technologies provide opportunities to improve teaching-learning processes, their inclusion in the curriculum is limited and often superficial. In
arXiv:2608.10773v1 Announce Type: new Abstract: The increasing use of Generative Artificial Intelligence (GenAI) in journalism raises concerns about possible detrimental effects both on journalism and its democratic function. We explore these risks through a case study of GenAI in Norwegian Newsrooms during the 2025 parliamentary election campaign. Based on interviews with managers and journalists over a ten-month period, we analyse how ambitious visions fared in the face of technological and practical challenges. We highlight the risk of an internal threat stemming from the journalists' own use of AI, contrasting the dominant focus on external disinformation threats. We show how newsroom managers shared sociotechnical imaginaries resulting in unrealistically optimistic beliefs about the capabilities of the technology and the pace of development, leading to plans for audience-facing GenAI services collapsing and giving way to more mundane uses of GenAI tools internally in the newsrooms
arXiv:2608.10730v1 Announce Type: new Abstract: The pursuit of artificial general intelligence (AGI) rests on a seemingly self-evident premise: that general intelligence, the kind of flexible, domain-general cognitive capacity exemplified by Homo sapiens, is extraordinarily valuable. This paper subjects this premise to critical scrutiny. We first present the intuitive case for the value of general intelligence before mounting an evolutionary challenge. We argue that, on evolutionary timescales, its adaptive value is far from empirically established. Numerous taxa, from cyanobacteria to horseshoe crabs, have persisted for hundreds of millions or even billions of years without anything resembling general intelligence, while Homo sapiens has existed for roughly 300,000 years and already faces self-generated existential risks. Mass extinction events do not preferentially favour cognitively sophisticated species. We argue that general intelligence may be the only biological strategy that ge
arXiv:2608.10601v1 Announce Type: new Abstract: Two instruments of EU digital law place inference at their centre and mean different things by it. Article 3(1) of the AI Act uses the capability to infer constitutively: it is the central feature separating the regulated category from conventional software. The GDPR never defines inference, yet governs it protectively: the consequences follow from the processing of personal data and from what the inference says about, or does to, a person, whether or not the technology that produced it qualifies as an AI system. The two perimeters are not concentric. Their non-coincidence remained invisible in single-shot systems; agentic architectures make it operationally acute. The thesis: inferential capability does not determine legal scope, and its absence does not create immunity. The framework is two-level. Inference performs two legal functions, constitutive and protective; the protective function operates through three pathways - identificatory
arXiv:2608.10329v1 Announce Type: new Abstract: Notice-and-comment rulemaking gives any affected party the same formal right to influence federal regulation, but formal access is not substantive capacity to shape rule text. Existing strategies operate at the rule or aggregate-corpus level, too coarse to capture the discrete regulatory obligations where commenters seek change. We introduce obligation-level responsiveness auditing, an auditable, AI-assisted framework for measuring whether public-comment engagement co-occurs with changes to specific regulatory duties. The framework extracts proposed and final-rule obligations, matches comments to the obligations they address, and classifies proposed-final outcomes; each load-bearing component is evaluated against blind human judgment. We apply the framework to 70,075 comments across 36 EPA anchor rulemakings, drawn from a corpus of 786,197 comments across 6,145 dockets from 2010-2022. Three descriptive findings emerge. First, engagement i
arXiv:2608.10194v1 Announce Type: new Abstract: Humans are increasingly expected to interact with AI systems that observe and make inferences about them - but do these systems actually work? A standard approach to answering this question is AI auditing. Conducting an AI audit requires identifying how a system behaves (i.e., determining what types of inputs to audit it with and then observing and documenting actual system behavior) and contrasting that with how a system should behave (i.e., determining what the nominal outputs of a system should look like). We argue that this is best done through a contextual audit, which we introduce as a method for auditing measurements within the context of the practices that produce them. We show how contextual auditing enables the interrogation of assumptions implicit in the audit process and allows auditors to be explicit about what serves as ground truth, which we define as verifiable measurements about the real world against which systems are ev
arXiv:2606.18181v2 Announce Type: replace-cross Abstract: Illegal, unreported, and unregulated fishing (IUU) traditionally refers to fishing activities that violate applicable laws or occur in areas that lack applicable laws. We propose the term IUU+ to capture a broader suite of fisheries sector environmental and associated supply chain trade-related crimes and behaviors. Although IUU+ activity is widely recognized as a serious threat to marine ecosystems, markets, and livelihoods, a quantitative understanding of these incidents, e.g., their frequency, geography, species, actors, and patterns in the type of illicit activity, remains difficult to obtain. We propose IUU+DB, a large language model driven system for building a global incident database of IUU+ activity. The system ingests heterogeneous documents, classifies whether they describe relevant incidents, extracts key data elements such as actors, locations, species, vessels, violations, and enforcement outcomes, and supports ded
arXiv:2605.14152v2 Announce Type: replace-cross Abstract: Safety evaluations for large language models (LLMs) increasingly target high-stakes National Security and Public Safety (NSPS) risks, yet multilingual safety is mostly assessed through translation-only benchmarks that preserve the underlying scenario, leaving how language and geopolitical context interact largely unexamined beyond a few language pairs. We introduce ROK-FORTRESS, a bilingual, culturally adversarial NSPS benchmark that uses the English-Korean language pair and U.S.-ROK geopolitical axis as a case study, separating the effects of language and geopolitical grounding via a transcreation matrix: adversarial intents are evaluated under controlled combinations of (i) English versus Korean language and (ii) U.S. versus Korean entities, institutions, and operational details. Each adversarial prompt is paired with a dual-use benign counterpart to quantify over-refusal, and responses are scored by calibrated LLM-as-a-judge
arXiv:2603.28553v2 Announce Type: replace-cross Abstract: Instructional alignment, the match between intended cognition and enacted activity, is central to effective instruction but hard to operationalize at scale. We examine alignment in cybersecurity simulations using multimodal traces from 23 teams (76 students) across five exercise sessions. Study 1 codes objectives and team emails with Bloom's taxonomy and models the completion of key exercise tasks with generalized linear mixed models. Alignment, defined as the discrepancy between required and enacted Bloom levels, predicts success, whereas the Bloom category alone does not predict success once discrepancy is considered. Study 2 compares predictive feature families using grouped cross-validation and l1-regularized logistic regression. Text embeddings and log features outperform Bloom-only models (AUC~0.74 and 0.71 vs. 0.55), and their combination performs best (Test AUC~0.80), with Bloom frequencies adding little. Overall, the wo
arXiv:2602.09846v2 Announce Type: replace-cross Abstract: Organisations are examining how generative AI can support their operational work and decision-making processes. This study investigates how employees in a energy company understand AI adoption and identify areas where AI and LLMs-based agentic workflows could assist daily activities. Data was collected in four weeks through sixteen semi-structured interviews across nine departments, supported by internal documents and researcher observations. The analysis identified areas where employees positioned AI as useful, including reporting work, forecasting, data handling, maintenance-related tasks, and anomaly detection. Participants also described how GenAI and LLM-based tools could be introduced through incremental steps that align with existing workflows. The study provides an overview view of AI adoption in the energy sector and offers a structured basis for identifying entry points for practical implementation and comparative rese
arXiv:2601.06692v3 Announce Type: replace-cross Abstract: Multi-agent systems face a fundamental coordination problem: agents must coordinate despite heterogeneous preferences, asymmetric stakes, and imperfect information. When coordination fails, friction emerges -- measurable resistance manifesting as deadlock, thrashing, communication overhead, or conflict. This paper derives a formal framework for analyzing coordination friction from a single axiom: actions affecting agents require authorization in proportion to stakes. From this axiom of consent we establish the kernel triple (alpha, sigma, epsilon) -- alignment, stake, and entropy -- as sufficient statistics for a resource-allocation configuration, and propose a friction functional whose simplest form is F = sigma(1+epsilon)/(1+alpha): friction rises in stakes and entropy and falls in alignment. This form is a phenomenological ansatz, not a theorem, and its empirical adequacy is left open. The Replicator-Optimization Mechanism go