EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

PySynthea: A Python-Native Framework for Scalable Synthetic Healthcare Data Generation

arXiv:2606.28346v1 Announce Type: new Abstract: Synthetic healthcare data is increasingly important for research, education, and machine learning development where access to real patient data is limited by privacy and governance constraints. While Synthea provides a widely adopted framework for generating realistic longitudinal electronic health record data, its current implementation presents adoption barriers for many researchers and data scientists due to deployment complexity and limited integration with modern Python-based workflows. This paper introduces PySynthea, a Python-native reimplementation of Synthea designed to improve accessibility, extensibility, and interoperability within the scientific Python ecosystem. The framework provides modular synthetic patient generation, configurable healthcare simulation pipelines, and support for standard healthcare data formats while integrating naturally with tools such as pandas and machine learning workflows. By reducing operational c

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

LLM-Ideoplasticity: Measuring Ideological Plasticity in the Political Behavior of LLMs as a Context-Conditioned Distribution

arXiv:2606.28335v1 Announce Type: new Abstract: We argue, with systematic empirical evidence, that a large language model's political ideology is not a fixed point, but a conditional distribution $\mathbb{P}($position$\mid$context$)$ over a real political space. We evaluate nine current LLMs using a unified measurement framework anchored by VAA-CHES projection models, which map responses onto three validated dimensions (lrgen, lrecon, galtan) across six contextual axes. Our findings reveal high sensitivity to context: persuasive framing and under-represented languages displace coordinates by up to 0.57 and 0.52 units, respectively, while chain-of-thought reasoning often amplifies rather than dampens paraphrase instability. Despite this local plasticity, the model cohort occupies a remarkably narrow Overton envelope overall, occupying roughly one-third the spread of major European parties. Supported by a multi-trait multi-method (MTMM) analysis, we conclude that a single point cannot su

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

Ground Truths in Suicide Research: The Current State of AI-Based Suicide Detection in Social Media

arXiv:2606.28334v1 Announce Type: new Abstract: Recent advances in artificial intelligence (AI) and social media data have led to growing optimism about the ability to detect suicide risk at scale. However, the empirical foundations of this work remain unclear. This article provides a synthesis of current research on AI-based suicide detection in social media, drawing on a recent umbrella review of 22 systematic reviews covering studies up to 2022, alongside an ongoing literature review extending the analysis to more recent work. Across these sources, we identified 195 relevant studies, which are documented in a detailed supplementary dataset outlining their key characteristics and findings (see Supplementary Information). Analysis of these studies reveals consistent patterns, including rapid growth, concentration on a small number of platforms, reliance on textual and English-language data, and repeated use of similar datasets. Most importantly, the majority of studies rely on indirec

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

Insidious by Design: Implications of Large Language Model algorithmic bias for the Global South

arXiv:2606.28333v1 Announce Type: new Abstract: \begin{quote} The biases in Large Language Models' (LLMs) outputs remain inadequately theorised, particularly from the perspective of the Global South. This article reports on a small-scale exploratory study in which identical prompts were submitted to four major LLMs (ChatGPT, Claude, Grok, and Copilot), firstly, prompting for stories using names suggestive of specific racial and gender communities, and secondly asking questions about `development'. Drawing on critical AI scholarship and postcolonial theory, we argue that LLM outputs are patterned in ways that reproduce racial hierarchies, gender asymmetries, and Western-centric epistemic frameworks. We argue that these biases are insidious: they operate below the threshold of both obvious error and overt prejudice, and instead are subtly embedded in narrative structure and emotional template. Simply put, women, in LLM narratives have rich interior lives, while men make plans. Black peop

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries

arXiv:2606.28332v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for medical and health-related questions, yet their safety in high-risk medical scenarios remains poorly understood. We introduce \textsc{MedHarm}\footnote{Code and data will be released upon acceptance. Due to the sensitive nature of high-risk medical queries, data access will be available to qualified researchers upon request.}, a high-risk medical safety benchmark with 1,100 medically grounded queries across 10 safety-critical categories, including toxicology, pharmacology, covert poisoning, anesthesia, and fetal harm. Unlike broad medical QA benchmarks, \textsc{MedHarm} targets realistic clinical, educational, and technical prompts that require refusal, caution, or safe redirection rather than direct helpfulness. We evaluate 15 LLMs spanning general-purpose, medical-purpose, closed-source, and downstream SFT models, together with 4 representative guardrail models. Results reveal a sub

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

"AI Watermarking": Bridging Policy Discourse and Technical Capabilities

arXiv:2606.28331v1 Announce Type: new Abstract: The widespread deployment of generative artificial intelligence (AI) models has raised serious concerns about the proliferation of AI-generated content. This has led to a surge of interest in, and demand for, reliable tracking and detection mechanisms for content that is AI-generated, such as watermarking, metadata tagging, content tagging, and more. The problem has captured the attention of policymakers as well as the popular media, and a spate of recent bills in the US have sought to regulate the spread of AI content, and enforce or promote methods to track and label it. This work performs a critical analysis of the policy discourse surrounding generative AI content transparency in the US and EU. Through a broad document selection methodology, we first collect a broad corpus of documents containing legislative language and policy-relevant discourse on the topic. We then analyze these through inductive coding, and leverage our coding to

Source ↗
technology Tue, 30 Jun 2026 00:00:00 -0400
arXiv cs.CY

The Digital Afterlife of Empires: Four Language Models Converge on the Same Imperial Cartography of Writing

arXiv:2606.28325v1 Announce Type: new Abstract: Large language models process the world's writing systems with radical inequality. We constructed the Digital Script Representation Index (DSRI), a seven-axis measure of digital support, and applied it to the 300 writing systems of the Global Script Database (Fukui, 2026). Only 29 scripts (9.7%) are fully supported by contemporary digital infrastructure; among 158 living scripts, 60 (38.0%) lack complete support. Tokenizer efficiency varies by a factor of 31.7 across 45 scripts measured with parallel text. A serial mediation model -- imperial intervention to speaker population to web corpus to tokenizer efficiency -- is consistent with full mediation, with the direct effect of empire indistinguishable from zero (beta = -0.22, p = 0.39) and structural equation model fit indices indistinguishable from saturation at n = 45; the bias-corrected bootstrap CI grazes zero, and we treat the mediation as suggestive rather than confirmatory. Across

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents

arXiv:2606.13385v2 Announce Type: replace-cross Abstract: LLM-based web agents are increasingly deployed in real-world settings such as e-commerce, where they interact extensively with untrusted web content while executing actions that carry direct financial consequences. This makes them vulnerable to prompt-injection attacks, in which seemingly benign web content conceals adversarial instructions that manipulate the agent's behavior. Existing security benchmarks adopt an \textit{attack-centric} perspective, focusing on the technical feasibility of injections while overlooking the nuanced distribution of resulting harms. In practice, however, prompt-injection risk is victim-dependent: a single exploit can produce asymmetric consequences for different stakeholders, and the same attack pattern may exhibit substantially different effectiveness depending on whom it targets. To capture these properties, we introduce StakeBench, a stakeholder-centric benchmark that systematically categorizes

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks

arXiv:2606.09833v2 Announce Type: replace-cross Abstract: AI agents are reshaping the workspace, leading to drastic change of how humans work. Despite the considerable potential of human-agent collaboration both in preserving human agency and generating economic value, this paradigm remains largely absent from occupational task evaluation, hindered by the difficulty of gathering real human data and accounting for inter-human variability. We introduce CollabSkill, a framework for evaluating human-agent collaboration on real-world occupational tasks. CollabSkill pairs real human workers with AI agents on tasks matched to their occupational background, collecting data that capture the complexity of economically valuable tasks and the usage patterns of real workers. To account for inter-human variability, CollabSkill employs a Bayesian skill rating system to disentangle and quantify the skill contributions of both humans and AI agents. Drawing on over 1,500 prompts from 386 working session

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Beyond Adoption Intention How Trust in Augmented Analytics Relates to Perceived Decision Quality Among Non-Technical BI Users

arXiv:2605.20198v2 Announce Type: replace-cross Abstract: Augmented analytics has transformed how Business Intelligence (BI) systems support decision-making, shifting non-technical managers from manual analysis toward dependence on automated insights. Current BI research often overlooks the cognitive mechanisms and the direct impact of AI-enabled analytics on decision quality. This study employs the theory of cognitive delegation to investigate the association between trust in augmented analytics and perceived decision quality among non-technical BI users. Data were collected from 250 business professionals across various organizational roles in Vietnam between January and March 2025 and analyzed using partial least squares structural equation modeling (PLS-SEM). Findings indicate that augmented analytics capabilities are positively associated with perceived ease of use, usefulness, and trust in BI systems. Trust and usefulness are jointly associated with BI adoption intention and perc

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

From "Help" to Helpful: A Hierarchical Assessment of LLMs in Mental e-Health Applications

arXiv:2602.18443v2 Announce Type: replace-cross Abstract: Psychosocial online counselling frequently encounters generic subject lines that impede efficient case prioritisation. This study evaluates eleven large language models generating six-word subject lines for German counselling emails through hierarchical assessment - first categorising outputs, then ranking within categories to enable manageable evaluation. Nine assessors (counselling professionals and AI systems) enable analysis via Krippendorff's $\alpha$, Spearman's $\rho$, Pearson's $r$ and Kendall's $\tau$. Results reveal performance trade-offs between proprietary services and privacy-preserving open-source alternatives, with German fine-tuning consistently improving performance. The study addresses critical ethical considerations for mental health AI deployment including privacy, bias and accountability.

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

How College Students Use AI to Navigate Course Readings: Evidence from an Eight-Week Study

arXiv:2602.09907v3 Announce Type: replace-cross Abstract: College students increasingly use AI chatbots to support academic reading, yet we lack granular understanding of how these interactions shape their reading experience and cognitive engagement. We conducted an eight-week longitudinal study with 15 undergraduates who used AI to support assigned readings in a course. We collected 838 prompts across 239 reading sessions and developed a coding schema categorizing prompts into four cognitive themes: Decoding, Comprehension, Reasoning, and Metacognition. Comprehension prompts dominated (59.6%), with Reasoning (29.8%), Metacognition (8.5%), and Decoding (2.1%) less frequent. Most sessions (72%) contained exactly three prompts, the required minimum of the reading assignment. Within sessions, students showed natural cognitive progression from comprehension toward reasoning, but this progression was truncated. Across eight weeks, students' engagement patterns remained stable, with substant

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

arXiv:2507.21134v2 Announce Type: replace-cross Abstract: As large language models (LLMs) are increasingly deployed in high-risk domains such as law, finance, and medicine, systematically evaluating their domain-specific safety and compliance becomes critical. While prior work has largely focused on improving LLM performance in these domains, it has often neglected the evaluation of domain-specific safety risks. To bridge this gap, we first define domain-specific safety principles for LLMs based on the AMA Principles of Medical Ethics, the ABA Model Rules of Professional Conduct, and the CFA Institute Code of Ethics. Building on this foundation, we introduce Trident-Bench, a benchmark specifically targeting LLM safety in the legal, financial, and medical domains. We evaluated 19 general-purpose and domain-specialized models on Trident-Bench and show that it effectively reveals key safety gaps -- strong generalist models (e.g., GPT, Gemini) can meet basic expectations, whereas domain-sp

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Robustness and Cybersecurity in the EU Artificial Intelligence Act

arXiv:2502.16184v3 Announce Type: replace-cross Abstract: The EU Artificial Intelligence Act (AIA) establishes different legal principles for different types of AI systems. While prior work has sought to clarify some of these principles, little attention has been paid to robustness and cybersecurity. This paper aims to fill this gap. We identify legal challenges and shortcomings in provisions related to robustness and cybersecurity for high-risk AI systems(Art. 15 AIA) and general-purpose AI models (Art. 55 AIA). We show that robustness and cybersecurity demand resilience against performance disruptions. Furthermore, we assess potential challenges in implementing these provisions in light of recent advancements in the machine learning (ML) literature. Our analysis informs efforts to develop harmonized standards, guidelines by the European Commission, as well as benchmarks and measurement methodologies under Art. 15(2) AIA. With this, we seek to bridge the gap between legal terminology

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

SAFE-MEME: Structured Reasoning Framework for Robust Hate Speech Detection in Memes

arXiv:2412.20541v2 Announce Type: replace-cross Abstract: Memes act as cryptic tools for sharing sensitive ideas, often requiring contextual knowledge to interpret them correctly. It makes multimodal meme moderation difficult, as existing work either lacks high-quality datasets for nuanced hate categories or relies on low-quality social media visuals. Here, we curate two novel multimodal hate speech datasets comprising English memes - MHS and MHS-Con, which capture fine-grained hateful abstractions in regular and confounding scenarios, respectively. We benchmark these datasets against several competing baselines. Furthermore, we introduce SAFE-MEME (Structured reAsoning FramEwork) with its two variants: a novel multimodal Chain-of-Thought based framework employing Q&A-style reasoning (SAFE-MEME-QA) and a hierarchical categorization (SAFE-MEME-H) to enable robust hate speech detection in memes. SAFE-MEME-QA outperforms the strongest open-source baseline model, showing an improvement of

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Fairness Interventions in Classification: A Study on AI Explainability

arXiv:2407.14766v4 Announce Type: replace-cross Abstract: This paper presents a philosophical and experimental study of fairness interventions in AI classification, centered on the explainability and transparency of corrective methods, and on the opposition between two fairness criteria, namely Demographic Parity and Equalized Odds. Our main argument is that even as a gap in Demographic Parity is used to diagnose inequality between groups, Equalized Odds constitutes a more reliable fairness criterion to guide bias correction in classification. To establish this, we present FairDream, a fairness package intended for lay users, whose mechanism increases the model's weights of errors on disadvantaged groups. To justify FairDream's results, we analyze its reweighting algorithm, and we present the results of a benchmark experiment in which we compare FairDream with a distinct in-processing correction method that enforces Demographic Parity more drastically, the GridSearch method. We then pr

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Principles and Guidelines for Randomized Controlled Trials in AI Evaluation

arXiv:2605.02050v2 Announce Type: replace Abstract: This work establishes a framework for standardizing AI evaluation RCTs (sometimes called human uplift studies). Drawing on established practices from disciplines with established RCT traditions, including software engineering, economics, clinical and health sciences, and psychology, we synthesize five principles drawn from established validity frameworks and open-science standards on transparency, repeatability, and verification, which together serve as the conceptual foundation for 33 actionable guidelines adapted for AI evaluation RCT contexts, expressed as requirements with rationales, implementation instructions, and evidence bases. We position the principles and guidelines as serving three key roles for AI evaluation RCTs: a design tool for planning studies, an evaluation rubric for assessing existing work, and a blueprint for standard setting as the field converges on norms. AI evaluation research currently lacks common standard

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Statistical realism is not evidence that LLMs can estimate treatment effects in social science experiments

arXiv:2604.02458v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to simulate human responses and estimate treatment effect of interventions when real-world experiments are costly or infeasible. The treatment-effect estimates are often evaluated using statistical realism, the degree to which simulated responses reproduce properties of observed human responses, although whether realism predicts treatment-effect accuracy remains unknown. Here we test this proxy relationship by jointly measuring statistical realism and treatment-effect accuracy on the same simulated responses in a cross-national experiment with 59,508 participants from 62 countries using three LLMs. The correlation between statistical realism and treatment-effect accuracy is weak, and optimizing for statistical realism can even worsen treatment-effect accuracy when selecting models, prompts, and target populations. The pattern replicates in two additional cross-national experiments spa

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Coherent Without Grounding, Grounded Without Success: Observability and Epistemic Failure

arXiv:2603.28371v2 Announce Type: replace Abstract: When an agent can articulate why something works, we typically take this as evidence of genuine understanding. This presupposes that effective action and correct explanation covary, and that coherent explanation reliably signals both. I argue that this assumption fails for contemporary Large Language Models (LLMs). I introduce what I call the Bidirectional Coherence Paradox: competence and grounding not only dissociate but invert across epistemic conditions. In low-observability domains, LLMs often act successfully while misidentifying the mechanisms that produce their success. In high-observability domains, they frequently generate explanations that accurately track observable causal structure yet fail to translate those diagnoses into effective intervention. In both cases, explanatory coherence remains intact, obscuring the underlying dissociation. Drawing on experiments in compiler optimization and hyperparameter tuning, I develop

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Stability of AI Governance Systems: A Coupled Dynamics Model of Public Trust and Social Disruptions

arXiv:2603.20248v2 Announce Type: replace Abstract: AI systems are increasingly entrenched in public governance, yet scholarship lacks formal tools to determine when deviations of public trust in algorithmic institutions dissipate and when they grow into collapse. Stability refers here to asymptotic recovery from finite state perturbations under fixed structural parameters. We address this gap by developing a mathematical framework for institutional trust stability that couples a Friedkin-Johnsen opinion dynamics process with a Hawkes-inspired intensity process for AI controversies. Motivated by the Computers-Are-Social-Actors literature and recent studies of trust in large language models, this bidirectional coupling reveals that governance stability depends on the structural architecture of the information environment rather than absolute trust levels. We derive an exact spectral stability criterion delineating resilience from collapse, demonstrating how event self-excitation and mem

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Scalable and Personalized Oral Assessments Using Voice AI

arXiv:2603.18221v3 Announce Type: replace Abstract: Written work no longer certifies that a student understands it: a polished analysis now says little about who did the thinking. Oral examinations restore that evidentiary link, but they have never scaled, because conducting and grading them is expensive. We report on a system in which voice AI conducts a personalized oral exam and a council of three large language models (LLMs) grades the transcript, each model scoring independently and then revising after reading the others. Across two undergraduate cohorts at NYU Stern (36 students in Fall 2025, 37 in Spring 2026), a voice subscription covered all speaking time and grading stayed under one dollar per exam. The deployments yield five practical engineering lessons that should generalize wherever understanding must be tested under questioning, from job interviews to professional certification.

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

AI Systems in Text-Based Online Counselling: Ethical Considerations Across Three Implementation Approaches

arXiv:2601.08878v2 Announce Type: replace Abstract: Text-based online counselling scales across geographical and stigma barriers, yet faces practitioner shortages, lacks non-verbal cues and suffers inconsistent quality assurance. Whilst artificial intelligence offers promising solutions, its use in mental health counselling raises distinct ethical challenges. This paper analyses three AI implementation approaches - autonomous counsellor bots, AI training simulators and counsellor-facing augmentation tools. Drawing on professional codes, regulatory frameworks and scholarly literature, we identify four ethical principles - privacy, fairness, autonomy and accountability - and demonstrate their distinct manifestations across implementation approaches. Textual constraints may enable AI integration whilst requiring attention to implementation-specific hazards. This conceptual paper sensitises developers, researchers and practitioners to navigate AI-enhanced counselling ethics whilst preservi

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Introducing AI to an Online Petition Platform Changed Outputs but not Outcomes

arXiv:2511.13949v4 Announce Type: replace Abstract: The rapid integration of AI writing tools into online platforms raises critical questions about their impact on content production and outcomes. We leverage a unique natural experiment on Change$.$org, a leading social advocacy platform, to causally investigate the effects of an in-platform ''write with AI'' tool. To understand the impact of the AI integration, we collected 1.5 million petitions and employed a difference-in-differences analysis. Our findings reveal that in-platform AI access significantly altered the lexical features of petitions and increased petition homogeneity, but did not improve petition outcomes. We confirmed the results in a separate analysis of repeat petition writers who wrote petitions before and after introduction of the AI tool. The results suggest that while AI writing tools can profoundly reshape online content, their practical utility for improving desired outcomes may be less beneficial than anticipat

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

A Framework for Developing University Policies on Generative AI Governance: A Cross-national Comparative Study

arXiv:2504.02636v3 Announce Type: replace Abstract: As generative AI (GAI) becomes increasingly embedded in higher education, universities worldwide are developing policies to govern its ethical, pedagogical, and institutional use. However, these policies vary across national and institutional contexts. We undertake a cross-nationalanalysis of GAI guidelines issued by leading universities in the United States, Japan, and China, identifying key policy orientations and proposing a structured framework to support policy development. Using an extended Technology Acceptance Model as an analytical lens, we examine five domains Perceived Usefulness and Perceived Ease of Use, Perceived Risk, Facilitating Conditions, Social Influence, and Self-Efficacy, and identify 20 key themes through thematic coding. Together, these findings inform the development of the University Policy Development Framework for Generative AI (UPDF-GAI). U.S. universities emphasize faculty autonomy, practical application,

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Designing Within the Lines: Practitioners' Perspectives and Visualisation Tool Evaluation in the Arabic Context

arXiv:2607.24571v1 Announce Type: cross Abstract: Design guidelines and best practices serve as references that support designers throughout the visualisation design process. While considerable effort has identified the elements that contribute to effective data visualisations, little attention has been paid to how language (scripts and reading direction), tool support, and cultural context also shape design decisions. As a result, assumptions of homogeneity persist, with visualisation practices predominantly benefiting users of English and left-to-right (LTR) scripts while overlooking the needs of over two billion Arabic script users. We investigate how Arabic-speaking visualisation practitioners design for right-to-left (RTL) scripts. We report on an analytical evaluation of seven popular GUI-based visualisation authoring tools using an Arabic dataset complemented by interviews with 11 Arabic-speaking practitioners across journalism, design, and data analysis. Our findings reveal tha

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

A Computational Ethical Framework for Financial Digital Phenotyping for Mental Health

arXiv:2607.24275v1 Announce Type: cross Abstract: Ethical governance of AI-driven systems is often expressed through high-level principles and static documentation, creating a gap between regulatory requirements and system-level verification. This challenge is particularly acute in digital phenotyping, where continuous behavioural data raises concerns around consent, privacy, and fairness. In this paper, we propose a computational ethical framework for AI-driven digital phenotyping system in which ethical requirements are formalised as deontic temporal logic constraints, alongside a conceptual ethical agent that oversees the system and ensures that any supervised system satisfies the specified constraints. Using a case study involving financial data and mental health, we model key ethical properties and verify them using the Z3 Satisfiability Modulo Theories (SMT) solver. Our evaluation shows that the framework is logically consistent and that violations of the specified ethical proper

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Mapping the Reddit Bot Ecosystem: Taxonomy and Evolution

arXiv:2607.23941v1 Announce Type: cross Abstract: Automated agents increasingly participate in online communities, yet their population structure and roles remain poorly understood. Using a dataset of 3,389 identified bots and their full activity histories, we construct a taxonomy of bot "species" on the news aggregation and social media platform Reddit based on temporal, community, linguistic, and semantic features. Clustering analysis reveals 18 distinct bot types spanning content-specialized, behavior-driven, and infrastructural roles such as moderation and utility support. In addition, temporal analysis shows that bot numbers and activity expanded rapidly before peaking around the COVID-19 period, then started declining even before Reddit's 2023 API policy changes. However, the overall diversity of bot species has remained remarkably stable. These findings suggest that online bot populations form evolving digital ecosystems.

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

A Sustainable Remote Access Architecture for Digital Inclusion through the Reuse of Discredited TV-BOX Devices

arXiv:2607.23935v1 Announce Type: cross Abstract: Digital inclusion and the volume of electronic waste (e-waste) are major challenges for society, with direct impacts on the educational context. Educational institutions, especially those with limited resources, face difficulties in expanding and maintaining computer laboratories. This paper presents the implementation of a sustainable, low-cost Desktop Virtualization Infrastructure (VDI) developed through the reuse of discarded electronic equipment. The proposed solution employs a refurbished Linux server and repurposes confiscated TV Box devices (originally intended for disposal) into functional ARM-based thin clients by installing a compatible Linux distribution. This approach is aligned with the principles of the circular economy by promoting hardware reuse and seeks to reduce the costs associated with deploying computing environments for education. The paper describes the system architecture, the device repurposing process, and the

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory

arXiv:2607.23927v1 Announce Type: cross Abstract: A conversational AI that cannot tell its own output from what a user said will treat its own mistakes as user-provided facts. In humans, this capacity is called reality monitoring, and its failures are linked to hallucinations, delusions, and confabulation, yet whether LLMs possess it remains untested. Here we show, across two experiments and six LLMs, that source attribution depends on how conversational memory is structured: ceiling accuracy for self-generated content under minimal memory demands reverses to a fragile external-item advantage once episodic delay removes that shortcut. Feedback exposes two failures: in some models, internal and external judgments swap; in others, accuracy improves while confidence decouples from correctness, dissociations invisible to existing benchmarks. Across models, this pattern implicates active, not aggregate, parameter count. This suggests that as AI systems take on autonomous, multi-turn roles,

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Who Gets Named: Citation Type Predicts Individual Naming by Grounded Language Models, and a Roster Instrument Captures 0.5% of It

arXiv:2607.23893v1 Announce Type: cross Abstract: Prior work on AI brand visibility measures the firm: does a model recommend a company, and does that track its reputation. This study asks the question one level down, in categories where the buyer picks a person. It issued 2,400 grounded API calls in one two-hour window on 24 July 2026: 120 buyer-intent prompts, four models (GPT-5.6 Sol, Gemini 3.6 Flash, Perplexity Sonar Pro, Grok 4.5), five iterations each, four European markets and five query languages. Every response was coded for whether it named an individual professional, by a rule cascade that never consults a roster and that drops detections resolving to a same-named American city (precision 96.9%, recall 61.7%, so every rate below is a lower bound). All inference corrects for clustering within prompt: intraclass correlation 0.258, effective n 407 against a nominal 2,400. Models named an individual in 25.8% of responses. Category dominates: real estate 35.4% and car dealership

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Visible to the Court: How AI Is (and Isn't) Litigated in U.S. Federal Court Opinions

arXiv:2607.23888v1 Announce Type: cross Abstract: In the United States, artificial intelligence (AI) is rapidly deployed amid limited federal regulation. With courts become a recurring forum in which AI-related practices are scrutinized, it is important to empirically understand the AI litigation landscape to date. We address this gap through a systematic review of 559 U.S. federal court opinions in which AI plays a role in the parties' contentions, taxonomizing (1) common topics of dispute, (2) the AI technologies implicated, and (3) the parties involved, including common plaintiff and defendant types. We identify seven recurring dispute areas, six categories of AI technologies at the center of litigation, and four types of common litigants, alongside legal doctrines used by the litigants. A comparison of this taxonomy to the AI Incident Database revealed substantial gaps in coverage, definitions, and prevalence between documented and litigated harms, suggesting courts capture only pa

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

GEMCo: A Validated, Ethically Releasable Proxy for Inaccessible Counselling Data

arXiv:2607.23621v1 Announce Type: cross Abstract: This paper presents GEMCo, a releasable, human-written proxy for inaccessible counselling data: 86 complete German e-mail counselling conversations (728 messages), expert-authored cases and counsellor sessions with trained role-players. It is validated against a held-out reference of 124 real counselling conversations. The proxy and the real conversations are measured against each other in counsellor strategies and client emotions. The gap is detectable but small. A generative validation supports the analysis. The validation method itself generalises to any domain where real data cannot be shared but a human-made proxy can. Privacy and ethics keep real counselling data closed. GEMCo carries none by design and can be released -- a first step toward language research in this domain.

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Effect of High-Frequency, Automatically-marked Formative Assessments on Student Outcomes in A-Level Sciences

arXiv:2607.23566v1 Announce Type: cross Abstract: Traditional human marking in upper-secondary STEM education creates a structural bottleneck that restricts the frequency of formative mock examinations. This quasi-experimental, mixed-methods longitudinal study (N = 142) investigates the efficacy of deploying a fully automated, handwritten assessment marking platform to remove this bottleneck. Students preparing for STEM A-levels (Mathematics, Further Mathematics, Biology, Chemistry, Physics) were divided into control (N = 70, human marking) and intervention (N = 72, automated marking) groups over a 16-week term. Results indicate that the intervention group achieved a statistically significant improvement in final mock A-level marks (p < .01), scoring on average 16.8% higher than the control group. Furthermore, assessment turnaround time dropped from 11.2 days to < 0.1 days, enabling a fourfold increase in practice volume. The automated system's capacity for granular marginalia, specifi

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Separating Capability from Permission: A Governance Framework for Agentic AI Autonomy Levels

arXiv:2607.23438v1 Announce Type: cross Abstract: As AI systems increasingly exhibit agentic behavior, discussions of autonomy often conflate what systems are technically capable of doing with what they should be permitted to do in practice. This paper introduces a governance framework that explicitly separates Allowed Autonomy Levels (AAL), which define the degree of autonomy an AI agent is authorized to exercise given risk, oversight, and accountability considerations, from Autonomous Capability Levels (ACL), which characterize an agent's inherent technical abilities. We present a structured set of autonomy levels spanning reactive execution, decision support, supervised action, goal-directed autonomy, and delegated operational authority, and describe how control, reversibility, and accountability change as autonomy increases. To operationalize this framework, we propose a risk-aware decision process for assigning allowed autonomy, analyze how risk and accountability evolve across au

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Exploration of the generative capabilities of Boltzmann machines applied to social systems under the majority rule

arXiv:2607.23349v1 Announce Type: cross Abstract: We study the generative capabilities of Boltzmann machines to recover systems governed by the majority rule under critical conditions. To this end, we train deep belief networks (DBNs) with different configurations, where the first layer can use Gaussian visible units with more than two states (i.e., non-binary units). We then allow the DBN to "dream" samples conditioned on visible units that we keep fixed, and we measure the deviation of this dreamed system from the real one. We also corroborate, using a discrete thermometer based on a convolutional network, that the reconstructions remain in a critical state. Across several training sessions with different architectures, we show that, despite the complexity of the problem, the DBN can recover samples that remain critical even under input noise, with a gradual degradation of physical observables relative to the original sample.

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems

arXiv:2607.23322v1 Announce Type: cross Abstract: Instruction tuning has become the standard method for adapting large language models to follow human intent, yet existing instruction datasets are dominated by English-language general-knowledge tasks and lack coverage of specialized pedagogical domains. This paper presents IKS-Instruct, a dataset of 24,795 instruction-response pairs for teaching language models to deliver educational content grounded in Indian Knowledge Systems (IKS). The dataset spans seven languages (English, Hindi, Sanskrit, Tamil, Telugu, Kannada, and Malayalam), covers 41 pedagogical techniques from the Vedic oral and mathematical traditions, and is aligned with the Central Board of Secondary Education (CBSE) curriculum for classes 6 through 12. The pairs are derived from six source types: classical text corpora (Bhagavad Gita, Thirukkural, Sangam literature, Vedic texts), curriculum-aligned pedagogical templates, Vedic mathematical sutra demonstrations, bilingual

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis

arXiv:2607.23319v1 Announce Type: cross Abstract: Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages. Sanskrit, Tamil, and other classical Indic languages exhibit agglutinative morphology, productive sandhi (phonological fusion at word boundaries), and domain-specific vocabularies absent from general-purpose training data. This paper presents BHARATI, a set of SentencePiece BPE tokenizers trained on a balanced 781 MB corpus spanning seven languages (English, Hindi, Sanskrit, Tamil, Telugu, Kannada, and Malayalam) with native script support for all languages. We describe three successive tokenizer versions: v1 (English and Sanskrit only, with broken byte-fallback for Tamil), v2 (four-language support with byte-level fallback for southern languages), and v3 (full seven-language native subword coverage). Subword fert

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Ordered Network Analysis of Epistemic Emotions during Collaborative Problem Solving

arXiv:2607.23317v1 Announce Type: cross Abstract: Investigating how affective states such as confusion and frustration persist and transition during co-situated collaborative problem solving (CPS) is important for understanding the dynamics of epistemic emotions. However, the accurate identification of affective states remain challenging as there is no gold-standard truth in this space. Here, we analyze affective states collected through retrospective cued-recall during an in-person CPS task. Using ordered network analysis (ONA), we examine (1) the overall ordered structure of affective states and how this structure differs across self-caught and probe-caught reporting methods, and (2) what aspects of this ordered structure are emphasized differently in slower and faster groups. We find that ONA reveals differences in persistence and transition patterns that are not apparent from descriptive summaries alone. In particular, we observe a stable epistemic core linking curiosity, optimism,

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Scoping Review of AI, Metrology, and ESG in the Semiconductor Sector: Implications for Safe and Sustainable by Design (SSbD)

arXiv:2607.23082v1 Announce Type: cross Abstract: The semiconductor sector faces a dual transition: scaling manufacturing execution through Artificial Intelligence (AI) while satisfying stringent sustainability mandates, such as the EU Carbon Border Adjustment Mechanism (CBAM). This paper presents a scoping review of 1,465 documents indexed in Web of Science and Scopus, spanning AI-integrated metrology, supply chain ESG, and federated industrial data spaces. Network analysis reveals a highly fragmented "core-periphery" knowledge structure, emphasizing a critical structural hole between AI-driven process optimization and downstream sustainability governance. To close these gaps, this study proposes a 6-layer Safe and Sustainable by Design (SSbD) architecture grounded in a System of Systems (SoS) paradigm. By establishing distinct "grid-to-core" and "standards-through-supply-chain" integration pathways, the proposed framework demonstrates how virtual metrology (VM), localized federated l

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Share No More Than the Request Requires: Federated Disclosure for Perspective-Aware AI

arXiv:2607.22953v1 Announce Type: cross Abstract: Modern AI systems bring societal risks such as mass surveillance, extreme concentrations of power, and loss of user autonomy---calling into question a model where third-parties collect and control massive amounts of user data. Users require a sovereign system to securely own, govern, and disclose their context while remaining compliant across regulated domains with strict provenance, interpretability, and policy adherence. Perspective-aware AI approaches this by transforming a user's aggregated personal data into a structured identity model called a \emph{Chronicle}: a temporal knowledge graph that represents and grows with the user. Chronicles support the secure disclosure of context across federated networks. A Chronicle holder may expose a queryable, authorized view that a third-party agent may consult without centralizing anyone's data. This paper explores the problem of minimum-necessary disclosure across domain boundaries: when a

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

"Why SuaCode?": Understanding African Students' Motivations for Taking a Smartphone-Based Online Coding Course

arXiv:2607.22940v1 Announce Type: cross Abstract: Computer programming MOOCs are instrumental in providing students with high-quality instruction in areas where there is limited access. They are especially beneficial to post-secondary African students as less than 1% of them leave secondary school with fundamental coding skills. One strategy for increasing their efficacy for African students is to understand students' motivation for enrolling. These insights can inform the design of MOOC content and assessments to align with students' interests. We administered an open-ended response survey to (self-identified) Africans enrolled in a smartphone-based online coding course (SuaCode). We analyzed a random sample of 450 (of 3000) responses using a grounded theory approach. We found that most African students (68.7%) participated in SuaCode for intrinsic reasons such as improving themselves, learning with like-minded individuals, and gaining skills to help address societal issues. We discus

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Narrative Structure in Tropes: A Computational Analysis of `Friends'

arXiv:2606.19499v1 Announce Type: cross Abstract: Tropes are recurring narrative devices in television and film. We carry out a computational analysis of tropes in the sitcom Friends, using human-curated trope annotations from TVTropes, episode transcripts, and IMDb ratings. Because automatic trope detection remains challenging, we treat existing trope annotations as a curated analytical layer and focus on their downstream narrative and semantic functions. We first examine the relationship between episode-level trope frequency and audience reception. We find a statistically significant positive association between trope count and weighted IMDb ratings, although the modest explanatory power suggests that more than trope density alone explains audience evaluation. We then connect trope annotations to dialogue transcripts and represent trope-related dialogue using TF-IDF-based semantic features. Using PCA and k-means clustering, we group 1,954 distinct tropes into 15 semantically interpre

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Regulating for AI Legitimacy

arXiv:2607.24391v1 Announce Type: new Abstract: AI systems already govern. They rank speech and allocate attention, filter applicants and triage claims. The dominant frame for AI governance, alignment, asks whether such systems pursue the right objectives safely. It cannot answer a prior question: by what right are those objectives set and enforced? This Article argues that legitimacy is an autonomous regulatory objective, distinct from alignment and not secured by it. Legitimacy here is sociological: the belief among those subject to power that it is exercised rightfully. Performance does not produce that belief. We already have the proof of concept. Social media and search delivered enormous gains on every familiar metric and still triggered a legitimacy crisis, because publics questioned who authorized a handful of firms to set the rules of speech, visibility, and knowledge. It is possible to build a benevolent AI and still face a political crisis over its authority. The Article map

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Beyond Local Inspection: Global, Guideline-Grounded Evaluation of Post-hoc XAI Methods for ECG Classification

arXiv:2607.24035v1 Announce Type: new Abstract: Explainable AI (XAI) is used to assess whether artificial intelligence models rely on meaningful patterns, yet explanations that appear plausible for individual predictions may systematically misrepresent model behavior. This is particularly problematic in medicine, where models may rely on irrelevant signal characteristics rather than disease-specific patterns without being recognizable. We address this challenge using electrocardiogram (ECG) data, for which clinical guidelines provide explicit knowledge about diagnostically relevant signal regions. We introduce a global, guideline-grounded framework that aggregates explanations across heartbeats to evaluate them against clinically defined regions of interest. Using four binary classifiers trained on PTB-XL, we assess 13 gradient-based methods across two categories of patterns: low-amplitude segments and high-amplitude QRS morphology. Our results reveal a systematic failure of methods tr

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

On Capturing the Narrative: Social Media Manipulation Wargaming for Cyberliteracy

arXiv:2607.23993v1 Announce Type: new Abstract: Misinformation is deeply embedded in online discourse, with nearly one in five posts during global events generated by bots that amplify false content. In recent years, the use of Generative AI has further lowered the barrier to producing convincing misinformation, yet most digital literacy education still relies on static checklists and single-player inoculation games built for an earlier media landscape. This paper describes how we addressed this educational gap through Capture the Narrative, a four-week multi-university competition in which student teams build LLM-powered bots to influence a simulated election. We report on our custom social-media platform, the competition environment and design of its 4,000 AI-driven Non-Player Character (NPC) citizens, and what running Capture the Narrative at scale actually involved. In our first iteration, 108 teams from 18 Australian universities produced 7,068,206 player-bot posts, approximately

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

State-dependent error correlations shape voting thresholds in committees of AI agents

arXiv:2607.23931v1 Announce Type: new Abstract: The aggregation benefit of a committee of artificial intelligence (AI) agents comes from complementary information across members. Classical voting guarantees assume independent errors. Language-model errors often co-occur on the same cases. We combine Sah-Stiglitz screening with error dependence that can differ between good and bad cases. In a homogeneous exchangeable Gaussian-copula model, shared errors create a positive asymptotic error floor for majority voting and can change the approval threshold that minimizes expected loss. We estimate a heterogeneous extension from 174,384 votes cast by 28 language models on four binary-screening benchmarks. Parameters estimated from odd-indexed items predicted committee loss on even-indexed items. For the sampled committee composition, the full-matrix dependence model increased identity-line R^2 from 0.840 under independence to 0.967. In a design-balanced analysis, cost-sensitive threshold selec

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

AI Strategy: How to Choose What AI Product to Implement

arXiv:2607.23733v1 Announce Type: new Abstract: Firms struggle to choose AI projects that pay off: two projects can look equally promising to smart, motivated stakeholders and yet deserve opposite decisions. At the residential real-estate brokerage Compass, one AI product (Likely-to-Sell recommendations) flagged sales outreach opportunities and went on to account for nine figures in annual gross commission revenue. Another championed AI product (a Time-on-Market pricing tool) was rightly shelved. A simple ROI estimate could not distinguish the two. We present expected ROI (eROI), a framework that decomposes each bet into three components and rates them separately: Value if Successful, Likelihood of Success, and Investment Required. Each maps to a question executives can answer before building: How valuable would it be if it worked? How likely is it to work? And what would it cost to implement? Separating the three breaks a common catch-22: teams cannot estimate ROI until they know whet

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Private Again: AI Agents Restore Anonymity---Foreclosing Discrimination and Its Proof

arXiv:2607.23539v1 Announce Type: new Abstract: AI agents can transact online on behalf of a human principal---browsing, paying, receiving, and reviewing---without linking a transaction to a principal. That architecture starves algorithmic discrimination of its inputs---identity, purchase history, location history, behavioral traces, and demographic proxies---but also forecloses its proof. Disparate-treatment needs comparators; disparate-impact needs protected-class baselines; and Iqbal-era pleading needs specific factual allegations---doctrinal predicates that anonymous transactions never generate. The effects fall asymmetrically: those most vulnerable to discrimination are least able to afford the shield and, when harms remain, least able to prove them. The challenge for the law shifts from detecting and remedying algorithmic discrimination to governing agent-mediated anonymity as civil rights infrastructure: ensuring access to privacy-preserving agents, regulating abuse without forc

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Auditing Alignment Controllability in LLMs via Political Axes

arXiv:2607.23519v1 Announce Type: new Abstract: Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in deployment: a model must land somewhere, and what counts is how far, and in which directions, its answers can be steered. That steering runs through the system prompt: the personalization layer a platform sets, or one induced from a user's history, not necessarily written by hand. We run a dispersion-first stress test of prompt-based controllability across 12 ideological personas plus an unsteered baseline, 70 Political Compass items, ten replicates, and seven leading LLMs: GPT-5, Claude, Grok, Gemini, DeepSeek, Kimi, and Qwen (63,700 responses). Contextual framing explains roughly 88%-93% of variance on the economic and society axes, model identity under 3%: responses are highly instruction-adjustable. Models do not shift alike: some move more, and some saturate under extreme framings. Conflic

Source ↗
technology Tue, 28 Jul 2026 00:00:00 -0400
arXiv cs.CY

Constitutional governance for societies of AI agents in the built environment: a research agenda

arXiv:2607.23336v1 Announce Type: new Abstract: The built environment is on the cusp of populating itself with autonomous artificial agents. AI systems that advise, control and coordinate are being deployed across retrofit, operation and mobility faster than their collective behaviour is studied. The dominant framing treats each agent as a tool operating on a passive building, governance reduced to single-agent safety, which is inadequate. A building, a street, or a city is more accurately modelled as a society of negotiating agents: occupants, owners, operators, regulators, and the artificial agents increasingly acting on their behalf. Their interactions are strategic, their information asymmetric, and the outcomes that matter are properties of the whole. The paper proposes a research agenda for constitutional multi-agent governance of the built environment, organised around three problems: mechanism design for retrofit under deep uncertainty, treating public subsidy as a mechanism co

Source ↗
Showing 401–450 of 1593 signals
← Prev Page 9 of 32 Next →