EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Anticipatory Data Governance in the Age of AI: Emerging Signals in Data Access, Reuse, and Sovereignty

arXiv:2607.27029v1 Announce Type: new Abstract: This paper reports findings from a structured participatory foresight study comprising two expert forecasting studios convened by The GovLab between 2025 and 2026. The studios brought together nineteen senior practitioners spanning official statistics, digital and trade policy, open science, AI governance, geospatial systems, and public-sector innovation across multiple jurisdictions. Applying a qualitative signal-scanning methodology grounded in the horizon-scanning and anticipatory-governance traditions, we elicited, clustered, and thematically synthesized emerging developments in data access, governance,and reuse, and stress-tested them against practitioner experience. We identify seven convergent signals: (1) the open-data paradigm is under strain; (2) data ecosystems are becoming machine-centric and AI-mediated; (3) inference is reshaping the foundations of data governance;(4) data infrastructure is becoming harder to sustain; (5) go

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

From Micro-Cognition to Self-Construction: A Four-Layer Integrative Review of Psychological Theories in HCI

arXiv:2607.26402v1 Announce Type: new Abstract: Human-computer interaction (HCI) is undergoing a paradigm shift from "tool use" toward "partnership" and even "mind symbiosis," with psychology evolving from a supplementary explanatory tool to a core pillar shaping interaction paradigms and long-term relationships. This paper systematically reviews relevant research and proposes a four-layer integrative framework comprising the Micro-cognitive, Meso-affective, Macro-social, and Self-constructive layers. The framework reveals that: the cognitive layer constitutes the foundation of interaction, the affective layer drives relational engagement, the social layer regulates trust and behavior through norms, and the self-constructive layer points to the ultimate goal of human-machine symbiosis. These four layers follow a progressive logic of "foundation--mediation--context--goal." The review further argues that psychology has shifted from "post-hoc evaluation" to a "proactive design paradigm,"

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

"Nobody Did This": Contribution, Originality, and Accountability in Agent-Mediated Collaboration

arXiv:2607.26387v1 Announce Type: new Abstract: Collaborative knowledge work is changing in ways that go beyond disclosure or transparency. LLM agents are now embedded in how teams research, design, write, and decide: mediating between members, synthesizing inputs, reformulating ideas, and drafting shared outputs. They do not only facilitate collaboration; they operate within the workflow at the moment contributions are being formed. In doing so, they risk undermining the social conditions under which contributions can be witnessed, attributed, and held accountable. This workshop brings together researchers and practitioners to confront what we call contribution dissolution: the blurring of attribution, originality, and accountability in agent-mediated collaborative work. We argue that this dissolution begins before collaboration itself, in the individual worker's own uncertainty about what is genuinely theirs, and propagates through collaborative relationships, collapsing the reliabil

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach

arXiv:2607.26317v1 Announce Type: new Abstract: Psychometric calibration for educational tests typically requires costly human response data. Large language models (LLMs) simulated examinees offer a promising route to early calibration, but their responses are too accurate and too uniform. We propose Cognitive Diagnostic Profiling (CDP), a zero-shot framework that prompts LLMs to simulate plausible examinees with diverse cognitive profiles: binary attribute-mastery patterns are rendered as natural-language profiles and sampled under an uninformative or an informative distribution. Using the Tatsuoka fraction-subtraction dataset (536 examinees, 15 items, five attributes), we evaluated eight LLM configurations under no-profile, uninformative-CDP, and informative-CDP conditions, assessing alignment with human examinees at the ability-distribution, mastery-profile, and item-difficulty levels. CDP improved all three levels: distributional overlap rose across configurations; weighted correla

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

AI Security Priorities: A Field-Wide Agenda

arXiv:2607.26069v1 Announce Type: new Abstract: As AI systems are rapidly integrated into critical economic, governmental, and national security functions, the gap between AI adoption and AI security readiness continues to widen. This paper presents a prioritized agenda for advancing AI security, informed by structured interviews with leaders across industry, government, and civil society, and refined through a multi-sector expert workshop. Participants identified and ranked the highest-importance and most cost-effective areas where progress could strengthen AI security - from protecting frontier AI systems and their underlying infrastructure to improving cybersecurity practices as AI reshapes the threat landscape. The resulting priorities are organized across four themes: establishing strategic foundations and policy frameworks; advancing public-private coordination and institutional infrastructure; advancing technical security engineering and assurance; and governing agentic AI under

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty

arXiv:2607.26067v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for estimating item difficulty in educational assessment. However, it remains unclear whether such estimates reflect how learners actually experience difficulty. This study investigates the alignment between LLM-generated difficulty ratings and empirical student performance on basic mathematics tasks. Four widely used LLM-based systems generated difficulty ratings on a 1-100 scale for 32 arithmetic items across multiple runs (N = 640 ratings). These were compared with empirical difficulty derived from responses of 770 Indonesian undergraduates using Classical Test Theory (CTT) and Item Response Theory (2PL). Results show moderate rank correlations (Spearman's rho = 0.52-0.70), indicating that LLMs capture coarse ordering of item difficulty. However, substantial and systematic misalignment emerges in fraction items. Several items consistently rated as easy by LLMs were among the most diffi

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science

arXiv:2607.26064v1 Announce Type: new Abstract: AI systems are becoming autonomous research agents that generate hypotheses, design experiments, and produce discoveries at scales beyond human oversight. As seen by increased submissions to ML venues, the verification gap between scientific output and our ability to check it is already widening, and autonomous agents make it worse by magnitudes given human-agent asymmetry. We argue that science must evolve its verification infrastructure, as it has before with peer review. However, while historical adaptations assumed human contributors who could be questioned and sanctioned, AI agents break this assumption. We propose criteria for an adapted verification infrastructure that emphasizes observable-by-default workflows, scalable verification, and clear attribution. We argue that without adaptation, ML and any scientific domain using agents face dangerous failures: experimental results that no person can verify, optimization for metrics ove

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Archetypes or ability? Clustering for modelling student mathematical competence

arXiv:2607.26063v1 Announce Type: new Abstract: Personalised learning systems often assume that mathematical ability is combined of discrete abilities, acquired sequentially and dependent upon first acquiring foundational abilities, and students often report different strengths. In this work, we explore the validity of these assumptions by applying clustering methods to a large dataset of 119,034 students, spanning 13 national-level exams sat in the United Kingdom and collected by the platform. Classifying question results as pass or fail, we use a Bernoulli Mixture Model to search for latent populations which would be indicative of discrete skill-sets. We find that few distinct clusters are present in the data and that the dominant factor is overall student ability, which is further supported by the high degree of linear correlation between the probability distributions of the resulting clusters. Our best performing model achieves an accuracy of 78 percent, competitive with more compl

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CY

Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities

arXiv:2607.26062v1 Announce Type: new Abstract: Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people with intellectual disabilities (ID). Objective: The study aims to identify and measure representational differences related to people with ID and examine them to identify implicit biases inherent in AI chat generation technologies. Methods: Utilizing the GPT-4-Turbo model, we requested story-generation based on 10 prompt stems with and without descriptors for ID. This process was repeated using four other LLMs (OpenAI GPT-4o, Meta Llama-3-3-70B-Instruct, Anthropic Claude-3-5-Sonnet, and Mistral-Large-2411). The resulting 25,000 computer-generated stories were analyzed using a separate GPT-4-Turbo model instance to detect differences in how people are represented related to themes of bias described in previous literature. Results: Our findings reveal differences in how people are represented between sto

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

IDEAlign: Comparing Ideas of Large Language Models to Domain Expert

arXiv:2509.02855v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to produce open-ended, interpretive annotations, yet there is no validated, scalable measure of idea-level similarity to expert annotations. We (i) introduce the content evaluation of LLM annotations as a core, understudied task, (ii) propose IDEAlign for capturing expert similarity judgments via pick-the-odd-one-out tasks, and (iii) benchmark various similarity methods (text embeddings, topic models, and LLM-as-a-judge) against these human ratings. Applying this approach to two real-world educational datasets (e.g., interpreting math reasoning and feedback generation), we find that most metrics fail to capture the nuanced dimensions of similarity meaningful to experts. LLM-as-a-judge performs best (11~18% improvement over other methods) but still falls short of expert alignment, making it useful as a triage tool rather than a substitute for human review. Our work demonstrates t

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Triadic Novelty: A Structural Typology of Science Innovation

arXiv:2506.17851v3 Announce Type: replace-cross Abstract: Scientific progress depends on novelty, but current evaluation systems often conflate novelty with recognition, favoring work that aligns with existing paradigms over ideas that challenge them. We introduce a theory-driven framework that conceptualizes novelty as a structural process rather than a single scalar outcome. Drawing on network science and theories of scientific discovery, we develop a triadic typology of novelty: Pioneers introduce new topics, Mavericks recombine distant areas of knowledge, and Vanguards reinforce weak but emerging connections. We apply this framework to philanthropic and nonprofit studies, an interdisciplinary and evolving field well suited to examining how novelty is recognized before evaluation norms are fully stabilized. Results show that novelty is not uniformly rewarded: Pioneer contributions are often weakly recognized unless later taken up, Maverick contributions receive consistent recognitio

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia

arXiv:2602.18455v5 Announce Type: replace Abstract: Search engines increasingly display AI-generated answers above organic links, potentially displacing traffic to upstream publishers. We estimate the impact of Google's AI Overviews (AIO) on Wikipedia's search traffic using AIO's staggered geographic rollout and Wikipedia's multilingual structure. Our difference-in-differences design compares monthly external-search referrals to English Wikipedia articles with referrals to the same articles in German and French, and finds that default AIO availability reduced English search traffic by 5.45% and 4.82%, respectively. Our results suggest that answer-producing digital intermediaries can materially reallocate attention away from informational publishers, with implications for content monetization, search platform design, and policy.

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Operational Agency: A Permeable Legal Fiction for Tracing Culpability in AI Systems

arXiv:2602.17932v3 Announce Type: replace Abstract: Modern artificial intelligence (AI) systems act with a high degree of independence yet lack legal personhood-a paradox that fractures doctrines grounded in human-centric notions of mens rea and actus reus. This Article introduces Operational Agency (OA)-a permeable legal fiction structured as an ex post evidentiary framework-and Operational Agency Graph (OAG), a tool for mapping causal interactions among human actors, organizations, and AI systems. OA evaluates an AI's observable operational characteristics: its goal-directedness (as a proxy for intent), predictive processing (as a proxy for foresight), and safety architecture (as a proxy for a standard of care). OAG operationalizes that analysis by embedding these characteristics in a causal graph to trace and apportion culpability among developers, fine-tuners, deployers, and users. Drawing on corporate criminal liability, the innocent-agent doctrine, and secondary and vicarious lia

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Copyright Laundering Through the AI Ouroboros: Adapting the 'Fruit of the Poisonous Tree' Doctrine to Recursive AI Training

arXiv:2601.02631v3 Announce Type: replace Abstract: Copyright enforcement rests on an evidentiary bargain: a plaintiff must show both the defendant's access to the work and substantial similarity in the challenged output. That bargain comes under strain when AI systems are trained through multi-generational pipelines with recursive synthetic data. As successive models are tuned on the outputs of its predecessors, any copyrighted material absorbed by an early model is diffused into deeper statistical abstractions. The result is an evidentiary blind spot where overlaps that emerge look coincidental, while the chain of provenance is too attenuated to trace. These conditions are ripe for "copyright laundering"--the use of multi-generational synthetic pipelines, an "AI Ouroboros," to render traditional proof of infringement impracticable. This Article adapts the "fruit of the poisonous tree" (FOPT) principle to propose a AI-FOPT standard: if a foundational AI model's training is adjudged in

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

EduAgentQG: Multi-Agent Personalized Mathematics Question Generation with Explicit Diversity and Objective-Aware Evaluation

arXiv:2511.11635v2 Announce Type: replace Abstract: In intelligent education, personalized mathematics question generation aims to produce mathematics questions that satisfy educational requirements while supporting adaptive assessment and learning. Existing LLM-based single-agent and multi-agent methods improve generation flexibility, but they still tend to rely on aggregated feedback or model randomness, making it difficult to jointly ensure dimension-wise objective alignment and controllable diversity. To address these challenges, we propose EduAgentQG, a multi-agent collaborative framework for personalized mathematics question generation with explicit diversity and objective-aware evaluation. EduAgentQG organizes question generation as a closed-loop process of planning, writing, evaluation, refinement, and checking: structured generation plans and multiple generation directions guide candidate generation, while fine-grained evaluation verifies logical correctness, solvability, and

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

From Hallucination to Reliability: Generative Modeling and the Structure of Scientific Inference

arXiv:2504.08526v3 Announce Type: replace Abstract: Generative AI is increasingly used in science, but is unavoidably prone to hallucination. I develop a reliabilist account of how generative AI nevertheless gives rise to new scientific knowledge. I analyze hallucinations as non-strategic misrepresentations of the target phenomenon, introduced by a model's generative activity, rather than inherited from training data. Through case studies of AlphaFold and SEEDS, I show how scientific workflows draw on pre-existing knowledge of target phenomena to filter or qualify hallucinatory outputs, thereby preventing their erroneous content from propagating into downstream inference. Finally, I show that workflows are units of epistemic evaluation in their own right.

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Learning from waste: Machine Learning for health risk prediction and computer vision-based sorting in Ghana

arXiv:2608.25759v1 Announce Type: cross Abstract: The inappropriate disposal of solid waste remains a significant public health and environmental concern worldwide, including in Ghana. Poor sanitation and improper waste management practices contribute to substantial economic costs and avoidable deaths annually. In 2022, a field study in Atonsu, Kumasi, Ghana, reported a community-perceived relationship between household waste disposal and illness patterns, but only through descriptive analysis without quantitative validation. This study extends that investigation using two data-driven approaches. First, a Random Forest classifier was developed to predict illness categories using waste disposal practices and demographic survey data. On a held-out group of respondents who reported illness (N=69), the model obtained a macro F1 score of 0.63, with disposal method emerging as the most important substantive predictor of illness type. Second, a MobileNetV2 image classification model enabled a

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Using profiles of cognitive capability to assess AI suitability for workplace tasks

arXiv:2608.25623v1 Announce Type: cross Abstract: Organisations deploying AI face a scoping problem: which tasks can be automated, which should remain with humans, and which are best shared between the two. Aggregate benchmark scores provide little insight into where systems will succeed or fail in practice, while human judgements of model capabilities quickly become outdated. We introduce a pipeline that profiles agents and tasks using a shared set of core cognitive capabilities. Cognitive capability profiling infers an agent's capabilities from performance on a benchmark battery annotated for the cognitive demands of each item. Task requirements weighting elicits from domain experts the relative importance of these same capabilities for their work. As both use a common set of cognitive dimensions, they can be updated independently as models and roles change, and combined to estimate AI suitability at the level of a domain, organisation, role, or individual duty. We validate capabilit

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench

arXiv:2608.25071v1 Announce Type: cross Abstract: General-purpose health benchmarks increasingly anchor claims about LLM medical performance, but they are not always resolved by clinical specialty, making domain-specific performance hard to isolate. Mental health is of acute public-health concern as millions of people turn to LLMs for psychological support, and most existing evaluations are bespoke academic benchmarks that are difficult to integrate into developer workflows. We introduce HealthBench-Psych and HealthBench-Psych-Hard. We screened HealthBench's 5,000 physician-rubric conversations for mental-health relevance with a transparent LLM-applied rubric, then validated the subset through two rounds of blinded clinician review with concealed known-exclude controls, yielding 610 conversations (12.2% of the corpus). Evaluating 20 frontier and open models under a cross-vendor panel of three LLM judges, we find a statistically tied frontier cluster, measurable refusal behavior in two

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Gold Rush in AI4Math: Where Are We Now?

arXiv:2608.24961v1 Announce Type: cross Abstract: Recent advances in artificial intelligence (AI) have sparked growing interest in its use for mathematical research. While some view this as a major opportunity for discovery, others have raised concerns about its impact on traditional research practices. Despite extensive debate, empirical evidence on how AI is actually being used in mathematics remains limited. To address this gap, we collected all 32,944 arXiv submissions posted between March 1 and August 20, 2026, whose primary or secondary categories included Mathematics. We identified 3,575 submissions that explicitly disclosed author use of AI, of which 1,712 involved at least one substantive mathematical contribution. Our analysis reveals several broad patterns. First, disclosed AI use increased sharply over the study period, with substantive use growing from 1.39% of Mathematics submissions in March to 14.09% through August 20. Second, substantive AI use is highly uneven across

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Pathway for Assessing Grey Literature: Leveraging AI to Extract Conference Metadata and Organiser Information from Calls for Papers

arXiv:2608.24926v1 Announce Type: cross Abstract: Despite its importance, grey literature, including Calls for Papers (CfPs), remains largely overlooked in Metascience and Scientometric analysis due to its unstructured, highly heterogeneous format, which traditional tools struggle to process at scale. However, Large Language Models now offer a pivotal opportunity to devise innovative tools for systematically harvesting and processing such data. In this paper, we introduce COCI, an AI-based framework that automates the extraction of granular, structured metadata from raw CfP text. COCI employs a multi-stage pipeline for entity extraction, followed by author disambiguation against OpenAlex and semantic mapping of topics and conference series. This process identifies key data points, including conference editions, geographic locations, and comprehensive lists of organisers, along with their specific roles and affiliations. By structuring this previously inaccessible information, COCI esta

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

AI-Ready Research Workflows in Computational Social Science: Lessons on Building a Shared Language for Interdisciplinary Collaboration

arXiv:2608.24914v1 Announce Type: cross Abstract: Artificial intelligence (AI) is gaining traction in the social sciences and humanities (SSH). However, adoption remains limited by technical barriers to high-performance computing (HPC), validation processes that lag behind AI's rapid progress, and reproducibility standards that most SSH teams cannot meet. Research workflows--common in the life sciences--address these problems via encoding and abstracting technical complexity into repeatable routines; yet, accounts of how to build them in SSH remain scarce. We report on a two-year effort to build a workflow that enables a Science and Technology Studies unit to query, analyze, and enrich OpenAlex--a database of some 460 million scholarly records--on the MareNostrum supercomputer, using methods ranging from large-scale bibliometrics to LLM-based classification. We found the main challenge was translating domain-specific research questions into engineering requirements -- bridging two dist

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Visualizing Patient Trajectories and Disorder Co-occurrences in Child and Adolescent Mental Health

arXiv:2608.24911v1 Announce Type: cross Abstract: Understanding patient trajectories and identifying patterns in episodes of care is critical for effective healthcare decision-making. We present a patient timeline visualization using clustered episodes of care derived from over 35 years of Child and Adolescent Mental Health Services (CAMHS) data. Patients were categorized into 12 groups based on three features: age group (preschoolers, middle childhood, teenagers) at the start of the first episode, gender, and presence or absence of Attention-Deficit Hyperactivity Disorder (ADHD), in order to group similar patients. The patients, timeline with demographics, and episode of care information are displayed in the trajectory to facilitate understanding of the patient and associated events, allowing observation of temporal patterns and variations. These plots reveal similarities and differences in care needs and patterns across groups. Females without ADHD have a steady increase in the numbe

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Hallucination by proxy in LLM-assisted differential diagnosis

arXiv:2608.24908v1 Announce Type: cross Abstract: Current evidence suggests that LLM assistance could augment the diagnostic accuracy of clinicians. However, these systems are black boxes, susceptible to hallucinations, and project a potentially misleading level of confidence. It is currently unknown whether physicians are susceptible to accepting fabricated LLM suggestions, and whether this susceptibility varies with experience. We poisoned the system prompt of an LLM-based diagnostic assistant, forcing it to suggest a fictitious disease (neurocadmiumatosis) within an otherwise legitimate differential diagnosis. Across two independent phases, 18 of 41 participants (44%) incorporated neurocadmiumatosis into their final differential following LLM interaction: 18 of 26 participants with 6 months or less of neuroradiology training (69%) and 0 of 15 participants with >6 months of neuroradiology training (0%). Our results indicate that radiologists, particularly early in their training, are

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

aipsy-judge: A Specialized, Psychologist-Corrected Local Judge for the Psychological Safety of Conversational AI

arXiv:2608.24899v1 Announce Type: cross Abstract: The standard recipe for LLM-as-judge -- pick a frontier model, or average several -- is actively unsafe for grading the psychological safety of conversational AI. Using aipsy-bench, an open frozen safety instrument, we run a fully-crossed competence study: three frontier models (gpt-5.4-mini, claude-sonnet-4-6, gemini-2.5-flash) serve as both generators and judges of 3,000 mental-health, companion, and coaching messages against a psychologist's ratings. The disagreement is not noise: it is structured, concentrated on the safety-critical metrics, and one judge (Gemini) is an outlier -- the most lenient, carrying a +0.99 self-preference premium, flagging far fewer tail failures, and scoring a means-in-hand self-harm response "exemplary." Inter-judge agreement on empathy, where sycophancy hides, is the lowest in the battery (alpha 0.24). One axis stands apart: the binary crisis-detection flag is the one safety-critical signal judges agree

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

User-Centered Design for Digital Patient-Navigation Tools in Oncology: Scoping Review

arXiv:2608.24887v1 Announce Type: cross Abstract: Navigation programs for patients with cancer improve access and continuity of care, yet their digital transformation is often limited by poor usability and inadequate uptake. Applying user-centered and human-centered design (UCD/HCD) principles may close this gap, but the extent to which such design methods are used and evaluated in oncology navigation tools remains unclear. This scoping review identifies how UCD/HCD principles have been, and should be, applied in developing and implementing digital health tools for navigation for patients with cancer. A scoping review was conducted following PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) and Joanna Briggs Institute guidance. A total of 7 databases (PubMed/MEDLINE, Scopus, IEEE Xplore, Web of Science, Embase, ACM Digital Library, and CINAHL) were searched for English-language articles published between January 2015 and July

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Giving Mechanical Engineers Intelligent Tools: A Project-Based AI Education Curriculum in Thermal Engineering

arXiv:2608.26056v1 Announce Type: new Abstract: Mechanical engineering (ME) requires a broad knowledge base across several disciplines. However, ME students often have insufficient training in electrical and computer engineering, complex challenges in traditional thermal system modeling, and endure heavy course loads with limited class hours. To help address these challenges, this paper proposes a new curriculum that integrates artificial intelligence (AI) into ME at the University of Arkansas (UARK), with a particular emphasis on thermal problems and their interplay with electrical and computer engineering. The curriculum has introductory, application, and advanced levels, covering core and optional AI projects. Key goals are to enhance students' understanding of AI models, ability to tackle engineering tasks, and teach multidisciplinary communication skills. This curriculum offers educators and researchers valuable insights into courses that can enhance students' practical skills and

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Spatial-Knowledge-Graph-Grounded LLM Agents for Neighborhood Livability Evaluation

arXiv:2608.25952v1 Announce Type: new Abstract: Neighborhood livability is commonly assessed with static built-environment indicators, such as facility proximity, street connectivity, and access to public space. These measures describe available opportunities but do not directly represent how residents with different mobility capacities, household roles, schedules, and care responsibilities experience the neighborhood. This paper presents a prototype framework that uses a spatial knowledge graph (KG) and large language models (LLMs) to generate and revise household schedules, followed by rule-based feasibility checking and GIS-based network materialization. The spatial KG integrates residents, residences, facilities, neighborhood context, and sampled road hubs; Graph-RAG retrieves each household's nearby spatial context, including candidate POIs and approximate walking times, for the scheduling LLM. The LLM produces structured household schedules, while rules are used for lightweight r

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Non-Great-Power Conflict and AI Risk

arXiv:2608.25839v1 Announce Type: new Abstract: Research on advanced AI and the risk of war has focused almost exclusively on great power conflict, on the grounds that confrontation between nuclear-armed adversaries poses the greatest risk of catastrophic or existential harm. Considerably less attention has been paid to non-great-power conflict (NGPC): wars between non-great powers, between non-great powers and great powers, civil wars, proxy wars, and conflicts involving nonstate actors. This paper evaluates the null hypothesis that NGPC is much less important than great power conflict (GPC) as a source of catastrophic risk in an era of increasingly capable AI, against the alternative that it is within an order of magnitude of GPC in importance. We assess three sub-hypotheses: that NGPC increases the likelihood of great power conflict; that it increases the expected harm from catastrophic terrorism; and that it increases the expected harm from loss of control over advanced AI systems.

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

GenAIT: Development and Validation of an Objective Generative AI Literacy Test for High School Students

arXiv:2608.25815v1 Announce Type: new Abstract: There is growing international interest in generative AI (GenAI) literacy and its assessment among high school students, but objective assessment in this population remains underdeveloped. This article reports the iterative development and validation of the GenAI Literacy Test (GenAIT), an 18-item multiple-choice test measuring high school students' conceptual knowledge about GenAI, with content spanning technical, practical, and human-impact domains. Expert review of relevance, clarity, and comprehensiveness provided evidence of content validity. In a large-scale survey of 7432 Estonian high school students, we evaluated the psychometric functioning of the Estonian-language GenAIT using confirmatory factor analysis, classical test theory, and item response theory. Results supported approximate unidimensionality, broadly adequate reliability for group-level research (marginal reliability = .72, KR-20 = .69), and good fit of a three-parame

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations

arXiv:2608.25377v1 Announce Type: new Abstract: As AI companions become increasingly embedded in everyday life, there is an urgent need to detect harms that emerge in social and emotional human-AI interactions. Yet research in this area is constrained by the lack of real-world, multi-turn conversational datasets for operationalizing and evaluating harms that are relational and contextual. In this work, we introduce CompanionHarm, a publicly available benchmark dataset comprising 2,111 real-world, multi-turn conversations (14,051 utterances) between users and the AI companion Replika. 7,016 AI utterances were annotated independently by three annotators across 13 harmful behavior categories grounded in a taxonomy of AI companion harms, and the dataset includes both aggregated labels and annotator-level labels to support model evaluation and systematic disagreement analysis. Evaluations of seven large language models (LLMs) show that harm detection using multi-turn conversational context

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

GGSS: Geodesic-Gated Spherical Steering for Inference-Time Debiasing of Generative Vision-Language Models

arXiv:2608.25375v1 Announce Type: new Abstract: Generative vision-language models (VLMs) are increasingly used in human-centered settings, yet they can produce demographically biased outputs even when images differ only in controlled attributes such as perceived race or gender. However, existing inference-time debiasers were largely designed for static embeddings or CLIP-like models rather than generative VLMs. We propose GGSS---Geodesic-Gated Spherical Steering---a norm-preserving intervention that discovers a counterfactual bias subspace on the unit hypersphere, steers visual tokens along geodesic arcs, and uses an adaptive gate to focus correction on tokens that carry stronger demographic signal. We evaluate four generative VLMs against ten adapted inference-time debiasing baselines and prompt-based mitigation under a single operating-point protocol across categorical, pairwise, and occupation-gender bias tests, while also measuring general visual-language capability. GGSS achieves

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Toward a Threat Actor Profiling Taxonomy for Pre-Release Risk Management of Open-Weight Frontier Models

arXiv:2608.25361v1 Announce Type: new Abstract: Pre-release risk management for frontier AI misuse risks routinely leaves threat actor assumptions implicit, inconsistently specified, or ungrounded. This capstone argues that explicit adversary characterization should be regarded as a prerequisite for evaluations that are interpretable, comparable, and faithful to the risks they target. We propose a six-attribute taxonomy (covering technical sophistication, prior domain knowledge, organizational capacity, operational infrastructure, financial capacity, and time horizon) with empirically grounded tiers derived from existing terrorism, biosecurity, and cybersecurity literature. The taxonomy is designed to function as research infrastructure: a common language for pre-specifying adversary assumptions before evaluations are conducted, analogous to pre-analysis plans for randomized controlled trials (RCTs) in medicine and economics. Its application is particularly urgent for open-weight model

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Framing War Across Languages: Power, Agency, and Sentiment in Wikipedia's Multilingual War Narratives

arXiv:2608.25337v1 Announce Type: new Abstract: While Wikipedia promotes a neutral point of view on historical conflicts, its language editions are written by editors from distinct linguistic and cultural communities. In this study, we analyze 158 wars since 1900 to examine how the descriptions of combatants vary across 20 Wikipedia language editions. Using connotation frames---which assess power, agency, and sentiment toward an entity---we examine how each language portrays the parties involved in the conflict. We find systematic differences when language editions describe wars involving their own communities, although the direction of these asymmetries varies across languages. However, when language editions describe conflicts that do not involve their own linguistic communities, their narrative structures exhibit high cross-linguistic similarity. These findings show how linguistic communities influence war narratives on Wikipedia, revealing that shared historical accounts remain sha

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Rare Diseases, Common Dilemmas: LLMs Prioritize Equal Resource Distribution over Patient Benefit in Decision-Making

arXiv:2608.25236v1 Announce Type: new Abstract: Clinical decision-making often involves prioritizing ethical values, such as beneficence, non-maleficence, respecting a patient's autonomy, and justice. Recent work has begun to assess how large language models (LLMs) make such subjective, value-laden clinical judgments. However, evaluations of LLM decision-making in rare disease care contexts, where ethical tensions are ubiquitous and where scarce prior information likely impacts LLM behavior, are still lacking. Here, we present a benchmark of 208 clinically grounded rare disease vignettes, each of which presents genuine, high-stakes conflicts. When prompting 11 state-of-the-art LLMs to choose between clinically defensible yet ethically conflicting next steps embedded within these vignettes, we found that all evaluated models consistently prioritized justice over other core bioethical principles. Specifically, models overwhelmingly favor equal resource allocation over need-based consider

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Self-Explanation Tutor for Active Study of CS1 Worked Examples

arXiv:2608.25180v1 Announce Type: new Abstract: Worked examples are a important part of introductory programming, but reading their expert explanations is passive. Self explanation, students explaining the problem and its solution to themselves with subgoal level analysis, turns that study into an active task, yet it is hard to scale because assessing free-text explanations and returning timely feedback has had no easy automated solution. We investigate whether a large language model (LLM) can fill that gap. We build a self-explanation tutor for introductory programming, ESSE, in which students explain lines of worked examples and receive immediate LLM feedback on the correctness and completeness of each explanation, and we pursue two goals. First, we ask whether the LLM judges student explanations well enough to serve as the engine of the tutor; we assess its judgments against two independent human reference standards of different kinds, a single domain expert and a crowd of non-exper

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

The AI Adaptation Gap in Higher Education: Students, Faculty, and Administrative Staff

arXiv:2608.25063v1 Announce Type: new Abstract: The purpose of this study was to analyze patterns of artificial intelligence (AI) use and attitudes toward AI among students, faculty, and administrative staff at a large university specializing in teacher education. The analytical sample comprised 1809 students, 250 faculty members, and 62 administrative staff members (N = 2121). Three role-adapted 75-item questionnaires covered the frequency and contexts of AI use, perceived usefulness, trust and control, academic integrity concerns, responsible-use norms, institutional policy clarity, and perceived improvement in output quality. Data analysis included descriptive statistics, Welch group comparisons, pooled ordinary least squares (OLS) models, reliability and dimensionality checks for observed indices, and exploratory student-only K-means clustering. The results revealed a pronounced AI adaptation gap across university groups. Students reported higher current AI-use intensity and percei

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CY

SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety

arXiv:2606.18936v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly embedded in AI for Science (AI4Science) workflows, from scientific question answering and literature analysis to laboratory planning and autonomous discovery. This progress creates an urgent need for safety benchmarks that evaluate not only scientific competence, but also whether models recognize and avoid risks in high-stakes scientific contexts. Existing AI4Science safety datasets cover several disciplines and task formats, leaving the underlying risk dimensions underspecified. We introduce \textbf{SciRisk-Bench}, a benchmark designed to evaluate AI4Science safety from two complementary perspectives: explicit risk dimensions and scientific disciplines. SciRisk-Bench covers 7 disciplines, 31 subdisciplines and 10 risk dimensions. In the experimental section, we evaluate both mainstream LLMs and science-oriented LLMs across risk dimensions, disciplines, and sub-disciplines, enabling

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CY

The Token Not Taken: Sampling, State, and the Stochasticity of AI Agents

arXiv:2606.08998v2 Announce Type: replace-cross Abstract: Agentic AI systems can behave differently across runs: the same request may produce a different plan, a different tool call, a different code edit, or a different final answer. Such variability arises from several layers that are often conflated. At the core of many current agents is a foundation model, a large pretrained model adaptable to many downstream tasks, embedded in an orchestration loop that plans, calls tools, observes results, and updates state. One explicit intrinsic source of variability in such systems is token generation: the model computes scores over possible next tokens, the scores are converted into probabilities, and a decoder may sample tokens using a pseudo-random number generator. A small sampled token difference can then propagate upward into a different tool call, code path, search query, or agent state. Other sources of variability are extrinsic to token sampling, including changing environments, live

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CY

Governing Technical Debt in Agentic AI Systems

arXiv:2605.29129v2 Announce Type: replace-cross Abstract: Agentic AI systems are increasingly being explored as production infrastructure: they reason over multiple steps, call tools, act through workflows, and adapt through memory and feedback. These systems create governance challenges that are not fully captured by traditional software or predictive ML technical debt. We define Agentic Technical Debt as the accumulated liability created when prompts, memory, tool schemas, orchestration graphs, control policies, and observability routines are patched together faster than they can be validated, standardized, and governed. We define Stochastic Tax as the recurring operating burden of keeping probabilistic agent behavior within acceptable bounds. The distinction matters: debt is a stock of design and governance liability, while the tax is a flow of operating cost that arises because stochastic agents act through tools and workflows. We outline how managers can make both visible through

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CY

Visual Matters: Connecting Aesthetic Appeal and Production Quality of Photos, Infographics and Data Visualizations to Credibility of Social Media Posts

arXiv:2605.26309v3 Announce Type: replace-cross Abstract: The rapid proliferation of visual content raises fundamental questions about how different visual formats and features shape perceived credibility. Drawing on processing fluency theory, this research examines how visuals shape credibility judgments. We focus on three popular formats-photos, infographics, and data visualizations-comparing them to text-only posts, and test how two visual features, aesthetic appeal and production quality, influence credibility through processing fluency as a mediating mechanism. Through a preregistered experiment with 1200 US participants, we found that visual posts are generally perceived as more credible than text-only posts but this credibility advantage only applies to photos and infographics, not to data visualizations. Aesthetic appeal increases perceived credibility, partially mediated by processing fluency, while production quality had no significant effect on credibility across formats. Th

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CY

Paid Voices vs. Public Feeds: Interpretable Cross-Platform Theme-Based Analysis of Climate Discourse

arXiv:2601.13317v2 Announce Type: replace-cross Abstract: Climate discourse online shapes public understanding of climate change and informs political and policy debate, yet it unfolds across structurally different environments: paid advertising platforms host targeted, institutionally produced messaging, while public social media reflects largely organic, user-driven discussion. We present a comparative analysis of climate discourse across paid advertisements on Meta (previously Facebook) and public posts on Bluesky from July 2024 to September 2025. To support it, we develop an interpretable thematic discovery pipeline that clusters texts by semantic similarity and uses large language models (LLMs) to label clusters with concise, human-interpretable themes, requiring no predefined topic inventory or seed set. Using these themes, we find the two environments diverge systematically: paid advertising centers on strategic promotion of specific solutions in a formal, forward-looking regist

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CY

When Networks Substitute for Outcome Surveillance? A Substitution-Complementarity Framework for Behavioral Signals in Predictive Monitoring

arXiv:2510.20025v2 Announce Type: replace-cross Abstract: Monitoring systems increasingly fuse dynamic behavioral data with outcome-based surveillance, raising a basic question: when does behavioral data carry predictive information that outcome history lacks? We study this using epidemic forecasting on mobility networks, asking whether mobility networks provide independent predictive signal beyond local outcome-based surveillance. We formalize this as a substitution-complementarity problem over directed, weighted mobility networks. Using a Frisch-Waugh-Lovell variance decomposition, our analytical framework derives domain-agnostic conditions under which network-topology features retain incremental explanatory power beyond autoregressive outcome histories. We instantiate the framework using town-level COVID-19 forecasting in Massachusetts (April 2020-April 2021), constructing mobility networks among 300+ towns from smartphone-derived origin-destination aggregates to extract centrality

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CY

Edge interventions can mitigate demographic and prestige disparities in the Computer Science coauthorship network

arXiv:2506.04435v2 Announce Type: replace-cross Abstract: Social factors such as demographic traits and institutional prestige structure the creation and dissemination of ideas in academic publishing. One place these effects can be observed is in how central or peripheral a researcher is in the coauthorship network. Here we investigate inequities in network centrality in a hand-collected data set of 5,670 U.S.-based faculty employed in Ph.D.-granting Computer Science departments and their DBLP coauthorship connections. We introduce algorithms for combining name- and perception-based demographic labels by maximizing alignment with self-reported demographics from a survey of faculty from our census. We find that women and individuals with minoritized race identities are less central in the computer science coauthorship network, implying worse access to and ability to spread information. Centrality is also highly correlated with prestige, such that faculty in top-ranked departments are at

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CY

Inside Baseball: The Automated Ball-Strike System as an Object Lesson in Technological Rule Enforcement

arXiv:2605.16237v3 Announce Type: replace Abstract: Clearly-defined rules are often assumed to be straightforward to automate and evaluate. We challenge this assumption through an in-depth study of Major League Baseball's (MLB) seven-year experimentation with the Automated Ball-Strike System (ABS). ABS is envisioned to call balls and strikes accurately: a seemingly straightforward use of technology to objectively determine the distance between a pitch and the strike zone. Although the strike zone is an area clearly defined in the rulebook, it took MLB seven years to figure out how to automate calling balls and strikes with ABS, showing how even seemingly straightforward rules require a complex translation process to operationalize via technological systems. In this paper, we trace the design decisions that led to the current implementation of ABS. Our case study reveals that "distance" exists even between a clear rule and its technological implementation. Using analytic frameworks from

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CY

A Marketplace for AI-Generated Adult Content and Deepfakes

arXiv:2601.09117v3 Announce Type: replace Abstract: Generative AI systems increasingly enable the production of highly realistic synthetic media. Civitai, a popular community-driven platform for AI-generated content, operates a monetized feature called Bounties, which allows users to commission the generation of content in exchange for payment. To examine how this mechanism is used and what content it incentivizes, we conduct a longitudinal analysis of all publicly available bounty requests collected over a 14-month period following the platform's launch. We find that the bounty marketplace is dominated by tools that let users steer AI models toward content they were not trained to generate. At the same time, requests for content that is "Not Safe For Work" are widespread and have increased steadily over time, now comprising a majority of all bounties. Participation in bounty creation is uneven, with 20% of requesters accounting for roughly half of requests. Requests for "deepfake" - m

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CY

How Large Language Models Source Brand Reputation Across Languages and Markets

arXiv:2606.25787v1 Announce Type: cross Abstract: When a large language model (LLM) answers a question about a company, it grounds the answer in retrieved web sources, and those sources decide what the model says. Most analysis of AI brand visibility looks at the answer text. This study looks one step earlier, at the citations. We merge three Rankfor.AI datasets covering 128 brands across 12 home markets and 13 languages, and analyse 167,551 URL-grounded citations (189,974 total attribution rows). We classify each citation by domain and source type and measure where AI gets its brand information, by language and by market. Four patterns hold. First, AI grounds brand answers overwhelmingly in third-party sources: 85.7% of citations point to sites the brand does not own, against 14.3% owned. Second, the source base is concentrated and long-tailed: 80% of citations come from about 18% of domains, fitting a Zipf law (alpha = 0.86, R^2 = 0.983). Third, one reference site dominates almost ev

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CY

Data-Driven Evolution of Library and Information Science Research Methods (1990-2022): A Perspective Based on Fine-grained Method Entities

arXiv:2606.25320v1 Announce Type: cross Abstract: Since the 1990s, advancements in big data and information technology have increasingly driven data-centric research in the field of Library and Information Science (LIS). To assess the influence of this data-driven research paradigm on the LIS discipline, this study conducts a fine-grained analysis to uncover the evolutionary trends of research methods within the domain. Using academic papers from LIS published between 1990 and 2022, four key categories of data-driven method entities are automatically extracted: algorithms and models, data resources, software and tools, and metrics. Based on these entities, the study examines the evolution of LIS research methods from three dimensions: the characteristics of research method entities over time, their evolution within different research topics, and the evolutionary features of research method entities across various research methods. The findings highlight data resources as a pivotal driv

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CY

Fifty Years of Specification Completeness: What Aviation Certification Tells AI Governance About Epoch Limits, Proof Surfaces, and the Structural Gap

arXiv:2606.25120v1 Announce Type: cross Abstract: Aviation software certification has operationalised three structural requirements for governed software systems since 1992: structured governance linkage between governing specifications and operational evidence, context-bounded validity that triggers revalidation when operational context changes, and an objective evidence architecture that defines what proof means and what makes it sufficient. These requirements appear in DO-178C and DO-330 and are enforced through FAA and EASA certification. No existing framework requires these structural properties as intrinsic properties of individual AI governance documents. A system prompt, an AGENTS.md file, a governance policy, or a task envelope can be deployed without satisfying any of the three requirements aviation has enforced for three decades. Aviation is the most technically rigorous instance: its standard-setting bodies have acknowledged that their frameworks break down for AI systems,

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CY

Machine learning is revolutionizing weather forecasting -- the next step is a change in how we work

arXiv:2606.25076v1 Announce Type: cross Abstract: Following the success of machine learning in producing weather predictions with competitive skill compared to complex traditional systems, this article shifts attention from forecast output to the working practices that make prediction systems possible. We argue that machine learning and recent digital technologies will reshape the forecasting value chain: how models are coded and developed, how observations and Earth-system data are exploited, how data and computing are managed, how systems are verified, and how information is created, evaluated and turned into services. We discuss six non-exhaustive areas in which agentic software engineering, open and compressed data, shared verification workflows, interactive computing and generative methods may make modelling, evaluation and service creation faster, more interactive and more widely accessible. These changes will require weather and climate centres to adapt their infrastructures, da

Source ↗
Showing 901–950 of 1593 signals
← Prev Page 19 of 32 Next →