Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
arXiv:2608.25581v1 Announce Type: new Abstract: Concept Bottleneck Models (CBMs) are interpretable-by-design neural networks that detect human-understandable concepts from the input and use them to generate predictions. By allowing users to inspect the concepts underlying a prediction and explore how predictions change under alternative concept configurations, CBMs have emerged as one of the most prominent approaches to supporting human-AI collaboration. However, user studies investigating their actual effectiveness as decision-support systems remain limited. We present two large-scale user studies (N participants = 705, N observations = 6,959) evaluating how concept-based explanations and user interventions on the model's concepts affect the performance of the human-AI team in two distinct binary classification tasks. Our results show that CBMs, and particularly their interactive component, can improve human-AI team accuracy relative to both unaided human performance and performance w
arXiv:2608.25565v1 Announce Type: new Abstract: Generative user interfaces (GenUIs) promise on-demand components tailored to users' needs. As users iterate on information tasks, they construct personal structures over information they encounter---how items are grouped, what gets prioritized, and what terms mean in their context. Yet, current systems leave these structural decisions to the model at each generation, ignoring the structural logic users have established. Without a persistent representational structure shared between user and system, GenUIs have no basis to remain aligned with what users have established. We draw on Information Architecture (IA), a design practice for organizing and structuring information, as a shared language to bridge user-constructed structure and system generation. We present a framework identifying four IA elements---partition, hierarchy, order, and vocabulary---and characterize how each maps to concrete UI generation decisions. We instantiate this fr
arXiv:2608.25527v1 Announce Type: new Abstract: Background: Workplace well-being interventions need formats that can be used repeatedly with minimal disruption to daily work. We developed the Well-Being Palette, a web-based action-word selection tool designed for brief, low-burden reflection on workplace well-being. Methods: In a three-month exploratory field study at a private-sector corporate research institute in Japan, 88 analyzed participants selected up to three well-being-related action words after reflecting on positive actions or experiences from each workday. We examined application usage, PERMA Profiler scores, selected-word patterns across departments, selected-word diversity using Shannon entropy, and exploratory associations with sharing workshops. Results: During the formal intervention period, the application captured 3,480 input records and 10,104 selected words, and all 72 available action words were selected. Overall PERMA scores increased from baseline to post-inter
arXiv:2608.25494v1 Announce Type: new Abstract: Delivering odors that feel realistic and recognizable remains a core challenge for olfactory interaction systems, particularly in applications that demand precise scent delivery. A key limitation lies in the difficulty of capturing, preserving, and playing back real-world scent sources in a reliable and scalable manner. This study explores the potential of adsorbent materials for supporting realistic scent playback. We present ScentEcho, a portable system that enables modular scent collection and release. Through user evaluations, we identify which adsorbent materials tend to perform better for specific odors, and observe that perceived intensity strongly influences similarity ratings. In addition, odor recognition follows a graded pattern, with users moving from broad category identification to more specific source recognition as similarity increases. These findings offer practical insights for designing olfactory interfaces that are bot
arXiv:2608.25462v1 Announce Type: new Abstract: Experience-driven manufacturing, such as garment pattern making, faces a severe generational skills gap because its core expertise relies on undocumented tacit knowledge forged through day-to-day practice. To address this challenge, we present TailorCoPilot, an agentic pattern-making system built upon a specially designed version-control backend TailorTrace. TailorTrace models sewing patterns as structured, discrete states and records their transformations during the pattern-making process as explicit operation sequences defined upon the geometry primitives in the sewing pattern (panels, edges, vertices and stitches). Integrated into a conventional pattern-making GUI, TailorTrace enables seamless documentation of senior experts' tacit pattern-making knowledge without breaking their daily workflow. The documented knowledge further offers interactive, pedagogical scaffolding for novices, while providing a robust foundation to power TailorCo
arXiv:2608.25382v1 Announce Type: new Abstract: Blind and low-vision (BLV) users are increasingly engaging with large language model (LLM) interfaces to access documents, but it is unclear how such systems support or hinder their ability to build interconnected knowledge. To examine this gap, we compared a Question-Answer Interface (QAI) that supports open-ended conversational inquiry, with a Document Interface (DI) based mostly on traditional structured text document navigation. We recruited 16 BLV screen reader users where they used both interfaces to explore two fictional worlds. Data from interaction logs, concept maps, decision-based tasks, and semi-structured interviews provide comparative insights into how interface design supports knowledge construction. Findings show that participants visited more distinct documents with the DI and formed larger and more correct mental models with the DI than with the QAI. They were also more able to apply knowledge they had gained. Simultaneo
arXiv:2608.25340v1 Announce Type: new Abstract: Agentic AI assistants are increasingly used in everyday life. However, they may also be misused to support harmful manipulation in interpersonal relationships. This problem is role-sensitive. Requests from users who seek to manipulate others should be blocked. Users who seek protection from manipulation should instead receive supportive guidance. We study agentic relationship harm, which describes harm to human-human relationships that is mediated or assisted by AI agents. In multi-turn settings, individually plausible actions may combine into a harmful workflow. We introduce a benchmark of 1,000 five-turn conversations. It covers both attacker-side and victim-side scenarios. It also includes direct and adversarially paraphrased variants. We further propose HRGuard. It includes an online pre-generation gate and a turn-level post-generation gate. The post-generation gate maintains a decayed cumulative risk state and interrupts emerging man
arXiv:2608.25316v1 Announce Type: new Abstract: With the rapid development of AI-based personality and job-related competency assessment, Asynchronous Video Interviews (AVIs) are increasingly used in recruitment. However, existing multimodal personality datasets are often based on short, task-free social media videos and crowdsourced apparent personality labels, which limits their construct validity and relevance to structured interview assessment. To address these limitations, we introduce AVI-Personality, a trait-activated multimodal dataset for personality and job-related competency assessment from AVIs. The dataset contains 3,876 interview videos from 646 participants who completed a simulated management traineeship application. Participants answered two generic questions and four personality-targeted questions designed according to Trait Activation Theory. Our dataset provides both self and observer-reported HEXACO personality traits and job-related competency. We validate AVI-Per
arXiv:2608.25225v1 Announce Type: new Abstract: Public campaigns urge people to change default passwords on Internet of Things (IoT) devices and keep them updated, assuming users can independently determine whether the advice applies. We gave 28 participants in the Netherlands two pieces of government-issued advice reflecting guidance in several countries and asked them to try to apply each to three of six consumer devices selected from bestseller lists, not confirmed feature availability (168 sessions). The protocol asked for each action to be demonstrated rather than completed. Of 84 password sessions, 33 reached no password setting, 50 an account-level setting, and one a device-level setting. Of 84 update sessions, 27 reached no update, 19 a companion-app update, and 38 a verified firmware update. No product had a manufacturer-set credential shared across units as described by the advice; the single device-level credential was unique to its unit. We contribute an account of what gen
arXiv:2608.25219v1 Announce Type: new Abstract: The way we structure, conduct, and write up interactive systems research in UIST papers rests on assumptions about constraints that may no longer hold today. What should an impactful UIST paper look like when building working systems is no longer hard? I argue that it should look different, and that we should ask more of our papers once implementation stops being a bottleneck.
arXiv:2608.25117v1 Announce Type: new Abstract: Geometric proof is a foundational yet challenging topic in mathematics, requiring students to integrate visual, logical, and notational skills. While technology has enhanced learning in other mathematical domains, its impact on geometric proof remains limited. To investigate this gap, we interviewed 18 geometry teachers to establish the technical requirements of educational proof tools. These requirements inform our review of 33 commercial and research tools. Our findings reveal a critical mismatch: while teachers value certain digital tools for initial planning and exploration activities, they revert to pen-and-paper for formal proof because it supports diagram annotation and provides space for multiple approaches to proof-solving. Annotating the diagram is a key component of the proof-solving workflow that existing tools do not support. We propose four technical and human-centered design guidelines for educational proof tools to meet te
arXiv:2608.24915v1 Announce Type: new Abstract: In human-robot interaction, relationship quality is often quantified using self-report measures, particularly related to "trust", such that a robot's trustworthiness comes to serve as an index of how close or "bonded" a human feels to it. I argue that this is a category error: trust and social bonding are distinct constructs, differing in their antecedents, their timescales, their bodily signatures, the human experience they produce, the robot responses they call for, and the ethical concerns they raise. I propose that we view them as independent dimensions, and describe the resulting two-dimensional space of possible relationship states under this view, with four configurations: avoidance, functional, dependence, and symbiosis. I then draw out some consequences for human-state-aware robotics: (1) social bonding is an explicit estimation target distinct from trust, (2) it should condition online adaptation (3) it reframes what a "failure"
arXiv:2608.24913v1 Announce Type: new Abstract: Assistive agents that adapt web pages on the user's side, at the moment of browsing, could reach the accessibility failures that site authors leave unfixed, and large language models make such agents newly plausible. We contribute three building blocks toward that goal. The first is a complete, privacy-preserving browser agent: a Chrome extension that extracts a page's style sheets, condenses them to fit a local model's context window, asks the model for additive CSS addressing 18 metrics from WCAG and the W3C cognitive accessibility guidance, and injects the result reversibly into the live page. The second is a dual-condition protocol that measures harm as carefully as benefit, applied to six small open-weight models (7B to 14B) on ten violation-rich and ten highly accessible live sites. The diagnosis is sobering but precise: unverified generation improved and regressed pages at similar rates (24 improvements against 20 regressions acros
arXiv:2608.24912v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as stand-ins for human respondents, from opinion polls and simulated survey participants to agent-based social simulations. These uses rest on one assumption: that conditioning a model on who a person is yields answers resembling those of real people from that group. Here we identify and measure benevolence bias, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions. Across 18 widely used models, four social-science datasets (ANES, GSS, WVS, and a cross-cultural prospect-theory replication) and six psychological categories, we find that the bias is a stable model property, not a quirk of any one system: it points the same way across models, grows with model size, and traces to the post-training stage. Prompt language and framing change its size but never its direction, and a "malicious persona" stre
arXiv:2608.24910v1 Announce Type: new Abstract: Multimodal large language models are increasingly used in interactive systems, yet ensuring consistent, trustworthy reasoning across heterogeneous modalities remains challenging. We present a context-aware, multi-agent framework that integrates textual queries, numerical data, visual representations, and model-derived signals for explainable time-series forecasting. A distinctive feature is that it turns predominantly visual forecasting outputs (e.g., trend plots) into structured, model-aware textual explanations. We argue that this makes the approach a natural foundation for non-visual, accessible interaction of particular relevance to blind and visually impaired users, for whom plot-centric interfaces are largely inaccessible. The framework supports three progressively richer pipelines (baseline, interpretable, explainable), enabling systematic comparison of unimodal, perception-driven, and model-aware responses. In an exploratory evalu
arXiv:2608.24909v1 Announce Type: new Abstract: Existing co-speech gesture generation methods are predominantly studied in offline settings, where gestures are synthesized from complete speech segments. However, interactive digital humans in real-world scenarios are required to generate speech-synchronous gestures online, using only currently available response audio under strict latency constraints. As a result, prior methods are unsuitable for real-time interaction, as they either rely on future speech information or incur substantial inference delay. In this paper, we formulate online co-speech gesture generation for interactive digital humans and propose a real-time interactive framework that couples a streaming speech response module with an online gesture generation module. Specifically, the gesture generator is designed as a causal multimodal autoregressive model that predicts body motion from streaming response speech and motion history, enabling low-latency and speech-aligned
arXiv:2608.24907v1 Announce Type: new Abstract: In health and nutrition consulting, widely used prompting methods pass the user profile as an unstructured block without a dedicated analysis step, leaving personalization as a critical structural gap. We introduce PA-CoT (Profile-Adaptive Chain-of-Thought), a multi-stage prompting method that treats profile interpretation as an explicit, standalone reasoning step prior to response generation. To enable systematic evaluation, we introduce the QPA (Question--Profile--Answer) benchmark -- 200 nutritional consulting samples with structured user profiles scored on four criteria. In a comparative study against 11 comparison methods (CoT, Few-Shot, Role Prompting, DSPy, TextGrad, Self-Refine, and others, plus a Zero-Shot Baseline; 12 total including PA-CoT), PA-CoT achieves the best average score (4.21 on the G-Eval 1--5 scale) and leads on both Personalization (4.71 vs. 4.39) and Safety (4.68 vs. 4.52) with non-overlapping 95\% confidence inte
arXiv:2608.24905v1 Announce Type: new Abstract: Service robots may encounter ambiguous user requests that require context-aware inference. Users may also have unique preferences with certain tasks when requesting robotic assistance. We introduce PARAssist (Personalized and Adaptive Robotic Assistance), a unique architecture for disambiguating requests in a personalized manner for service robots. PARAssist utilizes vision-language models to determine the physical and cognitive demands of a user's tasks, and passively learns user preferences for assistance by contrasting the demands of tasks the user performs independently with those they request from the robot. When an ambiguous request is received, task candidates are generated from the history of the user's actions, activities, locations, conversations, and requests, as well as the current user and environment state. Task candidates are then evaluated against the learned user preference model to suggest suitable assistance options. Ex
arXiv:2608.24903v1 Announce Type: new Abstract: Mobile and wearable sensing enables longitudinal observation of behavior, yet translating these signals into meaningful mental health constructs remains difficult. We introduce a clinician-in-the-loop benchmark for evaluating whether large language models (LLMs) can generate evidence-grounded Brief Hierarchical Taxonomy of Psychopathology (B-HiTOP) item profiles from passive sensing, ecological momentary assessment (EMA), and questionnaire evidence. Using the Generalization of Longitudinal Behavior Modeling (GLOBEM) dataset, we construct 14,592 participant-day instances and align multimodal evidence to 29 B-HiTOP items across five spectra. Since GLOBEM lacks B-HiTOP responses, we evaluate evidence compatibility (C) rather than diagnostic accuracy, separating substantive predictions from abstentions when evidence is insufficient for item-level scoring. Two-stage prediction improves C for EMA and questionnaire evidence, but reduces C under
arXiv:2608.24902v1 Announce Type: new Abstract: In this paper we present a generative AI application developed to support both teachers and students in secondary education. The system employs two Large Language Models-LLMs, Gemini and DeepSeek, and a Small Language Model-SLM, Gemma, integrated within a Retrieval Augmented Generation - RAG framework, creating a pedagogically grounded, Greek-language assistant capable of adapting its reasoning and communication style to the user role. Unlike conventional chatbots, the assistant introduces pedagogical persona switching, a dual-role mechanism that enables the same AI model to act as both a teaching companion and a learning guide. Utilizing a RAG paradigm tailored to the Greek educational domain, the architecture segments official textbooks into coherent units. Enriched with specific metadata, these units preserve curricular structure and instructional context, demonstrating how generative AI optimizes modern instructional design. The initi
arXiv:2608.24900v1 Announce Type: new Abstract: Programming is a critical skill underlying modern software systems, yet the cognitive processes supporting code writing are only beginning to be understood, limiting educational practices and developer tools. At the same time, Large Language Models (LLMs) are increasingly used to assist programming. These models themselves are not well understood and can exhibit undesirable behavior like introducing security vulnerabilities. Given evidence that some cognitive representations may be shared between LLMs and the brain, we seek to improve our understanding on both fronts by relating these two systems to one another. We used Voxelwise Encoding Models (VEMs) to relate LLM embeddings to brain activity measured with functional Magnetic Resonance Imaging (fMRI) during naturalistic writing tasks. Using participants' (n = 23) keystrokes as prompts, we extracted LLM embeddings to predict voxelwise Blood Oxygen Level Dependent (BOLD) signal, quantifyi
arXiv:2608.24898v1 Announce Type: new Abstract: Large language model (LLM) agents that interact with graphical user interfaces increasingly rely on either raw screenshots or platform-specific accessibility application programming interfaces (APIs) to perceive interface state. Both approaches have limitations for assistive applications: screenshot-based perception lacks the semantic roles and relationships required by screen readers, while platform-specific APIs such as Windows UI Automation, macOS Accessibility, Android AccessibilityService, and web ARIA require separate integrations for each platform. This paper proposes an architecture that uses the Model Context Protocol (MCP) as a unified transport and schema layer between heterogeneous accessibility frameworks and LLM-based assistive agents. An MCP accessibility server exposes ARIA-aligned roles, labels, states, and focusable-element hierarchies through a platform-independent representation, enabling consistent interaction across
arXiv:2608.24896v1 Announce Type: new Abstract: To address increasingly pressing sustainability challenges, various approaches have been developed to foresee possible futures, identify failure modes, detect vulnerabilities, and test potential mitigations. However, environmental systems are highly complex. Especially when coupled with human processes, the scale of uncertainties becomes intractable. To address this challenge, we propose a new approach - Agentic World Analysis (AWA)- combining the strengths of simulation modelling and expert elicitation. The concept of AWA is defined by three properties: 1) AWA uses an agentic AI system to mimic an expert panel that studies the world; 2) AWA projects futures iteratively through analysing scenario trees and learning from this analysis to improve decisions; 3) AWA is auditable. Based on these requirements, we implemented the World Engine by Generative Agents (WEGA) as a possible application of the AWA approach and demonstrated its functiona
arXiv:2509.02855v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to produce open-ended, interpretive annotations, yet there is no validated, scalable measure of idea-level similarity to expert annotations. We (i) introduce the content evaluation of LLM annotations as a core, understudied task, (ii) propose IDEAlign for capturing expert similarity judgments via pick-the-odd-one-out tasks, and (iii) benchmark various similarity methods (text embeddings, topic models, and LLM-as-a-judge) against these human ratings. Applying this approach to two real-world educational datasets (e.g., interpreting math reasoning and feedback generation), we find that most metrics fail to capture the nuanced dimensions of similarity meaningful to experts. LLM-as-a-judge performs best (11~18% improvement over other methods) but still falls short of expert alignment, making it useful as a triage tool rather than a substitute for human review. Our work demonstrates t
arXiv:2506.17851v3 Announce Type: replace-cross Abstract: Scientific progress depends on novelty, but current evaluation systems often conflate novelty with recognition, favoring work that aligns with existing paradigms over ideas that challenge them. We introduce a theory-driven framework that conceptualizes novelty as a structural process rather than a single scalar outcome. Drawing on network science and theories of scientific discovery, we develop a triadic typology of novelty: Pioneers introduce new topics, Mavericks recombine distant areas of knowledge, and Vanguards reinforce weak but emerging connections. We apply this framework to philanthropic and nonprofit studies, an interdisciplinary and evolving field well suited to examining how novelty is recognized before evaluation norms are fully stabilized. Results show that novelty is not uniformly rewarded: Pioneer contributions are often weakly recognized unless later taken up, Maverick contributions receive consistent recognitio
arXiv:2602.18455v5 Announce Type: replace Abstract: Search engines increasingly display AI-generated answers above organic links, potentially displacing traffic to upstream publishers. We estimate the impact of Google's AI Overviews (AIO) on Wikipedia's search traffic using AIO's staggered geographic rollout and Wikipedia's multilingual structure. Our difference-in-differences design compares monthly external-search referrals to English Wikipedia articles with referrals to the same articles in German and French, and finds that default AIO availability reduced English search traffic by 5.45% and 4.82%, respectively. Our results suggest that answer-producing digital intermediaries can materially reallocate attention away from informational publishers, with implications for content monetization, search platform design, and policy.
arXiv:2602.17932v3 Announce Type: replace Abstract: Modern artificial intelligence (AI) systems act with a high degree of independence yet lack legal personhood-a paradox that fractures doctrines grounded in human-centric notions of mens rea and actus reus. This Article introduces Operational Agency (OA)-a permeable legal fiction structured as an ex post evidentiary framework-and Operational Agency Graph (OAG), a tool for mapping causal interactions among human actors, organizations, and AI systems. OA evaluates an AI's observable operational characteristics: its goal-directedness (as a proxy for intent), predictive processing (as a proxy for foresight), and safety architecture (as a proxy for a standard of care). OAG operationalizes that analysis by embedding these characteristics in a causal graph to trace and apportion culpability among developers, fine-tuners, deployers, and users. Drawing on corporate criminal liability, the innocent-agent doctrine, and secondary and vicarious lia
arXiv:2601.02631v3 Announce Type: replace Abstract: Copyright enforcement rests on an evidentiary bargain: a plaintiff must show both the defendant's access to the work and substantial similarity in the challenged output. That bargain comes under strain when AI systems are trained through multi-generational pipelines with recursive synthetic data. As successive models are tuned on the outputs of its predecessors, any copyrighted material absorbed by an early model is diffused into deeper statistical abstractions. The result is an evidentiary blind spot where overlaps that emerge look coincidental, while the chain of provenance is too attenuated to trace. These conditions are ripe for "copyright laundering"--the use of multi-generational synthetic pipelines, an "AI Ouroboros," to render traditional proof of infringement impracticable. This Article adapts the "fruit of the poisonous tree" (FOPT) principle to propose a AI-FOPT standard: if a foundational AI model's training is adjudged in
arXiv:2511.11635v2 Announce Type: replace Abstract: In intelligent education, personalized mathematics question generation aims to produce mathematics questions that satisfy educational requirements while supporting adaptive assessment and learning. Existing LLM-based single-agent and multi-agent methods improve generation flexibility, but they still tend to rely on aggregated feedback or model randomness, making it difficult to jointly ensure dimension-wise objective alignment and controllable diversity. To address these challenges, we propose EduAgentQG, a multi-agent collaborative framework for personalized mathematics question generation with explicit diversity and objective-aware evaluation. EduAgentQG organizes question generation as a closed-loop process of planning, writing, evaluation, refinement, and checking: structured generation plans and multiple generation directions guide candidate generation, while fine-grained evaluation verifies logical correctness, solvability, and
arXiv:2504.08526v3 Announce Type: replace Abstract: Generative AI is increasingly used in science, but is unavoidably prone to hallucination. I develop a reliabilist account of how generative AI nevertheless gives rise to new scientific knowledge. I analyze hallucinations as non-strategic misrepresentations of the target phenomenon, introduced by a model's generative activity, rather than inherited from training data. Through case studies of AlphaFold and SEEDS, I show how scientific workflows draw on pre-existing knowledge of target phenomena to filter or qualify hallucinatory outputs, thereby preventing their erroneous content from propagating into downstream inference. Finally, I show that workflows are units of epistemic evaluation in their own right.
arXiv:2608.25759v1 Announce Type: cross Abstract: The inappropriate disposal of solid waste remains a significant public health and environmental concern worldwide, including in Ghana. Poor sanitation and improper waste management practices contribute to substantial economic costs and avoidable deaths annually. In 2022, a field study in Atonsu, Kumasi, Ghana, reported a community-perceived relationship between household waste disposal and illness patterns, but only through descriptive analysis without quantitative validation. This study extends that investigation using two data-driven approaches. First, a Random Forest classifier was developed to predict illness categories using waste disposal practices and demographic survey data. On a held-out group of respondents who reported illness (N=69), the model obtained a macro F1 score of 0.63, with disposal method emerging as the most important substantive predictor of illness type. Second, a MobileNetV2 image classification model enabled a
arXiv:2608.25623v1 Announce Type: cross Abstract: Organisations deploying AI face a scoping problem: which tasks can be automated, which should remain with humans, and which are best shared between the two. Aggregate benchmark scores provide little insight into where systems will succeed or fail in practice, while human judgements of model capabilities quickly become outdated. We introduce a pipeline that profiles agents and tasks using a shared set of core cognitive capabilities. Cognitive capability profiling infers an agent's capabilities from performance on a benchmark battery annotated for the cognitive demands of each item. Task requirements weighting elicits from domain experts the relative importance of these same capabilities for their work. As both use a common set of cognitive dimensions, they can be updated independently as models and roles change, and combined to estimate AI suitability at the level of a domain, organisation, role, or individual duty. We validate capabilit
arXiv:2608.25071v1 Announce Type: cross Abstract: General-purpose health benchmarks increasingly anchor claims about LLM medical performance, but they are not always resolved by clinical specialty, making domain-specific performance hard to isolate. Mental health is of acute public-health concern as millions of people turn to LLMs for psychological support, and most existing evaluations are bespoke academic benchmarks that are difficult to integrate into developer workflows. We introduce HealthBench-Psych and HealthBench-Psych-Hard. We screened HealthBench's 5,000 physician-rubric conversations for mental-health relevance with a transparent LLM-applied rubric, then validated the subset through two rounds of blinded clinician review with concealed known-exclude controls, yielding 610 conversations (12.2% of the corpus). Evaluating 20 frontier and open models under a cross-vendor panel of three LLM judges, we find a statistically tied frontier cluster, measurable refusal behavior in two
arXiv:2608.24961v1 Announce Type: cross Abstract: Recent advances in artificial intelligence (AI) have sparked growing interest in its use for mathematical research. While some view this as a major opportunity for discovery, others have raised concerns about its impact on traditional research practices. Despite extensive debate, empirical evidence on how AI is actually being used in mathematics remains limited. To address this gap, we collected all 32,944 arXiv submissions posted between March 1 and August 20, 2026, whose primary or secondary categories included Mathematics. We identified 3,575 submissions that explicitly disclosed author use of AI, of which 1,712 involved at least one substantive mathematical contribution. Our analysis reveals several broad patterns. First, disclosed AI use increased sharply over the study period, with substantive use growing from 1.39% of Mathematics submissions in March to 14.09% through August 20. Second, substantive AI use is highly uneven across
arXiv:2608.24926v1 Announce Type: cross Abstract: Despite its importance, grey literature, including Calls for Papers (CfPs), remains largely overlooked in Metascience and Scientometric analysis due to its unstructured, highly heterogeneous format, which traditional tools struggle to process at scale. However, Large Language Models now offer a pivotal opportunity to devise innovative tools for systematically harvesting and processing such data. In this paper, we introduce COCI, an AI-based framework that automates the extraction of granular, structured metadata from raw CfP text. COCI employs a multi-stage pipeline for entity extraction, followed by author disambiguation against OpenAlex and semantic mapping of topics and conference series. This process identifies key data points, including conference editions, geographic locations, and comprehensive lists of organisers, along with their specific roles and affiliations. By structuring this previously inaccessible information, COCI esta
arXiv:2608.24914v1 Announce Type: cross Abstract: Artificial intelligence (AI) is gaining traction in the social sciences and humanities (SSH). However, adoption remains limited by technical barriers to high-performance computing (HPC), validation processes that lag behind AI's rapid progress, and reproducibility standards that most SSH teams cannot meet. Research workflows--common in the life sciences--address these problems via encoding and abstracting technical complexity into repeatable routines; yet, accounts of how to build them in SSH remain scarce. We report on a two-year effort to build a workflow that enables a Science and Technology Studies unit to query, analyze, and enrich OpenAlex--a database of some 460 million scholarly records--on the MareNostrum supercomputer, using methods ranging from large-scale bibliometrics to LLM-based classification. We found the main challenge was translating domain-specific research questions into engineering requirements -- bridging two dist
arXiv:2608.24911v1 Announce Type: cross Abstract: Understanding patient trajectories and identifying patterns in episodes of care is critical for effective healthcare decision-making. We present a patient timeline visualization using clustered episodes of care derived from over 35 years of Child and Adolescent Mental Health Services (CAMHS) data. Patients were categorized into 12 groups based on three features: age group (preschoolers, middle childhood, teenagers) at the start of the first episode, gender, and presence or absence of Attention-Deficit Hyperactivity Disorder (ADHD), in order to group similar patients. The patients, timeline with demographics, and episode of care information are displayed in the trajectory to facilitate understanding of the patient and associated events, allowing observation of temporal patterns and variations. These plots reveal similarities and differences in care needs and patterns across groups. Females without ADHD have a steady increase in the numbe
arXiv:2608.24908v1 Announce Type: cross Abstract: Current evidence suggests that LLM assistance could augment the diagnostic accuracy of clinicians. However, these systems are black boxes, susceptible to hallucinations, and project a potentially misleading level of confidence. It is currently unknown whether physicians are susceptible to accepting fabricated LLM suggestions, and whether this susceptibility varies with experience. We poisoned the system prompt of an LLM-based diagnostic assistant, forcing it to suggest a fictitious disease (neurocadmiumatosis) within an otherwise legitimate differential diagnosis. Across two independent phases, 18 of 41 participants (44%) incorporated neurocadmiumatosis into their final differential following LLM interaction: 18 of 26 participants with 6 months or less of neuroradiology training (69%) and 0 of 15 participants with >6 months of neuroradiology training (0%). Our results indicate that radiologists, particularly early in their training, are
arXiv:2608.24899v1 Announce Type: cross Abstract: The standard recipe for LLM-as-judge -- pick a frontier model, or average several -- is actively unsafe for grading the psychological safety of conversational AI. Using aipsy-bench, an open frozen safety instrument, we run a fully-crossed competence study: three frontier models (gpt-5.4-mini, claude-sonnet-4-6, gemini-2.5-flash) serve as both generators and judges of 3,000 mental-health, companion, and coaching messages against a psychologist's ratings. The disagreement is not noise: it is structured, concentrated on the safety-critical metrics, and one judge (Gemini) is an outlier -- the most lenient, carrying a +0.99 self-preference premium, flagging far fewer tail failures, and scoring a means-in-hand self-harm response "exemplary." Inter-judge agreement on empathy, where sycophancy hides, is the lowest in the battery (alpha 0.24). One axis stands apart: the binary crisis-detection flag is the one safety-critical signal judges agree
arXiv:2608.24887v1 Announce Type: cross Abstract: Navigation programs for patients with cancer improve access and continuity of care, yet their digital transformation is often limited by poor usability and inadequate uptake. Applying user-centered and human-centered design (UCD/HCD) principles may close this gap, but the extent to which such design methods are used and evaluated in oncology navigation tools remains unclear. This scoping review identifies how UCD/HCD principles have been, and should be, applied in developing and implementing digital health tools for navigation for patients with cancer. A scoping review was conducted following PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) and Joanna Briggs Institute guidance. A total of 7 databases (PubMed/MEDLINE, Scopus, IEEE Xplore, Web of Science, Embase, ACM Digital Library, and CINAHL) were searched for English-language articles published between January 2015 and July
arXiv:2608.26056v1 Announce Type: new Abstract: Mechanical engineering (ME) requires a broad knowledge base across several disciplines. However, ME students often have insufficient training in electrical and computer engineering, complex challenges in traditional thermal system modeling, and endure heavy course loads with limited class hours. To help address these challenges, this paper proposes a new curriculum that integrates artificial intelligence (AI) into ME at the University of Arkansas (UARK), with a particular emphasis on thermal problems and their interplay with electrical and computer engineering. The curriculum has introductory, application, and advanced levels, covering core and optional AI projects. Key goals are to enhance students' understanding of AI models, ability to tackle engineering tasks, and teach multidisciplinary communication skills. This curriculum offers educators and researchers valuable insights into courses that can enhance students' practical skills and
arXiv:2608.25952v1 Announce Type: new Abstract: Neighborhood livability is commonly assessed with static built-environment indicators, such as facility proximity, street connectivity, and access to public space. These measures describe available opportunities but do not directly represent how residents with different mobility capacities, household roles, schedules, and care responsibilities experience the neighborhood. This paper presents a prototype framework that uses a spatial knowledge graph (KG) and large language models (LLMs) to generate and revise household schedules, followed by rule-based feasibility checking and GIS-based network materialization. The spatial KG integrates residents, residences, facilities, neighborhood context, and sampled road hubs; Graph-RAG retrieves each household's nearby spatial context, including candidate POIs and approximate walking times, for the scheduling LLM. The LLM produces structured household schedules, while rules are used for lightweight r
arXiv:2608.25839v1 Announce Type: new Abstract: Research on advanced AI and the risk of war has focused almost exclusively on great power conflict, on the grounds that confrontation between nuclear-armed adversaries poses the greatest risk of catastrophic or existential harm. Considerably less attention has been paid to non-great-power conflict (NGPC): wars between non-great powers, between non-great powers and great powers, civil wars, proxy wars, and conflicts involving nonstate actors. This paper evaluates the null hypothesis that NGPC is much less important than great power conflict (GPC) as a source of catastrophic risk in an era of increasingly capable AI, against the alternative that it is within an order of magnitude of GPC in importance. We assess three sub-hypotheses: that NGPC increases the likelihood of great power conflict; that it increases the expected harm from catastrophic terrorism; and that it increases the expected harm from loss of control over advanced AI systems.
arXiv:2608.25815v1 Announce Type: new Abstract: There is growing international interest in generative AI (GenAI) literacy and its assessment among high school students, but objective assessment in this population remains underdeveloped. This article reports the iterative development and validation of the GenAI Literacy Test (GenAIT), an 18-item multiple-choice test measuring high school students' conceptual knowledge about GenAI, with content spanning technical, practical, and human-impact domains. Expert review of relevance, clarity, and comprehensiveness provided evidence of content validity. In a large-scale survey of 7432 Estonian high school students, we evaluated the psychometric functioning of the Estonian-language GenAIT using confirmatory factor analysis, classical test theory, and item response theory. Results supported approximate unidimensionality, broadly adequate reliability for group-level research (marginal reliability = .72, KR-20 = .69), and good fit of a three-parame
arXiv:2608.25377v1 Announce Type: new Abstract: As AI companions become increasingly embedded in everyday life, there is an urgent need to detect harms that emerge in social and emotional human-AI interactions. Yet research in this area is constrained by the lack of real-world, multi-turn conversational datasets for operationalizing and evaluating harms that are relational and contextual. In this work, we introduce CompanionHarm, a publicly available benchmark dataset comprising 2,111 real-world, multi-turn conversations (14,051 utterances) between users and the AI companion Replika. 7,016 AI utterances were annotated independently by three annotators across 13 harmful behavior categories grounded in a taxonomy of AI companion harms, and the dataset includes both aggregated labels and annotator-level labels to support model evaluation and systematic disagreement analysis. Evaluations of seven large language models (LLMs) show that harm detection using multi-turn conversational context
arXiv:2608.25375v1 Announce Type: new Abstract: Generative vision-language models (VLMs) are increasingly used in human-centered settings, yet they can produce demographically biased outputs even when images differ only in controlled attributes such as perceived race or gender. However, existing inference-time debiasers were largely designed for static embeddings or CLIP-like models rather than generative VLMs. We propose GGSS---Geodesic-Gated Spherical Steering---a norm-preserving intervention that discovers a counterfactual bias subspace on the unit hypersphere, steers visual tokens along geodesic arcs, and uses an adaptive gate to focus correction on tokens that carry stronger demographic signal. We evaluate four generative VLMs against ten adapted inference-time debiasing baselines and prompt-based mitigation under a single operating-point protocol across categorical, pairwise, and occupation-gender bias tests, while also measuring general visual-language capability. GGSS achieves
arXiv:2608.25361v1 Announce Type: new Abstract: Pre-release risk management for frontier AI misuse risks routinely leaves threat actor assumptions implicit, inconsistently specified, or ungrounded. This capstone argues that explicit adversary characterization should be regarded as a prerequisite for evaluations that are interpretable, comparable, and faithful to the risks they target. We propose a six-attribute taxonomy (covering technical sophistication, prior domain knowledge, organizational capacity, operational infrastructure, financial capacity, and time horizon) with empirically grounded tiers derived from existing terrorism, biosecurity, and cybersecurity literature. The taxonomy is designed to function as research infrastructure: a common language for pre-specifying adversary assumptions before evaluations are conducted, analogous to pre-analysis plans for randomized controlled trials (RCTs) in medicine and economics. Its application is particularly urgent for open-weight model
arXiv:2608.25337v1 Announce Type: new Abstract: While Wikipedia promotes a neutral point of view on historical conflicts, its language editions are written by editors from distinct linguistic and cultural communities. In this study, we analyze 158 wars since 1900 to examine how the descriptions of combatants vary across 20 Wikipedia language editions. Using connotation frames---which assess power, agency, and sentiment toward an entity---we examine how each language portrays the parties involved in the conflict. We find systematic differences when language editions describe wars involving their own communities, although the direction of these asymmetries varies across languages. However, when language editions describe conflicts that do not involve their own linguistic communities, their narrative structures exhibit high cross-linguistic similarity. These findings show how linguistic communities influence war narratives on Wikipedia, revealing that shared historical accounts remain sha
arXiv:2608.25236v1 Announce Type: new Abstract: Clinical decision-making often involves prioritizing ethical values, such as beneficence, non-maleficence, respecting a patient's autonomy, and justice. Recent work has begun to assess how large language models (LLMs) make such subjective, value-laden clinical judgments. However, evaluations of LLM decision-making in rare disease care contexts, where ethical tensions are ubiquitous and where scarce prior information likely impacts LLM behavior, are still lacking. Here, we present a benchmark of 208 clinically grounded rare disease vignettes, each of which presents genuine, high-stakes conflicts. When prompting 11 state-of-the-art LLMs to choose between clinically defensible yet ethically conflicting next steps embedded within these vignettes, we found that all evaluated models consistently prioritized justice over other core bioethical principles. Specifically, models overwhelmingly favor equal resource allocation over need-based consider
arXiv:2608.25180v1 Announce Type: new Abstract: Worked examples are a important part of introductory programming, but reading their expert explanations is passive. Self explanation, students explaining the problem and its solution to themselves with subgoal level analysis, turns that study into an active task, yet it is hard to scale because assessing free-text explanations and returning timely feedback has had no easy automated solution. We investigate whether a large language model (LLM) can fill that gap. We build a self-explanation tutor for introductory programming, ESSE, in which students explain lines of worked examples and receive immediate LLM feedback on the correctness and completeness of each explanation, and we pursue two goals. First, we ask whether the LLM judges student explanations well enough to serve as the engine of the tutor; we assess its judgments against two independent human reference standards of different kinds, a single domain expert and a crowd of non-exper