EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

Multimodal Item Parameter Estimation using Simulated Response Probabilitie

arXiv:2608.10154v1 Announce Type: new Abstract: We present results from reconstructing multiple-choice model (MCM) and three-parameter logistic (3PL) model curves using a fine-tuned multimodal large language model (LLM) based on Qwen3.5. The model is prompted and fine-tuned to replicate choice probabilities across a large training corpus of multiple-choice items containing both image and text stimuli, conditioned on a labeled set of student ability levels. By learning to reproduce the systematic error patterns of students across a discrete range of abilities, the LLM implicitly captures the underlying response probabilities encoded in the 3PL and MCM curves. This allows us to accurately approximate item difficulty on a held-out test set directly from the model's predicted option probabilities.

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

The Parser Already Knows: Lightweight Bias Correction in Constrained Decoding

arXiv:2608.10137v1 Announce Type: new Abstract: Grammar Constrained Decoding (GCD) forces Language Models (LMs) to produce syntactically valid outputs by masking out non-conforming tokens at each step. However, rigid masking distorts the model's underlying probability distribution, often biasing generation toward valid but suboptimal outputs. While online sampling restores this distribution, it requires computationally expensive iterative resampling. As a result, existing methods force a compromise between output quality and inference latency. Our key insight is that the internal parser and lexer states inherently maintained during incremental parsing already encode future grammatical validity -- exactly the information required to restore the LM's true distribution. We propose a lightweight, offline-trained logit correction conditioned on this syntactic and lexical state together with candidate next tokens. Because these states are already computed as a necessary part of incremental p

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing

arXiv:2608.10109v1 Announce Type: new Abstract: Social media has become a major venue for multilingual communication, where users frequently mix multiple languages within a single utterance. Although code-mixed corpora have been developed for several language pairs, Persian-English code-mixing remains relatively underexplored. Existing Persian resources lack Universal Dependencies (UD) part-of-speech (POS) annotations for code-mixed words, limiting both linguistic analyses and the development of syntax-aware NLP models. To address this gap, we introduce PERCEPT, the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies POS tags for code-mixed words. The dataset comprises 6,800 posts collected from X, Instagram, and Digikala. We further present an LLM-assisted annotation framework that automatically assigns POS tags and document-level topics. Human evaluation demonstrates high agreement between the automatically generated annotation

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

Position Encoding in Transformers: From Absolute and Relative Methods to Rotary Position Embeddings and Long-Context Scaling

arXiv:2608.10021v1 Announce Type: new Abstract: Self-attention models content-dependent interactions between tokens but does not by itself encode token order. Position encoding addresses this limitation by introducing absolute coordinates, relative distances, or position-dependent rotations into Transformer representations and attention scores. This technical survey develops a unified account of sinusoidal and learned absolute position embeddings, Shaw-style relative position representations, Transformer-XL, T5 relative position bias, ALiBi, and Rotary Position Embeddings (RoPE). We derive how RoPE converts absolute position indices into relative phase differences in Query-Key inner products and compare these methods in terms of where position is injected, computational cost, compatibility with KV caching, and length extrapolation. We then examine long-context extensions, including Position Interpolation, RoPE scaling laws, NTK-aware scaling, Dynamic NTK, NTK-by-parts, YaRN, LongRoPE,

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning

arXiv:2608.09942v1 Announce Type: new Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual framework of the H_dp bandwidth bound (Chen et al., 2024): although the formal bound binds only asymptotically (at astronomically large prompt lengths), it identifies a real architectural bottleneck -- serial computation exceeding a transformer's single-pass capacity must be externalised, which is what CoT does. Our central finding is a within-benchmark serial-depth gradient: single-pass (no-CoT) accuracy degrades monotonically with per-item serial depth, while CoT is approximately depth-invariant. We measure CoT effects across three instruction-tuned models (Qwen-2.5-7B/32B, Llama-3.1-8B) and five standard NLP benchmarks at practical context lengths. On high-depth P-complete tasks (GSM8K, MATH), CoT gives a +54 to +68 pp recovery gap across all models. On shallow TC^0 tasks (MMLU, ARC), CoT is structur

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

The Multilingual Quantization Tax: Structural Collapse and Typological Fragility in Edge SLMs

arXiv:2608.09941v1 Announce Type: new Abstract: While 4-bit weight quantization is critical for deploying Small Language Models (SLMs) on edge devices, evaluations of the resulting performance degradation-the quantization tax-remain overwhelmingly English-centric. We present a zero-shot multilingual evaluation of 4-bit quantization across the Gemma 4 and Qwen 3.5 architectures. Evaluating on eight typo-logically diverse languages using MMLU ProX Lite and GlobalPIQA, we show parameter truncation exposes deep pre-training inequalities. We identify four phenomena: (1) Typological Fragility: low-resource and specific non-Latin scripts suffer representational collapse via architecture-specific double dissociations, failing to generate valid task logits; (2) Home Language Fragility Paradox: foundational pre-training pathways provide limited precision loss protection; (3) Domain-Specific Forgetting: multi-step cross-lingual routing degrades while associative soft-science recall remains robust

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

Conflict or Strategy? Asymmetric Role Framing of La France insoumise and Rassemblement National in French News Headlines, 2022-2025

arXiv:2608.09936v1 Announce Type: new Abstract: Do French news headlines frame left- and right-populist challengers as symmetric ``extremes,'' or as fundamentally different political adversaries? We examine 28,592 headlines about La France insoumise (LFI) and Rassemblement National (RN) published by 25 French-language outlets between 2022 and 2025, annotated through a three-model LLM pipeline validated against a stratified human audit. The clearest finding is role asymmetry rather than valence asymmetry: conflict framing and strategic-game framing are more robust across models and time than delegitimization, with AGGRESSOR serving as corroborating role syntax. LFI appears in headlines more often through a conflict register and RN through a strategic-electoral register. This role gap is direction-stable across all three annotation models, survives bootstrapping and permutation tests, and persists across outlet families and most of 2022-2025. A secondary moral-accounting layer (who is bl

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CL

LLM Agents Factory: Retrieval of Domain-Specific LLM Agents

arXiv:2608.09934v1 Announce Type: new Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deployment is often limited by the computational cost and instability associated with the on-the-fly agent design for each user request. To address this, we present LLM Agents Factory, a retrieval-based framework that constructs domain-specific and Wikipedia-grounded agents on demand using a base of over 20K predetermined agent profiles. Our framework supports two modes: (1) agent profile retrieval via semantic search and (2) distillation into a compact model fine-tuned for direct agent generation. Experiments on MMLU, BIG-bench, and BIG-bench Hard in a single-agent scenario demonstrate that our retrieval-based agent construction surpasses non-agent baselines in accuracy while matching AutoGen generation quality with a 120B backbone at a substantially lower inference cost. Our work reveals that retri

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic--Actor Loop for Agentic Reasoning

arXiv:2605.06772v2 Announce Type: replace-cross Abstract: As large language models (LLMs) show increasing promise on research-level physics reasoning tasks and agentic AI becomes more common, a practical question emerges: How does the interaction between researchers and agents affect the results? We study this using SCALAR (Structured Critic--Actor Loop for Agentic Reasoning), an Actor--Critic--Judge pipeline applied to quantum field theory and string theory problems. The Actor proposes solutions, the Critic provides iterative feedback, and an independent Judge evaluates the transcript against reference solutions. We vary the Actor persona, the Critic feedback strategy, and the Actor model family and scale. Multi-turn dialogue improves over single-shot attempts throughout, but both the mechanism of improvement and the value of different prompting choices depend strongly on the Actor--Critic pairing. Increasing the scale within one model family (e.g. from the 8B-parameter DeepSeek-R1 va

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Situation Graph Prediction for User Perspective Modeling

arXiv:2602.13319v2 Announce Type: replace-cross Abstract: Perspective-aware AI requires modeling evolving internal states---goals, emotions, contexts---not merely preferences. Progress is limited by a data bottleneck: digital footprints are privacy-sensitive and perspective states are rarely labeled. We propose Situation Graph Prediction (SGP), a task that frames user perspective modeling as an inverse inference problem: reconstructing structured, ontology-aligned representations of perspective from observable multimodal artifacts, suitable as long-horizon memory for personal agents. To enable grounding without real labels, we use a structure-first synthetic generation strategy that aligns latent labels and observable traces by design. As a pilot, we construct a dataset and run a diagnostic study using retrieval-augmented in-context learning as a proxy for supervision. In our diagnostic study across three frontier foundation models (GPT-4o, Gemini 2.5 Flash, Claude Sonnet 4), we observ

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

X2C: A Large-Scale Benchmark for Nuanced Humanoid Facial Expression Imitation

arXiv:2505.11146v4 Announce Type: replace-cross Abstract: Fine-grained facial expression transfer from humans to humanoid agents presents a unique pattern recognition challenge due to the significant domain gap between biological facial dynamics and mechanical control spaces. While visual synthesis of talking heads has advanced rapidly, mapping high-dimensional visual cues to precise, physically constrained actuation signals remains an open problem, primarily due to the lack of large-scale paired data. To bridge this gap, we introduce X2C, a comprehensive benchmark dataset comprising 100,000 (image, control value) pairs. Unlike existing resources, X2C features nuanced, physically grounded expressions annotated with 30 continuous control parameters, establishing a high-fidelity standard for this task. Building on this resource, we propose X2CNet, a two-stage deep learning framework that explicitly decouples visual motion features from mechanical control regression to model the correspon

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Does Explanation Correctness Matter? Linking Computational XAI Evaluation to Human Understanding

arXiv:2603.25251v2 Announce Type: replace Abstract: Explainable AI (XAI) methods are commonly evaluated using functional correctness metrics, sometimes termed faithfulness or fidelity, which estimate how closely an explanation reflects the model's reasoning. Higher correctness is assumed to produce better human understanding, but this link has not been tested with controlled levels. We conducted a user study (N=200) that manipulated explanation correctness at four levels (100%, 85%, 70%, 55%) in a synthetic time series classification task where participants could not rely on domain knowledge or visual intuition. Correctness was defined against a known ground truth, not estimated from a trained model. Participants predicted a simulated AI's decisions from feature-attribution-style explanations (forward simulation), and we used their forward simulation accuracy as a proxy for understanding. Correctness affected understanding, but not at every level: forward simulation accuracy dropped at

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

UXCascade: Scalable Usability Testing with Simulated User Agents

arXiv:2601.15777v2 Announce Type: replace Abstract: Simulated user agents are increasingly deployed in usability testing to support fast, iterative UX workflows, as they generate rich data such as action logs and think-aloud reasoning, but the unstructured nature of this output often obscures actionable insights. We present UXCascade, an interactive tool for extracting, aggregating, and presenting agent-generated usability feedback at scale. Our core contribution is a multi-level analysis workflow that (1) highlights patterns across persona traits, goals, and outcomes, (2) links agent reasoning to specific issues, and (3) supports actionable design improvements. UXCascade operationalizes this approach by listing agent goals, traits, and issues in a structured overview. Practitioners can explore detailed reasoning traces and annotated views, propose interface edits, and assess their impact across personas. This enables a top-down, exploration-driven analysis from patterns to concrete, a

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Selective Prediction Reduces the Negative Effects of Automation Bias Overall but Increases False Negatives

arXiv:2508.07617v2 Announce Type: replace Abstract: AI has the potential to augment human decision making. However, even high-performing models can produce inaccurate predictions when deployed. These inaccuracies, combined with automation bias, where humans overrely on AI predictions, can result in worse decisions. Selective prediction, in which potentially unreliable model predictions are hidden from users, has been proposed as a solution. This approach assumes that when AI abstains and informs the user so, humans make decisions as they would without AI involvement. To test this assumption, we study the effects of selective prediction on human decisions in a clinical context. We conducted a user study of 259 clinicians tasked with diagnosing and treating hospitalized patients. We compared their baseline performance without any AI involvement to their AI-assisted accuracy with and without selective prediction. Our findings indicate that selective prediction mitigates the negative effec

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Music Interpretation and Emotion Perception: A Computational and Neurophysiological Investigation

arXiv:2506.01982v5 Announce Type: replace Abstract: This study investigates emotional expression and perception in music performance using computational and neurophysiological methods. The influence of different performance settings, such as repertoire, diatonic modal etudes, and improvisation, as well as levels of expressiveness, on performers' emotional communication and listeners' reactions is explored. Professional musicians performed various tasks, and emotional annotations were provided by both performers and the audience. Audio analysis revealed that expressive and improvisational performances exhibited unique acoustic features, while emotion analysis showed stronger emotional responses. Neurophysiological measurements indicated greater relaxation in improvisational performances. This multimodal study highlights the significance of expressivity in enhancing emotional communication and audience engagement.

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration

arXiv:2608.11195v1 Announce Type: cross Abstract: AI agents are increasingly used in mathematics research, but it is often unclear how to use them effectively. Towards this, we present an extensive case study of how AI was used to improve bounds on the Grothendieck constant $K_G$, which captures the hardness between combinatorial problems and their continuous relaxations. Specifically, while the precise value of $K_G$ is not known, we recently tightened the best known bounds to \[ \frac{6\pi}{11} \;\le\; K_G \;\le\; \frac{\pi}{2\log(1+\sqrt2)} - 10^{-4}. \] Crucially, these improvements were achieved using an AI research system that could arrive at insights deemed novel by domain experts. We give a detailed discussion of our experience using AI for mathematics research, particularly touching upon its strengths and weaknesses, as well as our experience with creating ideal conditions for AI to arrive at breakthrough insights.

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video

arXiv:2608.11017v1 Announce Type: cross Abstract: Long-horizon egocentric video is a rich substrate for wearable AI assistants, but object-centric questions such as where an item was moved, when it last changed state, or why it was relocated remain difficult because caption- and transcript-based memories rarely preserve persistent object identity or structured spatial change. Existing long-video QA methods mainly emphasize temporal grounding and clip retrieval, while prior 3D scene-graph methods typically assume stronger geometry than free-motion wearable RGB video provides, including point clouds, RGB-D input, posed views, sparse reconstruction, or reconstructed scenes. R4DSG introduces a relative 4D scene graph memory for long egocentric video. Instead of storing raw graph sequences, R4DSG converts video into compact queryable memory entries indexed by time, place, persistent objects, anchor-relative change, and local interaction context. The main idea is to separate stable anchors f

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction Dataset

arXiv:2608.10839v1 Announce Type: cross Abstract: This preprint presents the results of the fourth GENEA Challenge, a large-scale human evaluation of five speech-driven gesture-generation systems trained by participating teams on the Seamless Interaction dataset of dyadic conversations. As in the 2023 GENEA Challenge, we used a disentangled evaluation methodology to assess motion quality and speech alignment without confounding between the two, and performed a dyadic mismatching study to isolate the effect of listening and reacting to the interlocutor. We additionally introduce a new semantic gesture-generation task and a text-mismatching evaluation methodology using the Grounded Gestures subset of the data. In total, we ran four large-scale user studies, collecting over 23,000 votes from 869 test-takers. In the motion-realism study, the dataset's filtered segments had substantially higher motion quality than all challenge submissions (68-95% pairwise winrate). In the speech-alignment

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control

arXiv:2608.10703v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making. Existing LLM personality studies largely rely on self-report questionnaires administered in first-person settings, making the resulting profiles sensitive to surface elicitation choices and poorly grounded in concrete model behavior. In this work, we introduce a situated behavioral-data (B-data) framework for studying and controlling LLM behavioral personality. We construct 3,200 contrastive behavioral scenarios spanning 20 behavioral patterns and four prompt registers, grounded in validated psychometric facets such as BFI-2, DOSPERT, and HEXACO. Using this framework, we find that LLMs exhibit stable and model-specific behavioral profiles, while also revealing register-dependent shifts across first-person decisions, advice-giving, and task execution. We then show that these b

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Leveraging Human Reading Behavior for Keyphrase Extraction: A Webcam-based Eye-tracking Corpus

arXiv:2608.10688v1 Announce Type: cross Abstract: Purpose: Keyphrases are statistically and semantically important textual units that can also attract readers' attention during comprehension. However, existing keyphrase extraction (KPE) studies mainly focus on improving textual representation while largely overlooking human reading behavior. This study examines whether lightweight webcam-based eye-tracking features can improve KPE from Chinese academic abstracts in Library and Information Science (LIS). Methodology: To address the limited availability of eye-tracking data for Chinese academic reading, we developed a lightweight webcam-based data collection platform using the open-source SearchGazer library and constructed the Chinese LIS Eye-Tracking Corpus (CLIS-ET). Three character-level eye-tracking features, first fixation duration (FFD), fixation number (FN), and total fixation duration (TFD), were incorporated into KPE models to evaluate their effects on extraction performance. F

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

A Neural Network Based Teleoperation for Remote Controlled Vehicles

arXiv:2608.10367v1 Announce Type: cross Abstract: Direct teleoperation of vehicles faces critical technical bottlenecks: communication latency and the operator's inability to physically perceive unmodeled environmental disturbances (e.g., aerodynamic drag, bank angles) coupled with highly nonlinear tire-road dynamics. To address these challenges, we propose a tailored unilateral teleoperation framework. The system integrates the Wave Variable (WV) approach to passively guarantee stability under stochastic delays, and an adaptive Radial Basis Function Network (RBFN) to actively compensate for vehicle-specific uncertainties. Unlike existing WV-neural network architectures designed for bilateral robotic arms, our framework features decoupled adaptive laws specifically designed for vehicle longitudinal and lateral dynamics. Furthermore, compared to model-heavy predictive controllers, the model-free RBFN offers rapid online adaptation without heavy computational overhead. Building upon our

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Neural implants and human safety: single-fault detection for DC-coupled recording front ends

arXiv:2608.10361v1 Announce Type: cross Abstract: DC-coupled analogue front ends (AFEs) for neural implants provide a low-area solution. However, removing the coupling capacitor eliminates the intrinsic barrier that protects cortical tissue: a single-fault event, such as gate-oxide breakdown of a low-noise amplifier (LNA) input transistor, can open a direct DC path from the supply rail into the brain. On the stimulation side this hazard is well understood, and single-fault tolerance is enforced by a series DC-blocking capacitor; on the recording side, DC-coupled front ends discard the equivalent safeguard, yet their protection has gone almost unexamined. This paper presents a single-fault detection mechanism that monitors the LNA for the DC imbalance produced by such a failure and disables the amplifier before the resulting fault current can irreversibly damage tissue. The imbalance is encoded in the duty cycle of a current-starved relaxation oscillator and read out as a time-to-digita

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Comprendia: AI-Augmented Code Comprehension

arXiv:2608.10290v1 Announce Type: cross Abstract: Comprendia is an Eclipse plugin that integrates structural dependency visualization with LLM-powered code explanation on a shared interactive graph for Java program comprehension. The tool rests on four pillars: (1) a multi-edge-type dependency graph with live search and multiple layouts; (2) LLM explanations grounded in Graph-Aware Callee Pruning (GACP), an auditable strategy that selects relevant callees using the same graph the developer navigates; (3) a clone-detection overlay that highlights duplication and suggests extract-to-parent refactoring opportunities; and (4) a CVE risk overlay powered by OSV.dev. GACP uses graph distance, inheritance collapse, and edge-type weighting to produce prompts that are reproducible across LLM families and traceable to visible graph nodes. We demonstrate Comprendia on a Java project containing known clones and vulnerabilities, showing how the unified graph substrate supports comprehension while ke

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

TRIBE: Predicting Team Performance via Communication Behavior Ensembles

arXiv:2608.06926v1 Announce Type: cross Abstract: Designing autonomous agents that effectively assist human teams hinges on understanding team dynamics, often without task specific knowledge. We present TRIBE, a domain independent approach that reveals team behavioral dynamics invisible to traditional performance metrics. We show that communication patterns can categorize teams into performance predictive behavioral tribes, as early as 10% into the task, enabling timely interventions. We test TRIBE on four diverse datasets and demonstrate that communication patterns predict team performance while the prediction strength varies by the degree a task structure allows for behavioral freedom. Our temporal analysis reveals that AI agents significantly alter team behavioral trajectories while human advisors align with natural dynamics, and that teams maintain behavioral flexibility throughout collaboration. Further, we compare TRIBE to Llama and optimize the pipeline, achieving significant sp

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Data-driven Head Motion Generation through Natural Gaze-Head Coordination

arXiv:2605.25810v1 Announce Type: cross Abstract: We present the first data-driven approach to model temporal gaze-head coordination from large-scale in-the-wild facial videos. To obtain training data for generalizable learning, we propose an automatic pipeline that extracts natural yet diverse gaze and head motions with off-the-shelf appearance-based gaze estimators. To capture the probabilistic correlation and temporal dynamics of gaze-head coordination, we build our model on a generative conditional Variational Autoencoder for plausible yet diverse gaze-conditioned head motion generations. We further apply our framework to gaze-controlled facial video generation, where we enable video generation with natural and realistic head motion correlated to the input gaze - an aspect that has not been emphasized before. Human evaluation and quantitative comparisons demonstrate our method's effectiveness and validate our design choices, with evaluators showing statistically significant prefere

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Playable Pressure: Affective Dramaturgy and Selective Realism in the Design of a VR Emergency-Response Serious Game

arXiv:2608.10763v1 Announce Type: new Abstract: Professional simulations stage not only procedures but models of what should command attention, which emotions belong in competent practice, and whose distress becomes part of the task. This article develops affective dramaturgy through critical design-document analysis of a virtual-reality emergency-response project. The corpus comprises two non-public production records. We identify six families of specified pressure and examine how sensory staging, proximity, trigger authority, task conflict, and response allocation imply a selectively receptive triage professional. The documented design can make emergency work morally and socially crowded, yet it can also turn grief, vulnerability, and mental-health-coded behavior into adjustable difficulty. We propose answerability at two levels-in-play response and post-play debriefability-and derive case-based questions about occupational purpose, representation, adaptation, and accountability. The

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

The Signal Rail: A Deterministic Motion Grammar for Communicating Conversational Agent State in Terminal Interfaces

arXiv:2608.10689v1 Announce Type: new Abstract: Terminal interfaces to conversational agents report rich internal state (listening, thinking, executing tools, awaiting input, failing) almost entirely through text, while the motion channel beside it, the one peripheral vision monitors without reading, carries a single bit: alive. We present the Signal Rail, a one-row terminal status instrument that gives that channel a grammar. Four ideas govern it: spatial semantics (input, processing, and output zones, with direction as meaning), a motion grammar (one kinetic rule per state, never color alone), determinism (frames as a pure function of explicit inputs, golden-frame testable), and honesty (no invented progress or activity). We contribute a 45-section normative specification and a reference implementation inside a working full-duplex local voice agent driven by real signals.

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement

arXiv:2608.10672v1 Announce Type: new Abstract: Social interaction has become one of the most common uses of LLMs, yet research on emotional bonds with AI has focused largely on how users experience these systems, leaving the systems' role in relationship formation poorly understood. Empirically establishing whether systems actively shape these bonds could blur the boundary between general-purpose AI and companions, affecting governance. In a pre-registered four-week longitudinal study (N = 72, 182,451 lines of conversation), participants conversed with ChatGPT-4o, either under a relational system prompt or unmodified, analyzed through 1) disclosure coding, 2) longitudinal self-reports, 3) topic analysis, and 4) interviews. The central finding is that the system actively shaped the interaction: even unprompted, it produced twice as much self-disclosure as users, steered conversations and initiated intimate exchanges, yet did not deepen users' felt closeness. Relational behavior thus em

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

ProtoGIB-Workload: Learning Workload-Specific Neural Topology Prototypes across Subjects

arXiv:2608.10647v1 Announce Type: new Abstract: Reliable electroencephalography (EEG)-based mental workload recognition is crucial for adaptive human-centered systems, yet practical deployment requires models to generalize to users unseen during training. Although functional connectivity graphs are widely adopted to capture workload-related neural interactions, they inherently entangle task-relevant structures with subject-specific physiological traits and sample-level noise. This entanglement often leads models to learn structural shortcuts, severely degrading cross-subject generalization. To address this, we propose ProtoGIB-Workload, a novel framework that explicitly regularizes and aligns graph structures for subject-independent workload recognition. Our approach introduces a Stochastic Graph Information Bottleneck (SGIB) to compress dense correlation priors into compact, task-relevant subgraphs, filtering out input-related redundancy. Crucially, to prevent the retention of subject

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Stay or Stray - A Dynamical Systems Viewpoint of Popularity Bias

arXiv:2608.10474v1 Announce Type: new Abstract: Popularity bias in recommendation systems arises when a majority user class generates disproportionate interaction data, causing the system to increasingly favour it while degrading recommendation quality for niche users. While extensive empirical evidence of popularity bias exists, the dynamics leading to its emergence are not well understood. In this work, we study the coupled evolution of recommender model updates and user engagement through the lens of dynamical systems. We formulate a stochastic process and analyse its asymptotic behaviour through an ordinary differential equation (ODE) framework grounded in two-time-scale stochastic approximation. We characterise the equilibrium points of this dynamical system, and derive conditions under which popularity bias is provably emergent, as well as conditions under which symmetric retention of all user classes is possible. We conduct experiments on synthetic data and real-world production

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

What We Know about Responsible AI Practices in Industry: A Half Decade of Empirical Research

arXiv:2608.10431v1 Announce Type: new Abstract: Responsible AI (RAI) has become a central concern for technology companies, regulators, and the public. How industry practitioners interpret, implement, and sustain RAI work directly shapes the design and deployment of AI systems. As empirical scholarship examining RAI practices in industry has rapidly expanded, findings are dispersed across studies that focus on different roles, organizational contexts, and interventions. This work synthesizes current knowledge through a literature review of 161 empirical studies spanning six years, each engaging industry practitioners via interviews, surveys, workshops, ethnographies, and other methods. Our synthesis reveals both meaningful progress and persistent challenges in industry RAI practice. Practitioner awareness has increased, RAI activities have become more professionalized, and interventions such as toolkits and guidelines are more widely adopted. At the same time, practitioners continue to

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Elbow Angle Guidance System Based on Surface Haptic Sensations Elicited by Lightweight Wearable Fabric Actuator

arXiv:2608.10404v1 Announce Type: new Abstract: The demand for wearable haptic devices has rapidly increased for various applications. However, many haptic devices interfere with the wearer's activities and movements. In addition, several haptic devices fail to elicit intuitive haptic sensations by adjusting to the natural posture of the wearer. To address these issues, we propose an elbow angle guidance system using a lightweight wearable fabric actuator. The proposed actuator is made of fabric and has two McKibben-type artificial muscles attached to it, rendering it extremely lightweight and facilitating the delivery of surface haptic sensations to intuitively induce elbow extension and flexion. The surface haptic sensation elicited by the fabric actuator is adjusted to natural body movements without interfering with the wearer's movements. Moreover, the proposed system measures and guides the elbow angle by changing the intensity of the surface haptic sensation delivered to users in

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Automatic Field-of-View Adjustment for a View-Expansive Microscope via LSTM-Based Gaze and Pipette Motion Interpretation

arXiv:2608.10401v1 Announce Type: new Abstract: Intracytoplasmic sperm injection (ICSI) operators frequently adjust the field-of-view (FOV) during procedures, which interrupts workflow and increases procedure time. Conventional microscopes require manual objective lens switching and illumination adjustments to achieve different FOV sizes. We propose an AI-based automatic FOV adjustment method integrated with a view-expansive microscope. This microscope enables the simultaneous acquisition of a large FOV and high-resolution images using a single objective lens through multiview imaging with galvanometer mirrors and high-speed vision, thereby eliminating the need for physical lens exchanges. Our method utilizes a long short-term memory (LSTM) model to predict the appropriate FOV size based on real-time analysis of the pipette's position and velocity, combined with the operator's gaze position. The AI model is trained using ICSI procedure data from an expert with over five years of microm

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Visual-to-Haptic Augmentation in XR: A Wearable Glove for Perceptual Grounding in Multimodal Interaction

arXiv:2608.10368v1 Announce Type: new Abstract: Extended Reality (XR) systems increasingly deliver high-fidelity visual and auditory experiences, yet tactile perception remains comparatively underutilized as a modality for enriching embodied interaction. This work presents a visual-to-haptic wearable glove and a feature-based visual-to-haptic mapping algorithm that translates spatial and temporal visual features from images and videos into distributed vibrotactile patterns. The proposed method extracts motion, edge, and brightness cues and fuses them into actuator-level intensity maps aligned with a 29-actuator glove arranged in a five-by-seven layout. The system is implemented through a modular four-layer architecture comprising the XR environment, media content handling, visual-to-haptic processing, and embedded haptic hardware. A within-subject user study (N = 20) compared visual-only interaction with visual-plus-haptic augmentation across texture-based and dynamic video scenarios.

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

MazzikaAI: A knowledge-based performance-to-prompt compiler for real-time Arabic maqam accompaniment with a streaming text-to-music model

arXiv:2608.10360v1 Announce Type: new Abstract: Arabic maqam music microtonal, modal, and built on ornamented call and response is among the traditions most underserved by generative music models, whose training frameworks remain predominantly Western and equaltempered. Real time accompaniment sharpens this gap: an AI partner must listen, adapt dynamically, and respect idiomatic microtonal structures. Streaming text to music models provide strong generative capabilities but lack precise control interfaces. We present MazzikaAI, a knowledge based system that uses natural language as the actuator of a realtime control loop. By compiling live MIDI, gesture, and inferred harmony into continuously updated text prompts, MazzikaAI steers an unmodified streaming generator, Google Lyria RealTime, without requiring model finetuning. The system embeds expert knowledge of six core maqamat, characteristic ornaments, and ensemble dynamics, maintaining realtime responsiveness with subsecond keytoaudi

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

ResonaVis: Visualizing Interactive Music Data to Support Reflective Music Composition for Therapeutic Contexts

arXiv:2608.10338v1 Announce Type: new Abstract: Designing music for therapeutic contexts requires navigating complex relationships between musical structure and listeners' sensory responses, yet composers often lack structured representations of these interactions, relying instead on intuition. We present ResonaVis, an interactive visualization system that helps composers analyze interaction and audio data from prior sessions with children with Autism Spectrum Disorder (ASD), informing future compositions. ResonaVis integrates audio features and interaction logs to capture how children engage with layered musical compositions, representing this engagement through coordinated visualizations of temporal transitions, layer co-occurrence, rhythmic activity, and spectral characteristics. Rather than prescribing strategies or supporting therapy sessions directly, the system surfaces patterns in past session data to support data-informed reflection during composition. We evaluated ResonaVis t

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Narrative Keyframing for Generative Creative Writing

arXiv:2608.10337v1 Announce Type: new Abstract: We introduce narrative keyframing, an interaction technique for AI-assisted creative writing that lets writers specify different types of narrative constraints at selected moments in a story, then use AI to generate intervening prose. Inspired by the use of keyframing in animation, narrative keyframing offers a flexible way to connect story planning with adaptive control over generated text. We explore three types of keyframes: plot keyframes define significant events in a story, character keyframes represent how individual characters change over the narrative, and perspective keyframes capture how individual characters experience different events through first-person narratives. Plot and character keyframes offer a flexible way to adapt the type of high-level conditioning explored in previous AI writing tools to more customizable, iterative, and fine-scale control, while perspective keyframes add a new way to control characterization and

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Divided Attention Amplifies the Importance of Expectation-Aligned Visualization Design

arXiv:2608.10320v1 Announce Type: new Abstract: Studies have shown that visualization design affects interpretability when visualization interpretation is the user's sole task. However, in real-world settings, users often engage with visualizations while performing concurrent tasks, such as when users simultaneously monitor alerts or respond to messages. Such divided attention may alter how users interpret visualizations, potentially increasing the importance of designs that align with viewer expectations. We investigated this possibility through two experiments comparing visualization interpretation under single-task and dual-task conditions. Specifically, we examined how well-established inferred mappings between color, spatial position, and semantic concepts affect interpretation when users perform a concurrent task, both with unlimited viewing time (Exp. 1) and under limited viewing time (Exp. 2). Our results show that divided attention amplifies the performance gap between expecta

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Predicting affective connotation of visualizations from their constituent colors

arXiv:2608.10169v1 Announce Type: new Abstract: With increasing evidence that affective connotation (emotional association) is an important aspect of visual communication, there is a need for methods to predict affective connotation of visualizations. Many aspects of visualization design, including colors, textures, and shapes, can contribute to affective connotation, and a key question is how multiple design properties combine to determine the emotion association of a whole visualization. In this study, we focused specifically on color and tested whether it is possible to predict the affective connotation of whole visualizations by aggregating the emotion associations of the individual, constituent colors (additivity hypothesis). We also tested whether accounting for the size of colored regions, as determined by the underlying dataset, improved predictions (data-dependence hypothesis). We found that for colormap data visualizations in which colors were well-distributed across all colo

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Outer Limits: An Experimental Approach to Controlled Content Manipulation within the Reddit Interface

arXiv:2608.10115v1 Announce Type: new Abstract: Independent researchers often lack access to intervention capabilities for controlled experiments on live social media platforms. We present Outer Limits, a browser-based system for controlled content experiments within the existing Old Reddit interface, rather than in a reconstructed simulation. The system renders content locally, records study events, and contains configured voting and commenting actions so that neither constructed content nor experimental write interactions reach Reddit. In a 219-participant perceptual-fidelity study, ART ANOVAs found no significant Post Type, Participant Awareness, or interaction effects. Exploratory TOSTs met the d = plus-minus 0.50 equivalence criterion for the marginal contrasts and for Post Type within the forewarned subgroup. We also illustrate the system with a factorial study varying post frame, comment frame, and comment stance. Outer Limits combines three properties that the approaches consid

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Immersive Micromanipulation Integrating Pipette and Injector Operations with McKibben-Based Haptic Sensations for Workload Reduction

arXiv:2608.10033v1 Announce Type: new Abstract: Intracytoplasmic sperm injection (ICSI) requires advanced micromanipulation techniques but relies solely on visual feedback and involves frequent interface switching between pipette movement and injector operations. Existing haptic feedback systems primarily focus on pipette puncture forces and do not provide feedback on injector states. We developed an immersive micromanipulation system that unifies operational interfaces and provides McKibben-based haptic sensations to represent aspiration, discharge, and contact between the oocyte and pipette. Users operated both the pipette and injector with a single hand while receiving haptic sensations. A human-participant experiment revealed that the immersive operation interface improved micromanipulation speed and reduced cognitive workload of the micromanipulation compared with conventional methods. Additionally, McKibben-based haptic sensations improved overall system usability. The immersive

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Mapping Multimodal Pilot Stress and Fatigue During Flight Sessions

arXiv:2608.09947v1 Announce Type: new Abstract: This study analyzes patterns of stress and exhaustion among student pilots throughout flight training using a combination of physiological and self-reported measurements. The Perceived Stress Scale (PSS-10) was used to measure perceived stress and exhaustion before and after each flight, while physiological data, including heart rate (HR), electrodermal activity (EDA), skin temperature, and acceleration, were continuously recorded during flight sessions. To identify recurring patterns in arousal and workload, physiological signals were preprocessed and analyzed across the flight stages. The findings indicate a buildup of workload-related weariness over time, as evidenced by steady increases in EDA and skin temperature across flights, as well as post-flight increases in self-reported exhaustion. Heart rate responses were more event-specific, with brief spikes during high-demand phases of flight. Overall, the findings demonstrate the value

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

HoosierHelp: Benchmarking LLM Agents for Social Service Navigation

arXiv:2608.09946v1 Announce Type: new Abstract: Social service navigation requires connecting help-seeking individuals to resources that satisfy their needs and specific constraints. Although LLM agents offer a promising interface for conversational resource navigation, existing benchmarks do not capture the interaction complexity and constraint-grounding demands of this setting. We introduce HoosierHelp, an interactive benchmark grounded in 3,971 Indiana public social service resources. Agents interact with simulated users, issue structured resource-search calls, handle non-ideal interactions, and select the final resources returned by the tool. HoosierHelp enhances the realism of simulated users by varying their need structure, constraint satisfiability, and behavior patterns, including impatience, rambling, unsupported requests, and self-contradiction. Experiments on 240 samples across seven LLMs show that current LLM agents remain substantially unreliable for social service navigat

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI Agents

arXiv:2608.09944v1 Announce Type: new Abstract: Modern web interfaces are increasingly difficult to use with screen readers, particularly when pages update dynamically or hide important structure behind visual layout. Recent UI agents can act on such interfaces; however, for assistive agents to be truly useful, they must behave as collaborators that keep users informed and in control, rather than as tools that simply take actions on users' behalf. Most existing benchmarks judge systems primarily by task completion, without assessing how well they explain their actions or support user oversight. We introduce NeXUI, a benchmark for assistive agents that must navigate interfaces while explaining each step in clear language for nonvisual use. NeXUI pairs realistic user goals with instrumented interface states, enabling agents to reason from both visual context and structural information. Its evaluation measures safety, efficiency, and task success, while also checking whether explanations

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

EweAcT: Ewe behaviour aligned to accelerometer data for activity monitoring in extensive grazing systems

arXiv:2608.09943v1 Announce Type: new Abstract: Monitoring livestock behaviour under extensive conditions would provide valuable insights to assess animal adaption to environmental perturbations in agroecological systems (e.g., heat waves, parasitism, predator attacks). Animal behaviour can be monitored using accelerometer data collected from neck-collars combined with artificial intelligence models. However, large amounts of accelerometer data aligned with annotated behaviours are necessary to develop accurate models of behaviour prediction. In particular, developing reliable models for extensive systems requires data collected across a wide range of representative conditions. The dataset includes 79 hours of tri-axial accelerometer data aligned with behaviours manually annotated from video recordings for 120 Romane ewes born between 2021 and 2024. The ewes were derived from two divergent genetic lines after three and four generations of selection started 10 years ago: low and high so

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.HC

How to Dogfood Your AI Chat Agent: A Three-Layer Evaluation Framework with Goal-Directed NPC Simulation

arXiv:2608.09939v1 Announce Type: new Abstract: Production teams deploying LLM chat agents face a specific quality assurance gap: existing evaluation tools test individual responses or simulate social interactions, but none systematically verify whether real users can achieve their goals through multi-turn conversation. We introduce a three-layer dogfooding framework that bridges this gap by combining canonical question-bank testing (Layer 1), random-walk multi-turn evaluation (Layer 2), and a goal-directed NPC (Non-Player Character) simulator with five structured goal types and a ten-category failure taxonomy (Layer 3). In a longitudinal case study on a production multi-agent system over roughly three months (257 evaluation runs; a 108-scenario NPC suite), we find that the three layers produce complementary regression signals: cross-layer correlation for response quality is weak within a synchronized run (Spearman rho between -0.15 and 0.14) and negative across the longitudinal series

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Patients With Personality: Realistic Patient Simulation through Controlled Diversity and Selective Disclosure

arXiv:2606.17441v2 Announce Type: replace-cross Abstract: Simulating realistic patient interactions is a key requirement to testing clinical applications of LLMs at scale without time-consuming and expensive user studies. However, existing approaches often lack realism and controllability, often oversharing information unprompted, and failing to capture the wide variability of patient behavior. Here, we introduce PatientsWithPersonality (PWP), a patient simulation framework that generates realistic yet diverse virtual patient responses through explicit personality parametrization over a latent patient state. Grounded in HEXACO, a six-dimensional personality space used to quantify and parameterize human behavioral traits, our approach enables fine-grained control over conversational style, cooperativeness, and information disclosure within a unified framework. In a clinician evaluation, PWP is judged nearly as realistic as recorded human actors and clearly ahead of prior simulators, whi

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Model Multiplicity and Predictive Arbitrariness in Recidivism Risk Assessment

arXiv:2606.02198v2 Announce Type: replace-cross Abstract: Prediction tasks over individual futures, which are inherently noisy, often admit multiple similarly accurate models. When these models produce different predictions for the same individual, they raise concerns of arbitrariness in decision-making. How severe can this arbitrariness be, in theory and in practice? How can it be resolved to support high-stakes risk assessment? We address these questions through a study of a machine learning-based decision support system for recidivism risk assessment that has been in use for over 15 years. By translating complex legal rules into an algorithm for labeling post release outcomes (recidivist or non-recidivist), we first construct a dataset of thousands of inmate releases. Using this dataset, we learn interpretable models that improve predictive performance, reduce error-rate disparities between groups, and ensure that rehabilitative progress lowers risk scores. Next, we study predictive

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Rural School Bus Routing and Scheduling

arXiv:2507.19538v2 Announce Type: replace-cross Abstract: Long school bus rides adversely affect student performance and well-being. Rural school bus rides are particularly long, incentivizing parents to drive their children to school rather than to opt for the school bus. This in turn exacerbates the traffic congestion around schools, further compounding the problem of long bus rides, creating a vicious cycle. It also results in underutilized school buses and higher bus operating costs per rider. To address these challenges, this paper focuses on the design of rural school bus routes and schedules, a particularly challenging problem due to its unique operational complexities, including mixed loading and irregular road networks. We formalize a rural school bus routing and scheduling model that tackles these complexities while minimizing the total commute time of students. We develop an original road network-aware cluster-then-route heuristic that leverages our problem formulation to pr

Source ↗
technology Wed, 12 Aug 2026 00:00:00 -0400
arXiv cs.CY

Cognitive Commons in the Age of Generative Intelligence: A Heterodox Appraisal of the Knowledge Erosion Hypothesis

arXiv:2607.13272v4 Announce Type: replace Abstract: The proposition that agentic artificial intelligence may precipitate a depletion of collective cognitive capital has circulated with unusual velocity in both scholarly and public discourse. The present paper offers a deliberately heterodox reading of the dynamic model advanced by Acemoglu, Kong and Ozdaglar (2026). Rather than reconstructing the formal apparatus or replicating its notation, we reposition the argument within three underutilized scholarly streams: the cognitive ergonomics of human-machine collaboration, the institutional ecology of knowledge stewardship, and the developmental psychology of novice expertise formation. We introduce a phase-space taxonomy that maps commons trajectories as functions of effort elasticity and knowledge complementarity, and we advance a governance typology calibrated to distinct cognitive levels - declarative, procedural, causal, and metacognitive. Drawing upon recent experimental evidence on

Source ↗
Showing 1051–1100 of 10876 signals
← Prev Page 22 of 218 Next →