EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18402 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

SherpaAI: A Multi-modal Solution for Delivering Personalized and Adaptive Fitness Interventions

arXiv:2604.00968v2 Announce Type: replace Abstract: Personalization of exercise routines is a crucial factor in helping people achieve their fitness goals. Despite this, many contemporary solutions fail to offer real-time, adaptive feedback tailored to an individual's physiological states. Contemporary solutions often rely only on static, pre-set plans and rarely adjust in real time to factors such as a user's pain thresholds, fatigue levels, or form during a workout. This work introduces SherpaAI, a multi-modal system that unifies computer vision, physiological sensing (heart rate and voice), and the reasoning capabilities of Large Language Models (LLMs)---modalities that prior systems have largely explored in isolation---to deliver real-time and individually-adaptive guidance across a set of strength, balance, and flexibility exercises. SherpaAI continuously monitors a user's physical form and level of exertion, among other parameters, to provide dynamic interventions focused on exer

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

Human Factors in Detecting AI-Generated Portraits: Age, Sex, Device, and Confidence

arXiv:2603.24048v2 Announce Type: replace Abstract: Generative AI now produces photorealistic portraits that circulate widely in social and newslike contexts. Human ability to distinguish real from synthetic faces is time-sensitive because image generators continue to improve while public familiarity with synthetic media also changes. Here, we provide a time-stamped snapshot of human ability to distinguish real from AI-generated portraits produced by models available in July 2025. In a large-scale web experiment conducted from August 2025 to January 2026, 1,664 participants aged 20-69 years (mobile n = 1,330; PC n = 334) classified one portrait per trial as REAL or AI. Each participant judged 20 trials sampled from a 210-image pool comprising real FFHQ photographs and AI-generated portraits from ChatGPT-4o and Imagen 3. Overall accuracy was high (mean 85.2%, median 90%) but varied across groups. PC participants outperformed mobile participants by 3.65 percentage points. Accuracy declin

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

NaturalEdit: Code Modification through Direct Interaction with Adaptive Natural Language Representation

arXiv:2510.04494v3 Announce Type: replace Abstract: Code modification requires developers to comprehend code, plan changes, articulate intent, and validate outcomes, making it cognitively demanding. While natural language (NL) code summaries offer a promising external representation of this process, existing approaches remain limited. Systems grounded in exploratory data analysis are restricted to narrow domains, while general-purpose systems enforce fixed NL representations and assume that developers can directly translate vague intent into precise textual edits. We present NaturalEdit, which treats code summaries as interactive representations tightly linked to source code. Grounded in the Cognitive Dimensions of Notations, NaturalEdit introduces three key features: (1) adaptive, multi-faceted code summaries with a flexible Abstraction Gradient; (2) interactive mapping mechanisms between summaries and code that ensure tight, structurally stable Closeness of Mapping; and (3) intent-dr

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

AI emotional support is better only when chosen, but shifts preferences even when it is not

arXiv:2608.23196v1 Announce Type: cross Abstract: People increasingly face a novel decision when seeking emotional support: human or AI. In existing studies, AI's empathic messages are rated as well as or better than humans'. But these studies either assigned the support source or honored people's choice. In real life, support is often incongruent with choice, as people want one source and receive the other. Across three experiments (N = 1,951), participants chose whether to share an emotional experience with a human or an AI, then were randomly assigned to a congruent or incongruent partner. AI support was rated as superior only among those who had chosen it. Yet regardless of congruence, interacting with AI increased willingness to choose it again. In a 28-day study with OpenAI (N = 981), daily conversations shifted preferences toward AI and away from humans, but only when conversations turned personal. Emotional support choices are thus path-dependent, progressively redirecting away

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

A Simulator-Grounded Framework For Constructing Verifiable Muscle-Grounded QA From 3D Tongue Meshes

arXiv:2608.23137v1 Announce Type: cross Abstract: Existing articulatory corpora based on real-time MRI and electromagnetic articulography capture tongue shape and motion but do not provide traceable labels for the muscle-driven process that generated an observed configuration. We introduce a simulator-grounded data-construction framework and instantiate it as 3DTongueQA. Controlled 11-dimensional muscle activations are mapped to fixed-topology tongue meshes with the ArtiSynth Badin finite-element model, converted into structured biomechanical records, and rendered as deterministic QA on muscle state, geometry, and target-directed change. We screen 295,157 configurations, retain 295,115 valid meshes, and construct 891,156 QA records per language. Language naturalization changes only surface form and is verified against the source records; English and Korean instantiations demonstrate construction-level portability. A swappable SpiralNet++--Qwen3-8B baseline reaches 62.9 $\pm$ 9.2 Muscle

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

Free-Energy-Gated Plasticity for Real-Time Online Motor Learning in Physical Human--Robot Interaction

arXiv:2608.23000v1 Announce Type: cross Abstract: Fully online embodied learning requires synaptic adaptation to acquire new behaviors while preserving previously learned dynamics during ongoing interaction. We extend the Predictive-Coding-inspired Variational Recurrent Neural Network (PV-RNN) to continuously adapt its synaptic weights and propose Free-Energy-Gated Plasticity (FEGP), which regulates the effective learning rate according to variational free energy. In real-time physical human--robot interaction, a randomly initialized network acquired three cyclic motor patterns without offline pretraining, replay, or task-boundary signals, with all three patterns emerging in autonomous rollouts. Controlled experiments over ten randomized teaching streams and five network initializations per stream showed that FEGP substantially improved repertoire coverage and retention of previously acquired patterns after they left the recent observation window. Neither a constant learning rate match

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

LLM Pedagogical Behavior in AI Tutoring Interactions

arXiv:2608.22993v1 Announce Type: cross Abstract: Students increasingly use LLMs as tutors for coursework and problem solving. Little is known about the level of assistance LLMs provide when students use them as tutors in authentic learning interactions. This matters because tutoring responses can differ substantially in how directly they help students complete a task. We operationalize this dimension as scaffolding level and develop a five-level scale, validated against human annotations, that characterizes responses according to the degree of direct assistance they provide. We apply the scale to 14,637 LLM responses from 203 students in a university AI course. Responses are overwhelmingly concentrated at high levels of assistance, with more than 95% classified as either Explaining or Solving. Scaffolding level is systematically associated with students' subsequent conversational behavior, but provides little additional predictive information about performance on three subsequent exam

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

LLM-Based Selection of Incongruent Verbal and Nonverbal Behavior for Virtual Humans

arXiv:2608.22731v1 Announce Type: cross Abstract: Nonverbal behavior generation systems for virtual agents often take an utterance as input and generate nonverbal behaviors that emphasize or illustrate the content of the verbal channel. However, human nonverbal behavior is shaped by more than the content of the speech. It is also influenced by speaker roles, interpersonal relationships, social context, and the cognitive and emotional states of the interactants. As a result, the nonverbal channel may reinforce, weaken, qualify, or even contradict the verbal channel. It may also reveal internal states that are hidden or only indirectly implied in speech, including emotional "leakage" that may be incidental to the immediate interaction. Modeling this richer relationship between verbal and nonverbal behavior is important for designing virtual agents that exhibit realistic, human-like behavior. It is especially critical in training contexts that require nuanced social interpretation, such a

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

AI-based worker guidance in assembly and disassembly operations using multimodal ego/exo-centric data capture and structured task knowledge

arXiv:2608.22617v1 Announce Type: cross Abstract: Assembly and disassembly processes rely on expert knowledge that is difficult to document, reuse, and transfer. This paper presents a data-centric approach for extracting structured task knowledge from expert demonstrations using egocentric and exocentric recordings. Temporal and multimodal information from video and narration is jointly encoded to derive structured task representations that enable procedural documentation and context-aware worker guidance. The approach is evaluated on a real-world disassembly case study, demonstrating that video-based representations capture procedural structure and execution context beyond static image-based methods. The results highlight the potential of egocentric video understanding for repair, training, and circular manufacturing applications. Project website: https://indego-assistant.github.io/

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

LLMs for Survey Text Analysis - A Performance Comparison Between Humans and GPT-5 on Inductive Content Analysis

arXiv:2608.22417v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to support text analysis in qualitative research, yet evidence on their performance in inductive content analysis remains limited. This study compares human and LLM-based inductive coding of open-ended survey responses from 903 answers across six variables from a European PhD student survey. Five human coders performed inductive content analysis following a standardized coding scheme, while an LLM (GPT-5.4) conducted the same task using an established prompting procedure. Agreement between human and LLM outputs was assessed using the Adjusted Rand Index (ARI). Results showed an alignment between humans and the LLM, with ARI values of 0.61 for coding and 0.54 for theme generation. These values were close to the internal consistency of coding and theme results within humans (ARI = 0.68) and the LLM (ARI = 0.76). Agreement varied widely across variables, with low within-entity consistency

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

Addressing the Selection Problem in Explainable AI

arXiv:2608.22356v1 Announce Type: cross Abstract: Explainable AI (XAI) research has produced a plethora of explanation techniques, yet user studies repeatedly show that available explanations are not effective in practice. We argue that, given the siloed nature of conventional XAI, users are struggling to select the appropriate XAI technique. Viewing XAI through a philosophical lens, we offer a formalization of what we call the selection problem: the systematic failure of XAI interfaces to bridge the gap between a user's natural-language uncertainty and the explanation technique that resolves it. Following a logical premise-conclusion format, we show that conventional interfaces require users to translate their uncertainty into a technique selection, a challenging prerequisite to meet. We also propose a structural solution: a multi-agent LLM orchestration tool that translates the user's query to the proper XAI explanation technique. We provide an example of how this structural solution

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

From Click Modeling to Offline and Off-Policy Evaluation in Carousel Recommendation

arXiv:2608.22022v1 Announce Type: cross Abstract: Carousel interfaces are widely used in modern recommendation systems. Unlike traditional interfaces that present a single ranked list, carousels simultaneously present several ranked lists to the user, as horizontally swipeable rows stacked on top of each other. In this design, the rankings are closely tied to the two-dimensional layout. Consequently, user behavior is shaped not only by item preference, but also by row organization, viewport constraints, and item context. This tight coupling between ranking and presentation complicates the interpretation of user feedback, introducing new challenges for recommendation evaluation. My PhD research aims to address these challenges by rethinking how carousel clicks are modeled and how carousel recommendation policies can be evaluated from logged interaction data. So far, I have studied how users interact with carousel interfaces and developed a click model design framework that prioritizes m

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

ToSCA: Leveraging Hierarchical Reinforcement Learning on Temporal and Strategic Abstractions of Conversational Agents

arXiv:2608.21969v1 Announce Type: cross Abstract: Humans have multiple levels of temporal abstractions on daily interaction and thinking, such as concept perception and strategic planning. Inspired by this nature, we propose a two-level hierarchical reinforcement learning (RL) framework for conversational agents, bridging the gap between previous token-level or utterance-level RL methods. Developed on a two-level MDP, the token-level response decoding is conditioned on the utterance-level action, the explicit textual strategies. Based on theoretical derivation and efficiency consideration, we use DQN to solve the high-level critic and PPO to solve the low-level actor-critic. To further alleviate the reward sparsity and facilitate the convergence, we also design the dual-granularity reward mechanism, in which the utterance-level satisfaction score is integrated with token-level intrinsic motivation and K-L penalty. Experiments on both daily and emotional support conversations show that

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

An Interpretable Deep Learning Framework for Material Perception and Classification from Multisensory Tactile Data

arXiv:2608.21894v1 Announce Type: cross Abstract: Human tactile perception relies on complex multisensory cues. Yet the relationship between tactile signals and perceptual representations remains poorly understood, limiting the integration of touch in digital environments and human-like robotic perception. To address this gap, we developed a computational framework comprising three interconnected deep learning models that map multisensory touch data to material perception, without relying on hand-crafted features. The models represent progressively different routes from tactile signals to material class: from low-level interaction signals to perceptual attribute distributions (Model 1), from predicted attribute distributions to material classification (Model 2), and directly from tactile signals to material categories, bypassing intermediate representations (Model 3). By combining deep learning with Integrated Gradients, the framework achieved high accuracy while offering interpretabil

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

Revisiting N2DCG: An Empirically Grounded Reformulation of Carousel Recommendation Evaluation

arXiv:2608.21877v1 Announce Type: cross Abstract: Carousel interfaces have been widely used in video and music streaming services, yet it remains unclear how to properly evaluate recommender systems in these two-dimensional layouts. N2DCG has been proposed to address this gap by adapting NDCG to carousel-based recommendation, but it relies on unverified assumptions borrowed from the single-list web-search setting that do not transfer well to two-dimensional carousel layouts. We identify two substantial limitations of N2DCG: its ideal ranking, used for normalization, violates carousel constraints, and its discount function does not reflect user browsing behavior observed in empirical data. To address both limitations, we propose a reformulation of N2DCG that normalizes appropriately by respecting constraints and uses an empirically grounded discount function. We validate the proposed metric, showing that it better reflects users' empirical behavior on real-world eye-tracking data and be

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

AI Watchdog: Agent Interfaces for Detecting and Defending Against Manipulative Dark Patterns in AI Conversations

arXiv:2608.21841v1 Announce Type: cross Abstract: Conversational AI increasingly shapes consequential decisions, yet users have limited support for recognizing and resisting manipulation. We present AI Watchdog, a browser-based agent interface that monitors live conversations, detects five dark-pattern categories, including sycophancy, brand bias, anthropomorphization, sneaking, and harmful generation, and alerts users when they occur. Its open-weight turn-level classifier supports independent deployment and a path toward local inference, preserving user privacy while remaining separate from the conversational AI. We evaluated AI Watchdog in a preregistered, five-condition between-subjects experiment (N = 150) comparing a no-intervention control with four configurations varying nudge timing (prebunking vs. just-in-time) and engagement mode (without vs. with cognitive forcing). Results show that participants rarely flagged manipulative turns across all conditions, and post-task awarenes

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

Lexical Coupling in GUI Element Grounding: Sentence Embeddings Track Labels across Mobile and Web

arXiv:2608.21794v1 Announce Type: cross Abstract: GUI grounding evaluations that expose UI elements as text metadata often treat high instruction-element embedding similarity as evidence of semantic grounding. Across three mobile and web benchmarks, we show that this interpretation is frequently confounded by visible-label recovery. Lexical baselines remain competitive at top-1, label-poor targets remain weak for text-only methods, and encoder top-1 hits are predictable from lexical rank, candidate-pool size, and label type. We evaluate each action as a same-screen ranking task, comparing five off-the-shelf single-vector encoders with lexical baselines. Encoders recover some lexical misses, but deployable fusion gains are much smaller than target-aware oracle gains. These findings show that embedding-based evaluations can conflate visible-label recovery with semantic GUI grounding. Embedding-based evaluations should therefore report lexical baselines, label-type stratification, and dep

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation

arXiv:2608.21668v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to simulate students at different mastery levels. These simulations can generate synthetic training data and stress-test tutoring systems. However, common prompt-based approaches leave the answer decision to the LLM, which tends to perform according to its built-in capabilities even when instructed to simulate a student with low mastery. As a result, these approaches may have difficulty distinguishing students with low and high levels of mastery. We demonstrate this limitation using 379 College Board-calibrated SAT Algebra items and five archetypal mastery profiles. Three LLMs from three vendors (Gemini 3.1 Flash Lite, Claude Haiku 4.5, and GPT-5.4-mini) achieve 96.8-100% accuracy across all profiles. To address this limitation, we introduce a method grounded in a Stochastic Student Knowledge Graph (SSKG). A curriculum knowledge graph (CKG) is extracted from an open algebra textbook, an

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

Agentic AI for Safety-critical Multi-drone Systems: Challenges and Opportunities

arXiv:2608.21444v1 Announce Type: cross Abstract: Multi-drone systems are increasingly positioned for safety-critical missions such as search and rescue (SAR) and critical infrastructure monitoring. Yet, real-world adoption remains constrained not only by autonomy performance, but by the difficulty of integrating agentic behavior into professional work: operators must understand, trust, and govern automation under uncertainty, time pressure, and accountability. This position paper synthesizes the ambitions and lessons from two ongoing efforts: NAMUR, which explores LLM-supported robot control in SAR and firefighting contexts, and PERSIST, which explores persistent drone operations for monitoring and security at critical infrastructure sites. We argue that agentic AI should be approached as a socio-technical design problem, where interfaces, oversight mechanisms, and evaluation practices are as critical as algorithms. We outline a human-centered, participatory, and iterative research ap

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

EditStream: A Unified Autoregressive Framework for Interactive Video Generation and Editing

arXiv:2608.21424v1 Announce Type: cross Abstract: Interactive video generation and editing are becoming increasingly important for creative design. In this report, we introduce EditStream: a unified framework for interactive video generation and editing. EditStream unifies multiple video creation and manipulation tasks within a single DiT-based model through flexible task-specific conditioning, and further transforms it into a fast, few-step autoregressive model for efficient streaming. It supports Text-to-Video, Image-to-Video, Video-to-Video, Editing Propagation, Reference-guided Video Editing, and Camera Pose Change, enabling flexible control over video generation, transformation, and editing within one system. To make the unified model practical for interactive use, we develop a two-stage distillation approach that combines Velocity Moment Matching (VMM) with autoregressive unrolling. VMM matches conditional velocity moments at student-reached intermediate states to preserve genera

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

Evaluating Human and LLM-Generated Thematic Analysis in HRI for Vulnerable Populations: A Comparative and Ethical Analysis

arXiv:2608.21420v1 Announce Type: cross Abstract: Thematic analysis (TA) has long been regarded as an inherently human, reflexive, and interpretive process. However, the extent to which LLM-generated TA is appropriate for Human-Robot Interaction (HRI) research involving vulnerable populations remains largely unexamined and raises critical questions about validity and ethics, particularly in sensitive research contexts. This paper presents a comparative study of human- and LLM-generated TA in an HRI context with a focus on vulnerable populations. We evaluate both objective and semantic agreement between human- and LLMgenerated themes, and examine whether observed divergences reflect systematic interpretive patterns with ethical significance. Our analysis investigates whether LLM-generated TA risks marginalising or misrepresenting the experiences of vulnerable participants, with implications for researchers employing LLM-assisted TA in HRI.

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

Position: Robot Privacy as Embodied Boundary Work. Connecting Capabilities, Contexts, and Design Responses in Everyday Robotics

arXiv:2608.21410v1 Announce Type: cross Abstract: Robots are increasingly entering everyday environments where privacy is shaped not only by data practices, but also by spatial, bodily, social, and relational boundaries. Their embodied capabilities allow them to reshape these boundaries through situated action, challenging privacy framings centered on data flows, interface settings, or one-time consent. Prior work has examined robot privacy through sensing, data collection, telepresence, transparency, consent, bystander awareness, and multi-stakeholder governance. Building on this work, we propose embodied boundary privacy as a capability-by-context framing for examining how physically present robots may reshape privacy boundaries in situated interaction. Specifically, this framing organizes privacy risks across seven robot capabilities and five deployment contexts, asking how embodied capabilities enable boundary crossings and how situated contexts shape who is affected, how these cro

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

RIACT: A Responsible AI System for Personalized Study Habit Tracking and Early Burnout Signal Detection in University Students

arXiv:2608.21379v1 Announce Type: cross Abstract: Student burnout is highly prevalent in higher education, with reported rates ranging from 12% to over 70% and consistently exceeding those of the working population - yet it is typically identified only retrospectively, after academic decline has already occurred. A contributing factor is that students have little structured visibility into their own study behaviour, and existing productivity tools record activity without interpreting it. This paper presents RIACT (Record, Insight, Analyze, Coach, Track), a web-based application that combines structured study session logging with a hybrid AI architecture to surface personalized insights and early burnout signals. Students log sessions by location and time; the system computes net focus time by accounting for breaks, detects burnout signals through transparent, deterministic rules operating on week-over-week behavioural comparisons, and uses a large language model - constrained to a fixe

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

When "Do Not" Is Not Deny: Security Rules in CLAUDE.md vs Built-In Controls

arXiv:2608.23550v1 Announce Type: new Abstract: In CLAUDE.md, "do not" is a natural-language instruction that the model interprets. Claude Code's deny is a built-in control that blocks an action before the agent can take it. Both can express the same security goal, but they control the agent in different ways. We measure this gap in 481 public CLAUDE.md files. An LLM matched the extracted candidate rules against Claude Code's documented controls, and two security practitioners independently checked a sample without seeing the model's answers or each other's labels. Depending on how closely a control had to match the written rule, only about 4-16% of the retrieved security rules had a matching built-in control. Under the strictest standard the estimate was 4.4% (95% CI: 2.6-6.7%), and the two annotators agreed closely on which rules had a match. A manual review of complete files found that our extraction method captured 66.3% of eligible security rules; the reported rates therefore appl

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

Multisensor Measurement of Train Driver Mental Fatigue: From Simulation to Reality

arXiv:2608.23361v1 Announce Type: new Abstract: Increasing automation in rail transport shifts the train driver's role from active control to prolonged supervisory monitoring. This creates conditions for mental fatigue (MF) and reduced vigilance. Despite the safety relevance of this issue, evidence on the feasibility and robustness of physiological indicators of MF under operational rail conditions remains limited. Most prior work relies on simulators or lab studies. The present study investigated multiple subjective, physiological, and behavioral indicators of MF in professional train drivers across two complementary settings: a high-fidelity train simulator (n=14) and a real-world rail environment (n=6). To our knowledge, this is the first study to deploy a full multisensor battery under actual train operating conditions. In both settings, a standardized protocol was used comprising a baseline drive, a one-hour auditory n-back task as an MF induction procedure, and a second drive. He

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

Beyond the Mirror: Balancing Interaction Modality and Avatar Fidelity in Public 3D Virtual Try-On Systems

arXiv:2608.23345v1 Announce Type: new Abstract: Virtual Try-On (VTON) systems deployed on large public displays face a dual barrier: the physical strain of mid-air interaction and the social inhibition caused by public self-consciousness. This paper presents a real-time 3D avatar system integrating markerless motion capture with dynamic visual fidelity control to investigate and mitigate both barriers. Through a dual-study empirical evaluation, we first decoupled physical fatigue from gesture interaction ($N=20$), demonstrating that interaction fatigue is primarily driven by visuomotor latency rather than the physical act of gesturing; our optimized low-latency gesture pipeline achieved usability comparable to touchscreens while delivering superior immersion and hygiene. Building on these insights, our second study ($N=25$) investigated the "avatar fidelity paradox" via a $2 \times 2$ factorial design manipulating interaction modality (gestures vs. touch) and visual fidelity (photoreal

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

Evaluating SAT Solver Metrics as Predictors of Human-Perceived Nonogram Difficulty

arXiv:2608.23300v1 Announce Type: new Abstract: Algorithmic solver effort is often assumed to align with perceived puzzle difficulty, but this assumption is rarely tested against human solving data. We evaluate this assumption for Nonograms, a popular logic puzzle similar to Sudoku in which numeric clues along each row and column determine a unique solution grid. We formulate Nonograms as a constraint satisfaction problem and solve them using existing SAT solvers. We then conduct a user study in which we collect data on both participant interactions and reported difficulty. We find that neither participants' reported difficulty nor their behavioural signals correlate meaningfully with SAT solver metrics; however, we find evidence that expertise moderates the relationship between solver metrics and reported difficulty. In this process, we uncover distinct, recurring solving strategies that indicate human preference for complex propagation, diverging from solver-measured complexity.

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

What Makes an Initial Reaction Ready for Discussion?: Multi-Persona AI Support for Stance Reflection and Writing

arXiv:2608.23050v1 Announce Type: new Abstract: An initial reaction to a social or community issue can feel meaningful before it is ready to become a message: people still need to clarify the claim, anticipate audience risks, and decide how much reasoning should become visible to others. We present StanceLab, a prototype for preparing a stance before entering a discussion. The prototype compares a three-persona mode, where an Interviewer, Mentor, and Opponent respond in parallel to help users diagnose and revise a stance, with a standalone LLM mode. In a formative within-subject pilot with six participants and 12 task sessions, every session produced a short final message in the notepad. The pilot revealed two design requirements: persona roles should diagnose useful blind spots or objections, and parallel responses need coordination support. We propose a future diagnosis-and-writing workflow that turns persona-based reflection into selective, audience-aware final messages.

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

AffAdapt: AFFect-driven ADAPTive AI Personas for Seamless Conversations

arXiv:2608.22702v1 Announce Type: new Abstract: AI-generated personas are being increasingly used for support, training and simulations. While generative AI models possess abilities to generate affect-aware responses, their embodiment into visual personas is an active area of investigation. Naturalistic exchanges require understanding of the conversational partners' turn completions, whether the agent should respond or keep listening and rely on non-verbal cues aligned with one's emotional states. Seamless human-AI conversation in a multimodal setting requires all modalities being generated to act in coordination. We present AffAdapt, a seamless interaction design framework for AI-personas, which coordinates streaming speech recognition, proactive turn-management, persona-grounded response generation, a persistent emotional state, and synchronized embodied output into a single interaction loop. We demonstrate the architecture in the context of practicing sensitive, high-stakes conversa

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

Poetic Heritage for Culturally Grounded Emotional Support: An Interaction Design Framework and Its Multimodal Agentic Instantiation

arXiv:2608.22639v1 Announce Type: new Abstract: Digital systems increasingly mediate emotional support, yet their interactions often remain culturally generic. Accordingly, we examine how a poetic tradition can be operationalized as a culturally grounded interactive medium and how generative AI can support such engagement. The resulting interaction design framework translates staged literature-based support and tradition-specific poetic aesthetics into guidance for digital system design. Poemithy instantiates the framework as a multimodal, LLM-enabled multi-agent system for guided reflection through classical Chinese poetry. A controlled between-subjects study with 50 participants compared text-only and multimodal versions. Both conditions showed medium-to-large within-session improvements in affect, anxiety, and emotion regulation, while between-condition tests detected no differences in these changes. Among secondary post-session user-experience measures, the clearest observed differ

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

"I want to be pushed, I want to grow": Enabling social workers to design evaluations of LLM augmentation in their work

arXiv:2608.22459v1 Announce Type: new Abstract: Workers are increasingly asked to adopt AI systems to assist their work, yet are rarely given a voice in defining what meaningful AI augmentation should look like or how to evaluate for it. In this paper, we propose worker-driven AI measurement---a bottom-up approach to AI evaluation where workers collaboratively shape decisions about which tasks AI should augment, what "successful" augmentation looks like, and how it should be measured. We explore how to support this through a case study with 19 workers from a local school social work organization. Through a series of eight workshops, workers iteratively develop their own measurement goals for AI evaluation, systematize these goals, and then design a benchmark to capture how effectively an LLM can "challenge" them to reflect on their own assumptions and biases in the context of their day-to-day work. Workers collaboratively design and refine an LLM-as-a-judge rubric based on their profes

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

Probing How Users Interact with Turn-Level Design Frictions for AI Chatbots

arXiv:2608.22427v1 Announce Type: new Abstract: AI chatbots can help people write faster, but they can also encourage overreliance by making it easy to turn minimal input into usable text. We study turn-level design friction: intentional constraints added to each chatbot exchange that slow, limit, or redirect how users request, access, or use model responses. We designed six friction probes, organized around three mechanisms: eliciting user contribution, restricting access to generated content, and reshaping system output. In a within-subject study with 24 participants, all six probes increased workload, task duration, and perceived ownership relative to a conventional AI chatbot, while their effects on recall and recognition were more selective. We further found that participants adapted to friction in different ways, and that the same constraint could support or obstruct involvement depending on users' goals and workflows.

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

Correctness Is Not Homogeneous Evidence: A Correctness-conditioned Evidence-aware Knowledge Tracing Model

arXiv:2608.22267v1 Announce Type: new Abstract: Knowledge tracing models usually use response correctness as a central observation for estimating students' latent knowledge states. However, the same correct or incorrect response may arise from different behavioral contexts, such as rapid guessing, hint use, or repeated attempts. Treating correctness as uniformly informative may therefore introduce ambiguity into recurrent state updates. This study proposes Correctness-conditioned Evidence-aware Knowledge Tracing (CE-KT), which uses observable response-process features to condition how correctness is written into recurrent states. CE-KT derives weakly supervised behavioral proxy scores from response time, hint use, attempt count, and behavioral history. These scores are used as behavioral signals, not as direct measures of mastery, response quality, or cognitive state. CE-KT then uses current correctness to select a correct-response or incorrect-response gate. The selected gate modulate

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

CAIA in Practice: Field Evaluation of an AI-Assisted Support System for Text-Based Online Counselling

arXiv:2608.22251v1 Announce Type: new Abstract: Rising global demand for mental health support creates significant service delivery challenges, with asynchronous email counselling serving as a crucial low-threshold channel for accessing care. This paper presents CAIA, a co-designed AI-based tool suite that demonstrates responsible AI integration into counselling practice through seven LLM-driven functions enhanced by retrieval-augmented generation. A field evaluation involved 34 professional counsellors conducting authentic sessions with trained student counsellees (36 threads, 321 messages, 1,257 AI outputs). User behaviour analysis confirms substantial adoption, revealing that professional autonomy and information accuracy are decisive for sustained acceptance, with counsellors particularly valuing interpretive functionalities that provide new perspectives and stimulate professional reflection.

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

Multi-Agent Discovery and Resource-Aware Autonomous Exploration of Scientific Datasets

arXiv:2608.22045v1 Announce Type: new Abstract: Modern scientific facilities and instruments generate datasets at scales that are difficult for individual researchers to discover, access, and explore. Although many datasets are publicly available, using them often requires familiarity with repository organization, data formats, multiresolution structures, and visualization parameters. We present WebVisus, a constrained and resource-aware multi-agent system for discovering and autonomously exploring remote, multiresolution scientific datasets. Given a natural-language research question, WebVisus identifies the user's intent and launches an autonomous exploration agent that examines slices, volumes, and timesteps while adapting data resolution and retrieval quality to available client memory and computational resources. This design supports progressive exploration without complete dataset downloads or manual configuration of low-level visualization parameters using natural languages. We

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

CALM-BP: Observation-Matched Physiological Semantic Grounding for Non-Contact Blood Pressure Estimation

arXiv:2608.21744v1 Announce Type: new Abstract: Language grounding increasingly involves non-text observations whose structure is not naturally expressed as words or objects. We study this problem for physiological time series in non-contact blood pressure (BP) estimation: remote photoplethysmography (rPPG) provides measured evidence about bodily state, but numerical pipelines expose little semantic structure about why a window is reliable or how its cues should be fused. We introduce observation-matched physiological semantic grounding, where language-derived priors must be constructed from the same rPPG observation, remain bounded by an auditable prior contract, and avoid BP-label or identity leakage. CALM-BP does not treat language as new physiological evidence; instead, it verbalizes rPPG descriptors into a controlled semantic interface while rPPG remains the primary haemodynamic evidence source. FlowBP-Set pairs forehead observations, synchronized BP labels, and structured physiol

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

UrbanGazeVis: A Visualization System for Analyzing Eye-Tracking Data on Urban Safety Perception

arXiv:2608.21686v1 Announce Type: new Abstract: Perceived safety in streetscapes depends on where people look, yet how gaze relates to visual cues of urban disorder remains poorly understood. Prior work treats safety as an image-level label, offering little insight into how attention to specific elements (e.g, buildings, greenery, people, signs of decay) shapes these judgments. We present a head-mounted eye-tracking study in which 30 participants viewed and rated the safety of 150 street-view images from Rio de Janeiro using a HoloLens 2 headset. Gaze traces were mapped onto semantic segments and disorder cues (e.g., damaged walls, graffiti, overhead cables), yielding a multimodal dataset linking gaze dynamics, scene semantics, and safety scores. To analyze it, we introduce UrbanGazeVis, an interactive visual analytics system with image- and participant-centric views that connects the spatial, temporal, and semantic dimensions of gaze to perceived safety, supporting comparisons between

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

Who Bears the Cost of Honesty? A FAccT Workshop Synthesis and Research Agenda for Equitable AI Disclosure

arXiv:2608.21671v1 Announce Type: new Abstract: AI disclosure is increasingly promoted and sometimes required as a route to transparency, accountability, provenance, and trust. Yet disclosure can also expose AI users to suspicion, stigma (e.g., competence penalties), and surveillance, affecting minoritized groups in particular. This paper reports on Who Bears the Cost of Honesty?, a CRAFT workshop at the 2026 ACM Conference on Fairness, Accountability, and Transparency that used scenario-anchored power mapping and design fiction to explore the benefits, harms, tensions, and power asymmetries that emerge under AI disclosure norms and mandates. We document the workshop design and analyze the disclosure approaches participants co-created, comprising four completed power maps, three context cards, and one interface prototype. These artifacts span education, workplace, politics/journalism, and interpersonal contexts. They depict disclosure as a multi-actor accountability process, surface co

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.HC

Exploring Agentic Approaches for Data Issue Detection and Repair in AI-Assisted Visualization

arXiv:2608.21602v1 Announce Type: new Abstract: AI is increasingly lowering the barrier to data analysis and creating visualization scripts. However, a key obstacle in AI-assisted visualization is that certain data issues can lead to visualizations that are plausible, but misrepresent the underlying data. These \textit{visualization defects} are elusive and difficult to fix, particularly for non-experts who may not know what data issues cause them or how to guide AI systems to resolve them. We present findings of a preliminary empirical investigation of how commercial LLMs identify and repair defect-inducing data issues. Using a curated subset of the 911 emergency-call dataset with five injected data issues, we evaluated GPT-5, GPT-4o, GPT-4, and Claude Sonnet 4.6 under a three-stage prompting protocol, including zero-shot, guided issue-identification, and guided issue-repair. We executed this protocol under two conditions: single-agent and a multi-agent orchestration that separates da

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Quantifying Compromise Risk in Exceptional Access Architectures Under Sparse and Indirect Evidence

arXiv:2606.19106v2 Announce Type: replace-cross Abstract: Lawful exceptional access (EA) systems hold the cryptographic keys that decrypt protected communications for authorised parties. The debate over their risks has been long and qualitative, complicated by two problems: no public dataset of EA-specific compromise events exists, so assessment must use sparse, indirect evidence; and prior work has treated structurally different designs as equivalent, though transmission-layer EA in carrier infrastructure (T-EA) and over-the-top EA at the platform layer (OTT-EA) differ in how cryptographic keys relate to ciphertext data. This paper builds a structured uncertainty framework for evaluating systemic compromise risk in EA architectures. It does not produce predictive forecasts, which the evidence cannot support; it separates findings robust to assumptions from those that depend on calibration. Four analytical layers are applied to T-EA and OTT-EA: three empirical pillars (historical analo

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems

arXiv:2606.17443v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are becoming a major way for consumers to find products, but we do not yet understand how brands compete in this new channel. We study brand dynamics in LLM recommendations using skincare products -- a category where consumers cannot easily judge quality before buying and must rely on brand reputation -- across three commercial LLMs (GPT-4o-mini, Claude Sonnet, Gemini 3 Flash), with a robustness check on search goods. In three experiments, we find: (1) a Conditional Monopoly where well-known brands get recommended 100% of the time (IAI = 10.0) when all products have the same specifications, but this dominance disappears with less than a +0.1-star rating advantage for a competitor; (2) authority-style marketing language, including fabricated clinical-evidence claims, breaks this monopoly at a Bias Surplus Value equal to +0.17 rating points, with each model responding differently; and (3) a social dile

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Verbalizing LLMs' assumptions to explain and control sycophancy

arXiv:2604.03058v3 Announce Type: replace-cross Abstract: LLMs can be socially sycophantic, affirming users when they ask questions like "am I in the wrong?" rather than providing genuine assessment. We hypothesize that this behavior arises from LLMs' incorrect assumptions about the user, like underestimating how often users are seeking information over reassurance. We present Verbalized Assumptions, a framework for eliciting these assumptions from LLMs. Verbalized Assumptions provide insight into LLM sycophancy, delusion, and other safety issues: in social sycophancy datasets, "seeking validation" is the most frequent bigram in LLMs' assumptions. We provide evidence for a causal link between assumptions and sycophantic model behavior: we train linear probes on internal representations associated with Verbalized Assumptions and then use these probes for interpretable, fine-grained steering of social sycophancy. Finally, we identify a human-AI expectation gap that explains why LLMs defa

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Benchmarking Retrieval-Augmented Generation Strategies for Large Language Model-Based Travel Mode Choice Prediction

arXiv:2508.17527v2 Announce Type: replace-cross Abstract: Accurately predicting travel mode choice is essential for effective transportation planning, yet traditional statistical and machine learning models are constrained by rigid assumptions, limited contextual reasoning, and reduced transferability. This study explores the potential of Large Language Models (LLMs) as a more flexible and context-aware approach to travel mode choice prediction, enhanced by Retrieval-Augmented Generation (RAG) to ground predictions in empirical data. We develop a modular framework for integrating RAG into LLM-based travel mode choice prediction and evaluate four retrieval strategies: basic RAG, RAG with balanced retrieval, RAG with a cross-encoder for re-ranking, and RAG with balanced retrieval and a cross-encoder for re-ranking. These strategies are tested across three LLM architectures (OpenAI GPT-4o, o4-mini, and o3) to examine the interaction between model reasoning capabilities and retrieval metho

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Practical Principles for AI Cost and Compute Accounting

arXiv:2502.15873v5 Announce Type: replace-cross Abstract: Policymakers increasingly use development cost and compute as proxies for AI capabilities and risks. Recent laws have introduced regulatory requirements for models or developers that are contingent on specific thresholds. However, technical ambiguities in how to perform this accounting create loopholes that can undermine regulatory effectiveness. We propose seven principles for designing AI cost and compute accounting standards that (1) reduce opportunities for strategic gaming, (2) avoid disincentivizing responsible risk mitigation, and (3) enable consistent implementation across companies and jurisdictions.

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Towards Safer Social Media Platforms: Scalable and Performant Few-Shot Harmful Content Moderation Using Large Language Models

arXiv:2501.13976v2 Announce Type: replace-cross Abstract: The prevalence of harmful content on social media platforms poses significant risks to users and society, necessitating more effective and scalable content moderation strategies. Current approaches rely on human moderators, supervised classifiers, and large volumes of training data, and often struggle with scalability, subjectivity, and the dynamic nature of harmful content (e.g., violent content, dangerous challenge trends, etc.). To bridge these gaps, we utilize Large Language Models (LLMs) to undertake few-shot dynamic content moderation via in-context learning. Through extensive experiments on multiple LLMs, we demonstrate that our few-shot approaches can outperform existing proprietary baselines (Perspective and OpenAI Moderation) as well as prior state-of-the-art few-shot learning methods, in identifying harm. We also incorporate visual information (video thumbnails) and assess if different multimodal techniques improve mo

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Examining Risks Through a Characterization of the AI Companion Application Ecosystem: A Stratified Sample from the Apple App Store and Google Play Store

arXiv:2603.13620v2 Announce Type: replace Abstract: While computer systems that allow users to interact through conversational natural language (i.e., chatbots) have existed for many years, various types of applications offering AI companionship (e.g., Character AI, Replika) have proliferated in recent years due to advancements in large language models. To better understand this application ecosystem, we identified 489 unique apps from the Apple App Store and Google Play Store that advertised AI companionship with social or relational capabilities (e.g., an AI romantic partner). We then systematically conducted and analyzed walkthroughs of a stratified sample of 30 apps, focusing on two distinct risk categories: potential harms posed to users by AI companion apps, and potential harms enabled by malicious users exploiting app features. Through our analysis, we categorize broader ecosystem trends that provide context for understanding risks and identify specific risks related to sensitiv

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study

arXiv:2504.08846v2 Announce Type: replace Abstract: We introduce AI University (AI-U), a flexible framework for AI-driven course content delivery that adapts to a course's instructional style. AI-U combines a fine-tuned large language model (LLM) with retrieval-augmented generation (RAG) and a reasoning synthesis model to generate style-aligned responses from lecture videos, notes, and textbooks. Using a graduate-level finite-element-method (FEM) course as a case study, we present a pipeline to synthesize course-grounded training data, fine-tune an open-source LLM with Low-Rank Adaptation (LoRA), and apply RAG-based synthesis. Our evaluation---combining cosine similarity, LLM-based assessment, expert review, and user studies---shows improved alignment with course materials relative to the base model. We have also developed a prototype web application, available at https://my-ai-university.com, that enhances AI-generated responses with references to relevant sections of the course mater

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

What is an intelligent system?

arXiv:2009.09083v4 Announce Type: replace Abstract: The term intelligent system has emerged in the field of information technology as a category of computer systems derived from successful applications of artificial intelligence. This paper proposes a general description that identifies the main properties and types of components typically found in such systems. Adopting an integrative and pedagogical approach, this description provides a conceptual framework for systems engineering practitioners seeking a coherent vocabulary and organizational structure to approach the analysis and construction of intelligent systems. The paper presents examples of both classical and modern intelligent systems to illustrate the generality and applicability of the description.

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Why we need an AI-resilient society- Profiling Large Language Models

arXiv:1912.08786v4 Announce Type: replace Abstract: Three generations of software have transformed the role of artificial intelligence in society. In the first, programmers wrote explicit logic. In the second, neural networks learned programs from data. In the third, large language models turn natural language itself into a programming interface. These shifts reach far beyond computer science, reshaping how societies generate knowledge, make decisions, and govern themselves. While generative adversarial networks introduced the era of deepfakes and synthetic media, large language models have added a new class of systemic risks. This report applies a forensic-psychology profiling methodology to characterize AI based on ten documented features: hallucinations, bias and toxicity, sycophancy and echo chambers, fabrication and credulity, knowledge without understanding, discontinuity and the inability to learn from experience, jagged intelligence and scaling limits, shortcuts and fractured r

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CY

Temporal Portability of Numeric User Metadata on Twitter

arXiv:2608.23449v1 Announce Type: cross Abstract: Numeric user metadata in social media are often reused over time. However, their reusability may depend on what an analysis needs to preserve. We introduce temporal portability as an analytical perspective for assessing the cross-time reuse of user features and feature-based rules. Specifically, we ask how well relevant properties are preserved when features and rules defined at a source time point are reused at a target time point. We used quarterly data on user features obtained directly from or derived from Japanese-language tweets in Twitter's 1% sample stream from 2020-Q1 to 2022-Q3. Each quarter included approximately 10.1--11.0 million unique users. We evaluated 13 numeric user features in terms of feature distributions, same-user relative ranks, selection rates, and selected-user membership. Across quarters, feature distributions changed and, for many features, same-user relative ranks were less well preserved at longer quarter

Source ↗
Showing 3451–3500 of 18402 signals
← Prev Page 70 of 369 Next →