EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18298 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Wed, 26 Aug 2026 00:00:00 -0400
arXiv cs.HC

MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation

arXiv:2608.24189v1 Announce Type: cross Abstract: Memory systems for conversational LLMs are conventionally evaluated by direct, fact-seeking questions about prior dialogue (Direct QA): can the model recall fact X from a prior conversation? We tested whether higher Direct QA accuracy correlates with higher user satisfaction in a 4-month deployment (40 users, 1,872 sessions, 7 memory conditions). Existing-benchmark Direct QA varies from 19.7% to 70.1% across the 7 conditions, but satisfaction does not change. We hypothesize that existing benchmarks and user satisfaction are tracking different capabilities: benchmarks measure elicited retrieval (recall when asked), while conversation requires natural integration (detecting relevance and naturally weaving prior context into a response). To examine this, we introduce MemUse, a set of real user-cued memory moments drawn from the deployment, scored by an integration-aware judgment of the natural conversational response. Holding the model and

Source ↗
technology Wed, 26 Aug 2026 00:00:00 -0400
arXiv cs.HC

ViSculpt: Visual-Centric Agentic Geometry Editing

arXiv:2608.24169v1 Announce Type: cross Abstract: 3D geometry editing is a critical yet labor-intensive part of the graphics pipeline, requiring artists to translate creative intent into precise operations in complex professional software. Large language models (LLMs) have shown promise for script-based 3D creation, but script generation is less suited to perception-driven editing of arbitrary existing meshes, where execution must remain visually grounded and untouched regions should be preserved. We present a \emph{visual-centric}, training-free multi-agent system that edits existing 3D meshes directly in Blender by emulating the iterative workflow of human artists. Rather than generating scripts or regenerating geometry, our system operates through the Blender GUI: multimodal LLM agents observe the viewport, reason about the current mesh state, and execute localized edits through simulated user interactions. Experiments on a curated benchmark provide initial evidence that this agenti

Source ↗
technology Wed, 26 Aug 2026 00:00:00 -0400
arXiv cs.HC

Bridging Teacher Expectations and Robot Learning via Coupling Dynamics

arXiv:2608.23994v1 Announce Type: cross Abstract: Human-robot teaching focuses on enabling nontechnical experts to customize robots according to their needs after deployment. With recent advances in machine learning, human-robot teaching is no longer confined to offline learning where the data gathering step from a human teacher is separated from when the robot learns. Instead, more recent approaches for human-robot teaching focus on coupling human teaching with robot learning. This coupling impacts the structure, timing, and content of the teaching and learning interaction. However, it is currently unclear how such coupling dynamics affect humanrobot teaching effectiveness and human perceptions towards the teaching process. Informed by human learning theories, in this paper we propose a new scale for classifying human-robot teaching interactions according to coupling dynamics present between the human teacher and robot learner. We apply this scale to a subset of the human-robot teachi

Source ↗
technology Wed, 26 Aug 2026 00:00:00 -0400
arXiv cs.HC

When Youth Enter The Chat: An Epistemic Shift in the Validation of LLM-Based Measures of Student Talk

arXiv:2608.23780v1 Announce Type: cross Abstract: LLMs are being used increasingly to measure aspects of student discourse (e.g. talk moves, collaboration, equity of voice) at scale. Typically, LLM-based measures of student talk use transcriptions of classroom conversations that only include verbal contributions, which de-contextualize student language. Common practices for validating these measures include comparing outputs against expert annotations by adults, using held out evaluation sets and F1 scores. We argue that these approaches are insufficient to ensure that such measures are meaningful and equitable for teaching and learning, particularly for racially and linguistically marginalized youth. In order to center the youth whose talk is being analyzed, re-contextualizing these classroom conversations and engaging youth in the research process is necessary. Sharing epistemic authority with youth, ultimately, centers their point of view and adds crucial nuance to the analysis of t

Source ↗
technology Wed, 26 Aug 2026 00:00:00 -0400
arXiv cs.HC

AI Agents Push Humans Out of the Loop

arXiv:2608.23642v1 Announce Type: cross Abstract: AI agents pose significant risks as they are granted increasing autonomy. A commonly proposed solution is human oversight and keeping a ''human in the loop'', but this is not a simple solution: Not only do current approaches to AI agent design impede effective human oversight, but the cognitive capacities required for it are also themselves degraded by extended use of AI systems. This position paper argues that current approaches to the development and deployment of AI agent systems do not support effective human oversight -- they contribute to its degradation. To address this, a top priority in the advancement of AI agents should be supporting the situated goals and cognitive requirements of effective human oversight, treating the human needs of overseers at the same level of importance as AI agent capability. To put this idea into practice, we connect work on automation and human-computer interaction to AI agent processes, outlining d

Source ↗
technology Wed, 26 Aug 2026 00:00:00 -0400
arXiv cs.HC

Ten Years Later: Replicating Two Color Discrimination Studies

arXiv:2608.24789v1 Announce Type: new Abstract: Color discrimination is a fundamental aspect of visualization as it influences how people interpret visual encodings. Many visualization guidelines are informed by perceptual studies, yet relatively few have been replicated. Acknowledging that the interaction between human perception, visual tasks, and display technology can change over time, we replicate two crowdsourced color discrimination studies conducted 10 years earlier. Specifically, we replicated a visualization-focused color discrimination task (N=144) and a more general perceptual discrimination task (N=394). In both studies, our results reproduced the original perceptual effects. We further use the replication to investigate whether color-related practice influences color discrimination. Specifically, we extended our replication studies by adding questions about participants' engagement with color practices. We then examined whether diverse color-related practices (e.g., artis

Source ↗
technology Wed, 26 Aug 2026 00:00:00 -0400
arXiv cs.HC

StrokeGuard: A Multi-Agent Guided System for Prehospital Stroke Assessment

arXiv:2608.24555v1 Announce Type: new Abstract: Prehospital stroke assessment aims to accurately identify stroke symptoms and make rapid decisions through standardized procedures within an extremely narrow time window, thereby saving valuable time for subsequent treatment. In clinical practice, FAST-based scales are widely used for prehospital stroke assessment by issuing instructions that guide subjects to perform specific actions to screen facial, arm, and speech functions. However, in home and community settings, non-clinical users often encounter challenges such as inaccurate descriptions, incomplete symptom observation, and difficult operational procedures, which may lead to inaccurate or biased assessment results. To address these challenges, this paper presents StrokeGuard: a multi-agent guided system designed for prehospital stroke assessment that makes mobile FAST screening more standardized and executable. Specifically, to overcome the limitations of traditional single-agent

Source ↗
technology Wed, 26 Aug 2026 00:00:00 -0400
arXiv cs.HC

Latent-surrealism: Revisiting surrealism and its aesthetics in relation to contemporary AI-Generated cultural production

arXiv:2608.24367v1 Announce Type: new Abstract: This chapter examines AI-media objects - creative and artistic outputs generated through generative artificial intelligence in the form of text-to-X tools - in relation to three avant-garde movements of the twentieth century: Dadaism, Surrealism, and Conceptual Art. Drawing on Lewis Carroll's Through the Looking-Glass as an early precursor to these three movements and to anti-rationalist aesthetics, and on three case studies in AI-generated conceptual architecture - Matias del Campo's "Deep House" and Hassan Ragab's "Post-Pharaonic Architecture" and "A State of Decay" - the chapter develops the concept of latent-surrealism. Latent-surrealism includes a set of aesthetic and methodological conditions inherent to creative AI-media objects. These include the use of readymade datasets reassembled through collage-like processes, the absurd as an aesthetic quality of machine hallucinations, and the decoupling of craft from artistic value. The ch

Source ↗
technology Wed, 26 Aug 2026 00:00:00 -0400
arXiv cs.HC

How Do Professional Editors Evaluate the Editing Quality of AI-Generated Cinematic Video Ads?

arXiv:2608.24329v1 Announce Type: new Abstract: On social media, we often encounter short-form video ads that employ cinematic editing techniques to evoke an emotional response. While AI tools are beginning to generate such cinematic ads automatically, we lack a fine-grained framework for evaluating these ads. In this paper, we first characterize social media video ad formats and identify cinematic ads as a recurring format in our corpus. We then analyze the duration, shot structure, audio and text elements, and editing techniques of cinematic ads to inform a two-stage generation pipeline in which an LLM first generates a shot plan and a video generation model renders the video. Using this pipeline, we generated 70 cinematic ads for 35 real brands and recruited professional video editors to critique their editing choices. From their critiques, we derive six dimensions of editing quality: narrative progression, audiovisual coordination and sound design, visual composition and graphics,

Source ↗
technology Wed, 26 Aug 2026 00:00:00 -0400
arXiv cs.HC

When AI "Works," When Does Help Begin?: Intergenerational Support Around Older Adults' LLM Usage

arXiv:2608.24297v1 Announce Type: new Abstract: LLMs are becoming part of everyday life, including for older adults (OAs). OAs often learn digital technologies with younger family members, who have traditionally served as "warm experts" providing trusted and personalized operational help. LLMs expand this role: family supporters may also help OAs judge appropriate uses, consider what information to disclose, assess the credibility of outputs, and decide when AI-generated advice is safe to act on. We conducted a formative qualitative study with six OAs and seven younger adults (YAs), using semi-structured interviews and scenario-based think-aloud activities. OA participants described using LLMs to lighten their recurring reliance on family, while preserving family as a selectively invoked support channel. However, because LLMs rarely produced visible operational breakdowns, YAs had limited signals for when support was actually needed. Instead, YAs relied on OAs' partial disclosures and

Source ↗
technology Wed, 26 Aug 2026 00:00:00 -0400
arXiv cs.HC

Aura: Dynamic Intra-Turn Emotion-Aware Adaptation of Large Language Model Responses

arXiv:2608.24224v1 Announce Type: new Abstract: Effective human-AI interaction requires systems that dynamically adapt to a user's behavior and evolving understanding. When users interact with Large Language Models (LLMs), these models typically respond to prompts without sensing the user's immediate reactions. This lack of communicative synchrony can lead to information overload or leave confusion unresolved in real time. In this paper, we introduce Aura, a framework that enables LLM systems to dynamically modulate output based on a user's evolving emotions. Aura's Perception Module continuously estimates the user's emotional state from facial expressions. Our Policy Module then selects interventions through a probabilistic belief model. Finally, Aura's Generation Module uses parameter-efficient Low-Rank Adaptation (LoRA) adapters to produce contextually tailored responses mid-turn during response generation. We evaluated Aura in a within-subjects user study (N=20) on information-seek

Source ↗
technology Wed, 26 Aug 2026 00:00:00 -0400
arXiv cs.HC

Balancing Evidence and Interpretation: Historical Grounding Ratio as a Design Parameter for AI-Generated Urban Storytelling

arXiv:2608.24157v1 Announce Type: new Abstract: Location-aware generative systems can now select historical archives and real-time contextual information based on a user's surroundings to automatically generate narratives for urban heritage walks. Yet when multiple sources jointly inform generation, existing systems provide neither a clear representation of how much content from each source actually appears in the output nor an operational means of measuring it. We introduce the Historical Grounding Ratio (HGR), defined as the proportion of claim-bearing information units in a generated narrative that are supported by historical archives. HGR turns the realized share of historical evidence in a narrative into a directly measurable design parameter. In GeoDrama, a mobile narrative system, we created three conditions that used a common retrieval procedure and comparable evidence-bundle sizes while varying the allocation of information from different sources during generation. We evaluate

Source ↗
technology Wed, 26 Aug 2026 00:00:00 -0400
arXiv cs.HC

Negotiating Ontological Boundaries in User-Authored Personal Sensing Systems

arXiv:2608.24058v1 Announce Type: new Abstract: Designed artifacts are ontological, shaping, and at times limiting, what becomes possible or imaginable. One path toward mitigating such foreclosures is giving people power over how systems are designed and built. Despite decades of scholarship around systems that enable such authorship, these systems are often evaluated on whether or not they are usable, useful, or technically feasible, leaving questions of ontological boundary negotiation, unexamined. We design two open-ended probes that utilize a Wizard of Oz technique to enable the experience of training a personalized machine learning system on phenomena people define themselves. In a week-long exploratory study, participants use one of two probes in the course of their everyday lives. We identify four sites where ontological boundaries were negotiated; the boundaries of a phenomena, the subject as part of relations, what is signal and what is noise, and the objectivity of data. We o

Source ↗
technology Wed, 26 Aug 2026 00:00:00 -0400
arXiv cs.HC

Who Chooses How Preferences Are Aggregated? Auditing Aggregation-Rule Authority in LLM-Based Group Recommendation

arXiv:2608.23966v1 Announce Type: new Abstract: AI systems increasingly make joint recommendations for users with conflicting preferences. However, when reasonable aggregation rules support different actions, a further question arises: who may choose how those preferences are combined? We study this interaction-level problem as aggregation-rule authority. Using synthetic preference profiles and profiles constructed from empirical ratings, we conduct a controlled behavioral audit of three LLMs under three authority conditions: unspecified, explicitly retained by users, and delegated to the model. In cases where two witness rules supported different actions, models almost never committed when users retained authority, but committed in every delegated case. All three models executed both witness rules perfectly when directly instructed. Yet when authority was unspecified or delegated, their aggregation-consistent outcome distributions differed across models and preference settings. Togeth

Source ↗
technology Wed, 26 Aug 2026 00:00:00 -0400
arXiv cs.HC

ColorA11Y: Enhancing Creative Design Workflows with Just-in-Time Color Accessibility Recommendations

arXiv:2608.23852v1 Announce Type: new Abstract: Effective color contrast in visual design is essential for content accessibility. While existing tools can identify contrast issues, they often operate in isolation from design workflows or are used as an afterthought. We present ColorA11Y, a system that supports designers in creating accessible content by providing just-in-time feedback and actionable recommendations throughout the authoring process to meet accessibility color contrast guidelines. Our system analyzes the visual properties of text and background elements and offers recommended changes, including text color adjustments, background modifications, and opacity changes. Through two user studies, we evaluate ColorA11Y's effectiveness. A user preference study (n=40) revealed varying effectiveness of different recommendations based on design context, while a qualitative study with designers (n=8) indicated a more seamless workflow experience in comparison to a baseline using a co

Source ↗
technology Wed, 26 Aug 2026 00:00:00 -0400
arXiv cs.HC

From Anonymous Shapes to Named Places: A Tool for Braille and Place-Semantic Annotation of Tactile Maps

arXiv:2608.23820v1 Announce Type: new Abstract: On a 3D-printed tactile map, a building felt under the finger is an anonymous shape: touch alone cannot tell which footprint is which, and a spoken description cannot reliably point to one shape at one place. We present a web-based tool that lets a sighted helper click to add on-shape Braille labels to an already-generated map model, downstream of the geometry generator so that whoever knows the reader and the local Braille standard does the labeling. The tool offers click-based OpenStreetMap matching, hand-editable abbreviation that shrinks a name to fit a footprint, and print-safe dot geometry with a review step that catches anomalies before printing. We demonstrate it on five printed maps of different place types, from a downtown core to a college campus and a small dining mall. In formative sessions in which ten BLV readers compared an unlabeled print with an annotated one, four read Braille fluently, so we treat Braille as one output

Source ↗
technology Wed, 26 Aug 2026 00:00:00 -0400
arXiv cs.HC

Technology Caregiving: Reframing How Older Adults Are Supported in Everyday Digital Activities

arXiv:2608.23751v1 Announce Type: new Abstract: Transformed by digitization, everyday activities-paying bills, shopping, managing transportation-increasingly require older adults to navigate digital systems. To accomplish these digital activities of daily living (DADLs), older adults often rely on help that looks less like IT support-institutional, episodic, and product-oriented-and more like caregiving: relational, ongoing, and aimed at preserving their functional independence. We argue that this practice is technology caregiving and introduce a framework characterizing it along four dimensions: why support is needed, who provides it, when it occurs, and how it is delivered. Applying this framework, we then systematically review the literature on how older adults are supported in DADLs. From 3,381 unique records, 36 articles met the inclusion criteria. Findings show that technology caregiving involves burden, like traditional care, but is distinctly shaped as much by digital systems a

Source ↗
technology Wed, 26 Aug 2026 00:00:00 -0400
arXiv cs.HC

The Ordinal Annotation Game: How Construct Abstraction Shapes Crowdsourced Consensus

arXiv:2608.23727v1 Announce Type: new Abstract: Inter-annotator disagreement in real-time affect annotation is widely treated as stochastic noise. We challenge this view by modelling ordinal annotation as an implicit game-theoretic coordination process against an internalised population prior under a post-hoc majority vote. We present the Ordinal Annotation Game, a conceptual scaffold in which the mapping from individual effort to collective consensus is governed by the semantic abstraction of the target construct. We evaluate it across two experiments sharing identical interface software and a uniform sensitivity threshold: a controlled sensory tracking study and an in-the-wild engagement study. Sensory annotation yields a consensus-dominant regime where active updates reinforce agreement, whereas engagement annotation inverts into an effort-limited regime where more labelling penalises consensus. The payoff slope reverses sign under identical processing, showing that ordinal disagree

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

Quantum Cinema: An Interactive Cinematic Exploration of Quantum Computing Hardware via Generative World Models

arXiv:2606.17102v2 Announce Type: replace-cross Abstract: Quantum computing promises transformative advances across science and industry, yet the physical hardware that enables these computations remains invisible to the public: quantum processors operate inside sealed dilution refrigerators at temperatures near absolute zero, making direct observation impossible. This "imagination gap" between quantum computing's growing societal impact and the public's ability to visualize it represents a significant barrier to quantum literacy and workforce development. We present Quantum Cinema, an open-source, browser-based interactive application that closes this gap by transforming invisible quantum hardware into explorable, cinematic experiences using generative world models. Quantum Cinema guides users through a four-act narrative -- from the foundational Nobel Prize-winning science of quantum entanglement, through curated video introductions to three major quantum computing architectures (tra

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

Lect\=uraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching

arXiv:2606.16428v2 Announce Type: replace-cross Abstract: Effective personalized AI-assisted learning demands systems that can not only generate accurate learner-specific educational materials, but also dynamically adapt their instruction to diverse learners. However, existing educational agents have primarily focused on lecture content automation and simulations, which often fall short of modelling multimodal and embodied instructional methods tailored for the individual learner. To this end, we propose Lect\=uraAgents - a multi-agent framework that enables personalized learning through end-to-end adaptive embodied teaching. At its core, Lect\=uraAgents mirrors a professor-student relationship, in which a ProfessorAgent leads a collaborative team of specialized subordinate agents through research, planning, review, and embodied delivery of lecture contents that adapt to a learner's needs. The framework offers three main contributions: (1) a hierarchical multi-agent architecture for en

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

Face versus Body Tracking for Human-Robot Interaction: An Egocentric Dataset

arXiv:2606.03694v2 Announce Type: replace-cross Abstract: Meaningful human-robot interaction (HRI) requires a robot to continuously assess user engagement through persistent user tracking. However, state-of-the-art Multi-Object Tracking models are heavily optimized for surveillance or autonomous driving. A social robot faces distinct egocentric challenges, such as humans moving in unpredictable nonlinear patterns, obstructing each other, or leaving and reentering the scene. These dynamics trigger frequent identity switches (IDSW), causing the robot to lose its footing mid-conversation. To address this, we introduce a focused, custom-annotated egocentric dataset collected via the Furhat robot. We present a systematic evaluation isolating detection errors from tracking logic, comparing face versus body tracking, and assessing the impact of extended memory and appearance re-identification (ReID). Results indicate that increasing temporal memory mitigates prolonged occlusions but fails on

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

Lightweight Test-Time Adaptation for EMG-Based Gesture Recognition

arXiv:2601.04181v2 Announce Type: replace-cross Abstract: Reliable long-term decoding of gestures from surface electromyography (EMG) is hindered by signal drift caused by electrode displacement, muscle fatigue, and/or posture changes. Although modern models achieve high intra-session accuracy, their performance often degrades substantially across recording sessions. Existing approaches to mitigate this problem typically rely on large training datasets or computationally intensive pipelines that are unsuitable for energy-efficient wearable devices. We propose a lightweight test-time adaptation framework for EMG decoding. The framework includes three complementary adaptation strategies: (i) causal adaptive batch normalization for online statistical alignment, (ii) Gaussian Mixture Model alignment with experience replay to mitigate forgetting, and (iii) meta-learning for rapid few-shot calibration. We evaluate these methods on the multi-session NinaPro DB6 dataset. All approaches substan

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

MILE: A Mechanically Isomorphic Hand Exoskeleton and Visuotactile Robotic Hand for Data Collection in Dexterous Manipulation

arXiv:2512.00324v4 Announce Type: replace-cross Abstract: Dexterous robotic hands are expected to perform complex, contact-rich object manipulation, but learning such skills remains challenging because high-dimensional hands require high-fidelity demonstrations. Imitation learning provides a practical route for acquiring dexterous manipulation skills from human demonstrations, yet collecting synchronized multimodal demonstrations with accurate hand actions and tactile observations remains a key bottleneck. We present MILE, a teleoperation-based data-collection system comprising the human-first MILE exoskeleton and the mechanically corresponding MILE-Tac robotic hand. The system integrates custom-designed and fabricated modular joint encoders and compact MILE fingertip visuotactile sensor modules. The exoskeleton is informed by human-hand anatomy and ergonomic constraints, while the robotic hand is co-designed to preserve the selected four-finger kinematic topology. This correspondence

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

Multimedia and Visual Analytics in the Agentic Era

arXiv:2504.06138v3 Announce Type: replace-cross Abstract: Professional users need tools to help them gain actionable insights from large multimedia collections. Foundation models and AI agents have rapidly changed the playing field, and improving their accuracy, trustworthiness, and reasoning capabilities are active topics in the computer vision, machine learning, and multimedia communities. Most current research focuses on benchmark driven algorithmic improvements. The multimedia community is the place to go beyond algorithms and consider complete multimedia analytics systems that support professional users in their complex tasks and achieve a true teaming of humans and AI. Supporting users with machine learning and visualizations has been studied for decades in the visual analytics field. In this paper, we propose a framework to bring multimedia and visual analytics together and indicate how it could impact current and new multimedia analytics solutions. Additional information can be

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

AI-Driven Analytics of Team-Teaching Talk: Acoustic Patterns across Experience, Cohorts and the Learning Design

arXiv:2606.09831v2 Announce Type: replace Abstract: As classroom cohorts expand, team teaching is increasingly used to integrate the expertise and pedagogical perspectives of multiple teachers. Yet, there is limited empirical understanding of how team teaching unfolds in practice, particularly regarding differences in teachers' contributions across experience levels, student cohorts, and learning task design. Prior research on team teaching has largely relied on retrospective self-reports or small-scale observations, offering limited insight into the micro-level processes through which team teaching is enacted. Teacher talk offers a scalable lens on these processes. While research in individual teaching contexts shows that acoustic features of speech (e.g., voice quality, intonation, and loudness) can shape student learning, evidence from team-teaching settings remains scarce. Moreover, capturing such features through manual observation or transcription is especially challenging in tea

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

MyoInteract: A Framework for Fast Prototyping of Biomechanical HCI Tasks using Reinforcement Learning

arXiv:2602.15245v2 Announce Type: replace Abstract: Reinforcement learning (RL)-based biomechanical simulations have the potential to revolutionise HCI research and interaction design, but currently lack usability and interpretability. Using the Human Action Cycle as a design lens, we identify key limitations of biomechanical RL frameworks and develop MyoInteract, a novel framework for fast prototyping of biomechanical HCI tasks. MyoInteract allows designers to setup tasks, user models, and training parameters from an easy-to-use GUI within minutes. It trains and evaluates muscle-actuated simulated users within minutes, reducing training times by up to 98%. A workshop study with 12 interaction designers revealed that MyoInteract allowed novices in biomechanical RL to successfully setup, train, and assess goal-directed user movements within a single session. By transforming biomechanical RL from a days-long expert task into an accessible hour-long workflow, this work significantly lower

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

"GenAI Defaults to Bias!" Gamify AI Literacy Through Reflections on Prompts

arXiv:2509.13679v2 Announce Type: replace Abstract: As Generative AI (GenAI) becomes widespread, it is increasingly important for the public to understand the model's behaviors and biases. However, existing AI literacy efforts miss opportunities to engage the general public to reflect on enduring GenAI bias and behaviors (e.g., how GenAI defaults to its internal bias in response to ambiguous or challenging prompts). In this work, we introduce ImaginAItion, a multiplayer game to help adults better reflect on GenAI bias and understand GenAI behaviors. ImaginAItion is grounded in reflective play to surface GenAI limitations by encouraging players to manipulate prompt specificity (e.g., an underspecified prompt "CEO" defaults to a white man). From ten sessions (n=30), we find that the game significantly improved players' understanding of GenAI behaviors by 35% in accuracy. Qualitative analysis showed how game mechanisms supported player reflections, including on prompting strategies to mit

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

From Inquisitorial to Adversarial: Using Legal Theory to Redesign Online Reporting Systems

arXiv:2506.07041v2 Announce Type: replace Abstract: User reporting systems play a central role in how online communities address interpersonal conflict and harassment, especially in private spaces such as direct messages, voice chats, and end-to-end encrypted messaging. These settings complicate evidence collection for community moderators while heightening users' concerns about procedural justice and privacy. To examine these challenges, we draw on adversarial legal frameworks from offline judicial systems and apply them to community-level reporting systems, using Discord as a research site. We find that online community reporting systems often follow an inquisitorial model, in which moderators lead evidence collection and case development, rather than an adversarial model, which gives users greater control over how evidence is presented and contested. Although adversarial practices can strengthen procedural justice and protect privacy, they can also introduce new risks of abuse, unde

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

Assessing Distribution Shift in Human Activity Recognition for Domain Generalization

arXiv:2606.24781v1 Announce Type: cross Abstract: While the field of Human Activity Recognition (HAR) continues to draw interest from researchers and advance in important ways, some key challenges remain. One of the most difficult aspects of building HAR models that show good performance in real-world settings is dealing with data diversity from device and sensor heterogeneity, and contextual changes that are intrinsic to real-world applications. While data diversity in HAR has been well-acknowledged in the literature, there remains a gap in understanding the effect of various types of distribution shifts on HAR models and the domain generalization problem that arises. Towards that end, this paper systematically evaluates 4 different types of distribution shifts, including variations in device type, sensor placement, sampling rate, and user behavior. Quantifying their effects, we illustrate that diversity shifts predominantly define all types of shifts, indicating the existence of uniq

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

Task Decomposition for Efficient Annotation

arXiv:2606.24734v1 Announce Type: cross Abstract: High-quality annotations of structured representations are expensive to collect over large corpora. Manual annotation of structure is laborious, and model-based annotation, although cheaper to generate, requires expensive validation and potentially significant supervision to ensure that the annotation quality is strong enough to be useful downstream. In traditional annotation workflows, annotation of each complete example is performed end-to-end by a single annotator. However, structured annotation is complex, and each aspect of the task represents a unique challenge with an associated inferential load for a given annotator. Modern annotation projects can incorporate heterogeneous groups of annotators, including both models and human annotators with varying domain and linguistic expertise. It remains unclear, however, how to redesign annotation tasks in this setting, where efforts are discriminately allocated across heterogeneous annota

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

Measuring User's Mental Models of Speech Translation in Human-AI Collaboration

arXiv:2606.24644v1 Announce Type: cross Abstract: Millions of people use machine translation (MT) tools daily, yet little is known about their perception of what systems can and cannot do. This paper studies users' mental models of speech translation systems through a new framework based on cross-lingual question answering, where users either accept MT output or request professional re-translation to answer questions based on the information presented in a foreign language. By analyzing user behavior and accuracy trends across varying translation qualities, we examine to what extent they can predict where the system is likely to be wrong, and how this mental model evolves. Users develop stronger mental models with practice, especially when they have some knowledge of the source language, primarily by relying on surface-level error cues. Moreover, providing speech transcriptions can help users develop better mental models. Our results show the promise of cross-lingual question answering

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback

arXiv:2606.24622v1 Announce Type: cross Abstract: Training safe Reinforcement Learning (RL) systems is inherently challenging, with no guarantee of avoiding unwanted behaviors. The most effective defenses against this are (i) transparency through explainability and (ii) alignment via human feedback. While both show promising results, no publicly available framework currently combines them. To address this, we introduce Themis, an XAI-enabled testing and evaluation framework for Reinforcement Learning from Human Feedback. Themis supports over 200 widely used environments and is easily configurable for experiments in RL, transparency, and alignment. Our results show that Themis can train reward models that match or outperform the environment's true reward signal using human preferences. We also provide a cloud-based platform for collecting human feedback and managing experiments. It is user-friendly, auto-scalable, and supports large participant groups across multiple experiments without

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

Reinforcement Learning for Computer-Use Agents with Autonomous Evaluation

arXiv:2606.24515v1 Announce Type: cross Abstract: Computer-Use Agents (CUAs) execute high-level user goals by perceiving and acting directly within graphical user interfaces. However, reinforcement learning for CUAs remains difficult because open-ended desktop environments rarely provide scalable, machine-readable reward signals: task success is often visually grounded and hard to specify with handcrafted reward functions or dense manual labels. We propose an RL fine-tuning framework that uses autonomous vision-language evaluation as a scalable supervision signal for GUI agents. Given a final screenshot and the original instruction, a Vision-Language Model judges task completion and provides terminal feedback without task-specific heuristics or manual labels during policy optimization. Because autonomous evaluators are imperfect, we model their feedback as a noisy binary reward channel and derive a noise-corrected reward estimator for Proximal Policy Optimization. Experiments across ma

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

Real-Time Interactive Music Generation via Data-Free Streaming Consistency Distillation

arXiv:2606.24307v1 Announce Type: cross Abstract: Interactive music and live performance relies on real-time human expression, but modern generative music AI remains largely absent from this domain due to its prohibitive inference latency and offline rendering paradigm. To provide pioneer musicians with a novel medium for interactive composition, we should fundamentally change these static models into dynamic, playable instruments. In this paper, we propose a framework that bridges this gap. To achieve the low latency required for live interaction without sacrificing structural coherence, we formulate distillation within a streaming autoregressive latent space. Our approach gets rid of the need for expensive paired audio-latent datasets by utilizing prompt-only inputs to synthesize teacher-guided, chunk-wise trajectories on the fly. Because live instruments require high acoustic fidelity, we introduce music-aware consistency objectives, which combine latent, spectral, and temporal-diff

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

Dialogue to Discovery: Attribute-Aware Preference Elicitation for Conversational Product Search Assistants

arXiv:2606.24194v1 Announce Type: cross Abstract: Conversational product search assistants offer a more expressive, natural, and interactive alternative to traditional keyword-based product search. With limited screen space, showing only a few items increases the need for precise preference elicitation, which can prolong conversations, leading to user frustration and session abandonment. Conversely, rushing to recommend items without a clear understanding of preferences risks poor matches and a degraded user experience. We present Dialogue to Discovery (D2D), an attribute-oriented preference elicitation framework that dynamically exploits the structure of product attributes to efficiently steer conversations toward the user's desired item. D2D adaptively prioritizes the most informative queries and strategically times product recommendations, reducing premature or off-target suggestions that harm engagement. To evaluate D2D, we curate three datasets from the Amazon Reviews corpus. In s

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

Aspect-Based Sentiment Evolution and its Correlation with Review Rounds in Multi-Round Peer Reviews: A Deep Learning Approach

arXiv:2606.24188v1 Announce Type: cross Abstract: Mining sentiment information from the textual content of peer review comments offers valuable insights into the scientific evaluation process. However, previous studies are often constrained by coarse-grained analysis and the lack of differentiation across review rounds. Notably, the dynamic shifts in reviewers' focus and sentiment tendencies throughout multiple review stages remain underexplored. To address this gap, the present study investigates the distribution and evolution of aspect-level sentiments and examines their correlation with the number of review rounds. We begin by segmenting the multi-round review comments of 11,063 accepted papers from Nature Communications and identifying fine-grained review aspect clusters. A manually annotated corpus of approximately 5,000 review sentences is then constructed. Using this dataset, we train a series of deep learning-based aspect sentiment classification models. Among them, the LCF-BER

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

Ten Digits on a Train: AI-Assisted Verification of Two Eigenvalue Problems

arXiv:2606.23821v1 Announce Type: cross Abstract: Accurate numerical eigenvalues are often difficult to certify, especially in singular or non-normal settings. This article reports a human--AI collaboration on two such computations. For a singular self-adjoint Schr\"odinger operator, a verified zero count and Dirichlet--Neumann bracketing certify the complete negative spectrum to ten decimal places. For a delicate non-normal atom--molecule benchmark, a previously unresolved resonance pair is separated, with each member enclosed to ten digits. The second result is achieved not by increasing the precision of one-way shooting, but by reformulating the problem as a global matching system for projective solution lines. The infinite tail is encoded as uncertainty in the terminal projective data, and a componentwise, tail-robust Krawczyk--Brouwer inclusion supplies the certificate. This gives a reusable architecture for analytic boundary-value systems with ill-conditioned propagation and unce

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

EvidenceLens: A Claim-Evidence Matrix for Auditing Financial Question Answering

arXiv:2606.23724v1 Announce Type: cross Abstract: Large language models are increasingly used to answer questions over annual reports, earnings decks, and analyst notes, yet their outputs remain difficult to verify in high-stakes financial workflows. A fluent answer can blend directly grounded statements, weak synthesis, and unsupported claims across narrative text, tables, and charts. We present EvidenceLens, a visual analytics prototype that treats financial question answering as a claim-evidence alignment problem. The system decomposes an answer into atomic claims, summarizes support composition and confidence, support gaps, and coordinates claim-level inspection with source passages, table cells, and chart regions. Its core visual representation is a multimodal claim-evidence matrix that makes coverage, contradiction, and modality imbalance immediately visible. To support reproducibility, we also specify a JSON-based artifact schema, a lightweight multimodal alignment pipeline, and

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

Zero-Shot Neural Priors for Generalizable Cross-Subject and Cross-Task EEG Decoding

arXiv:2606.23706v1 Announce Type: cross Abstract: The development of generalizable electroencephalography (EEG) decoding models is essential for robust brain-computer interfaces (BCI) and objective neural biomarkers in mental health. Conventional approaches have been hindered by poor cross-subject and cross-task generalization, owing to high inter-subject variability and non-stationary neural signals. We address this challenge with a zero-shot cross-subject decoding framework on the large-scale Healthy Brain Network dataset, benchmarking a convolutional neural network baseline, a hybrid LSTM, and a Transformer-based foundation model. To adapt the Transformer for regression while averting catastrophic forgetting, we propose a novel progressive unfreezing strategy. The baseline yielded an nRMSE of 0.9991, whereas our fine-tuned Transformer achieved 0.9799 on unseen subjects. This work advances scalable, calibration-free EEG decoding for computational psychiatry and behavioral prediction.

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

Evaluating LLM Usage for Efficient and Explainable Numerical and Classified Implicit Sentiment Analysis of Product Desirability

arXiv:2606.23701v1 Announce Type: cross Abstract: Qualitative product feedback can reveal nuanced user experiences, but its implicit sentiment is difficult to measure. This paper presents a scalable and interpretable framework that uses large language models (LLMs) to quantify product desirability from such data. Using two Product Desirability Toolkit (PDT) datasets from ZORQ and CARMA comprising 106 respondent term groupings with gold-standard human annotation, zero-shot continuous numerical sentiment scoring and categorical sentiment classification are evaluated without relying on explicit review scores. Across the datasets, LLMs generated numerical sentiment scores directly from qualitative responses and closely matched expert labels, achieving Pearson correlations up to 0.97 and classification accuracy up to 94%. LLMs maintained robustness even when handling data presented in multiple forms and consistently expressed high confidence. In contrast, lexicon-based and transformer basel

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

A Geometry-Informed Computer Vision Method for Detecting and Examining Overtaking Vehicles From A Bicycle

arXiv:2606.23699v1 Announce Type: cross Abstract: Instrumented bicycle studies have produced direct field evidence on vehicle passing behavior, but extracting overtaking events from continuous rear-facing video has remained dependent on manual, frame-by-frame annotation. This bottleneck constrains sample sizes and limits naturalistic cycling safety research. We present a geometry-informed computer vision pipeline that automates overtaking event detection from a single bicycle-mounted camera without multi-sensor configurations or explicit camera calibration. The system combines RT-DETR object detection with ByteTrack multi-object tracking through a three-stage geometric validation module enforcing bearing angle trend, apparent size growth, and spatial confirmation criteria derived from perspective projection principles. Validated on 315 manually annotated real-world overtaking events from urban roads in Ann Arbor, Michigan, the pipeline achieved 97.8% recall with zero false positives. T

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

"Zooming In" on Agentic Web Browsers as Assistive Technologies: A Case Study with a Low-Vision Technology Expert

arXiv:2606.24870v1 Announce Type: new Abstract: Agentic Web Browsers (AWBs), powered by Large Language Models (LLMs), are emerging as autonomous systems capable of navigating the Web on behalf of users. Beyond enhancing productivity, they could also offer significant promise as Assistive Technologies (ATs) for visually-impaired individuals, transforming web interaction into a fluid conversational exchange. In this paper, we present a case study with a low-vision technology expert, examining how AWBs can support visually-impaired users in web navigation. The findings show that, despite the current limitations, the navigation experience is notably fluid and flexible, underscoring the strong potential of AWBs to enhance accessibility and reduce barriers in web interaction, with implications that may extend beyond accessibility to agentic UX more broadly.

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

It's Complicated: On the Design and Evaluation of AI-Powered AAC Interfaces

arXiv:2606.24854v1 Announce Type: new Abstract: Artificial intelligence (AI) can enhance what people who use augmentative and alternative communication (AAC) are able to do with their systems. However, evaluating AI-powered AAC interfaces can be difficult. People are intersectional beings and current evaluation metrics can struggle to capture the multifaceted and nuanced desires people may have for their AAC. We explore the complicated nature of six AAC problem spaces, explore how AI might be used in these spaces, and suggest more robust methods of evaluation that take the intersectional nuances of people into account. We also discuss broader issues that arise across these problem spaces and how they could be addressed using our proposed evaluation methods.

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

Virtual Simulation for Mental Health

arXiv:2606.24826v1 Announce Type: new Abstract: Poorly designed interventions or those deployed without adequate safeguards can harm the communities they aim to serve, thus exacerbating existing vulnerabilities and leaving individuals unsupported. This is especially the case for the mental health context, where there is a growing trend of relying on technological interventions due to their accessibility and ability to deliver large-scale support. However, the mental health context is also particularly sensitive to change and risks of failure are dire; at their worst, failures in mental health interventions can result in lasting negative outcomes for individuals and tragic losses as people fall through the cracks. Thus, enabling safe ways to experiment in the mental health context is vital to allow both individuals and communities to engage with new interventions without risk of their real-world consequences. Virtual simulation, which uses virtual environments to replicate real-world in

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

SciFi-VIS: Way Out There -- How SciFi and Visualization Influence Each Other

arXiv:2606.24731v1 Announce Type: new Abstract: We propose a hybrid half-day workshop at IEEE VIS 2026, calling for participation from visualization researchers and science fiction creators in order to develop a systematic understanding of the two-way relationship these communities have long shared. We invite submissions of creative formats showcasing connections and inspiring future research. Our workshop plan includes a keynote, lightning talks, brainstorming, cross-community critique, affinity mapping, and discussion around identified themes.

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

SupplyNet: Supporting Visual Exploratory Learning in Supply Chain via Contextual Multi-Agent Simulation

arXiv:2606.24694v1 Announce Type: new Abstract: Simulation has long supported supply chain management instruction by letting learners observe network behavior and test decision strategies. Recent progress in LLM-driven agents opens new possibilities for richer, more adaptive simulations, but many existing systems still present abstract, opaque data that overwhelms learners and discourages active exploration. We introduce \textit{SupplyNet}, a gamified visual simulation system built on a contextual graph-based LLM multi-agent framework that models interdependent supply chain dynamics and provides responsive feedback through tiered challenges. \textit{SupplyNet} turns the simulation into a manipulable decision space by integrating an interactive network view of system state, a branching timeline for "what-if" exploration and comparison, and a task-oriented analysis console for structured performance breakdowns. Together, these visual components support counterfactual exploration, causal

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

Optimizing Visual Analytics Workflows: From Theory to Practice

arXiv:2606.24454v1 Announce Type: new Abstract: The principle of visual analytics (VA) is to provide integrated workflows where human-centric processes (e.g., visualization and interaction) and machine-centric processes (e.g., statistics and algorithms) complement each other. To implement this principle in practice, it is necessary to reason about the trade-offs among different processes and make optimal use of them in a workflow. Building on an existing ontology of the methodology for analyzing such trade-offs information-theoretically and for optimizing VA workflows systematically, we investigate ways to transform this methodology from theory to practice. In particular, we adopted the action research method. Through case studies in different application domains, VA researchers with different background knowledge and experiences offered their answers to several hypotheses about using the methodology in practice and proposed ways forward. In this paper, we present our collective analys

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

Average Rankings Mask Per-Subject Optimality: A Friedman-Nemenyi Benchmark of EEG Motor-Imagery BCI Decoders

arXiv:2606.24394v1 Announce Type: new Abstract: Electroencephalography (EEG) is the dominant non-invasive modality for brain-computer interfaces (BCIs), yet reliable decoding of motor imagery is hampered by inter- and intra-individual variability. A recurring claim is that one decoding pipeline, most often a spatial or Riemannian method, is broadly preferable. We test the weakest version of that claim under the most favourable conditions. Using the Mother of All BCI Benchmarks (MOABB) framework, we evaluated 1,056 decoding configurations (feature extractor x scaler x classifier), >340,000 subject-level model fits, across three public left-versus-right motor-imagery datasets (PhysionetMI, 109 participants; Cho2017, 52; Zhou2016, 4) and two frequency bands (8-15 Hz, 8-30 Hz). Every model is fit and tested within a single session of a single participant, the easiest regime, giving every pipeline its best chance. We apply the statistics standard for multi-classifier comparison: Friedman om

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

A Dynamic Coupling Theory of Expertise Through Thinking Flow and Workflow Evolution

arXiv:2606.24197v1 Announce Type: new Abstract: Expertise has long been explained through tacit knowledge, deliberate practice, skill acquisition, and expert performance. While these perspectives have advanced understanding of expertise, they often describe its conditions or outcomes rather than the cognitive architecture through which expertise continuously emerges and evolves. This paper proposes Workflow Cognition as a theoretical framework for explaining expertise as a dynamic cognitive phenomenon. Workflow Cognition is defined as the cognitive architecture emerging from the recursive coupling of Thinking Flow and Workflow Evolution. Thinking Flow refers to ongoing processes of perception, interpretation, judgement, decision-making, and reflection; Workflow Evolution refers to the continuous adaptation of actions, task structures, and operational strategies within situated practice. Through their coupling, expertise is not treated as a static accumulation of knowledge or skill, but

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.HC

Human-Centered Design: The Disclosure of Generative Artificial Intelligence for Emerging Professionals

arXiv:2606.24136v1 Announce Type: new Abstract: As the Human centered design continues to grow, generative AI has the potential to streamline the research process by iterating tasks within established workflows to increase efficiency. However, integrating AI raises concerns surrounding ethical bias, complexity, and the lack of prioritization of humanistic values. Emerging professionals represent a cohort with the opportunity to learn Human Centered Design principles, yet without this foundation AI becomes more of a crutch than a tool, leading to reduced experience with deep work, decreased autonomy, and deskilling of key foundations. Disclosures are a common method to self report AI usage, but they provide little clarification on appropriate implementation and may encourage omission to avoid consequences. This paper reflects on experiences in the Human Centered Design course ITIS8300, which emphasized optimizing user experience, enhancing innovation and collaboration, and improving eff

Source ↗
Showing 51–100 of 1631 signals
← Prev Page 2 of 33 Next →