EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Monitoring

arXiv:2509.07260v5 Announce Type: replace-cross Abstract: Mobile and wearable healthcare monitoring play a vital role in facilitating timely interventions, managing chronic health conditions, and ultimately improving individuals' quality of life. Previous studies on large language models (LLMs) have highlighted their impressive generalization abilities and effectiveness in healthcare prediction tasks. However, most LLM-based healthcare solutions are cloud-based, which raises significant privacy concerns and results in increased memory usage and latency. To address these challenges, there is growing interest in compact models, Small Language Models (SLMs), which are lightweight and designed to run locally and efficiently on mobile and wearable devices. Nevertheless, how well these models perform in healthcare prediction remains largely unexplored. We systematically evaluated SLMs on health prediction tasks using zero-shot, few-shot, and instruction fine-tuning approaches, and deployed t

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

When Should AI Follow? Task Structure and Joint Adaptation by Human and AI Agents

arXiv:2504.20903v4 Announce Type: replace-cross Abstract: How should organizations divide and sequence decision tasks between human and artificial agents? We develop a computational model of joint sequential adaptation in which two agents differ in a single, precisely specified way: the memory regime governing how past decisions shape subsequent ones. A recency-weighted regime, motivated by behavioral evidence on human adaptation, privileges recent outcomes; a uniform-memory regime, motivated by the scale-free consistency of algorithmic updating, weights a window of past outcomes equally. Situated in the lineage of NK/NKC models but developed on its own terms as a sequential-adaptation model, the framework varies task scope (N), within-task coupling (K), and cross-agent coupling (C) across modular and sequenced task structures. Three mechanisms organize the results. First, threshold dynamics create absorbing high- and low-payoff regimes, so adaptation compounds whatever it inherits. Se

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

Designing a Visualization Atlas: Lessons & Reflections from The UK Co-Benefits Atlas for Climate Mitigation

arXiv:2604.20781v2 Announce Type: replace Abstract: This paper reports on the process of designing the UK Co-Benefits Atlas, which communicates and publicizes data for climate mitigation. Visualization atlases--an emerging type of platform to make data about complex topics comprehensive through interactive visualizations and explanatory content--pose challenges beyond traditional visualization projects. Atlases must address diverse and often uncertain audiences and use cases, support both explanatory and guided exploration, and accommodate complex, evolving data. Over 10 months, our team of visualization and domain experts conducted 8 design workshops, iterative prototyping, 15 stakeholder onboarding sessions, and continuous reflection. These intertwined processes informed the development of the Atlas, comprising over 400 pages of visualizations and explanations. They also enabled a deeper understanding of how stakeholders may critically engage with the atlas in practice, in terms of i

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

Facial-Expression-Aware Prompting for Empathetic LLM Tutoring

arXiv:2604.15336v2 Announce Type: replace Abstract: Large language models (LLMs) enable increasingly capable tutoring-style conversational agents, yet effective tutoring requires sensitivity to learners' affective and cognitive states beyond text alone. Facial expressions provide immediate and practical cues of confusion, frustration, or engagement, but remain underexplored in LLM-driven tutoring. We investigate whether facial-expression-aware signals can improve empathetic tutoring responses through prompt-level integration, without end-to-end retraining. We build a scalable simulated tutoring environment where a student agent exhibits diverse facial behaviors from a large unlabeled human facial expression video dataset, and compare four tutor variants: a text-only LLM baseline, a multimodal baseline using a random facial frame, and two Action Unit estimation model (AUM)-based methods that either inject textual AU descriptions or select a peak-expression frame for visual grounding. Ac

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

Ask don't tell: Reducing sycophancy in large language models

arXiv:2602.23971v4 Announce Type: replace Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an alignment failure, particularly in high-stakes advisory and social contexts. While prior work has documented conversational features correlated with sycophancy, we lack a systematic understanding of what provokes or prevents AI sycophancy. Here, we present a set of controlled experimental studies where we first isolate how input framing influences sycophancy, and second, leverage these findings to develop mitigation strategies. In a nested factorial design, we compare questions to various non-questions where we vary three orthogonal factors: epistemic certainty (statement, belief, conviction), perspective (I- vs user-perspective), and affirmation vs negation. Measuring expressed sycophancy, how sycophantically a model phrases its free-text response, we show that (1) sycophancy is substantially higher

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

Grating haptic perception through touchscreen: Sighted vs. Visually Impaired

arXiv:2511.10026v3 Announce Type: replace Abstract: Providing haptic feedback via smartphone touch screen may potentially offer blind people a capability to understand graphs. This study investigated the discrimination performance of haptic gratings in different frequencies, in both visually impaired (VI) and sighted (S) individuals. 6 VI participants and 10 S participants took part in two experiments designed to compare their ability to interpret grating images with a finger swiping across a smartphone touchscreen without vision. The swipe gesture activates phone vibration temporally synchronized with the black stripes. Their tasks were: (1) determining whether a grating pattern is presented on the touchscreen, (2) comparing two different grating frequencies and determining the wider one. Results demonstrated that the VI group exhibited superior tactile sensitivity compared to the S group, as evidenced by their significantly better performance in Experiment 1 (accuracy of 99.15\% vs.

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

AI LEGO: Scaffolding Cross-Functional Collaboration in Industrial Responsible AI Practices during Early Design Stages

arXiv:2505.10300v2 Announce Type: replace Abstract: Responsible AI (RAI) efforts increasingly emphasize the importance of addressing potential harms early in the AI development lifecycle through social-technical lenses. However, in cross-functional industry teams, this work is often stalled by a persistent coordination challenge: how technical roles hand off technical intent, how teams establish shared structures for collaboration, and how non-technical roles are supported in systematically evaluating harms. Through literature review and a semi-structured interview study with 8 practitioners, we unpack how this challenge manifests---technical design choices are rarely handed off in ways that support meaningful engagement by non-technical roles; collaborative workflows lack shared, visual structures to support mutual understanding; and non-technical practitioners are left without scaffolds for systematic harm evaluation. Existing tools like JIRA or Google Docs, while useful for product

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

APEX-Accounting

arXiv:2607.27189v1 Announce Type: cross Abstract: We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set comprises 160 tasks, split across 10 worlds. Each world contains an accounting system, as well as spreadsheets, PDFs, and other files. Every task was authored and solved by experts in accounting and bookkeeping, who also wrote grading rubrics. Across nine frontier models, Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3, ahead of Muse-Spark-1.1 (xHigh) at 52.6%. No model scores more than 2.6% Pass^8 (GPT-5.6-Sol (Max+Pro)) and the highest Pass@8 is 21.5% (Muse-Spark-1.1 (xHigh)). We experiment with increasing the token budget from $1 to $50 and observe an instance of Simpson's paradox: scores increase as the token budget increases but within a given budget-cons

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork

arXiv:2607.27177v1 Announce Type: cross Abstract: Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc teamwork (AHT) approaches assume that agents will collaborate on a single, fixed task and that the partner's capabilities, their ability to successfully execute the desired action, are already known. In reality, a partner's true capabilities are often hidden, and human collaborators may act sub-optimally on tasks with multiple valid strategies. To address these limitations, we extend ad-hoc teamwork into a multi-task setting by re-framing it as a problem of joint planning with decentralised execution under hidden partner capabilities. We introduce CE-CM (Capability Estimation via Contextual Models), an approximate Bayesian method that infers task-invariant capability vectors. By using simulation-based sampling, the agent estimates capabilities and induces a contextual Multi-agent Markov Decision Processes for planning. T

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

arXiv:2607.27155v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding. The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a privacy-preserving process. On average, these tasks require 2.32 hours of human labor to complete. An important feature of the benchmark is that each task is paired with two economic signals: human labor time and task price proxy. These signals enable direct comparisons between human costs and LLM inference costs, as well as value-weighted evaluation. To support stable evaluation, we develop code-based verifiers from fine-grained rubrics. We evalua

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

An Attention-Based Framework for Alzheimers Disease Classification Using Resting-State fMRI

arXiv:2607.26746v1 Announce Type: cross Abstract: Accurate identification of Alzheimers disease (AD) using resting-state functional magnetic resonance imaging (rs-fMRI) remains challenging due to the high dimensionality, noise, and complex inter-regional dependencies inherent in functional brain connectivity, which limit the effectiveness of traditional approaches based on handcrafted connectivity features or conventional machine learning models. In this work, we present an attention-based deep learning framework for Alzheimers disease classification that operates directly on rs-fMRI functional connectivity matrices by treating brain regions as tokens and employing a Transformer-inspired self-attention mechanism to model long-range and global functional dependencies across distributed brain networks. The proposed framework learns discriminative functional representations without reliance on manual feature engineering and is evaluated on a longitudinal cohort from the Alzheimers Disease

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

Contrastive ESA: Human Evaluation of Multiple Translations at Once

arXiv:2607.26640v1 Announce Type: cross Abstract: Current human evaluation of machine translation typically assesses single outputs in isolation, a paradigm that suffers from high annotator noise and cost. We introduce Contrastive Error Span Annotation (cESA), a protocol that presents multiple translations of the source input (text, video, audio, image). In cESA, the annotator sees multiple translations of the same document, marks major and minor error spans, and then assigns a score from 0% to 100% on absolute scale. By allowing annotators to access the shared context across multiple outputs, cESA facilitates more consistent and efficient judgments. We validate cESA using a large-scale human evaluation of English->Japanese translations of 12 models, demonstrating reductions in annotation time and noise compared to standard pointwise evaluation. Unlike existing contrastive ranking methods, cESA yields absolute quality judgments that enable simple, interpretable non-parametric model ran

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants

arXiv:2607.26611v1 Announce Type: cross Abstract: AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities that recur in user-specific ways across tasks and sessions. Existing disambiguation methods typically address each ambiguous request in isolation within the current coding session, often through eliciting additional clarification. However, whether resolved session history from the same user can serve as memory for resolving recurring personalized ambiguity in a newly opened session remains underexplored. We formulate personalized ambiguity adaptation as a new task: given a user's previously resolved coding sessions and a new ambiguous request, an assistant should identify the recurring ambiguity pattern, produce the intended executable solution, and minimize clarification. To benchmark this task, we introduce CAPA, which characterizes personalized coding ambiguity through six mechanisms and injects thes

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

(Im)Paired Programming: Coding Agents Improve Productivity but Harm Understanding

arXiv:2607.26375v1 Announce Type: cross Abstract: Coding agents (e.g., Cursor) improve developer productivity by optimizing task completion, but shifting users from writing code to prompting and reviewing may harm their understanding, impeding oversight, learning, and communication. To probe this, we have 54 students create a website with one of two AI systems: an agent that edits user code; or a chatbot where users write code alone or adapt generic code snippets. We test understanding via comprehension questions and a task where users extend their code without agents, showing: (1) While agents aid initial task completion, they harm users' code comprehension and thus do not prepare users to extend their code; (2) Low-effort agent interaction types, like copy+paste prompts and auto-accepted edits, are linked with lower comprehension; and (3) Despite self-reported weaker understanding, users still prefer coding agents because they are quick and easy to use. While users stay in the loop f

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

AgentGUI: An Interface for Observing and Steering Long-Running AI Agents

arXiv:2607.26300v1 Announce Type: cross Abstract: AI agents are increasingly adept at tackling complex, long-running tasks. With the rapid surge of autonomous capabilities, human oversight is systematically lagging behind due to limited human-centered interfacing. Aiming to address this, we introduce AgentGUI, a user-friendly, locally hosted GUI for seamlessly observing and steering AI agents amid multiple concurrent, long-running sessions. AgentGUI features 1) rich agent trajectory visualizations, 2) effective manual and automated steering, and 3) integration with and coordination between open-source and frontier agent frameworks. A controlled user study demonstrates statistically significant reduction in the time it takes to identify key elements from agent traces (38% faster, p = 0.023). In a preliminary experiment, AgentGUI's automated drift prevention feature raises the task completion rate of small local agents by as high as 34pp across a 0.8B--9B model ladder (N=50 runs per mode

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

FloDR: An invertible dimensionality reduction method based on a normalising flow

arXiv:2607.26278v1 Announce Type: cross Abstract: It is common for two-dimensional embeddings of high-dimensional data to be read far beyond what they can support. Distances in and between clusters, the meaning behind empty spaces, and the amount of structure hidden at each point are generally invisible in the output of methods such as t-SNE and UMAP. This is because the information that could support the meaning of these properties is discarded during the optimisation process. Here, we present FloDR, a dimensionality reduction method that embeds data through an invertible normalising flow. While FloDR only uses the first two output coordinates to create a two-dimensional embedding, it retains the remaining coordinates rather than discarding them. In addition to the embedding, an exact inverse and an exact density are properties of a trained mapping, which enable diagnostic visualisations that are computed from the exact inverse of the model that drew the layout rather than from an app

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

The Attention-Directing Ability of Teams

arXiv:2607.26109v1 Announce Type: cross Abstract: Why do some teams consistently mobilize collective effort and achieve superior performance while others struggle to coordinate action? We introduce Attention-Directing Ability (ADA), a latent team capability capturing how effectively members' interaction signals elicit engagement and coordinated responses from others. Extending the attention-based view, we conceptualize attention direction as an emergent coordination capability embedded in patterns of interaction rather than as a cognitive state or an outcome. Teams differ in the extent to which attention-directing signals trigger collective responses, and these differences shape how teams mobilize effort and perform. We examine ADA in a distributed innovation effort involving 2,233 participants collaborating asynchronously in 79 self-organized teams across 165 public Slack channels, generating over 30,000 messages. We model the causal responsiveness among interaction signals and derive

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

Reproducibility in Recommender Systems: A Survey

arXiv:2607.26074v1 Announce Type: cross Abstract: Reproducibility has become a cornerstone of credible recommender systems research, driven by growing concerns about the reliability and generalizability of experimental results. In response, the ACM RecSys conference introduced a dedicated Reproducibility Track in 2020 to encourage rigorous, transparent, and repeatable research. This paper presents a structured analysis of the track from 2020 to 2025, covering 51 accepted papers. We classify contributions by type and analyze common patterns in datasets, algorithms, frameworks, and evaluation practices, with the goal of understanding how reproducibility is operationalized in practice within the community. Our findings reveal three main trends. First, the track has expanded in scope, evolving from a focus on reproduction and replication to include benchmarking, resources, and methodological contributions. Second, reproducibility papers exhibit a consistent methodological profile, relying

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

GraphQAG: A Knowledge-Graph-Guided Visual Analytics Framework for Question-Answer Pairs Generation

arXiv:2607.27182v1 Announce Type: new Abstract: Question-answer (QA) pairs are widely used in knowledge base construction, question-answering systems, and the post-training of large language models (LLMs). However, important knowledge in long documents is often distributed across multiple paragraphs and connected through complex entity relationships. Such fragmented and relational knowledge poses substantial challenges for existing QA generation methods, which often fail to adequately cover core document content, cross-paragraph semantic connections, and multi-entity relationships. We present GraphQAG, a knowledge graph-guided visual analytics framework for generating high-quality QA pairs from long documents. GraphQAG follows a three-stage workflow. First, it constructs a document knowledge graph by segmenting the document into paragraphs and extracting salient entities and relations. Second, it builds a graph-based generation space from entities, relations, and multi-hop paths to con

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

TactiPlay: Multi-Granularity Tactical Parsing and Video-Anchored Match Review for Amateur Badminton Players

arXiv:2607.27125v1 Announce Type: new Abstract: Amateur badminton players increasingly record matches, yet existing tools provide only aggregate statistics or generic summaries, leaving most unable to extract tactical insights without expert guidance. A formative study (N=8) reveals the need for multi-granularity, video-anchored tactical analysis centered on rallies. We derive a taxonomy of performance issues from national-level athletes' annotations and present TactiPlay, an interactive system that instantiates an expert-taxonomy-guided, rally-level, video-anchored review workflow. The system's analytical pipeline organizes match events into taxonomy-grounded feedback, while its interface links structured reports to rally summaries, video evidence, and court visualizations. A within-subjects study (N=16) shows that TactiPlay elicits more frequent, concrete, actionable, and appropriate reflections than a report-and-statistics baseline. These findings show how organizing reviewed match

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

AI as Friction for Reflection Support in Ideation

arXiv:2607.26827v1 Announce Type: new Abstract: Generative AI tools for creative work tend to be designed around the goal of removing friction, on the assumption that smoother iteration and faster output translate into more value for the designer. We argue, however, that this framing leaves out something important about how design ideation works, namely reflection-in-action. The act of accepting, rejecting and reworking candidate ideas is both a path to a final outcome and the process through which designers develop the rationale that allows them to think with their ideas and to communicate them to others. This becomes particularly important in group ideation, where ideas need to be expressed and explained to others to allow the group to extend, reject or combine them further. We suggest that AI in design ideation might be more usefully thought of as a friction agent for reflection rather than as a smoothing agent for output. This reframing opens up a different role for AI in design id

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

Affective Tools for Thought: Towards Shared Attention and Affective Reorienting in AI-Supported Thinking

arXiv:2607.26731v1 Announce Type: new Abstract: Current Tools for Thought (TfTs) treat affect as either friction that slows cognitive progress or a signal to optimise it. Drawing on enactive cognitive science, we argue that affect is constitutive of cognition: it reshapes the trajectory of thinking, not just the speed. We identify two core barriers for Affective TfTs: the lack of Shared Attention (caring, directed attention to the user's mode of engagement) and the lack of Affective Reorienting (the capacity to use emotional moments to open new trajectories rather than reinforcing predetermined ones), and propose three design strategies that address both: Chain of Emotion X Chain of Thought, Affective Mirror, and Prompted Reorienting. The strategies are grounded in empirical findings from a study of a touch-aware conversational agent for embodied craft learning, and are oriented as provocations for future design.

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

Galvanic Vestibular Stimulation in Latent Space

arXiv:2607.26659v1 Announce Type: new Abstract: Galvanic vestibular stimulation (GVS) is widely used to modulate self-orientation, balance, and motion perception; the discriminability of frequency-encoded cues further suggests its potential as a standalone modality for embodied feedback. However, synthesizing GVS waveforms congruent with target events or bodily states remains challenging. GVS waveforms combine current direction, intensity, duration, and onset and offset transitions, yet how these parameters jointly shape users' perceptual and associative responses remains underexplored. To address this gap, we contribute a dataset linking GVS waveforms to free-form experience descriptions, as well as a retrieval-guided generative model for synthesizing candidate waveforms from target descriptions. The dataset comprises 100 GVS waveforms and 1,526 valid free-form sensation descriptions collected from 16 participants. Semantic analysis revealed diverse motion- and force-related sensation

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

A Design Study on Voice-based Interaction for Immersive Network Visualization and Analysis

arXiv:2607.26526v1 Announce Type: new Abstract: Visual network analysis leverages network visualization authoring techniques to facilitate sensemaking, serendipitous discovery, and hypothesis verification on network data. However, transferring the same paradigm to immersive environments is non-trivial due to insufficient UI affordance for authoring operations. Researchers have studied combining multiple modalities for interactions, but the high learning curve of such input systems limits their adoption by typical data analysts, let alone for network analytics. In this work, we investigate the advantages and limitations of voice as the primary input modality with a research-through-design (RtD) study, in which we design a system that supports voice-based interactions for immersive network visualization facilitated by Large Language Models (LLMs). Through a user study on social network data analysis with participants from social science and computer science backgrounds, we find that voic

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

Repair as Representational Work: Integration Bottlenecks in AI-Assisted Development

arXiv:2607.26517v1 Announce Type: new Abstract: Research on AI-assisted programming has concentrated on the gulf of execution -- how users write successful prompts. We report a candidate phenomenon, an integration bottleneck, that lies in Norman's gulf of evaluation: a repair-relevant contribution reaches the user and fails to become actionable at the point of receipt. Two cases in an eighteen-case corpus of publicly shared AI-assisted-development accounts report this, from a peer and from the system's own output; both fall at evaluation's interpretation stage, and a third, which would fall at comparison, is reached only on an inferential reading and reported as a boundary case. A within-case contrast is consistent with actionability turning on whether the contribution can be restated as an instruction without an intervening judgement. We report this as a candidate warranting dedicated study, not an established regularity; its evidence base is retrospective author self-reports. On the

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

How Do Researchers Manage Visualization Experiment Stimuli?

arXiv:2607.26443v1 Announce Type: new Abstract: Visualization experiments need a set of "good" stimuli that effectively address research questions and hypotheses. Creating, managing, and deploying stimuli are often challenging, as these tasks require tremendous care. Inappropriate stimuli can make the outcome invalid or uninteresting, wasting both researchers' and participants' resources. As the speed of science increases, better support for stimuli-related tasks is essential, yet we lack a closer look at how visualization researchers deal with them. To understand the experiences of visualization experimenters and guide future improvements, we interviewed 19 visualization researchers with diverse backgrounds and experiences. Our findings describe practices and challenges across the life cycle of stimuli, from exploration and selection through shipment, deployment, and analysis. For example, stimuli management and deployment require tedious manual effort, which does not scale for experi

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

FleetScape: A Mixed Reality Sandtable for Spatial Supervision and Control of Scalable Drone Fleets

arXiv:2607.26423v1 Announce Type: new Abstract: As autonomous drone deployments scale from individual units to coordinated swarms, the human operator's role shifts from direct piloting to high-level supervision. Current interfaces often treat multi-drone control as a scaled-up version of single-drone operation. We instead investigate how reframing fleet supervision as spatial interaction can better support the spatial, temporal, and safety demands of complex missions. We present FleetScape, a Mixed Reality (MR) sandtable system that externalizes layered real-time mission, safety, and environmental data while enabling fluid transitions between manual intervention and autonomous supervision. We developed a high-fidelity building inspection simulation that generates and streams synchronized multi-drone and environmental data for MR visualizations. We used this prototype to conduct a user study with six experienced drone pilots managing fleets of up to 15 drones. Our findings show that Fle

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

Sensor-Placement-Agnostic Sonomyography: Toward Continuous High-Dimensional Control by Users with Tetraplegia

arXiv:2607.26401v1 Announce Type: new Abstract: Sonomyography (SMG) enables continuous device control via ultrasound-measured muscle deformation signals, but existing SMG interfaces generally require substantial user- and sensor-location-specific training data and provide only one proportional signal or task-specific classification. We present a real-time, sensor-placement-agnostic SMG control system based on sparse optical flow tracking that enables continuous 1-DOF control after minimal calibration (3 pose definitions). We also present a preliminary expansion of this method that augments this algorithm with a short computer-aided calibration to enable 2-DOF control. We evaluate both 1- and 2-DOF systems' performance for a preliminary cohort of 3 cervical spinal cord injury survivors and 6 uninjured individuals across 6 sensor placements spanning the arm, neck, and upper torso. As assessed by a cursor trajectory tracking task, all participants achieved continuous 1-DOF control at all

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

Designing Needs- and Attention-Aware AI Learning Tools for Engineering Education: Insights from Psychological Outcomes

arXiv:2607.26338v1 Announce Type: new Abstract: Artificial Intelligence (AI) is transforming higher education, but its benefits can vary depending on where, how, and how often it supports learning. While prior research emphasizes cognitive and academic outcomes, this study examines how AI chatbots support the psychological needs and motivational states of engineering students. A survey of college engineering students (n = 206) examined perceived effects of AI chatbots on autonomy, relatedness, and relief from competence frustration. Structural equation modeling with latent interaction effects examined how baseline autonomy, competence frustration, relatedness, and personal agency contributed to perceived AI outcomes. Results indicate that students perceived that AI provided the greatest benefits as relief from competence frustration, smaller benefits for autonomy, and the weakest benefits for relatedness. Baseline motivational states mattered more than demographic factors, and inattent

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

Pragmatic Reasoning in Design

arXiv:2607.26322v1 Announce Type: new Abstract: People can often understand and use novel artifacts after only a few interactions, suggesting that design choices communicate underlying affordances and causal structure. We propose a formal account of this process by framing cooperative, user-centered design as a cooperative game in which the user is the principal and the designer is an assistant. Inspired by prior work on pragmatic communication (e.g. RSA), our model treats a designer's design decisions as communicative signals and predicts user judgments via recursive mentalizing: designers make design decisions to trade off informativeness about the artifact with efficiency, and users infer the true model of the artifact by inverting this cooperative designer model. We evaluate the model in a design game where designers place visually identical keys on trays to help a user infer which keys unlock which doors in grid-world layouts. We find that pragmatic designer and user models better

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

User-Reported Misinformation Exposure Across Social Media Platforms

arXiv:2607.26218v1 Announce Type: new Abstract: In this study, we surveyed users for their perception of misinformation exposure across social media platforms. Such perceived exposure is important because individuals' beliefs about how often they encounter false information can shape their trust in institutions, platforms, and even their friends. In a survey of 1,010 United States residents, we found that perceived exposure to misinformation varies substantially across platforms and is only moderately correlated with the frequency of platform use. A much larger percentage of participants also reported being exposed to misinformation from the public feed than from known contacts. Based on these results, we propose governance strategies across three categories of platform types: discovery, interpersonal, and discourse. This work offers insight into users' perceptions of social media misinformation and a corresponding research agenda for governance.

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

Reading Between the Curly Braces: On Textual Data Serialization Format Usability

arXiv:2607.26211v1 Announce Type: new Abstract: Textual data serialization formats, such as JSON or XML, are ubiquitous, supporting tasks like software configuration and data tabularization. Despite their prominence, little is known about their usability. What makes one good or bad? Is there a best one for cognitive efficiency? We explore these questions via a (N=215) crowd work study and a (N=9) semi-structured interview study. We find that format distinctions (like indentation versus curly braces) do not consistently translate into substantial usability differences. While HJSON and YAML performed better than other formats in certain modification tasks, these advantages disappeared in more realistic settings where task complexity was either trivial or highly demanding. Instead, usability appears driven by sociotechnical ecosystems: the tooling, documentation, and community practices surrounding a format matter more than syntax.

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

How Wrangling Tools Shape Wrangling: A Technical Dimensions Analysis

arXiv:2607.26198v1 Announce Type: new Abstract: Wrangling consumes a disproportionate share of the effort associated with any data project. While a variety of tools support it, relatively little is known about how their differing interface forms shape the way people actually wrangle. We conduct a between-subjects (N=40) observational study of data cleaning tasks performed in tools spanning distinct interface paradigms: Jupyter (notebook), Excel (spreadsheet), ChatGPT (conversational AI), and OpenRefine (visual wranglers). We situate our observations within the Technical Dimensions of Programming Systems framework, which we use as a conceptual scaffold for comparing across interface paradigms. Within the context of our study, the results suggest that tool affordances steer user strategies but do not determine outcomes. There is no consistent advantage of any single tool, nor convergence of results within tools observed across our outcome measures. Instead, we identify trade-offs and con

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.HC

Hakka Kitchen: Engagement with Culinary Cultural Heritage Through Immersive Game Play

arXiv:2607.26183v1 Announce Type: new Abstract: Intangible Cultural Heritage (ICH) experiences are difficult to share with the public because they are essentially processes that rely on physical interactions with embodied, tacit, and situated phenomena in specific cultural contexts. We consume non-interactive media such as videos and books to learn about culinary ICH experiences, but they do not allow us to grasp actual interactive procedures that embody the cultural knowledge. To engage people in a traditional cooking experience, we created a gamified VR experience Hakka Kitchen, where players are guided by a chef of Hakka cuisine through a modeled physical process of making the traditional dish of stuffed bitter melon. Compared against watching a video in VR providing the same information in a between-subjects study (N=40), Hakka Kitchen led to increased sensory, imaginative engagement, positive affect, and willingness to transmit awareness for the culinary ICH. Heritage recognition

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

A Simulator-Grounded Framework For Constructing Verifiable Muscle-Grounded QA From 3D Tongue Meshes (extended version)

arXiv:2608.23137v2 Announce Type: replace-cross Abstract: Existing articulatory corpora based on real-time MRI and electromagnetic articulography capture tongue shape and motion but do not provide traceable labels for the muscle-driven process that generated an observed configuration. We introduce a simulator-grounded data-construction framework and instantiate it as 3DTongueQA. Controlled 11-dimensional muscle activations are mapped to fixed-topology tongue meshes with the ArtiSynth Badin finite-element model, converted into structured biomechanical records, and rendered as deterministic QA on muscle state, geometry, and target-directed change. We screen 295,157 configurations, retain 295,115 valid meshes, and construct 891,156 QA records per language. Language naturalization changes only surface form and is verified against the source records; English and Korean instantiations demonstrate construction-level portability. A swappable SpiralNet++--Qwen3-8B baseline reaches 62.9 $\pm$ 9.

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Learning to Prompt: Improving Student Engagement with Adaptive LLM-based High-School Tutoring

arXiv:2606.20138v2 Announce Type: replace-cross Abstract: LLMs can personalize education, although current static-prompt tutoring systems struggle to adapt to diverse academic disciplines. We develop and test a system with subject-aware prompting, based on 14 pedagogical features (e.g., tutor scaffolding, student understanding) extracted from raw transcripts. We first train a prompt routing model in a simulation environment, and then deploy it for online adaptation with actual high-school students. The simulation benchmark shows the router outperforming two static baselines ($0.694$ vs. $0.647$ and $0.64$, $p<0.001$). A/B testing ($N=656$ conversations from 359 students) shows sim-to-real transfer where the model switches from analytical to scaffolding learning strategies. Our adaptive prompt selection mechanism improves instructional efficiency, maintains pedagogical quality and reduces interactions by around 3 turns ($p=0.007$). While a greedy router achieves a comparable exercise co

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Predicting Time Pressure of Powered Two-Wheeler Riders for Proactive Safety Interventions

arXiv:2601.03173v4 Announce Type: replace-cross Abstract: Time pressure critically influences risky maneuvers and crash proneness among powered two-wheeler riders, yet its prediction remains underexplored in intelligent transportation systems. To address this gap, we propose MotoTimePressure (MTPS), a deep learning model combining convolutional preprocessing, dual-stage temporal attention, and Squeeze-and-Excitation feature recalibration, achieving 91.53% accuracy and 98.93% ROC AUC, outperforming six baselines, with only 172K parameters, 0.66 MB model size, and 0.21 ms inference on CPU. To validate and benchmark MTPS, we present a dataset of 129,209 feature windows from 153 simulator sessions by 51 experienced male PTW riders under No, Low, and High Time Pressure conditions. Each sequence captures 63 features spanning vehicle kinematics, control inputs, behavioral violations, and environmental context. Our empirical analysis shows High Time Pressure induces 48% higher speeds, 36.4% gr

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Helping the Helper: LLM-Assisted Problem Articulation for Older Adults Seeking Technology Support

arXiv:2601.10018v2 Announce Type: replace Abstract: Older adults often struggle to articulate technology support needs due to unfamiliar technical terminology and age-related cognitive changes. We explore how large language models (LLMs) can facilitate this problem articulation process. Through a diary study (n = 27), we identified four communication barriers in older adults' queries: verbosity, incompleteness, over-specification, and under-specification. To mitigate these barriers, we developed an LLM pipeline that clarifies context and paraphrases unstructured queries. LLM-rephrased queries significantly improved automated solution accuracy (69% vs. 35%). Furthermore, younger adults (n = 48) acting as technology helpers understood LLM-rephrased queries better (93.7% vs. 65.8%) and reported greater ease in providing support. Older adults (n = 34) also found the resulting solutions highly actionable (94.7%). Finally, we contribute the first synthetic dataset of older adults' technology

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Automated Healthcare Thematic Analysis using Multi-Agent Large Language Model: Algorithm Development and Evaluation

arXiv:2512.16063v2 Announce Type: replace Abstract: Understanding patients experiences is essential for advancing patient-centered care. Qualitative thematic analysis is widely used to explore these experiences, however, the process remains labor-intensive, subjective, and difficult to scale. This study aimed to develop and evaluate Collaborative Theme Identification Agent (CoTI), a multi-agent large language model framework designed to support manual thematic analysis by rapidly generating supporting excerpts, initial codes, and themes. CoTI consists of three agents: Instructor, Thematizer, and CodebookGenerator. The Instructor refines instruction prompts, the Thematizer extracts supporting excerpts and generates initial codes for each transcript, and the CodebookGenerator groups similar codes across all transcripts into a codebook with themes. We evaluated CoTI primarily using 12 heart failure patient transcripts. CoTI-generated outputs were compared against the reference standard de

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Human learning is an understudied but promising lever for boosting human--AI synergy

arXiv:2512.13253v3 Announce Type: replace Abstract: Humans collaborating with artificial intelligence (AI) hold the promise of achieving superior outcomes compared to either acting alone (i.e., human--AI synergy). However, the conditions that facilitate such synergy when humans are advised by AI are not well understood. A recent meta-analysis showed that, on average, human--AI combinations do not outperform the better individual agent. We argue that this pessimistic conclusion arises from insufficient attention to human learning in experimental designs. To substantiate this claim, we re-analyzed all 74 studies included in the original meta-analysis and found that most previous research overlooked design features that foster human learning (e.g., outcome feedback to participants). Our re-analysis further revealed that studies providing outcome feedback show tentatively higher synergy than those without outcome feedback. Crucially, feedback paired with AI explanations was associated with

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

When Do Reactive Notebooks Fail to React?

arXiv:2511.21994v2 Announce Type: replace Abstract: Computational notebooks are convenient for programmers, but can easily become confusing and inconsistent due to the ability to incrementally edit a program that is running. Recent reactive notebook systems, such as Ipyflow, Marimo and Observable, strive to keep notebook state in sync with the current cell code by re-executing a minimal set of cells upon modification. However, each system defines reactivity a different way. Additionally, within any definition, we find simple notebook modifications that can break each system. Overall, these inconsistencies make it difficult for users to construct a mental model of their reactive notebook's implementation. This paper proposes Rex, a fine-grained test suite to discuss and assess reactivity capabilities within reactive notebook systems. We evaluate Rex on three existing reactive notebook systems and classify their failures with the aims of (i) helping programmers understand when reactivity

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

People readily follow personal advice from AI but it does not improve their well-being

arXiv:2511.15352v4 Announce Type: replace Abstract: People increasingly seek personal advice from large language models (LLMs), yet whether humans follow their advice, and its consequences for their well-being, remains unknown. In a longitudinal randomised controlled trial with a representative UK sample (N = 6,474), we found that up to 79% of participants who had a 20-minute discussion with one of three AI chatbots (GPT-4o, LLama-3.3-70B, Gemini 3 Pro) about health, careers or relationships subsequently reported following its advice. Advice-following remained above 65% even for high-stakes recommendations, suggesting that users only weakly calibrate their reliance on AI advice to potential consequences. Based on autograder evaluations of chat transcripts, LLM advice rarely violated safety best practice. However, when queried 2-3 weeks later, participants receiving personal advice from AI showed no sustained well-being benefits compared to a control group who discussed hobbies and inte

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching

arXiv:2608.26094v1 Announce Type: cross Abstract: Existing action quality assessment (AQA) datasets and methods rely primarily on visual inputs such as RGB and pose, overlooking physiological dynamics such as muscle mechanics and often modeling actions as monolithic patterns. These limitations hinder fine-grained, biomechanically grounded feedback. We introduce MyoMechanix, a multimodal ecosystem for weight-loaded actions that aligns motion with muscle activity. Expert-annotated, it contains 7,500+ samples of 20 actions from 38 subjects, with synchronized multiview RGB video, 3D pose, sEMG, and additional physiological signals, forming the largest multimodal AQA benchmark to date. We further construct the Fitness Knowledge Graph (FKG), which organizes expert annotations into structured relationships among actions, phases, key steps, errors, and corrective feedback, enabling compositional scoring and interpretable assessment. Building on these representations, we develop CUBIST (Composi

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Simultaneous Digital Communication and Deformation Sensing over a Single Stretchable Interconnect

arXiv:2608.25801v1 Announce Type: cross Abstract: Stretchable hybrid electronics integrate rigid solid-state electronics with stretchable materials and structures to achieve both high deformability and stable electronic performance. However, most existing systems treat stretchability only as a mechanical attribute without exploiting device deformation to encode its own mechanical state. This problem arises from adapting conventional rigid circuit architectures to stretchable substrates, affording a loss in compatibility with the sensors required for strain measurement. This study addresses this issue by proposing a communication-integrated deformation sensing architecture for stretchable hybrid devices. In the proposed approach, standard universal asynchronous receiver-transmitter digital signals transmitted between rigid nodes are amplitude-modulated by strain-induced resistance changes in stretchable liquid metal interconnects. By reading both amplitude changes and digital patterns,

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Development of a Voice-Controlled Tendon-Driven Bionic Hand

arXiv:2608.25222v1 Announce Type: cross Abstract: The impairment of the hands can seriously affect the abilities of every individual to perform the every-day activity, so the design of stable and controllable support devices is a significant field of study. This paper is about the design and implementation of an automated bionic hand which is dedicated to the coordinated finger movement through the simplified and efficient actuation mechanism. The method that the proposed system was designed on is the tendon-based method whereby the servo motors generate the movement of the fingers, with assistance of the angular control which is calibrated. An actuation is controlled by a microcontroller that will be programmed by use of an Arduino-based microcontroller to carry out programmed gestures that include open hand, fist, pinch and half flexion. It has an interface that is voice command enabled to make it easy to interact with a Bluetooth based sender receiver architecture which offers an op

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Longitudinal Robot Learning from Demonstration with Care Providers in a Home Environment

arXiv:2608.25196v1 Announce Type: cross Abstract: Learning from demonstration (LfD) methods enable non-expert end users to teach robots novel skills without explicit programming. However most evaluations of the usability of LfD with non-experts has been conducted in controlled laboratory environments with a robotics experimenter present. In this work we identify non-expert end users' key barriers when teaching robots via demonstration without live robotics expert feedback in a home environment. In our human subjects experiment we support the non-expert end users through two forms of demonstrator guidance developed in prior work: pre-training and adaptive feedback. Towards the ecological validity of the evaluation, we conduct this experimentation over multiple visits, with a population of care providers. Finally, we propose to open source the resulting LfD dataset of care providers teaching a robot assistive tasks over multiple visits to a home environment.

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Semantic Variability of Replies Across LLMs: Implications for Designing Conversation-Based Assessment

arXiv:2608.24920v1 Announce Type: cross Abstract: This study examines whether LLM-generated replies remain semantically consistent when the underlying LLM changes. Using messages from real collaborative conversations, we compared the semantic similarity of generated replies across LLMs under two conditions: with and without preceding chat history. Results show that model choice and conversational context both affect response similarity and alignment with human replies. These findings indicate that prompting and conversational context alone may not be sufficient to preserve response consistency across LLMs, highlighting the need for infrastructure and design strategies that can maintain stable and comparable responses amid the rapid and continuous evolution of LLMs.

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Dynamic Influence-Weighted Distillation for Single-IMU Activity Recognition

arXiv:2608.24904v1 Announce Type: cross Abstract: Inertial sensors at multiple body locations can improve activity recognition, but requiring every sensor at inference increases the deployment burden. We study whether four synchronized IMUs available during training can improve a student that uses only the right-arm IMU during fitting and inference. A frozen four-IMU teacher provides logit and feature targets. Fixed-weight knowledge distillation applies each target with the same strength to every fitting sample, although the student may not benefit equally from them. We introduce dynamic influence weighting (DIW), which tests a one-step candidate update on separate fold-internal training participants. DIW then assigns separate sample-wise gates to the logit and feature losses. On WEAR, we evaluate 19 labels and 68,298 complete windows from 22 participants using subject-disjoint five-fold cross-validation. Pooled out-of-fold macro-F1 is 0.561820 for Supervised and 0.571623 for Fixed-wei

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores

arXiv:2608.24901v1 Announce Type: cross Abstract: A decodable "empathy" direction is routinely read as a causal lever, conflating decodability, automated-metric control, and human-perceived change. We test this for two EPITOME-derived facets -- Recognition (cognitive) and Resonance (affective) -- in three instruction-tuned LLMs, scoring every intervention with two LLM judges and a discriminative EPITOME classifier, each gated by an emotional-vs-neutral positive control. The control passes for the affective facet across all automated instruments, but cognitive range is inconsistent across them. Both facets remain decodable after residualizing against a sentence-embedding-derived surface score, and steering can substantially rewrite the text. Yet adding the Resonance direction raises the affective score only partially -- in Qwen by +0.29 (approximately 26% of the natural gap). A direct between-direction contrast confirms the shift is facet-specific in Qwen and Llama (not Gemma); we do no

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.HC

From Producing to Validating: How AI Is Deskilling Freelancers

arXiv:2608.26089v1 Announce Type: new Abstract: Generative AI is promoted as a way to enhance knowledge work, yet its benefits and drawbacks fall unevenly across the workforce. Freelance and gig workers, who commonly lack the upskilling pathways available to traditional employees, face heightened risks to both skill development and job security as AI adoption advances. We review empirical evidence on AI's impact on knowledge-worker workflows and upskilling, then predict the primary and downstream effects of AI adoption among clients and workers in the freelance economy. We anchor this in two cases of the same shift, machine-translation post-editing and software development. We argue that freelancers are the leading edge of a change that also reaches salaried HCI practitioners, and we close with questions for the platforms and clients that mediate this work, and for HCI researchers.

Source ↗
Showing 851–900 of 1631 signals
← Prev Page 18 of 33 Next →