EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.HC

From Agent Behaviour to Agent-Friendly Documentation: An Empirical Study of How Coding Agents Discover, Read, and Write Technical Documentation

arXiv:2608.20195v1 Announce Type: cross Abstract: Technical documentation is written for human developers, but an increasing share of software changes is now authored by autonomous coding agents. Which documents they consult, when, and what follows remain unknown. We conduct a behaviour-grounded study of agent-documentation interaction across two public datasets: 557 agentic coding sessions from SWE-chat, yielding 94,813 development events including 3,033 documentation interactions; and 33,097 agentic pull requests from AIDev, with 690,260 classified file-level change records. Four findings challenge current documentation practice. First, agents' documentation work is dominated by agent-facing artefacts: instruction files and working notes account for 60.5% of all documentation interactions, versus 10.6% for classical technical documentation and 1.3% for API references. Second, the link between consultation and code editing is unresolved: the adjacent transition probability is 0.002 an

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.HC

Generating Diverse Personas for User Simulators to Test Interview Dialogue Systems

arXiv:2608.19549v1 Announce Type: cross Abstract: This paper addresses the issue of the significant labor required to test interview dialogue systems. While interview dialogue systems are expected to be useful in various scenarios, like other dialogue systems, testing them with human users requires significant effort and cost. Therefore, testing with user simulators can be beneficial. Since most conventional user simulators have been primarily designed for training task-oriented dialogue systems, little attention has been paid to the personas of the simulated users. During development, testing interview dialogue systems requires simulating a wide range of user behaviors, but manually creating a large number of personas is labor-intensive. We propose a method that automatically generates personas for user simulators using a large language model. Furthermore, by assigning personality traits related to communication styles when generating personas, we aim to increase the diversity of comm

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.HC

APPROVE: Visual End-User-in-the-Loop Robot Programming with LLMs

arXiv:2608.19281v1 Announce Type: cross Abstract: Programming robots remains challenging for non-experts, as traditional methods require expert knowledge and even block-based interfaces often lack flexibility. Recent work has explored Large Language Models (LLMs) to automatically generate robot programs from natural language, but these systems remain limited by a lack of transparency, missing mechanisms to ensure alignment with user intent, and little support for reuse. We present APPROVE (AI-Powered Programming for Robots with Visual End-User Feedback), an LLM-based multi-modal end-user programming framework that integrates natural language input with a block-based interface and an explicit user confirmation step. Generated programs are visualized using a block-based interface in Blockly, allowing users to confirm, modify, or reject them before execution. Confirmed functions are stored in a library for reuse, gradually building a set of reliable program components. Our approach contri

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.HC

Navigating and Retrieving Information in Immersive Model-Based Design Reviews: An Exploratory Study

arXiv:2608.20128v1 Announce Type: new Abstract: Digital engineering uses many models from different perspectives, creating a connected set of digital artefacts across a product's life cycle. Designers seeking a holistic view must navigate numerous models and views, requiring domain-specific software, languages, and representations. This can lead to getting lost in scattered information and the cognitive burden of mentally integrating details across diagrams. To overcome these issues, we developed the virtual environment GraphXplore. GraphXplore enhances perceptual and conceptual integration by linking all relevant visual items from different perspectives into an interactive, layered 3D graph displayed in virtual reality, providing a holistic view of the system. We compared GraphXplore with a conventional on-screen setup using a PowerPoint slide deck with model screenshots viewed on a desktop PC. In an experiment with N=33 volunteers (mainly industrial product design postgraduates and p

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.HC

What Do Visualization Instructors Want Students to Learn? Introducing a Concept Inventory for Visualization Design

arXiv:2608.20090v1 Announce Type: new Abstract: The term "visualization design" encompasses multiple concepts and skills that go well beyond current assessments of graphical perception and visualization literacy. In the context of education, what exactly should a student be able to do if they "know" visualization design? To answer this question, we draw on existing methodology from the field of education to propose a concept inventory for visualization design, i.e., a theoretical model capturing the most important concepts and skills commonly associated with visualization design. We initially draft the concept inventory using a qualitative analysis of course objectives from visualization course syllabi. Then, we iteratively refine the concept inventory by soliciting feedback from instructors through semi-structured interviews. Based on our experiences in developing the concept inventory, we reflect on open questions and future research directions in visualization education, such as dev

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.HC

Evaluating Smart Home Device User Responses to their (Un)Confirmed Privacy Expectations

arXiv:2608.19873v1 Announce Type: new Abstract: Users of smart home devices are often unaware of how their devices handle personal data. We examine how revealing these data practices influences user trust, satisfaction, and coping behaviors, including decisions to block device communications. Using Expectation-Confirmation Theory, we conducted two complementary studies to balance ecological validity with experimental control. An in-situ field study used network monitoring to reveal actual device traffic, and an online experiment presented simulated reports with manipulated levels of advertising-related communications. Across both studies, when data practices aligned with user expectations, satisfaction increased, strengthening intentions to continue using the device. Defensive responses, however, followed different pathways: satisfaction predicted willingness to block in the in-situ field study, whereas collection concerns were the primary predictor of blocking in the experiment. Toget

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.HC

Dancing Through Soundscapes: Designing a Low-Cost, Sound-Based Device for Sensing and Interpreting Movement and Dance

arXiv:2608.19827v1 Announce Type: new Abstract: When we move through space, we often rely on multiple senses beyond vision to perceive and act in that environment: we ``feel'' the presence of others; we build internal representations and models and recall them to navigate the environment. We also leave traces and impressions that others pick up on. The traces include echoes, heat, the displacement of objects such as furniture or footprints, air movement close to the face of another, smells such as perfume, but also the immediate sounds we make when we move and breathe. Movement is a spatial and temporal activity, and dance as a form of movement practice requires coordination of oneself in relation to others, the space and a potential score. When rehearsing dance, dancers have to relate to others often not just by looking but more often by feeling and imagining or remembering where others are based on experience and shared practice. So how can we approach technology-mediated movement an

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.HC

Calming Robot Pitches? Exploring the Influence of Robot Voice Pitch on Children's Stress Levels

arXiv:2608.19826v1 Announce Type: new Abstract: This study examined whether variations in robot speech pitch influence children's stress levels during a robot-guided game. Although lower-pitched voices have been shown to facilitate stress regulation in human communication, it remains unclear whether this effect generalizes to synthetic voices in child-robot interactions. Twenty-seven Dutch children aged 8-12 years were randomly assigned to interact with a Zenbo Junior II robot using either a lower-pitched or a higher-pitched voice. The interaction consisted of an introduction followed by a timed LEGO-building game. Stress levels, measured with an adapted version of CAM-S, increased during the game, confirming the stress-inducing nature of the task. No differences emerged between pitch conditions. These findings suggest that the benefits of lower pitch in reducing stress may not directly translate to child-robot interactions. Possible explanations include children's developing sensitivi

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.HC

Grounding Mindfulness in Embodied Tangibles: A Scoping Review & Theoretical Framework for HCI Design

arXiv:2608.19673v1 Announce Type: new Abstract: Embodied and tangible devices are increasingly used to support mindfulness practices across meditation, yoga, and everyday routines. However, existing HCI research lacks a coherent theoretical foundation for explaining how such systems support distinct mindfulness processes and outcomes. First, we report findings from a scoping review of tangible devices (n=65) for mindfulness in HCI based on the mechanisms of action of mindfulness. The review found that most systems primarily target attentional regulation and body awareness, while emotion regulation and change in perspective on the self remain comparatively underexplored. Also, the evaluation methods used to assess the effectiveness of tangible systems for mindfulness-related outcomes were found to be fragmented and weakly grounded in theory. Building on these findings, a theoretical framework grounded in the Self-Awareness, Self-Regulation, and Self-Transcendence (S-ART) framework is pr

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.HC

IRIS: Navigating and Reflecting on Writing Traces Using Intelligent Document Histories

arXiv:2608.19614v1 Announce Type: new Abstract: Much of the text produced throughout the lifetime of a document is impermanent. In this paper, we explore how writing activity traces can be made visible and interactive to help writers navigate their document histories and understand their writing processes. Using the Flower and Hayes cognitive process model of writing, IRIS infers writing process states from keystroke logs and presents them using an AI-enhanced version history. IRIS provides three primary interactions: revision highlighting that shows local process histories in-situ, conceptual filters that constrain the version history by process type or topic, and natural language inquiry that lets writers pose reflective questions about their writing and process. Following a formative and a longitudinal study, we find that writers use the interfaces to locate specific revisions and understand the progression of their writing. They use system outputs as interpretive material, relating

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.HC

Localized Ecological Momentary Assessment for Mental Health Research in China: An Implementation-Oriented Framework and Preliminary Case Application

arXiv:2608.19588v1 Announce Type: new Abstract: Background: Ecological momentary assessment (EMA) is increasingly used in mental health research, but research-grade deployment requires platforms supporting protocol configuration, automated delivery, participant management, and data export. In China, these requirements are not consistently supported. Objective: We aimed to identify workflow gaps affecting localized EMA deployment, develop an implementation-oriented framework for platform assessment, and assess Huixin EMAI. Methods: We reviewed EMA platforms reported in Chinese mental health studies in CNKI and Wanfang. A multidisciplinary panel of 6 experts developed the Multi-dimensional EMA Platform Evaluation Framework (MEPEF) and benchmarked 7 platforms across 43 indicators in 6 domains. MEPEF was then applied to Huixin EMAI using deployment logs from 48 participants, questionnaires from 44 participants, and semistructured interviews with 6 researchers. Results: We identified 66 emp

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.HC

Delegating or Doing? Understanding User Behavior in Hybrid Human-Agent Interfaces

arXiv:2608.19551v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly embedded into applications, allowing users to complete tasks either through direct manipulation or by delegating actions to conversational agents. However, little is known about how users balance these modalities when both are available. We present a web-based content management system augmented with an LLM agent through the Model Context Protocol (MCP), enabling users to perform CRUD tasks through a graphical interface, a conversational agent, or both. We conducted a between-subjects study (N=73) comparing three interaction modes: Traditional-Only, AI-First, and Hybrid. Across sixteen scenarios, we analyzed task completion time, interaction logs, and delegation behavior. AI-assisted interaction significantly reduced clicks, page navigations, and scrolling indicating lower interaction effort. Surprisingly, these reductions did not translate into faster task completion, as task duration did not

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.HC

Does Listening Matter? Backchanneling and Nodding in AI Clone

arXiv:2608.19527v1 Announce Type: new Abstract: AI clones that imitate a specific person typically reproduce what the person says and how they sound, but not how they listen. We investigate whether adding multimodal listening behaviors gives such a clone more presence and authenticity. We integrated verbal backchannels and head nodding, driven by real-time prediction models, into an AI clone equipped with voice cloning and LLM-based responses. In a within-subjects study (N=35), adding these behaviors significantly improved the perceived attentiveness of the avatar, the sense of talking with the real person, and the feeling of co-presence. These results indicate that AI clone fidelity should extend beyond voice and response content to include interactive listening behavior.

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.HC

Cyber-Physical Systems for Accessibility and Ability Augmentation: Bridging Diverse Communities

arXiv:2608.19422v1 Announce Type: new Abstract: The powerful convergence of wearables, robotics, extended reality, and smart environments is expanding the design space for cyber-physical systems (CPS) that support and augment human abilities in daily life. By sensing real-world contexts, modeling user needs, and providing situated assistance, these systems can improve accessibility for people with disabilities while enhancing broader human abilities such as perception, memory, learning, and mobility. However, realizing this potential requires addressing key challenges in context sensing, user modeling, adaptive interaction, privacy, and evaluation to ensure that CPS are reliable and effective in real-world contexts. This workshop will bring together researchers and practitioners across HCI, AI, wearables, robotics, XR, smart environments, accessibility, and ability augmentation to examine shared strategies and challenges for designing accessibility- and ability-centered CPS. Through pa

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.HC

Scientific Visualization as a Collaborative Data Infrastructure

arXiv:2608.19413v1 Announce Type: new Abstract: Scientific visualization is an active site of infrastructuring with many layers of data and evidential claims. This paper reflects on collaborative Mars geoscience research conducted at the NASA Jet Propulsion Lab, which produced the PIXLISE spectroscopic analysis platform, through a retrospective analysis of a scientific discovery made by the team using PIXLISE. The history of infrastructuring in scientific visualization is exceedingly rich, which presents an opportunity for archival research. Simultaneously, the field of visualization aspires to become a science of communication, opening the door to future collaborations. Finally, with regard to visualization practitioners, we believe that opportunities for infrastructuring in both science and science communication are actively emerging.

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.HC

MultiVerse: A Creator-Centered Approach to Steering Context-Adaptive Lyrics

arXiv:2608.19350v1 Announce Type: new Abstract: Generative AI may enable new forms of context-aware creative expression by dynamically tailoring media content to its consumption context. For instance, AI systems could adapt song lyrics to the listener and their current activity. However, existing media adaptation systems primarily optimize for audience experience, often neglecting artists' intent, style, and preference. We address this challenge by introducing a novel creator-centered approach to adaptive media authoring and present MultiVerse, a system that instantiates this approach for steering adaptive lyrics. Our approach allows creators to explicitly author controls based on their intent, lyric structure, and audience context, and uses rule-based validations to ensure controls are followed. We conducted a study with 10 songwriters, comparing MultiVerse with a prompting-based workflow for composing adaptive lyrics. The comparison revealed that creators preferred to author how lyri

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

Zero2Skill: Bootstrapping Robot Skills through Autonomous Data Collection, Training, and Deployment

arXiv:2607.14047v2 Announce Type: replace-cross Abstract: Autonomous data collection governs the volume and quality of real-world trajectories for manipulation policy learning. Existing pipelines reduce human effort via self-resetting, VLM verification, or language-guided correction, yet episode-scoped fixes must be reissued whenever the same failure recurs, so oversight cost grows with session length rather than with the number of distinct problems. We present Zero2Skill, a human-robot symbiotic agentic system in which corrections are retained and reused across rounds. The collection loop collects, verifies, and resets autonomously, pausing for a remote operator only when a phase exhausts an explicit retry budget. An LLM parser maps each natural-language utterance to a structured adjustment stored in Corrective Memory, so addressed failure modes typically need not be corrected again under the same conditions. On a real-robot desktop-clearing testbed, Zero2Skill matches teleoperation e

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

EEG-based AI-BCI Wheelchair Advancement: Transformer-Based Learning with Motor Imagery for Brain Computer Interface

arXiv:2509.25667v3 Announce Type: replace-cross Abstract: This paper presents an Artificial Intelligence (AI) integrated approach to Brain-Computer Interface (BCI)-based wheelchair development, utilizing a motor imagery right-left-hand movement mechanism for control. The system is designed to simulate wheelchair navigation based on motor imagery right and left-hand movements using electroencephalogram (EEG) data. A pre-filtered dataset, obtained from an open-source EEG repository, was segmented into arrays of 19x200 to capture the onset of hand movements. The data was acquired at a sampling frequency of 200Hz. The system integrates a Tkinter-based interface for simulating wheelchair movements, offering users a functional and intuitive control system. We propose TFormerEEG, a Transformer-driven deep learning architecture, for motor imagery EEG classification. The model achieves a test accuracy of 93.04% compared with various machine learning baseline models, including XGBoost, EEGNet, a

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

Supporting Multimodal Data Interaction on Refreshable Tactile Displays: An Architecture to Combine Touch and Conversational AI

arXiv:2602.15280v2 Announce Type: replace Abstract: Combining conversational AI with refreshable tactile displays (RTDs) offers significant potential for creating accessible data visualization for people who are blind or have low vision (BLV). To support researchers and developers building accessible data visualizations with RTDs, we present a multimodal data interaction architecture along with an open-source reference implementation. Our system is the first to combine touch input with a conversational agent on an RTD, enabling deictic queries that fuse touch context with spoken language, such as "what is the trend between these points?" The architecture addresses key technical challenges, including touch sensing on RTDs, visual-to-tactile encoding, integrating touch context with conversational AI, and synchronizing multimodal output. Our contributions are twofold: (1) a technical architecture integrating RTD hardware, external touch sensing, and conversational AI to enable multimodal

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

JEEVHITAA -- An HCAI Ecosystem to Support Collective Care

arXiv:2512.06364v4 Announce Type: replace Abstract: Current mobile health platforms are predominantly individual-centric and lack the support for coordinated, auditable multi-actor workflows. However, in many settings worldwide, health decisions are enacted through multi-actor coordination rather than individual users. We present JEEVHITAA, a cross-platform mobile system enabling role-aware sharing and verifiable information flows within permissioned care circles. JEEVHITAA ingests platform and device data, builds layered profiles from sensors and tiered onboarding, and enforces fine-grained, time-bounded access control across care graphs. Data stays secure both within the application and the cloud. Integrated retrieval-augmented Large Language Models produce structured, role-targeted summaries and action plans, offer evidence-grounded verification with provenance and confidence scores, and support advanced insights on reports. We describe the architecture, connector abstractions, and

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

Harmonious Color Pairings: Insights from Human Preference and Natural Hue Statistics

arXiv:2508.15777v3 Announce Type: replace Abstract: While color harmony has long been studied in art and design, a clear consensus remains elusive, as most models are grounded in qualitative insights or limited datasets. In this work, we present a quantitative, data-driven study of color pairing preferences using controlled hue-based palettes in the HSL color space. Participants evaluated combinations of thirteen distinct hues, enabling us to construct a preference matrix and define a combinability index for each color. Our results reveal that preferences are highly hue dependent, thereby challenging classical color harmony theories. Yet, when averaged over hues, statistically meaningful patterns of aesthetic preference emerge, with certain hue separations perceived as more harmonious. Strikingly, these patterns align with hue distributions found in natural landscapes, pointing to a statistical correspondence between human color preferences and the structure of color in nature. Finally

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

Can AR Embedded Visualizations Foster Appropriate Reliance on AI in Spatial Decision-Making? A Comparative Study of AR X-Ray vs. 2D Minimap

arXiv:2507.14316v5 Announce Type: replace Abstract: Artificial Intelligence (AI) and indoor sensing increasingly support decision-making in spatial environments. However, traditional visualization methods impose a substantial mental workload when viewers translate this digital information into real-world spaces, leading to inappropriate reliance on AI. Embedded visualizations in Augmented Reality (AR), by integrating information into physical environments, may reduce this workload and foster more appropriate reliance on AI. To assess this, we conducted an empirical study (N = 32) comparing an AR embedded visualization (X-ray) and 2D Minimap in AI-assisted, time-critical spatial target selection tasks. Surprisingly, evidence shows that the embedded visualization led to greater inappropriate reliance on AI, primarily as over-reliance, due to factors like perceptual challenges, visual proximity illusions, and highly realistic visual representations. Nonetheless, the embedded visualization

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

Beyond Interestingness: Semantic and Context-Aware Natural Language Query Recommendations for Visual Data Analysis

arXiv:2201.04868v3 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have made natural language interfaces (NLIs) widely accessible for data exploration, yet analysts who have a broad analytical objective still face the challenge of decomposing it into effective step-by-step queries, especially over unfamiliar, multi-table relational databases. Rather than generating high-level analytical agendas, we investigate how to augment an NLI with semantic- and context-aware next-step query recommendations that act as analytical scaffolding for relational database exploration. Our approach goes beyond interestingness-only methods by jointly integrating semantic relevance, data interestingness, and context coherence to guide exploration toward coherent, topic-focused analyses and potentially insightful subsets. We evaluate QRec-NLI with NL2SQL benchmarking, LLM-enhanced description validation, agentic comparisons against interestingness-only and LLM-based prompting

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

teLLMe Why (Ain't Nothing but a Jam): Exploratory Causal Analysis of Urban Driving Data

arXiv:2607.15254v1 Announce Type: cross Abstract: Traffic agencies now have access to large volumes of video-derived data for studying safety and congestion. Most of these data are observational and collected without interventions, which makes causal questions such as "How would rain change traffic density?" difficult to answer. We present teLLMe, a system for exploratory causal analysis of urban driving datasets. The system starts from a structured event table built from dashcam annotations and combines causal structure learning with the PC algorithm, bootstrap-based stability checks, and query-specific effect estimation using linear regression and DoWhy. Natural-language questions are mapped to structured causal queries through a schema-aware LLM, enabling users to specify treatments, outcomes, and subpopulations. teLLMe returns a "Causal Card" that summarizes effect estimates, adjustment sets, DAG support, and assumptions, followed by a short natural-language explanation. Case studi

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

Divergent Gaze Patterns in Artistic Viewing: Spatial and Temporal Signatures of Attention Across Autistic Individuals, Artists, and Neurotypical Observers

arXiv:2607.15227v1 Announce Type: cross Abstract: How different populations visually explore artworks bears on cognitive science and on accessibility design, yet most eye-tracking work in autism has used social scenes rather than art, and has analysed where the eyes land while ignoring when and in what order. We present a comparative free-viewing study across three groups, autistic adults (ASD), trained artists, and neurotypical observers, who each viewed 30 paintings for 15s. We introduce a directed, metric-grounded framework that compares groups along two complementary axes: a spatial axis, in which one group's fixation-density map predicts another's fixations under six saliency metrics (AUC-Judd, NSS, CC, SIM, KL, Information Gain); and a temporal axis, in which individual scanpaths are compared with MultiMatch, ScanMatch, a foveal-disc IoU score (FDISS), and dynamic time warping (DTW). Fixations are extracted uniformly for all groups with a dispersion-threshold algorithm. Three res

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

Self-Evolving Human-Centered Framework for Explainable Depression Symptom Annotation

arXiv:2607.15202v1 Announce Type: cross Abstract: Annotation quality is a major bottleneck in building reliable and explainable artificial intelligence (XAI) systems for mental health research. In depression-related datasets, labels are often assigned without structured evidence, symptom-level justification, or traceable alignment with the criteria of the Diagnostic and Statistical Manual of Mental Disorders, Fifth Edition, Text Revision (DSM-5-TR), limiting both transparency and downstream model interpretability. We propose a self-evolving, expert-in-the-loop annotation framework for Major Depressive Disorder (MDD) that combines large language model (LLM)-assisted labeling with expert verification. The framework is intended to support the construction of explainable, DSM-5-TR-aligned datasets rather than to perform clinical diagnosis. It operates in three stages: candidate evidence selection from textual records, criterion-level DSM-5-TR analysis, and case-level synthesis that produce

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy

arXiv:2607.15176v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualization literacy assessment test, a standardized SciVis literacy assessment comprising 49 items based on 18 scientific visualizations and illustrations, spanning 8 techniques and 11 task types. We evaluate three closed-source and three open-source models under a closed-world protocol and compare their performance using data from 485 human participants. Results show that current MLLMs do not exhibit uniform SciVis literacy. Gemini is the strongest model overall, exceeding the human mean across the evaluated subsets, whereas the open-source models remain below the human baseline. Performance is highly uneven across techniques and tasks: models perform best on scientif

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

Setup Complete, Now You Are Compromised: Weaponizing Setup Instructions Against AI Coding Agents

arXiv:2607.15143v1 Announce Type: cross Abstract: AI coding agents set up projects by reading documentation and installing the dependencies it lists, without verifying their names, sources, or known vulnerabilities. By editing only a README, requirements file, or Makefile, an attacker can redirect the agent to an untrusted registry, a known-vulnerable version, or a wrong-but-plausible name: documentation becomes a vector for code execution. We present the first systematic evaluation of package-install-time supply-chain attacks delivered through ordinary project-setup documentation across production coding-agent harnesses, probing frontier models on twelve scenarios in five attack classes, grounded in documented incidents. The same model catches an attack through one harness and installs it through another: install-time security rests on the harness-model combination, not the model alone. Agents catch blatant typosquats reliably, but plausible separator-confusion names (azurecore for az

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

Catch, Throw, Repeat: Planning for Human-Robot Partner Juggling

arXiv:2607.15129v1 Announce Type: cross Abstract: Dynamic object exchange between humans and robots remains a challenging problem due to uncertainty in perception, timing, and contact-rich interaction. Human-robot juggling represents a particularly demanding instance of this problem, requiring precise real-time coordination, predictive motion planning with feedback control, and robustness to variability in human motion. Enabling such skills is of interest for advancing physical human-robot interaction and shared autonomy. We present a real-time planning and control architecture for human-robot partner juggling that enables a robot to reliably catch and throw balls in synchronized multi-ball patterns with a human partner. The system integrates predictive ball tracking, adaptive online trajectory optimization using a multiple-shooting formulation, and a state-machine-based coordination logic to enable synchronized multi-ball human-robot partner juggling. In a user study with 8 participan

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

The Misclassification of Autistic Writing as AI-Generated

arXiv:2607.14729v1 Announce Type: cross Abstract: Recent findings suggest that detection models for artificial intelligence (AI) cannot accurately identify AI-generated text and may exhibit bias against certain minority groups. In the present study, anecdotal claims that autistic writers more often have their work flagged as AI-generated are examined empirically. A corpus of approximately 60,000 Reddit posts split into "likely-autistic" and "general-Reddit" subcorpora is used to compare the distribution of probabilities output by the OpenAI GPT-2 detection model. Differences in textual features between subcorpora are observed and compared to reported features of AI-generated text. Results showed that while less than two-percent of either subcorpus was flagged as AI-generated by the model, significantly more texts from the likely-autistic subcorpus were flagged. Connections between features of text with likely-autistic authors and AI-generated text were not straightforward. The widespre

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications

arXiv:2607.14673v1 Announce Type: cross Abstract: Evaluations (Evals) are a deployment bottleneck for real-world AI applications: public benchmarks rarely match a team's users, context, or policies, and human review is often tedious to scale. Motivated by our work with AI applications in the public sector, this project addresses recurring evaluation challenges encountered when applications must satisfy local policy and governance requirements. We present Kaleidoscope, an integrated workflow for contextual functional evaluation that links persona-based test generation, contextualized rubrics, and human review for reliability-gated automated scoring. Generated test cases are scored against application-specific rubrics; human annotations provide reviewable labels; and LLM judges automate scoring only when their agreement with those labels meets a configured threshold. Kaleidoscope is therefore a practical, inspectable, iterative workflow for product teams. We report early evidence from a

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

Volition Elicitation: Operational Semantics for People and Their Machines

arXiv:2607.14138v1 Announce Type: cross Abstract: The most prevalent distributed systems today include people and their personal machines (smartphones). In such systems, computations are driven by people's volitions: a payment when a person wishes to pay someone, befriending when two people wish to become friends, etc. Volition-Guarded Multiagent Atomic Transactions were proposed as an abstract specification language for such systems, in which each agent consists of a person and their machine, and a transaction can be guarded by both the machine states and the personal volitions of its participating agents. Here, we define the programming language volition-guarded GLP (vGLP), which extends GLP with volition-guarded clauses, and define its operational semantics as an instance of Communicating Volitional Agents. As the semantics requires the person to will a volition-guarded clause reduction, a correct implementation must elicit the person's volitions: finding out ``what's in the person'

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

UzWordnet and Generative AI for Learning Uzbek by Game Playing

arXiv:2607.14104v1 Announce Type: cross Abstract: This paper presents an educational system architecture that enables learners to practice the Uzbek language through game-playing. The architecture integrates UzWordnet and the largest currently available orthographic dictionary for Uzbek as core lexical resources, together with generative AI as a fundamental component for learning support. We design four educational games to facilitate Uzbek language learning and propose a game-based methodology for improving UzWordnet as a direct by-product of game dynamics. Our approach combines game design and lexical resources to address objectives that are at the same time educational (language learning) and lexical (improvement and enrichment of a lexical resource).

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

All Polarized but Still Different: a Multi-factorial Metric to Discriminate between Polarization Behaviors on Social Media

arXiv:2312.04603v1 Announce Type: cross Abstract: Online polarization has attracted the attention of researchers for many years. Its effects on society are a cause for concern, and the design of personalized depolarization strategies appears to be a key solution. Such strategies should rely on a fine and accurate measurement, and a clear understanding of polarization behaviors. However, the literature still lacks ways to characterize them finely. We propose GRAIL, the first individual polarization metric, relying on multiple factors. GRAIL assesses these factors through entropy and is based on an adaptable Generalized Additive Model. We evaluate the proposed metric on a Twitter dataset related to the highly controversial debate about the COVID-19 vaccine. Experiments confirm the ability of GRAIL to discriminate between polarization behaviors. To go further, we provide a finer characterization and explanation of the identified behaviors through an innovative evaluation framework.

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

When AI Blurs the Boundaries of Contribution: An Empirical Study of Authorship Calibration

arXiv:2607.15006v1 Announce Type: new Abstract: The broad adoption of Artificial Intelligence (AI), especially Generative AI, raises pressing questions about how users interact with these systems to produce new content. In this paper, we introduce the concept of authorship calibration, defined as users awareness of their actual authorship when interacting with AI. Using the CoAuthor dataset, we empirically examine how authorship calibration varies across users and how it relates to their frequency of AI use. Our results reveal high variability: users relying heavily on AI tend to misjudge their authorship, whereas those using AI less frequently exhibit more accurate authorship calibration. These findings suggest that AI can obscure users perception of their own authorship. In learning contexts, miscalibration can affect metacognitive monitoring and learning strategies, ultimately impacting learning outcomes. Fostering authorship calibration then appears essential for promoting responsi

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

Authoring Narrative Visualization in Motion: Visual Storytelling in Swimming Videos

arXiv:2607.14924v1 Announce Type: new Abstract: We investigate how to support authoring narrative visualizations in motion in sports videos, drawing on automated data preparation, systematic analysis, technology probe design, and evaluation, using swimming races as a case study. Sports videos are widely broadcast and shared across social media, where content creators increasingly seek to present and explain complex events to general audiences. Visualization in motion has been explored as an efficient way to embed data into videos and to move with the data referents, providing additional information and helping audiences understand races. However, existing approaches primarily focus on embedding visualizations in videos, lacking exploration of how to support authoring narratives that coordinate views, data, and temporal progression to explain the unfolding races. To address this gap, we use swimming videos as an ideal case for exploration, as swimming is a sport with rich, dynamic data

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

Understanding of Task-specific and Subject-specific Components in Surface EMG

arXiv:2607.14744v1 Announce Type: new Abstract: Surface electromyogram (sEMG) signals are widely used in human-machine interfaces for gesture recognition and user identification, but existing models often struggle to generalize across individuals due to subject-specific neuromuscular characteristics. This study introduces a disentanglement model that separates task-specific and subject-specific components from sEMG signals, thereby improving the generalization and interpretability of gesture recognition and user identification systems. Experimental results demonstrate that the disentangled components significantly improve the accuracy of both gesture classification and user identification across subjects and days, outperforming conventional methods under the same experimental conditions. Further analysis reveals that the task-specific components capture consistent activation patterns associated with the same gestures across individuals. In contrast, the subject-specific components refl

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

Dendrite: A Real-Time Python Application for Online Brain-Computer Interface Research and Development

arXiv:2607.14655v1 Announce Type: new Abstract: Online brain-computer interface research requires software that can acquire multimodal physiological data, train and update decoders, run live inference, and preserve the full experimental provenance in a reproducible workflow. We present Dendrite, a real-time brain-computer interface application in Python that brings signal acquisition, decoder training, and live inference together in a single, ready-to-run application that stays modifiable. Dendrite records several signal streams at once, each at its native rate, and executes multiple processing modes concurrently against them. A decoder can start from a previously trained model or be fit mid-session while the pipeline keeps running, and the same recordings feed offline training in the same application. Each recording, decoder, and training run is tracked in a database, and every decoder records the configuration and the source recordings it was trained from, so a deployed decoder trace

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

Memory-Driven Self-Disclosure and Relational Turning Points: A Longitudinal Multimodal Study of Human-AI Interaction

arXiv:2607.14593v1 Announce Type: new Abstract: As conversational AI systems are designed for repeated use, a central question is how a series of interactions becomes a relationship. We present a longitudinal multimodal study of a memory-augmented conversational agent (24 participants x 10 sessions), in which participants rated five relational constructs -- familiarity, self-disclosure, perceived memory, conversational quality, and enjoyment -- after each session. Two complementary dynamics emerge. First, conversational quality strongly shapes how enjoyable a session feels in the moment but does not carry forward across sessions, whereas perceived memory is relationally conditioned -- predicted by prior relational state rather than reflecting system capability alone -- and it shapes later enjoyment indirectly, via subsequent self-disclosure. Second, relationships are punctuated by discrete turning points -- crashes and surges -- that are partially traceable in multimodal behavior and o

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

Conversational Tactile Data Interfaces: Co-Designing Accessible Data Experiences with Blind Users Using Refreshable Tactile Displays and Conversational AI

arXiv:2607.14588v1 Announce Type: new Abstract: Combining refreshable tactile displays (RTDs) with conversational AI offers a promising approach to accessible data visualization for people who are blind or have low vision (BLV). However, it remains an open question how these modalities should be integrated to support accessible data experiences. We address this through a co-design process with three BLV co-designers. Building on our prior Wizard-of-Oz study, we created a conversational tactile data interface (CTDI) that combines an RTD with an LLM-powered conversational agent, refined through four workshops over eight months. In addition to the resulting system, Graphy, we contribute design knowledge and recommendations for CTDIs. Co-designers used touch as the primary sensemaking channel for spatial understanding of the data's shape, trends, and relationships, reserved the agent for what touch could not resolve (e.g., calculation and analysis), and used the chart on the RTD to verify

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

Skeleton: Visual Authoring of Non-visual Data Experiences

arXiv:2607.14579v1 Announce Type: new Abstract: When sighted practitioners author accessible data visualizations, they build navigation structures (the nodes, edges, and input bindings that govern how assistive technologies traverse an interface) entirely in code, with no visual representation. Without a representation to react to, practitioners cannot develop judgment about what makes navigation good or bad, and the quality ceiling of non-visual experiences is set by the absence of a feedback loop. We address this problem through longitudinal co-design with practitioners across cartography, design systems, and open-source visualization, and make three contributions. First, we introduce an Inspector that renders navigation graphs as interactive node-link diagrams, and a Dimensions API that expresses navigation in terms of data dimensions rather than explicit graph construction. Second we present Skeleton, a direct-manipulation authoring environment in which the properties of an accessi

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

Not Just Pockets: Understanding Phone-Carrying Behaviors of Wheelchair Users for Mobile Context-Awareness

arXiv:2607.14437v1 Announce Type: new Abstract: Smartphone-based context-awareness holds significant promise for wheelchair users -- from detecting everyday accessibility barriers to enabling ability-based adaptations. Such capabilities often build on passive context inference through mobile sensing, yet their accuracy hinges on how and where phones are carried and the resulting signal quality. While prior work documents phone-carrying behaviors in the general population, patterns specific to wheelchair users remain underexplored. Through a mixed-methods approach combining a survey of 91 and interviews with 15 wheelchair users, we systematically investigate their phone-carrying locations and influencing factors. Our findings reveal distinct patterns extending beyond pocket storage to diverse wheelchair-mounted accessories and around-body placements, shaped by the interplay of physical ability, wheelchair design, and everyday contexts, including social, activity, and device factors. Gro

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

From Product-Centred Retrieval to Experience-Led Commerce:Twelve Candidate Design Principles for Fashion E-Commerce User Experience

arXiv:2607.14429v1 Announce Type: new Abstract: This paper proposes twelve candidate Experience-Led Commerce design principles for high-constraint, relational fashion e-commerce, surfaced through design-led induction while building VogueDrop, a multi-vendor prototype. The principles address multi-entry discovery, experience continuity, relational exploration, preference sovereignty, evidence-scoped correspondence, recommendation-time feasibility, customer-compatible commercial ranking, adaptive but stable workspaces, attributable transaction authority, outcome-linked learning, shared composition authoring, and accountable human to agent handoff. The paper uses fashion as an intentionally bounded domain in which fit, composition, material, identity-sensitive preference, seller fragmentation, visual correspondence, and delivery timing make the interaction breakdown observable; it does not claim universal applicability across e-commerce. Each candidate principle is paired with a prespecif

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.HC

"Trust Junk" Leads to Unjustified Support for Highly Discriminatory Predictive Models

arXiv:2607.14152v1 Announce Type: new Abstract: The persuasive power of data visualizations can go awry: for instance, in an explainable AI (XAI) context, visualizations can produce over-trust of predictive models. In this paper, we use a crowdsourced study to show that providing accurate (but superfluous or irrelevant) data in a model explanation can, in fact, result in unjustified trust and other positive beliefs about a model, even when the model is patently discriminatory and unfair. Our results suggest that XAI designers and developers need to consider the implicit or explicit rhetorics of their work, and beware of the potential of visualizations to imbue models with unearned trust.

Source ↗
technology Fri, 14 Aug 2026 00:00:00 -0400
arXiv cs.HC

Agentic Neurosymbolic Collaboration for Mathematical Discovery: A Case Study in Combinatorial Design

arXiv:2603.08322v2 Announce Type: replace-cross Abstract: We study mathematical discovery through the lens of neurosymbolic reasoning, where an AI agent powered by a large language model (LLM), coupled with symbolic computation tools, and human strategic direction, jointly produced a new result in combinatorial design theory. The main result of this human-AI collaboration is a tight lower bound on the imbalance of Latin squares for the notoriously difficult case $n \equiv 1 \pmod{3}$. We reconstruct the discovery process from detailed interaction logs spanning multiple sessions over several days and identify the distinct cognitive contributions of each component. The AI agent proved effective at uncovering hidden structure and generating hypotheses. The symbolic component consists of computer algebra, constraint solvers, and simulated annealing, which provides rigorous verification and exhaustive enumeration. Human steering supplied the critical research pivot that transformed a dead e

Source ↗
technology Fri, 14 Aug 2026 00:00:00 -0400
arXiv cs.HC

Bayesian Distributional Models of Executive Functioning

arXiv:2510.00387v4 Announce Type: replace-cross Abstract: This study uses controlled simulations with known ground-truth parameters to evaluate how Distributional Latent Variable Models (DLVM) and Bayesian Distributional Active LEarning (DALE) perform in comparison to conventional Independent Maximum Likelihood Estimation (IMLE). DLVM integrates observations across multiple executive function tasks and individuals, allowing parameter estimation even under sparse or incomplete data conditions. To establish known-ground truth, we uniformly sample individual sessions from a neural network learned latent space and map them to distributional cognitive performance across different tasks. The individual test-items are then sampled from these distributions using either DALE, random procedure or a standard fixed battery approach. When given the same set of observations, DLVM consistently outperformed IMLE, especially under smaller amounts of data, and converges faster to highly accurate estimat

Source ↗
technology Fri, 14 Aug 2026 00:00:00 -0400
arXiv cs.HC

The Impact of Generative AI on Collaborative Open-Source Software Development: Evidence from GitHub Copilot

arXiv:2410.02091v4 Announce Type: replace-cross Abstract: Generative artificial intelligence (AI) facilitates content production and enhances ideation, with potentially important implications for developer productivity and participation in software development. To explore its impact on collaborative open-source software (OSS) development, we investigate the role of GitHub Copilot, a generative AI pair programmer, in OSS development where multiple distributed developers voluntarily collaborate. Using GitHub's proprietary Copilot usage data, combined with public OSS project data obtained from GitHub, we find that Copilot use increases project-level code contributions by 5.9%. This gain is accompanied by a 3.4% increase in developer coding participation and a 2.1% increase in individual code contributions. However, Copilot use is also associated with an 8% increase in coordination time and more code discussions. This reveals an important tradeoff: While AI expands who can contribute and h

Source ↗
technology Fri, 14 Aug 2026 00:00:00 -0400
arXiv cs.HC

Harmonizing Safety and Speed: A Human-Algorithm Approach to Enhance the FDA's Medical Device Clearance Policy

arXiv:2407.11823v4 Announce Type: replace-cross Abstract: The United States Food and Drug Administration's (FDA's) 510(k) pathway allows manufacturers to gain medical device approval by demonstrating substantial equivalence to a legally marketed device. However, the inherent ambiguity of this regulatory procedure has been associated with high recall among many devices cleared through this pathway, raising significant safety concerns. In this paper, we develop a combined human-algorithm approach to assist the FDA in improving its 510(k) medical device clearance process by reducing recall risk and regulatory workload. We first develop machine learning methods to estimate the risk of recall of 510(k) medical devices based on the information available at the time of submission. We then propose a data-driven clearance policy that recommends acceptance, rejection, or deferral to FDA's committees for in-depth evaluation. We conduct an empirical study using a unique dataset of over 31,000 subm

Source ↗
technology Fri, 14 Aug 2026 00:00:00 -0400
arXiv cs.HC

Vibe to Code: Elucidating Strategic Oscillation of Tacit Knowledge in Generative AI Design Workflows -- An Exploratory Qualitative Study

arXiv:2607.23126v2 Announce Type: replace Abstract: The rapid adoption of generative AI tools has created new literacy demands for designers who must verbalize tacit knowledge through natural language prompts. Yet the micro-level cognitive processes by which designers externalize implicit intentions during iterative AI dialogue remain underexplored. This exploratory qualitative study examined five expert designers (11-20 years UI/UX experience, M = 15.4 years) using think-aloud protocols. We identified "Strategic Oscillation" -- experts' intentional return to vague language (Vibe) after progressing toward operational specifications (Code), leveraging AI's probabilistic nature for creative exploration. We observed shifts from "instruction" to "consultation" mode, deep reflection triggered by semantic collisions, and one participant's disengagement revealing conceptual alignment as a tentative boundary condition. We propose the ECRT cycle (Expectation-Collision-Reflection-Transformation)

Source ↗
technology Fri, 14 Aug 2026 00:00:00 -0400
arXiv cs.HC

Psychological Benefits and Costs of Diversifying Algorithmic Recourse

arXiv:2605.11793v2 Announce Type: replace Abstract: Algorithmic recourse provides counterfactual action plans that help people overturn unfavorable AI decisions. While diverse recourse sets may improve transparency and motivation, they may also impose cognitive load and negative emotions by increasing counterfactual reasoning demands. To examine this trade-off, we conducted a between-subjects controlled experiment (N=750) that manipulated recourse-set diversity and size, and evaluated these effects on psychological benefits and costs. Results show that diversification enhances psychological benefits (e.g., willingness to act) for small sets without incurring additional psychological costs, whereas for large sets, it makes cognitive load more salient. These findings suggest that naively diversifying recourse can burden decision subjects, underscoring the need for new diversification methods that incorporate human cognition and psychology to mitigate such costs.

Source ↗
Showing 1451–1500 of 1631 signals
← Prev Page 30 of 33 Next →