EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.HC

FIERO: Empowering Creative Writing Through Collaborative Game Play

arXiv:2607.11837v1 Announce Type: new Abstract: Creativity often flourishes in collaboration, such as when designers brainstorm a new app together, or storytellers collectively build a world with elements of each person's narrative. However, collaborative storytelling can have challenges for its participants, such as when they disagree about the plot proposed, or when different ideas become fragmented when voiced individually. While current tools for creative collaboration focus on synchronous online text sharing, they often neglect the social dynamics of in-person collaboration critical to creative synergy. To address this, we created FIERO, a multiplayer web-based card game. Physical cards provide tangible scaffolding and social interaction, while the digital interface generates contextual visuals, facilitate group decisions, ensure narrative coherence, and synthesize different idea contributions using generative AI. Compared against online collaborative writing alone, the game signi

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.HC

Supporting Reflection in LLM-based Exploratory Search

arXiv:2607.11810v1 Announce Type: new Abstract: Large Language Models (LLMs) can make exploratory search more efficient but may undermine the reflection and iterative sensemaking needed in unfamiliar domains. Existing LLM tools often prioritize rapid answers over supporting users in tracking how their understanding evolves and how well their strategies align with their goals. We present TrailLM, a system that helps users reconstruct and revisit their exploration paths to support reflection and metacognitive engagement during information seeking. By aligning LLM assistance with users' sensemaking workflows, TrailLM aims to preserve the benefits of LLM-based search while enhancing opportunities for critical reflection on one's own search process.

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.HC

HandPad: A Bimanual Hand Interface for Fluid Window Interactions in VR

arXiv:2607.11807v1 Announce Type: new Abstract: Virtual Reality (VR) offers potential for productivity work by creating expansive displays anywhere, yet current systems often rely on external input devices that limit the on-the-go use of mobile VR. We introduce HandPad, a suite of bare-hand interaction techniques that leverage the benefits of asymmetric bimanual coordination and self-haptic support. HandPad assigns the non-dominant hand (NDH) to establish spatial frames and interaction contexts, while the dominant hand (DH) performs fine-grained manipulation. Users can use NDH gestures as an input modifier to change the mode and target of DH interactions, including multi-window navigation, in-window content interaction, and window management. The palm surface of the NDH also serves as a physical touch surface, providing passive haptic feedback for effective DH touch interaction. Both hands and their interactions are spatially remapped to the window surface, enabling comfortable and dir

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.HC

"We are all in big trouble! *Shock Emoji": Personal Narratives in Expressing Emotions, Opinions, and Data Regarding Climate Change in TikTok Short Videos

arXiv:2607.11803v1 Announce Type: new Abstract: Climate change is a source of anxiety about the future. Understanding how people express themselves about climate change enables us to address such concerns. To study climate change expression on social media, we analyzed 200 TikTok videos tagged with #climatechange, identifying four categories of content: expression-feelings, views-appeals, news-information, and trend-hijacking. We found that creators use humor to package sharp critiques, avoiding direct confrontation. They replace complex discussions with life stories, such as adopting a vegetarian lifestyle or deleting emails. They borrow from news media to present fragmented information as scientific interpretations, creating a perception of scientific credibility, balancing scientific accuracy with emotionality. Analysis of viewer responses showed they engaged empathetically, reshaping interpretations of videos. These interactions risk reinforcing existing views but help build commun

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.HC

Toward Inclusive Avatar Design with Limb Differences Through Artificial Intelligence

arXiv:2607.11512v1 Announce Type: new Abstract: As extended reality becomes more popular for social interaction and entertainment, 3D avatars must represent the full diversity of body types. Most 3D avatar systems only support normative bodies and do not accurately depict people with limb differences, amputations, or other morphological variations. This paper reviews emerging technical approaches for inclusive 3D avatar customization for this group and current guidelines that promote respectful and accurate representation. We highlight persistent challenges, including the scarcity of diverse datasets and the limitations in animation for non-normative anatomies. This paper positions artificial intelligence as a promising path to overcoming these limitations and advancing inclusive 3D avatar generation.

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.HC

ManiScope: LLM-Assisted Visual Analytics of Cryptocurrency Manipulation Risk

arXiv:2607.11451v1 Announce Type: new Abstract: Cryptocurrency markets are vulnerable to trade-based manipulation, such as wash trading, which can distort price signals and mislead investors. Prior research has mainly focused on detecting manipulation using fixed rules or labeled examples, offering limited flexibility and interpretability for assessing potential risks. Existing visual analytics tools can reveal basic manipulation-related signals, such as token distribution, but still require substantial manual effort to integrate holder relationships, suspicious behaviors, and market dynamics for risk assessment. To address these limitations, we propose ManiScope, an LLM-assisted visual analytics system for analyzing trade-based manipulation risks in cryptocurrency markets. ManiScope provides coordinated views of token distributions, holder relationships, detailed holder behaviors, price dynamics, and suspicious trading patterns. To further enhance user analysis, ManiScope introduces a

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.HC

DiffLens: A Visualization System to Explore Local Differences in Graph Sampling

arXiv:2607.11424v1 Announce Type: new Abstract: Graph sampling techniques have been widely used to simplify network computation and visualization, which also results in inevitable differences between the sampled networks and the original networks in terms of nodes, edges and structures. Investigating such differences can inform graph sampling technique users of the pros and cons of different techniques and select the appropriate one, and can also help graph sampling developers evaluate their own technique. However, there are still no systematic ways to achieve such a goal. This paper fills this research gap by first proposing systematic and generic quantitative measures to quantify three categories of graph differences (i.e., neighbor-based, path-based, and structure-based). Built upon this, we further propose DiffLens, a novel visualization system to help graph sampling developers and users intuitively explore local differences at different regions of their interest within a sampled g

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.HC

Same Stories, Different Journeys: From Social Comparison to Sensemaking in AI-Mediated Peer Career Exploration

arXiv:2607.11039v1 Announce Type: new Abstract: Young job seekers frequently turn to social media to compare themselves with peers and make sense of career possibilities. However, passive feed browsing creates a paradox: the authentic peer content that provides emotional grounding also triggers potentially detrimental upward social comparison and cognitive overload. Previous work has either structured online user-generated content to reduce noise without changing the passive browsing modality, or built AI-powered career exploration systems that disregard authentic human experiences. To address this gap, we developed JobMate, an interactive system that transforms real social media career posts into persona-grounded conversational AI agents, shifting the interaction from passive scrolling to active, personalized dialogue. We conducted a between-subjects study ($N$ = 24, three disciplines) comparing JobMate with native RedNote browsing. Our study shows that JobMate's AI-mediated dialogue

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.HC

When Context Dominates: Multimodal Signatures of Takeover Readiness Under Varying Hazard and Cognitive Load Conditions

arXiv:2607.10945v1 Announce Type: new Abstract: Semi-automated driving systems promise to reduce crashes by assisting with perception and control, yet they simultaneously introduce additional human factors challenges by requiring drivers to monitor automation and rapidly resume control when failures occur. Prolonged passive monitoring can degrade vigilance, delay reactions, and increase takeover risk, but the extent to which distraction, hazard context, and drivers' underlying cognitive and physiological states jointly shape takeover performance remains insufficiently understood. This study investigates these interacting factors using a controlled, within-subjects driving simulator experiment that crosses two hazard types (dynamic pedestrian and static crash events) with three levels of secondary task engagement (no task, conversation, and working memory load). Driver responses were assessed using a multimodal sensing framework that integrates vehicle-dynamics measures, subjective work

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.HC

What to Distinguish and How? Opportunities and Challenges of Augmenting Multiple, Cluttered Objects in Complex Scenes for People with Low Vision

arXiv:2607.10902v1 Announce Type: new Abstract: People with low vision (PLV) struggle to perceive complex scenes like busy kitchens and crowded streets, which contain many objects, visual clutter, and dynamic elements. Prior AR systems for low vision either enhance low-level visual features or augment task-relevant objects for single tasks in simple settings, leaving multi-object augmentation in complex scenes underexplored. Informed by a formative study characterizing important objects and their perceived importance for PLV, we built SceneGlance, a wearable AR system that recognizes important objects and visually distinguishes them by importance level. Through a controlled lab study with 12 PLV in a mock-up kitchen scene and a free-form think-aloud study with 13 PLV navigating an outdoor route, we found that AR distinction on object importance shifted PLV's attention toward objects of higher importance, and supported perception strategies such as building mental snapshots from the aug

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.HC

FaciliTrain: Practicing Facilitation Skills through AI-Simulated Group Dialogue

arXiv:2607.10850v1 Announce Type: new Abstract: Skilled facilitation supports inclusive small-group dialogue, but deliberate practice is hard to scale: it depends on expert coaches, live practice partners, and iterative feedback. We present FaciliTrain, a voice-based training system in which learners step into the facilitator role of an AI-simulated multi-participant conversation, apply five evidence-based techniques, and receive structured AI feedback to support reflection. We report findings from a mixed-methods study with 24 participants, conducted as a formative study (N = 12) and a controlled pilot (N = 12; 6 treatment, 6 control). Both conditions achieved comparable accuracy on a live evaluation task, though treatment participants' self-rated comfort declined significantly while control participants' comfort improved (p = .018). Reflexive thematic analysis identifies four themes: the taxonomy externalizes implicit facilitation intuitions; Making Connections is the most cognitivel

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.HC

Lottery and Sprint Arcade: Enabling Player-Driven Game Editing with Generative AI

arXiv:2607.10711v1 Announce Type: new Abstract: Large language models (LLMs) are shifting game generation from offline automation toward play-driven modification through natural language interaction. In this work, we present a play-driven game editing system that enables players to modify a retro Space Invaders - style arcade game through voice-based natural-language commands during play. Spoken instructions are interpreted by an LLM and translated into structured updates of internal configuration parameters, allowing iterative play - edit - feedback cycles in an invader-style game environment without exposing underlying system details. The game includes approximately 100 editable configuration fields controlling mechanics, visuals, interaction patterns, and audio behavior, enabling gameplay transformation through incremental parameter changes. To investigate how users experience play-driven AI-mediated editing (RQ1) and how emergent editing patterns relate to variations in player expe

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.HC

U-Lens: Supporting User Uncertainty Management in Long-Form LLM Responses

arXiv:2607.10604v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate long-form answers for knowledge-intensive tasks, but users often struggle to decide which parts of a response deserve scrutiny, why they may be unreliable, and what to do next. Prior work on uncertainty communication has largely focused on making uncertainty visible through cues such as confidence scores, leaving less support for the broader process of managing uncertainty distributed across a long response. Through a formative study, we examine how users manage such uncertainty across three stages: interpretation, evaluation, and decision. Based on these insights, we derive design guidelines that address both stage-specific and cross-stage needs: uncertainty target representation, evaluative explanation, response guidance, and interactive presentation. We instantiate these guidelines in U-Lens, an uncertainty-management support system that organizes uncertain information in l

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.HC

Motif: Discovering and Automating Personal Web Workflows

arXiv:2607.10531v1 Announce Type: new Abstract: Recent advances in LLMs and existing work on programming by demonstration have made it possible for end users to create automations by explicitly demonstrating their behavior to LLMs. However, these approaches rely on the assumption that users know what to automate and what is capable of being automated. Additionally, automation via LLM agents is often expensive compared with programs. We introduce Motif, a system that passively observes everyday browser activity to discover recurring interaction patterns that are programmable, makes recommendations to users whenever a pattern is discovered and generate a program to install after user confirmation. Users can review, and refine the program using natural language. We evaluated Motif in a multi-day study, comparing its ambient discoveries against automations users attempted to build via ``vibe coding.'' With eight participants, Motif discovered more automatable patterns than users recognized

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.HC

Spatula: Exploring On-Demand In-Situ Interfaces and Interaction for Attribute Control

arXiv:2607.10405v1 Announce Type: new Abstract: Controlling attributes is a critical step toward achieving the final creative outcome, yet current approaches fall short in supporting users in the iterative refinement of generative content. We propose Spatula, a proof-of-concept system that generates on-demand, in-situ attribute control interfaces and interactions for creating motion graphics. Building on a technical probe that automatically analyzes animation context and generates corresponding attributes and UI, we frame attribute control as an explorable landscape and explore the attribute control space along four key dimensions: Discoverability, Resolution, Scope, and Expandability. Findings from a user study (N=12) show that our system provides intuitive and convenient interactions while supporting diverse needs for fine-grained parameter control. Furthermore, our applications demonstrate that the plug-and-play design generalizes to other domains, such as web design and 3D modeling

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.HC

Learning behavior accounts for background-related advantage in AI-assisted education

arXiv:2607.10101v1 Announce Type: new Abstract: Generative AI has been found, and will likely be found increasingly, useful in education. However, existing AI-for-education studies provide inconsistent evidence on its average effects. More broadly, research on prior educational technologies shows that average effects often mask substantial heterogeneity across student populations. Motivated by this evidence, this study examines heterogeneity in students' learning behavior with AI, which students benefit from AI assistance, and how learner profiles and learning behavior shape these patterns. To this end, we recruited 318 university students to participate in structured learning experiments lasting up to 125 minutes. Our findings indicate that students' learning behavior is strongly associated with learning outcomes, with behaviors characterized by proactive and critical engagement, rather than limited engagement, associated with significantly better performance. These behavioral differe

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.HC

SyncSpace: Layout-Conditioned 3D Gaussian Splatting for Space Reskinning in Mixed Reality

arXiv:2607.10050v1 Announce Type: new Abstract: We present SyncSpace, a system that achieves both spatial alignment and visual consistency between a generated 3DGS world and physical space. We first scan the space via depth sensing to extract 3D bounding boxes, which we render into a layout-only panorama and feed as a geometric prior to a generative world model, producing a Gaussian splat scene in which objects are re-semantized to fit a target style without per-object control. We then align the generated scene to physical space with a coarse-to-fine registration algorithm, refined manually via pinch gestures when automatic registration does not converge. We demonstrate a hand-tracked engulfment interaction in which the virtual world rises to replace the physical space, and show a single space reskinned into multiple stylistically distinct worlds with its layout preserved.

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.HC

"Code Is Cheap. Show Me the Talk.": Lessons from Teaching and Managing AI Coding Tool Usage in a Visualization Course

arXiv:2607.09938v1 Announce Type: new Abstract: Generative Artificial Intelligence (GenAI) coding tools are transforming visualization education. They can assist with implementation and design, but they can also let students bypass intended learning trajectories. In this paper, we share our retrospective experience managing and teaching AI use in an upper-level visualization course. We implemented prompt injections, asked oral checkout questions, and taught two AI coding labs. Prior to our coding labs, at least half of the students had already used AI tools in their assignments. In both AI coding labs, refinement accounted for about half of students' prompting logs, and explanation was almost absent. In the lab where AI coding was optional, 44 of 78 (56.4%) submissions preferred the scaffolded instructions over designing their own prompts. Students' final projects were more polished than in our previous offering, but also more visually homogeneous. Our reflections point to the need for

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.HC

When LLM Tutoring Responses Work: Evidence from Student Programming Conversations

arXiv:2607.09919v1 Announce Type: new Abstract: As students increasingly use LLM tutors in computer science education, one question becomes especially important: what kind of response helps a student continue productively? Prior work has studied how students use LLMs in computer science education, but less is known about how tutoring response styles are associated with student follow-up across programming help-seeking contexts. This paper analyzes StudyChat (UMass, 2026), a public dataset of student and ChatGPT tutoring conversations from an artificial intelligence course. We transformed StudyChat into 16,851 assistant-response interactions from 203 students and 2,214 conversations. Using local LLM-assisted annotation with Gemma 4, we labeled student help-seeking situations, student state, assistant response style, and student next-turn outcome. Human validation showed 82\% agreement with the LLM-assisted labels (Cohen's $\kappa=.74$). We analyzed productive continuation and unresolved

Source ↗
technology Tue, 14 Jul 2026 00:00:00 -0400
arXiv cs.HC

The Individual-Targeting Assumption: A Systematic Review of Proactive Robots in Human Group Settings

arXiv:2607.09734v1 Announce Type: new Abstract: Proactive robots are increasingly deployed in public environments where people are encountered not as isolated individuals but as members of cohesive social groups. Yet whether the prevailing design paradigm in proactive human-robot interaction (HRI) accounts for the relational structure that defines a group as a social unit remains largely unexamined. Through a systematic review of 63 proactive HRI studies in group settings from 2000 to 2025, we identify a recurring tendency, the Individual-Targeting Assumption (ITA), in which robots treat co-present people as independent engagement targets. We find that ITA is present in 60.3% of the corpus, with group-aware approaches emerging almost entirely after the robot is already embedded in an ongoing interaction. Critically, how a robot should detect and negotiate entry into a pre-formed group before initiating contact remains unaddressed across the corpus. Three failure modes, engagement misde

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

Loom: Multi-Region Analysis of Spatial Transcriptomics with Local Neighborhoods and Global Trajectories

arXiv:2607.22505v2 Announce Type: replace-cross Abstract: We present Loom, a spatial transcriptomics (ST) visual computing system to support the analysis of pseudo-temporal trajectories, comparative investigation across samples and regions of interest, and the examination of spatially structured processes within local microenvironments. ST is a molecular profiling technology that measures gene expression directly within a thin tissue section while preserving its spatial organization. For practical application-driven analyses, the ST local microenvironment data needs to be integrated with cell reference datasets and temporal simulations of cell behavior. This integration is challenging due to multi-modal registration issues and the complexity of the pseudo-temporal patterns, spatial enrichment data, and gene expression dynamics. Loom leverages a novel glyph coupled with a computational backbone to facilitate the detailed pseudo-temporal exploration of local microenvironments, cross-samp

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

Persona Prompting in Multimodal Urban Perception: Descriptive Convergence and Interpretive Variation

arXiv:2605.29064v2 Announce Type: replace-cross Abstract: This study examines how persona prompting shapes language generated by two multimodal large language models in urban perception, a setting for examining subjective interpretations of shared visual evidence. We organize outputs into three functional levels: descriptive grounding (captions), intermediate semantic layer (perception tags), and interpretive framing (justifications). Using approximately 60,000 persona-conditioned annotations per model from Qwen3-VL-8B and Gemma-4-E4B-it, we find that captions converge strongly across persona profiles and show only small attribute-associated differences. Justifications vary substantially more: economic status produces the largest difference in both models, with political orientation and personality also prominent. Paired image-level comparisons confirm larger justification than caption differences for these three attributes. For perception tags, personas sharing the same attribute leve

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

Cognitive Energy Modeling for Neuroadaptive Human-Machine Systems using EEG and WGAN-GP

arXiv:2604.01653v2 Announce Type: replace-cross Abstract: Electroencephalography (EEG) provides a non-invasive insight into the brain's cognitive and emotional dynamics. However, modeling how these states evolve in real time and quantifying the energy required for such transitions remains a major challenge. The Schr\"odinger Bridge Problem (SBP) offers a principled probabilistic framework to model the most efficient evolution between the brain states, interpreted as a measure of cognitive energy cost. While generative models such as GANs have been widely used to augment EEG data, it remains unclear whether synthetic EEG preserves the underlying dynamical structure required for transition-based analysis. In this work, we address this gap by using SBP-derived transport cost as a metric to evaluate whether GAN-generated EEG retains the distributional geometry necessary for energy-based modeling of cognitive state transitions. We compare transition energies derived from real and synthetic

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

Knowing When to Ask: Resolving Uncertainty in Human-Robot Joint Planning via Explicit Dialogue and Implicit Intent Cues

arXiv:2603.07822v2 Announce Type: replace-cross Abstract: Effective human-robot collaboration in open-world environments requires joint planning under uncertainty about the task, the environment, and the human teammate. Communication is the most direct means of resolving such uncertainty, yet most existing systems support only one-way communication: robots listen and act, treating humans as passive supervisors rather than conversational teammates capable of two-way dialogue. We propose a unified human-robot joint planning system in which the robot actively resolves uncertainty through two complementary communication channels. When uncertainty is decision-critical, an uncertainty-mitigation joint planning module engages the human in clarification dialogue: it grounds ambiguous instructions via an LLM-assisted active elicitation mechanism, enumerates traversability hypotheses through a hypothesis-augmented A* search, and computes a cost-optimal querying policy via dynamic programming, so

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents

arXiv:2508.04412v3 Announce Type: replace-cross Abstract: The advent of large language models (LLMs) has sparked an evolution of autonomous web browsing agents: given a web browsing task and serialised user interface (UI) state, an LLM is expected to suggest input actions that incrementally solve the given task. The central challenge lies in serialising UI state for LLMs. Web agents have increasingly relied on grounded graphical UI (GUI) snapshots - screenshots augmented with visual cues - favoured for their modest input token footprint. Document Object Model (DOM) snapshots, serialised as HTML, represent a compelling alternative that leverages previously demonstrated HTML interpretation capabilities of LLMs. Their excessive token footprint, however, has precluded reliable deployment with web agents to date. We propose D2Snap, an algorithm to downsample the DOM, premised on preserving actionability and actionability-discriminating features. We evaluate D2Snap-downsampled DOM snapshots

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

X2C: A Dataset Featuring Nuanced Facial Expressions for Realistic Humanoid Imitation

arXiv:2505.11146v3 Announce Type: replace-cross Abstract: Fine-grained facial expression transfer from humans to humanoid agents presents a unique pattern recognition challenge due to the significant domain gap between biological facial dynamics and mechanical control spaces. While visual synthesis of talking heads has advanced rapidly, mapping high-dimensional visual cues to precise, physically constrained actuation signals remains an open problem, primarily due to the lack of large-scale paired data. To bridge this gap, we introduce X2C, a comprehensive benchmark dataset comprising 100,000 (image, control value) pairs. Unlike existing resources, X2C features nuanced, physically grounded expressions annotated with 30 continuous control parameters, establishing a high-fidelity standard for this task. Building on this resource, we propose X2CNet, a two-stage deep learning framework that explicitly decouples visual motion features from mechanical control regression to model the correspon

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

TreeHop: Efficient Embedding-Level Query Rewriter

arXiv:2504.20114v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) systems face significant challenges in multi-hop question answering (MHQA), where complex queries require synthesizing information across multiple document chunks. Existing approaches typically rely on iterative LLM-based query rewriting and routing, resulting in high computational costs due to repeated LLM invocations and multi-stage processes. To address these limitations, we propose TreeHop, an embedding-level framework without the need for LLMs in query refinement. TreeHop dynamically updates query embeddings by fusing semantic information from prior queries and retrieved documents, enabling iterative retrieval through embedding-space operations alone. This method replaces the traditional "Retrieve-Rewrite-Vectorize-Retrieve" cycle with a streamlined "Retrieve-Embed-Retrieve" loop, significantly reducing computational overhead. Moreover, a rule-based stopping criterion is introduced to fu

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

Transformer Explainer: Learning LLM Transformers with Interactive Visual Explanation and Experimentation

arXiv:2408.04619v2 Announce Type: replace-cross Abstract: The Transformer architecture underpins modern large language models powering state-of-the-art text generation and AI applications. However, its complexity makes it difficult for non-experts to learn. Existing resources often lack interactivity, rely on static descriptions of simplified architectures, or fail to reflect models' behavior with real data. To address this gap, we introduce Transformer Explainer, an interactive visualization tool for non-experts to learn Transformers. The tool integrates an overview illustrating the Transformer's data flow with on-demand explanations that gradually reveal mathematical details. Smooth transitions across abstraction levels highlight the interplay between high-level structures and low-level operations. Running a live GPT-2 instance directly in the browser, Transformer Explainer empowers learners to experiment with custom input and hyperparameters without setup, observing next-token predi

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

VisCanvas: A Node-Based Interface for Exploratory Visualization Authoring with LLMs

arXiv:2607.21886v2 Announce Type: replace Abstract: Visual data analysis involves both open-ended exploration and targeted question answering. Visualization authoring tools support this process by enabling users to create visualizations for these tasks. With the rise of large language models (LLMs), substantial effort has been devoted to developing visualization authoring tools that use natural language instructions. However, existing systems are typically based on a linear chat interface, which is not well suited to exploratory visual analysis workflows. In this paper, we introduce VisCanvas, a node-based interface for exploratory visualization authoring with LLMs. VisCanvas allows users to create, revise, branch, and merge visualizations in a non-linear way, enabling more efficient exploration of multiple analytical directions. We conducted a user study with 20 participants to evaluate the effectiveness of VisCanvas compared to a baseline chat-based interface. The results show that V

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

MIA: A Visual Analytics System for Multimodal Spectral Imaging Data

arXiv:2606.00874v2 Announce Type: replace Abstract: Hyperspectral bioimaging techniques such as infrared (IR) microscopy and laser ablation-inductively coupled plasma-mass spectrometry (LA-ICP-MS) produce high-dimensional, spatially resolved datasets that require sophisticated analysis to reveal chemically and anatomically meaningful structures. Existing software solutions are typically modality-specific and cover only parts of the analytical workflow, forcing researchers to transfer data across multiple tools and manually reconcile results. We present MIA (Multiscale Image Analysis), a modality-agnostic visual analysis environment that integrates the full exploratory workflow -- from spectral preprocessing and dimensionality reduction to interactive segmentation and spectral similarity analysis -- within a single, tightly coupled interface. MIA supports hierarchical and landmark-based embeddings to handle datasets of varying scale and complexity, interactive and automatic segmentation

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

Quieting the Cobwebs: Browser Interaction for Visual Floaters

arXiv:2605.12739v2 Announce Type: replace Abstract: Floaters, cobweb-like shadows that move around a person's visual field, impair vision for nearly 33% of the population, yet have limited treatment options. Floaters especially harm screen use, since they reduce contrast, introduce clutter, and add moving distractions. While existing high-contrast tools offer some help, few address the motion that makes screen use with floaters uniquely difficult. In this paper, we build a floater simulation inspired by the physics of the eye, use it to quantitatively assess text readability at varying levels of motion, and build a novel web extension that minimizes eye movement, maximizing the signal-to-noise ratio of performing browser tasks. Importantly, our tool works not only for text, but for all UI elements, requiring no modifications to existing websites.

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

Designing Youth Social Media through Problem Space Attunement

arXiv:2605.07018v3 Announce Type: replace Abstract: Social media is central to how young people maintain relationships, develop identity, and access communities, yet dominant platform designs often leave youth feeling disempowered rather than supported. My dissertation argues that youth social media design is shaped by three forms of problem-space misattunement. \textit{Conceptual misattunement} occurs when the language of ``social media'' anchors participants to existing platforms' interaction templates. I address this through a Fictional Inquiry design workshop that frees youth from preconceived notions of social media by having them brainstorm ways to ``magically connect with remote wizard friends'' rather than ideas for ``social media.'' \textit{Definitional misattunement} occurs when researchers define what ``better'' means on youth's behalf. I address this through a Discord-based asynchronous community that supports youth-led collective inquiry. \textit{Evaluative misattunement}

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

Beyond Screenshots: Evaluating VLMs' Understanding of UI Animations

arXiv:2604.26148v2 Announce Type: replace Abstract: AI agents operating on user interfaces must understand how interfaces communicate state and feedback to act reliably. As a core communicative modality, animations are increasingly used in modern interfaces, serving critical functional purposes beyond mere aesthetics. Thus, understanding UI animation is essential for comprehensive interface interpretation. However, recent studies of Vision Language Models (VLMs) for UI understanding have focused primarily on static screenshots, leaving it unclear how well these models handle dynamic UI animations. To address this gap, we created AniMINT, a novel dataset of 300 densely annotated UI animation videos. We systematically evaluate state-of-the-art VLMs on UI animation understanding, including their abilities to perceive the animation effects, identify animation purposes, and interpret animation meaning. Our results show that VLMs can reliably detect primitive motion. However, their high-leve

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

MAESTRO: Adapting GUIs and Guiding Navigation with User Preferences in Conversational Agents with GUIs

arXiv:2604.06134v2 Announce Type: replace Abstract: Modern task-oriented chatbots present GUI elements alongside natural-language dialogue, yet the agent's role has largely been limited to interpreting natural-language input as GUI actions and following a linear workflow. In preference-driven, multi-step tasks such as booking a flight or reserving a restaurant, earlier choices constrain later options and may force users to restart from scratch. User preferences serve as the key criteria for these decisions, yet existing agents do not systematically leverage them. We present MAESTRO, which extends the agent's role from execution to decision support. MAESTRO maintains a shared preference memory that extracts hard and soft preferences from natural-language utterances and provides two mechanisms. Preference-Grounded GUI Adaptation applies in-place operators (augment, sort, filter, and highlight) to the existing GUI according to preference strength, supporting comparison among options. Pref

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

DiSCo: Diffusion Sequence Copilots for Shared Autonomy

arXiv:2603.22787v2 Announce Type: replace Abstract: Shared autonomy combines human user and AI copilot actions to control complex systems such as robotic arms. When a task is challenging, requires high dimensional control, or is subject to corruption, shared autonomy can significantly increase task performance by using a trained copilot to effectively correct user actions in a manner consistent with the user's goals. To significantly improve the performance of shared autonomy, we introduce Diffusion Sequence Copilots (DiSCo): a method of shared autonomy with diffusion policy that plans action sequences consistent with past user actions. DiSCo seeds and inpaints the diffusion process with user-provided actions with hyperparameters to balance conformity to expert actions, alignment with user intent, and perceived responsiveness. We demonstrate that DiSCo substantially improves task performance in simulated driving and robotic arm tasks. Project website: https://sites.google.com/view/disc

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

Autonomy Reshapes How Personalization Affects Privacy Concerns and Trust in LLM Agents

arXiv:2510.04465v3 Announce Type: replace Abstract: LLM agents require personal information for personalization in order to effectively act on users' behalf, but this raises privacy concerns that can discourage data sharing, limiting both the autonomy levels at which agents can operate and the effectiveness of personalization. Yet the expanded design space of agent autonomy also presents opportunities to shape these effects, which remain underexplored. We conducted a $3\times3$ between-subjects experiment ($N=450$) to study how agent autonomy level influences personalization's effects on users' privacy concerns, trust, and willingness to use, as well as the underlying psychological processes. We find that risk-contingent autonomy, where the agent delegates control back to users upon detecting potential privacy leakage, improves users' perceived control. This in turn attenuates personalization's adverse effects: privacy concerns rise less and trust declines less. Our results suggest tha

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

HeedVision: Attention Awareness in Collaborative Immersive Analytics Environments

arXiv:2505.07069v4 Announce Type: replace Abstract: Group awareness--the ability to perceive the activities of collaborators in a shared space--is a vital mechanism to support effective coordination and joint data analysis in collaborative visualization. We introduce collaborative attention-aware visualizations (CAAVs) that track, record, and revisualize the collective attention of multiple users over time. We implement this concept in HeedVision, a standards-compliant WebXR system built with React Three Fiber that runs on modern AR/VR headsets, and complement it with proof-of-concept implementations covering the remaining three quadrants of our design space--varying presentation (embedded vs. separated) and situatedness (world space vs. camera space). Through a mixed-methods exploratory study where pairs of co-located analysts performed visual search tasks in a shared immersive AR environment, we investigate how attention revisualization affects collaborative coordination in immersive

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

WhichTok? Comparing Three TikTok Data Acquisition Tools

arXiv:2608.09917v1 Announce Type: cross Abstract: TikTok's global growth has made it a prime platform for both entertainment and political discourse, prompting increased social science research. However, this rapidly evolving research field faces a fundamental reproducibility crisis. TikTok's opaque algorithmic systems hinder researchers from drawing meaningful empirical inferences, while the lack of standardized data collection methods compounds these challenges. This study addresses these methodological gaps by systematically comparing three data collection tools - the official TikTok Research API, Pyktok, and Apify. We evaluated five endpoints: User, Hashtag, Keyword, Comment, and Related Video. Results show substantial cross-tool differences, especially for hashtag and keyword searches. The Research API uses back-end API calls, whereas Apify and Pyktok rely on front-end web scraping, producing systematic differences in the time periods and popularity levels represented in retrieved

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

Removing Infrastructure Barriers in Human-Robot Collaboration Through Wireless Reconfigurable Cells

arXiv:2608.09658v1 Announce Type: cross Abstract: Human-Robot Collaboration (HRC) plays a vital role in dynamic, high mix, low volume industrial scenarios such as remanufacturing, which frequently face workcell rearrangements. Traditional setups are constrained by power and data cabling, restricting modularity and reconfigurations, while the selection of commercial wireless devices suitable for real-time perception and safe collaboration are limited in availability. This paper presents a highly flexible, wireless, 5G-based system that serves as a versatile experimental testbed for applications including remanufacturing, operator training, and user studies. To eliminate infrastructure barriers, the workcell integrates a novel battery-powered, multi-sensor platform prototype. Additionally, to support operator safety and system adaptability across environmental shifts, the system integrates a computer vision module for object detection and pose estimation, further augmented for robust han

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

LITEWAY: LIghtweight HAR via Temporal Efficient highWAY

arXiv:2608.09421v1 Announce Type: cross Abstract: Wearable human activity recognition (HAR) remains challenging due to the computational and energy constraints of deep learning models on resource-limited devices. Existing lightweight approaches often rely on recurrent architectures (e.g., GRU and LSTM), limiting parallelism and increasing inference latency. We propose LITEWAY, a modality-agnostic, fully convolutional framework for multichannel sensor time series that replaces recurrent temporal modeling with structured convolutional decomposition. LITEWAY combines lightweight convolutional blocks, strided temporal processing, and convolution-attention pooling to efficiently capture temporal dependencies while reducing computational complexity. We evaluate LITEWAY on 16 HAR datasets against TinyHAR, TinierHAR, and MLP-HAR. LITEWAY achieves competitive macro F1 while reducing model size by 4.06x-9.52x (Light) and 3.87x-9.07x (Full) compared with TinyHAR and TinierHAR. Deployment experime

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

When Confidence Fails: Overconfidence in LLMs under Uncertainty and Missing Clinical Information

arXiv:2608.09080v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks. However, their reliability under uncertainty remains poorly understood which raises critical concerns for deployment in high-stakes clinical settings. In such environments, incorrect predictions are inherently risky, but confident incorrect predictions can be particularly harmful as they may mislead clinical decision-making. In this paper, we conduct a systematic behavioral analysis of LLMs under clinical information uncertainty. We propose an evaluation framework based on the MedMCQA dataset consisting of two complementary uncertainty settings. First, we introduce linguistic uncertainty cues through prompt modifications to simulate ambiguous clinical contexts. Second, we construct an answer removal setting, wherein the correct option is deliberately excluded mandating the model to recognize insufficient information

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

Can Webcam Gaze Constrain Mesa-Objectives in Driving Models? An Instrument Precision Analysis

arXiv:2608.08947v1 Announce Type: cross Abstract: Current hazard detection systems in autonomous driving may develop mesa objectives, learned internal goals that achieve high training performance through spurious correlations rather than genuine hazard recognition. We investigate whether human gaze patterns, captured via webcam-based eye tracking (WebGazer.js), can serve as privileged information to constrain mesa-objective formation. We collected 137,663 frame-level gaze samples synchronized with hazard annotations across 388 real dashcam clips, then test this hypothesis across two calibration protocols (9-point/45-click and 11-point/440-click), two model architectures (Random Forest and causal Transformer), and five random seeds per experiment with paired t-tests. No experiment yields a statistically significant improvement from gaze (p = 0.919, 0.578, and 0.667 respectively). A geometric analysis reveals the root cause: WebGazer's reported error (~130-257 px depending on configurati

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

SHRIMP: Iterative Refinement of Robot Task Plans

arXiv:2608.08884v1 Announce Type: cross Abstract: As collaborative robots have entered domains such as manufacturing, agriculture, and healthcare, programming or adapting robot behavior typically requires robotic expertise that most end users lack. Natural language lowers this barrier. Recent advancements in large language models (LLMs) have made it feasible to translate natural language into robot task plans. However, language-based task specification suffers from semantic ambiguity, and generative models lack transparency for how language instructions become robot actions, making it difficult for users to validate the plan before execution. To address these issues, we introduce SHRIMP, a system that allows users to automatically generate a hierarchical robot primitive plan using natural language and iteratively revise their plan through re-prompting and explicit correction. At each revision, SHRIMP allows users to validate their plan in simulation, and once satisfied, execute it on t

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

Preview-Based Relative-Motion Control of an Insertion Tool for Neural-Thread Placement in Pulsating Tissue

arXiv:2608.08860v1 Announce Type: cross Abstract: Robotic neural-thread placement requires regulating the insertion-tool tip relative to tissue that moves with cardiac and respiratory pulsation. This paper develops a preview-based relative-motion controller that estimates latency-delayed periodic surface motion, predicts it over a short horizon, and uses offset-free model predictive control to regulate relative placement while limiting actuator effort and lateral relative velocity. In MuJoCo, the 1-DOF controller achieves 12.0\um\ free-space and 1.9\um\ contact RMS relative-placement error, versus 18.3/176.8\um\ for delayed-feedback impedance and 286.1/275.5\um\ for lab-frame PD, at the cost of higher peak contact force (3.43 versus 2.00~mN) since offset-free tracking drives the tip fully to the commanded depth rather than yielding against the tissue. In 3 DOF, coupled preview reduces contact lateral shear from 1.34 to 0.50~mm/s with 2.1\um\ lateral RMS error. A feasibility-restored oc

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

Walking through Discussions: A Mobile Visual Analytics System for In-Situ Group Discussion Analysis

arXiv:2608.08617v1 Announce Type: cross Abstract: Group discussion-based teaching is widely used to foster collaborative learning, yet teachers in physical classrooms often struggle to simultaneously monitor multiple groups and quickly diagnose a target group before intervening. Existing visual analytics tools primarily support post-hoc analysis on desktop, providing limited support for in-situ walk-around teaching. To address this gap, we present MobileGroupVis, a mobile visual analytics system for in-situ analysis of classroom group discussions. MobileGroupVis integrates multi-group monitoring, single-group diagnosis, and instructional intervention into a concise analytical workflow tailored for small-screen touch interaction. The system is powered by a lightweight streaming analysis pipeline that converts group audio into structured discussion data and further extracts interaction patterns, topic progression, and topic deviation through a dialogue analysis module. To enable both gla

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis

arXiv:2608.08362v1 Announce Type: cross Abstract: Recent Text-To-Speech (TTS) systems have achieved strong naturalness and zero-shot voice cloning performance, but fine-grained control of expressive speech at the word or phoneme level remains challenging. We propose CtrlSpeech, a controllable, expressive TTS framework with coarse-to-fine control. Built on the DiTAR architecture, CtrlSpeech combines global speaker conditioning with phone-aligned pitch, loudness, and duration signals, enabling localized prosodic control while preserving the target speaker's timbre. This design allows users to adjust expressive attributes at a fine temporal granularity, making speech refinement more flexible and controllable. Experimental results show that CtrlSpeech achieves competitive zero-shot TTS performance and improves controllability over expressive attributes, demonstrating its effectiveness for flexible and practical expressive speech synthesis.

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

Privacy-Preserving Data Drift Detection and Recovery for Large-Scale LLM Applications via Proxy Representations

arXiv:2608.08245v1 Announce Type: cross Abstract: LLM applications deployed at scale face a fundamental challenge: privacy constraints prevent direct inspection of user interactions, making it difficult to obtain any representative evaluation dataset or to track the ongoing evolution of production traffic. We present ProxyDrift, a framework that (i) identifies and measures drift between production traffic and offline evaluation sets, and (ii) constructs and refreshes those evaluation sets accordingly; all without access to raw user data. Our approach operates entirely on non-PII proxy representations: structured, multi-dimensional descriptors derived from LLM-based classification of user interactions. We introduce (1) a chance-calibrated, redundancy-aware (RA) alignment score that aggregates per-dimension drift measurements via mutual information; (2) a conditional sampler that generates synthetic proxies respecting inter-dimensional dependencies; (3) a roundtrip consistency analysis t

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

Focus particles and scalar inferences across humans and language models

arXiv:2608.08227v1 Announce Type: cross Abstract: Focus particles such as "even" and "only" are central to formal semantic theories that posit structured representations over sets of alternatives. "Even" highlights unexpected or extreme alternatives, while "only" enforces exclusivity. If such scalar representations are robust and generalizable, they should give rise to consistent judgments across contexts and systems. In this work, we test whether humans and large language models (LLMs) construct stable scalar representations from sentences containing these particles. Using a dataset of approximately 100 items, participants and models were asked to make scalar judgments. Preliminary results suggest that similar outputs across humans and LLMs may arise from different underlying mechanisms.

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

Spatiotemporal Context-dependent Personalized Movement Compensation in Delayed Telemanipulation

arXiv:2608.08200v1 Announce Type: cross Abstract: Communication delay remains a central challenge in telerobotics, where it disrupts visuomotor coordination and reduces task precision. Motion scaling is an effective countermeasure to delay-induced overshoot, yet typical deployments rely on uniform gains that neglect individual and contextual variability. We propose a human-centered method that fits personalized delay-, direction-, and distance-specific scaling parameters for each participant. We conducted experiments with twenty participants who performed delayed reaching tasks in a virtual simulator. Scaling gains were computed to minimize mean overshoot in simulation in each combination of experimental conditions. Evaluation was done in simulation and on a telesurgical robot to evaluate assistance benefits. Performance was assessed across multiple delays, distances, and movement directions using overshoot, endpoint error, trajectory smoothness, economy of motion, and a composite erro

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.HC

Evaluating and Improving Weak Scalability Analysis of Visualization Algorithms

arXiv:2608.08166v1 Announce Type: cross Abstract: Research on visualizing large-scale datasets traditionally relies on empirical evaluation of scalability, determining the effectiveness of specific computation methods, algorithmic strategies, or implementations. Weak scalability, which assesses the algorithm's performance as problem size and computing resources increase, is a valuable indicator for a method's applicability at scale. However, sufficiently large data sets with increasing size are needed for weak scalability studies. To this end, it is customary to use simple scaling techniques to increase problem size by generating larger input data sets from a base data set. Nevertheless, many visualization algorithms' workload depends on factors beyond input size, such as input data complexity or output size, leading to inaccuracies in the attributed weak scalability. In this work, we highlight different common data scaling methods on multiple algorithms and data sets, recognizing that

Source ↗
Showing 551–600 of 1631 signals
← Prev Page 12 of 33 Next →