EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.HC

Preventive Care Recommendations by Large Language Models

arXiv:2608.06379v1 Announce Type: new Abstract: Preventive care services (PCS) extend life, yet physicians often underprioritize highly effective interventions such as lifestyle modifications (Zhang et al., JAMA Network Open 2020). We evaluated whether large language models (LLMs) replicate and augment physician prioritization of PCS under time constraints. Using Zhang et al.'s validated survey with two patients assessed during long and short visits, we compared seven LLMs with historical physicians. We generated 137 simulated physician personas matching cohort demographics and tested three prompts per model. Primary outcomes were concordance with physician rankings, measured by Spearman correlation, and Consensus-Stratified Agreement (CSA), the proportion of LLM selections rated 4 or higher that matched physician consensus across agreement strata. Secondary outcomes included life-years gained per prioritized choice (LYGPC), consistency, and selectiveness. Augmentation was assessed by

Source ↗
technology Mon, 10 Aug 2026 00:00:00 -0400
arXiv cs.HC

Multimodal Drivers' Emotion Recognition and Safety-Oriented Intervention for Intelligent Transportation Systems

arXiv:2608.06378v1 Announce Type: new Abstract: Driver emotions can affect risk perception, decision-making, and vehicle control under complex road conditions. Existing studies mainly focus on driver emotion recognition, while limited attention has been given to context-aware intervention that jointly considers driver emotion and road perception. This paper proposes a safety-prioritized multimodal driver assistance framework that analyzes speech-derived emotional cues and visual road conditions to generate structured driving interventions. The framework first provides road safety reminders and then generates emotion-aligned verbal support. We construct a multimodal dataset by aligning emotional speech signals with structured road environment descriptors and introduce the CARE (Context-Aware Road-Emotion Evaluation) score to jointly evaluate emotion recognition, risk identification, and intervention generation. Experimental results show that the proposed framework balances environmental

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why

arXiv:2606.00093v2 Announce Type: replace-cross Abstract: Whether a rubric-based LLM judge can replace human annotation is decided by its measured agreement with human labels. Yet the same verdicts can support wildly varying agreement numbers, depending on seemingly minor choices: the judgment scale, the retained cases, the handling of abstentions and invalid outputs, and the pooling of verdicts across items and rubric criteria. The statistics that settle these choices are established, but in psychometrics, econometrics, and corpus annotation rather than in the evaluation practice that needs them. We treat the choices as a measurement protocol that fixes what the reported number estimates before any metric is computed, assemble the relevant results into a single source-attributed analysis, and apply it to three published LLM-judge evaluations. For non-degenerate binary verdicts, Pearson's $r$, Spearman's $\rho$, Kendall's $\tau_b$, the phi coefficient, and the Matthews correlation coef

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Shaping Scientific Explanations to Expert Perspectives with Persona-Conditioned Reinforcement Learning

arXiv:2603.21846v2 Announce Type: replace-cross Abstract: Explainable AI is increasingly important to scientific discovery. However, existing methods largely ignore that explanation quality is not universal: experts differ in how they assess evidence, prioritize mechanisms, and construct explanatory narratives. We introduce perspective-conditioned explanations, a framework for adapting explanation generation to epistemic variation in expert judgment. Using knowledge graph reasoning paths in drug discovery, we show that preferences organize into coherent epistemic perspectives that can be captured by agentic personas, representations of how experts evaluate explanations. Persona-aligned rewards then guide reinforcement learning-based explanation generation without large-scale expert supervision. Expert user studies show that perspective-conditioned explanations are preferred over general-purpose explanations and improve perceived relevance and validity. Moreover, they match or exceed st

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Combining Large Language Models and Symbolic Reasoning for Multi-Robot Temporal Planning through Explainable Knowledge Bases

arXiv:2502.19135v2 Announce Type: replace-cross Abstract: We present PLANTOR, a framework for generating and executing multi-robot task plans from natural-language task descriptions through LLM-assisted knowledge-base construction. The approach uses large language models to synthesize a structured Prolog knowledge-base, applies consistency checks to detect and repair modeling errors, generates a high-level symbolic plan, refines it into low-level robot actions, and computes a temporally optimized schedule that is converted into an executable behavior tree. The framework is designed to preserve inspectability by exposing the generated knowledge-base, intermediate plans, and scheduling constraints. We evaluate the approach on scenarios inspired by the Blocks World and Grippers benchmark across multiple language models, and we report both the quality of generated knowledge-bases and the runtime of the planning pipeline. We further demonstrate end-to-end execution in a real multi-arm assem

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Generative AI in Action: Field Experimental Evidence from Alibaba's Customer Service Operations

arXiv:2603.29888v2 Announce Type: replace Abstract: In collaboration with Alibaba, we study how a generative AI assistant affects service performance in e-commerce after-sales operations. In a large-scale field experiment, human agents providing digital chat support were randomly assigned access to a gen AI assistant. The assistant drafts issue diagnoses and solution proposals in the opening stage only; agents can adopt, modify, or disregard them. Because of this discretion, we estimate the effects of both gen AI access and usage. On average, gen AI improves service speed and subjective service quality, measured by customer ratings, but has no significant effect on objective service quality, measured by customer retrials. These gains come from more than automation. Gen AI reshapes agent-customer interactions: treated agents respond faster and take a more proactive role, while customers provide less input; both patterns persist into later chat stages. These average effects, however, mas

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

arXiv:2607.29602v1 Announce Type: cross Abstract: Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic ice-breaker conversation. Every pair answers the same type of prompt, so only the manner of interaction can reveal the answer. Across text, audio, and video, we compare 26 models from seven companies against matched human panels over 96 balanced dyads. The best model and the human crowd are statistically indistinguishable on accuracy in every modality, but reach it differently: humans stay balanced across the two answers, while the strongest models lean toward "stranger"---a difference in effective prior, not discrimination. Richer channels help both unequally, and only humans gain from visible behavior on top of speech. We release the stimuli, human ratings, and model predictions.

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

STAGE: STyle-controllable Action GEneration for personalized autonomous driving

arXiv:2607.29517v1 Announce Type: cross Abstract: Driving style refers to the behavioral preferences that drivers maintain during driving, shaped by their diverse experiences, habits, and needs, and is typically reflected in varying levels of aggressiveness. If humans choose to use autonomous driving systems, they would expect the driving style of the systems to closely resemble their own habit. However, this is challenging for current industrial autonomous driving systems. To address this, we developed a style controllable action generation method, STAGE, for driving tasks. Its training process is based on imitation learning, incorporating both style value and latent value action modality encoding. Preference learning is then used to identify the user's driving style as a continuous, monotonic style value. And to reduce the cost of human involvement in the preference training process, we also developed a set of rules to compare driving style in data pairs. Then, during inference, the

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

YazSes: An Offline, Privacy-First, Cross-Platform Hold-to-Talk Voice-Dictation System

arXiv:2607.28878v1 Announce Type: cross Abstract: Cloud voice-dictation services deliver strong accuracy but require streaming a user's speech to a remote provider, an unacceptable trade-off in privacy-sensitive professions and offline or air-gapped settings; the leading on-device alternatives are either platform-locked or aimed at expert scripting rather than plug-and-play dictation. We present YazSes, an open-source (Apache-2.0) hold-to-talk voice dictation daemon that runs entirely on-device, with a single codebase targeting Linux, macOS, and Windows through a protocol-based platform abstraction. YazSes transcribes speech locally with faster-whisper (CPU, int8) and injects the result into the focused application; a fast regex command grammar, backed by an optional small-language-model router, maps utterances to editor and terminal actions. Nothing leaves the machine: recording is push-to-talk rather than always-listening, there is no telemetry, and an opt-in personalization loop kee

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Data Visualization Style Guides in Practice: Why They Emerge, How They Work, and When They Bend

arXiv:2607.29645v1 Announce Type: new Abstract: Visualization style guides play a crucial role in shaping how data is interpreted and trusted, yet they often receive little scrutiny in their creation and use. Understanding their impact requires looking beyond the specific rules that style guides prescribe and examining how they function within organizations to coordinate visual work, manage trade-offs, and support judgment under real constraints. Analyzing interviews with nine authors of twenty-six style guides across journalism, government, industry, and the public sector, we reveal how these guides reflect the specific challenges of their organizations, including consistency, training, governance, and accountability. Our study highlights the tensions between standardization and flexibility, guidance and discretion, and automation and human oversight. We propose PRISM, a socio-technical framework that characterizes visualization style guides by their Purpose, Rules & Mechanisms, Insti

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Exploratory Integration of EEG Spectral Features and Gaze Variability for Mild Cognitive Impairment Discrimination

arXiv:2607.29493v1 Announce Type: new Abstract: Early detection of mild cognitive impairment (MCI) is an important challenge in aging societies. Electroencephalography (EEG) and eye-tracking have independently been explored as potential biomarkers; however, their integrative effects remain insufficiently examined. This exploratory study investigated whether combining EEG spectral features with gaze variability may provide complementary information for MCI discrimination. EEG signals were recorded using the 10--20 system, and spectral power features were extracted. We compared three models: (a) high-dimensional EEG features, (b) L1-regularized feature selection (LASSO), and (c) integration of the selected EEG features with gaze variability. Performance was evaluated using leave-one-out cross-validation and area under the ROC curve (AUC). Model (a) yielded limited discrimination (AUC = 0.52). Feature selection increased AUC (0.64), and additional integration of gaze variability further i

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

SAVVY: Student Attention Visualization for Video-based Learning Analysis

arXiv:2607.29413v1 Announce Type: new Abstract: Video-Based Learning (VBL) has become a popular delivery medium of education in the past decade, ranging from online education to hybrid learning. Students' rising expectations for video quality have motivated teachers to enhance the design of instructional videos before releasing them. Analyzing the attention of pilot cohorts in advance has become a conventional optimization strategy to guide course improvement. However, existing attention quantification algorithms are highly susceptible to noise in real-world environments, degrading estimation accuracy. Moreover, even when attention data are available, teachers must still invest substantial effort in empirical revision attempts, limiting practical feasibility. To address these challenges, we first propose a novel attention modeling framework based on multimodal brain signals that enables stable tracking of student attention levels. We then develop SAVVY, a novel interactive visual analy

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

An Algorithmic Perspective on Information Visualization

arXiv:2607.29360v1 Announce Type: new Abstract: Information visualization is inherently a field that brings together various research domains. Roughly speaking, we may identify two perspectives: the design perspective, revolving around how to ensure that a human can work effectively with the visual representations of data and the tools that offer them, and the algorithmic perspective, focusing on how to automatically create such visual representations. Munzner's model for visualization design places design choices before algorithmic considerations. It offers predominantly a design perspective; as a consequence, applications of this model may consider the algorithmic perspective as an afterthought, bypassing a step that translates the design into the formalism necessary for algorithmic study. As a result, the design may be entangled with the algorithms used to compute a visualization. Focusing on layout algorithms, we explore the ramifications of this entanglement: quality often goes un

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

The persuasive power of large language models does not depend on their perceived national origin

arXiv:2607.29334v1 Announce Type: new Abstract: Conversational AI developed by geopolitical rivals reaches citizens worldwide, raising concerns that it could sway public opinion or be rejected as foreign propaganda, with consequences for democratic discourse and information sovereignty. Yet, whether an AI's perceived national origin shapes its persuasive power is unknown. In a preregistered randomized experiment, 403 adults from a nationally representative United States sample held a three-round debate with a chatbot introduced as either American ("DiscoveryAI") or Chinese ("ZhengheAI"), discussing a political or non-political topic. In all conditions, participants actually conversed with the same model (GPT-4o), instructed to argue against their initial position. We combined pre- and post-conversation self-reports of attitudes, trust, and collective narcissism with computational analyses of 1,209 participant turns, including LLM-coded stance and argumentative conduct, stance-sensitive

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Designing a digital word-learning intervention with neurodiverse children: Experiences and ideas from children with developmental language disorder

arXiv:2607.29113v1 Announce Type: new Abstract: Introduction - Developmental language disorder (DLD) is a neurodevelopmental condition often characterised by word-learning difficulties that can lead to significant social and academic challenges. The disorder shares some features with other neurodevelopmental conditions such as autism spectrum disorder (ASD). Despite affecting 7 percent of children, the condition has received little coverage in participatory design research. This paper addresses this by reporting on the emotional responses and design outputs of children with DLD following participatory sessions to inform a word-learning intervention. Method - Principles from learner-centred, cooperative, and accessible co-design approaches were integrated to tailor activities for four children with DLD. Design sessions were refined through ongoing monitoring of the children's experiences. The data that informed the findings included design artefacts such as children's drawings, structur

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth

arXiv:2607.28890v1 Announce Type: new Abstract: Evaluations of LLM-assisted qualitative coding almost universally measure model performance as agreement with human coders, a practice that presumes human coding is the standard to approximate. This study provides empirical evidence that the presumption fails in ways agreement metrics cannot detect. Five LLM systems and three trained human coders independently applied a 72-item hierarchical codebook to 2,560 educator messages from a K-12 AI platform. Beyond conventional agreement analysis, an independent domain expert judged 855 pairwise comparisons of code sets blind to source, treating human and machine sources symmetrically. The two evaluation approaches diverge in both directions. Human-LLM agreement (mean Jaccard 0.30) falls well below human-human agreement (0.52), which standard practice would read as inferior LLM coding, yet the blind verifier preferred human and LLM coding at indistinguishable rates (51.5% vs. 48.5%, p = 0.537), a

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Human-LLM Collaborative Inductive Coding for Conceptualizing K-12 Educator AI Use

arXiv:2607.28889v1 Announce Type: new Abstract: Qualitative researchers increasingly encounter interaction corpora whose scale exceeds what manual coding alone can address, and large language models (LLMs) are frequently proposed as analytic assistants. The open questions are not whether LLMs can participate in qualitative analysis but to what extent, in what phases, and under what safeguards. This article provides a detailed procedural account of a multi-phase human-LLM collaborative pipeline that adapted open, axial, and selective coding to develop a hierarchical codebook from 45,000 messages exchanged between K-12 educators and a generative AI platform. Across three phases, LLMs generated candidate labels and structured annotations at scale, while human researchers retained conceptual authority over category definitions, merging decisions, and interpretive frameworks. The resulting instrument was then tested through systematic human coding, in which three trained coders with educati

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Spatial Visual Analytics for Multi-Document Summary Verification

arXiv:2607.28853v1 Announce Type: new Abstract: Large language models increasingly generate summaries from collections of documents to support sensemaking and reporting, but verifying whether summary statements are grounded in source materials remains difficult. In multi-document summarization (MDS), evidence is distributed across many source documents and may be incomplete, conflicting, or missing. We present Summary Verification Space (SVS), a visual analytics system for verifying multi-document summaries through spatial document organization and coordinated provenance visualization. To support scalable verification, we investigate two alternative 2D canvas layouts: a SUMMARY-GUIDED layout that organizes documents by alignment with summary sentences, and a SOURCE-GUIDED layout that arranges documents by semantic similarity. Coordinated provenance visualization then makes relationships among summary content, source documents, and supporting evidence explicit, enabling users to trace s

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Guided Exploration of Iterative Schedule Modifications: A Design Study on Railway Traction Unit Scheduling

arXiv:2607.28694v1 Announce Type: new Abstract: Traction unit scheduling in large railway networks involves complex operational constraints: multi-objective optimization produces feasible circulation plans under ideal assumptions, while simulation is required to assess their robustness under realistic operating conditions. A critical refinement mechanism relies on crossing operations, in which co-located traction units exchange their remaining schedules to reduce delay propagation. The space of possible crossing sequences, however, grows exponentially. Existing tools provide limited support for identifying promising candidates, evaluating their impact, and managing the resulting exploration. We present an interactive visual exploration approach that tightly couples schedule visualization, simulation-based evaluation, and a three-level guidance mechanism to support the systematic exploration and interactive optimization of traction unit circulation plans. The system renders the circulat

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Unanticipated Effects of Generative AI on Expertise Pathways and Performance Perception in System Administration

arXiv:2607.28650v1 Announce Type: new Abstract: While industry discourse often emphasizes immediate productivity gains and frames GenAI primarily as a tool for automation, the integration of GenAI into system administration may involve deeper shifts in professional practice that are not yet fully understood. Drawing on 14 semi-structured interviews with IT professionals, this paper explores the lived reality of embedding GenAI into daily routines of troubleshooting, scripting, and system verification. Through inductive thematic analysis, we uncover two unanticipated socio-technical findings. First, we describe a "compression of traditional expertise pathways" where GenAI appears to function as both a mentor-like tutor and a "ladder-shortening" tool. While the tool can support faster task performance in unfamiliar domains, our findings suggest it may also reduce a practitioner's exposure to the foundational, hands-on cycles of building, failing, and debugging that historically served as

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention

arXiv:2607.28649v1 Announce Type: new Abstract: COSI-Lab presents a multimodal, multi-sensor dataset of an interdisciplinary scientific workshop containing 32 academics at an international conference. It captures ecologically valid social interactions in a weakly scripted setting consisting of two 30-minute mingling sessions with real professional and social consequences for the participants involved. We argue that future intelligent systems could be better equipped to handle subjective perceptions by modeling their multiplicity not as label noise but as a explainable perspective-driven reasoning process. We focus on the Apparent Intent Inference (AII) problem as determined by ex-situ observers and conceptualize intentions to be independent of manifest future outcomes. We contribute 1. a novel annotation process for AII that accounts for a perceiver's own interpretative tendencies, 2. quantitative and qualitative analyses of intent narratives with respect to diversity, grounding, and p

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Why It Hurts: Identifying the Drivers of Negative Thoughts in Emotional Support Conversations

arXiv:2607.28648v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used for emotional support tasks, such as negative thought reframing. This task relies on modifying cognitive appraisals, the subjective interpretation of events that elicit negative emotions, which is typically conceptualized along multiple discrete dimensions. Current LLM-based frameworks model cognitive appraisal by exhaustively evaluating all possible dimensions, but they fail to account for the varying saliency of these dimensions across different contexts. In this work, we investigate a vital yet overlooked question: "Can LLMs infer the salient appraisal dimensions from emotional support conversations?" To address this question, we introduce the AppraiSal benchmark, containing 996 emotional support conversations with human-annotated mental states, including salient cognitive appraisal dimensions. Furthermore, we propose PRISM, a multi-agent probabilistic framework grounded in Bayesian In

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

ConnectED: A Curriculum-Aligned AI System for Vietnamese Instructional Lesson Planning and Student Learning

arXiv:2607.28647v1 Announce Type: new Abstract: This paper presents ConnectED, a human-centered AI system that supports the full instructional lifecycle in Vietnamese education by linking curriculum-aligned lesson design, interactive student learning, and feedback-driven refinement. Built on VietEduQwen, a Vietnamese educational large language model trained via supervised fine-tuning and direct preference optimization, the system ensures academically accurate, pedagogically appropriate, and student-safe interactions. ConnectED operationalizes the ADDIE framework through structured prompt templates aligned with Official Dispatch No. 5512/BGDDT-GDTrH, where each phase serves as both a generation step and a teacher validation gate. The Evaluation phase further closes the loop by connecting student performance data with iterative lesson improvement. Beyond lesson generation, the system integrates a student-facing interactive environment, enabling continuous collection of learning signals t

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

"YES! YES! I absolutely love this insight!" Affirmative Narration as Interactional Strategy in Dialogues with LLM Chatbots

arXiv:2607.28646v1 Announce Type: new Abstract: This article analyses narrative mechanisms that are common in dialogues with LLM chatbots. In combination, these mechanisms produce an interactional strategy for maximising user engagement, which we call affirmative narration. Affirmative narration serves to convince users of the chatbot's utility. We analyse three narrative mechanisms that support affirmative narration in human-LLM dialogues: firstly, guiding the user to view the chatbot as an intelligent and reliable character; secondly, activating masterplots, culturally significant and recurring story templates; and thirdly, using characters and masterplots not only to affirm, but also to isolate the user. The case studies range from a journalist's unsettling chatbot experiment to cases where users have experienced delusions or even died by suicide after lengthy interactions with a chatbot. The analyses illustrate the worrying sides of affirmative narration, and the article thus concl

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation

arXiv:2607.28645v1 Announce Type: new Abstract: Recent multimodal large language models can convert visual designs directly into executable code, but real mobile products require multiple screenshots to become a buildable codebase with shared components and working navigation. This project-level setting exposes three limits of existing design-to-code benchmarks: they focus on single-page generation rather than complete codebases, cannot evaluate cross-page navigation, and do not measure project-wide maintainability. We introduce MobileForge, the first benchmark for project-level multi-screen mobile app generation, comprising real mobile apps, human-reviewed screens, structured page-relationship annotations, and navigation test specifications. MobileForge supports five-axis evaluation of build, navigation, visual fidelity, code maintainability, and efficiency. We also propose state-isolated navigation testing to avoid cascading failures in navigation evaluation and an anchor-referenced

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Seeing Differently: Modeling Interpretive Perspectives in Computational Creativity using a Four-World Framework

arXiv:2607.28644v1 Announce Type: new Abstract: Creativity in computational systems is often evaluated as an objective property of artifacts, with existing Computational Creativity (CC) frameworks assessing creative merit at the level of outputs or systems rather than interpretive context. However, artistic meaning is inherently perspective-dependent and can vary across viewers and critical traditions. This paper proposes a computational approach to modeling interpretive perspectives rather than treating creativity as a single measurable construct. The study adopts a twelve-trait creativity framework, organized across four conceptual domains, and operationalizes it through three evaluative personas: formalist, social-historical, and iconographic. Using 1,069 artworks from the SemArt dataset, the analysis generates 38,484 persona-based evaluations to examine how perspectives shape creativity assessment. Results show systematic divergence across perspectives, with traits such as Social R

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

To Facilitate or not to Facilitate: Human and LLM Facilitator Tendencies in Online Discussions

arXiv:2607.28643v1 Announce Type: new Abstract: Automating facilitation in online discussions is a long-standing social concern given the increasing time we spend on online spaces and the failure of content moderation approaches. While studies have been conducted on how to facilitate, none have answered the essential question of when to do so. A potential answer is using LLMs, which ostensibly make automated, large-scale intervention increasingly feasible. In this study, we examine when LLMs decide to facilitate by defining what facilitation is, observing when humans decide to facilitate, and comparing their decisions with those made by LLMs. To this end, we create PEFK, a corpus standardizing and aggregating all relevant facilitation datasets. We are the first to run a survey on facilitation timing, which we execute using expert facilitative participants and LLM-as-a-judge models. We discover that while humans are more cautious, LLMs are excessively eager to facilitate, although both

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

arXiv:2607.09306v3 Announce Type: replace-cross Abstract: Behavioural auditing asks whether a language model behaves as it claims, but detection scores are reported without separating two targets: whether a reply was produced under a behaviour-inducing condition (exposure) and whether the behaviour surfaced in it (manifestation). Scoring a compact 146-million-parameter auditor's frozen-representation read-out and a frontier judge against each label on the identical 720 replies, the gap between the instruments moves by roughly 0.2 AUROC when the target changes. Under the judge's deployed interface, a single verdict, the ranking reverses: the auditor leads on exposure, 0.804 against 0.718, and trails on manifestation, 0.690 against 0.811. Matching the output resolution from either direction, by asking the judge a target-specific question answered with a continuous confidence score or by thresholding the auditor's read-out, removes the reversal but not the interaction, which excludes zero

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

To Ban or not to Ban? How Open Source Projects Govern GenAI Contributions

arXiv:2603.26487v2 Announce Type: replace-cross Abstract: Generative AI (GenAI) is playing an increasingly important role in open source software (OSS). Beyond completing code and documentation, GenAI is increasingly involved in issues, pull requests, code reviews, and security reports. Yet, cheaper generation does not mean cheaper review - and the resulting maintenance burden has pushed OSS projects to experiment with GenAI-specific rules in contribution guidelines, security policies, and repository instructions, even including a total ban on AI-assisted contributions. However, governing GenAI in OSS is far more than a ban-or-not question. The responses remain scattered, with neither a shared governance framework in practice nor a systematic understanding in research. Therefore, in this paper, we conduct a multi-stage analysis on various qualitative materials related to GenAI governance retrieved from 67 highly visible OSS projects. Our analysis identifies recurring concerns across co

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

Sonic Stage: Auto-Generating Interactive Spatial Soundscapes to Facilitate Dialogue Video Comprehension for Blind Viewers

arXiv:2607.20835v2 Announce Type: replace Abstract: Audio description (AD) makes film and television accessible to blind and low-vision (BLV) audiences by narrating characters' actions. However, in scenes with lots of dialogue, AD often omits important actions because it is constrained not to overlap with speech. It is not yet known how to convey characters' actions during dialogue. We present Sonic Stage, a system that transforms dialogue videos into interactive spatial soundscapes, enabling BLV audiences to intuitively understand characters' actions and movements through immersive auditory cues. Sonic Stage conveys essential visual information during dialogue through three auditory techniques: (1) spatialized dialogue to represent spatial layout, (2) diegetic sound to convey character actions, and (3) interactive descriptions to provide context-specific visual details. Evaluation with 12 BLV viewers showed that Sonic Stage significantly improved video comprehension, spatial presence,

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

Creative Reading: Scaffolding Reading for Transformation

arXiv:2606.04308v2 Announce Type: replace Abstract: Reading augmentation systems increasingly help readers process text at scale. While these tools address real constraints of time and cognitive load, they often implicitly frame reading as information transmission, or "reading to discard," delegating interpretation and effort to the machine. Yet this delegation changes the outcome of reading. For example, in scholarly reading, deciding what a research text implies and why it matters is central to the work of scholarly production. We propose creative reading as an alternative goal: reading augmentation that supports readers in creating both readings and themselves as readers. By putting literary and narrative theories into conversation with scholarly sensemaking and creativity support, we present a provocation-oriented design space for valuing the process of reading as a way of preserving a plurality of readings and transforming readers over time.

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

StateScribe: Towards Accessible Change Awareness Across Real-World Revisits

arXiv:2604.23749v2 Announce Type: replace Abstract: Real-world environments evolve continuously, yet blind and low-vision (BLV) individuals often have limited access to understanding how they change over time. Unexpected or relocated objects, layout modifications, and content updates (e.g., price changes) can introduce safety risks and cognitive burden. While existing visual assistive technologies can describe immediate surroundings, they operate as one-off interactions and lack mechanisms to surface meaningful changes across revisits. Informed by a survey of 33 BLV individuals, we develop StateScribe, a system that supports accessible awareness of real-world changes across revisits. StateScribe employs a dual-layer memory architecture that integrates episodic scene memory and object-centric temporal memory to enable scalable and structured change tracking. It provides both live descriptions of the current scene, and descriptions of what has changed, when and where it occurred across r

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

Linking Heterogeneous Data with Coordinated Agent Flows for Social Media Analysis

arXiv:2510.26172v2 Announce Type: replace Abstract: Social media platforms generate volumes of heterogeneous data, capturing user behaviors, textual content, and network structures. Analyzing such data is crucial for understanding phenomena such as opinion dynamics, community formation, and information diffusion. However, discovering insights from this complex landscape is exploratory, conceptually challenging, and requires expertise in social media mining and visualization. Existing automated approaches, including large language models (LLMs), remain largely confined to structured tabular data and cannot adequately address the heterogeneity of social media analysis. We present SIA (Social Insight Agents), an LLM agent system that links heterogeneous multi-modal data, including raw inputs (e.g., text, network, and behavioral data), mined analytical results, and rendered visual artifacts, through coordinated agent flows. Guided by an insight-oriented taxonomy connecting insight types wi

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

AI-Assisted Data Extraction for Systematic Reviews in Education

arXiv:2501.11840v2 Announce Type: replace Abstract: Systematic reviews are time-consuming endeavors that require knowledgeable human reviewers to screen studies for relevance and extract data following a specific coding scheme before any analysis or synthesis can occur. Large language models (LLMs) hold promise for substantially accelerating this process and reducing reviewer workload, yet their application within the context of systematic reviews in the field of education remains underexplored. We address this issue in two ways: through empirical studies and the iterative development of an open-source software tool. First, we conducted two empirical studies examining the efficacy of using LLMs for data extraction using data from a published review on pedagogical agents. We extracted a variety of data types from 112 studies and compared the results to data extracted by human coding. Results indicate that LLMs struggled with extracting data accurately and therefore are not ready to be u

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

Personalizing Privacy Protection With Individuals' Regulatory Focus: Would You Preserve or Enhance Your Information Privacy?

arXiv:2402.17838v2 Announce Type: replace Abstract: In this study, we explore the effectiveness of persuasive messages endorsing the adoption of a privacy protection technology (IoT Inspector) tailored to individuals' regulatory focus (promotion or prevention). We explore if and how regulatory fit (i.e., tuning the goal-pursuit mechanism to individuals' internal regulatory focus) can increase persuasion and adoption. We conducted a between-subject experiment (N = 236) presenting participants with the IoT Inspector in gain ("Privacy Enhancing Technology" -- PET) or loss ("Privacy Preserving Technology" -- PPT) framing. Results show that the effect of regulatory fit on adoption is mediated by trust and privacy calculus processes: prevention-focused users who read the PPT message trust the tool more. Furthermore, privacy calculus favors using the tool when promotion-focused individuals read the PET message. We discuss the contribution of understanding the cognitive mechanisms behind regul

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

Snapshot plots: displaying summary tables as parallel univariate plots with consistent color highlighting

arXiv:2607.28302v1 Announce Type: cross Abstract: For empirical studies, social and health scientists give background characteristics of their sample and summarize them in the famous "Table 1". When treatment/ control groups are present, this table gives summary statistics by group to see whether the background characteristics differ by group. We propose snapshot plots --- parallel univariate plots with consistent highlighting --- to visualize such tables. Compared to "Table 1", such plots are designed to facilitate comparisons of background characteristics --- in particular among groups --- and give more detail on numerical variables. We provide a web app as well as a python implementation of snapshot plots. Snapshot plots arise as edge cases of hammock plots (parallel coordinate plots for mixed categorical/ numerical data). We demonstrate the usefulness of snapshot plots for two "Table 1"s.

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm

arXiv:2607.27851v1 Announce Type: cross Abstract: Emotional dialogue research includes two influential strategy traditions. Empathetic dialogue prioritizes understanding a speaker's emotional experience. Emotional support conversation selects and sequences support for the seeker's current needs. Sustained use introduces a further goal. Effective support should sustain users' capacities for emotion regulation, coping, self-endorsed decisions, and social connection across the interaction lifecycle. We propose capability-sustaining emotional dialogue (CSED) as a longitudinal research paradigm that aligns supportive strategy with this goal and organizes data, models, system design, evaluation, and governance around repeated use, non-use, transition, and termination. A targeted literature-and-corpus audit motivates this position. In a PRISMA-ScR-guided sample, 95% of 60 system-building papers pursue relief-oriented goals. None evaluates capability or longitudinal outcomes, and only 1 consid

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

Measuring Alignment With Reader Highlights Net of Position and Length

arXiv:2607.27739v1 Announce Type: cross Abstract: Context compression discards most of a document before a language model reads it, and is normally evaluated by downstream task accuracy - which makes another model the judge of what mattered. Naturalistic social highlighting offers a non-circular reference: many people independently marking passages on the same page. But the obvious metric, the fraction of crowd-marked sentences a compressor keeps, is confounded twice: crowd marks are front-loaded and crowd-marked sentences are longer, so any method favouring early or long sentences scores well regardless of readers. We remove both by matching each marked sentence against unmarked sentences of the same document at equal relative depth and equal within-document length rank, and we calibrate every estimator on synthetic nulls built from position and length alone - a step that matters, since depth-only stratification returns a false positive on 20-36% of nulls containing no effect. On 120

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

Recognition and Label-Free Adaptation Across Recording Sessions in Surface-EMG Gesture Decoding

arXiv:2607.27568v1 Announce Type: cross Abstract: Recognition accuracy obtained during a recording session does not persist when a user puts on the electrodes again after the electrodes had previously been removed. The electrodes may have moved slightly, the skin may be drier or wetter, or the elbow may be positioned differently; these factors all contribute to day-to-day variability and therefore represent a major obstacle to implementing successful pattern-recognition based myoelectric control systems in daily practice. However, simply recalibrating a user's hand for 20 min at every doff/don event is a clearly unrealistic expectation. A montage-agnostic encoder built for cross-user, cross-montage transfer is trained here using data collected during a particular recording session, and then applied to data collected later in a different recording session without adjusting anything, on the ten intact subjects of NinaPro DB6. The performance of this approach is compared to that of a per-

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

A Montage-Agnostic Encoder for Calibration-Light Cross-User Gesture Recognition from Surface Electromyography

arXiv:2607.27565v1 Announce Type: cross Abstract: Pattern-recognition control promises a myoelectric prosthesis that responds to many intended gestures rather than one or two, but the promise has stayed in the laboratory. A recogniser trained on one person rarely transfers to the next, and useful performance usually demands a fresh round of labelled calibration from the end user. A montage-agnostic encoder is introduced that reads each electrode with shared weights and locates it by its physical coordinate rather than its index, so one architecture ingests any channel count without montage-specific parameters. Trained across users, it exceeds a per-user Hudgins and linear-discriminant classifier by 0.234 macro-F1 on DB1 for every held-out subject and by 0.108 on DB2, and falls below it on the ten-subject DB5. Each of the encoder's three key components individually accounts for more than half of its 3-shot macro F1 in an otherwise budget-matched ablation study. A controlled subject-coun

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

FADEx: Feature Attribution and Distortion-based Explanation of Dimensionality Reduction

arXiv:2607.27463v1 Announce Type: cross Abstract: Dimensionality Reduction (DR) is a fundamental tool for high-dimensional data exploration, reducing the complexity of latent spaces of machine learning models, and assisting in the explanation of complex opaque models. However, non-linear DR techniques often function as opaque transformations themselves, making it challenging to understand how individual features influence instance positioning in the reduced space. This lack of transparency complicates the analysis and interpretation of structural patterns, hindering the ability to reason about the organization of high-dimensional data based on the projected layout. In order to address this challenge, dimensionality reduction explanation methods have shown promise in improving the understanding of the observed groups and cluster structures. Unfortunately, existing DR explanation approaches tend to suffer from limitations such as multiple attributions per feature and restricted applicabi

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

CrossAtlas: Evaluating Projection Techniques for Spatial Referencing in Cross-Reality Collaboration

arXiv:2607.28583v1 Announce Type: new Abstract: Cross-reality collaboration increasingly connects immersive and desktop users within synchronized workspaces, yet little is known about how bidirectional projection techniques between immersive 3D layouts and desktop 2D views influence communication. Spatial referencing depends on shared spatial understanding, but different mappings preserve and distort geometric relationships in different ways, altering perceived adjacency, orientation, and coverage across collaborators' views. We present CrossAtlas, a synchronized PC-VR collaboration platform that integrates multiple bidirectional projection techniques, including three planar projection variants and equirectangular, a spherical projection variant, across layouts of varying curvature. In a controlled study with 24 dyads, collaborators completed spatial referencing tasks under different projection-layout conditions while we collected performance and subjective measures. Our results show t

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

Multi-Session User Experience Assessments of Computationally Optimized Automated Vehicle Functionality Visualizations

arXiv:2607.28552v1 Announce Type: new Abstract: Understanding automated vehicles (AVs) is crucial to improving their acceptance. Numerous approaches to visualizing relevant traffic information to passengers have been proposed and empirically evaluated. As this is time-consuming, costly, and reduces the possible design parameters, we employed multi-objective Bayesian optimization to optimize the design of visualizations in AVs. In particular, we evaluated multi-session aspects involving iterative optimization. We optimized the design for passenger trust and perceived safety while minimizing cognitive load. Results from an online study (N=74) show that this method effectively identifies visualization design parameter values that improve trust, safety, and predictability while making the design process more efficient and scalable. However, shortcomings of the computational approach when optimizing for subjective measurements are highlighted and discussed.

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

Effects of Auditory Information for People With Visual Impairments in Highly Automated Vehicles

arXiv:2607.28544v1 Announce Type: new Abstract: Automated vehicles promise to improve accessibility and access to personal mobility for everyone. However, their design and current research trends in visualizing relevant information do not reflect this commitment to accessibility for users with visual impairments. Therefore, we designed and implemented a visual and auditory communication concept for people with visual impairments seated inside fully automated vehicles. Furthermore, in an online video-based study (N=35, 12 with visual impairments), we compared three levels of auditory information communication: low (safety-relevant information only), medium (additionally including vehicle control and route updates), and high (additionally including sightseeing and destination information). Results showed that trust and user experience significantly improved with additional information, with a corresponding, albeit less robust, effect on perceived safety. However, they also revealed that

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

Identifying a Level-up Pathway for AI-assisted Counterspeech through Elaboration

arXiv:2607.28239v1 Announce Type: new Abstract: Given the profound societal impact of vaccine-skeptical content on social media, community-driven counterspeech has emerged as a promising participatory response to contest and curb such objectionable content. Yet crafting effective counterspeech remains challenging for ordinary users, limiting their willingness and ability to engage constructively. We designed and evaluated three generative AI-assisted counterspeech writing systems that vary by assistance stage (co-writing vs. re-writing) and mode (guided vs. unguided) to support lay users' responses to vaccine-skeptical content. We ask whether AI can help users craft counterspeech perceived as both effective and authentic, which forms of AI support work best, and through what mechanisms. In a randomized controlled trial with social media users, participants wrote counterspeech responses to both statistical and narrative vaccine-skeptical content. Across evidence types, AI-assisted writi

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

Toward Annotation-Efficient Continuous Emotion Arousal Quantification via Group-Level EEG Dynamic Neural Synchrony

arXiv:2607.28204v1 Announce Type: new Abstract: Continuous emotional arousal quantification remains bottlenecked by time-consuming and labor-intensive manual annotation. This work investigates group-level EEG dynamic neural synchrony (DNS) as a principled signal for continuous arousal quantification that bypasses per-subject manual labeling. Using Correlated Component Analysis (CorrCA) with sliding-window computation across four EEG datasets spanning 142 subjects and over 207 hours, we systematically evaluate DNS as a group-level marker for emotional arousal dynamics. Three key findings emerge. First, DNS exhibits significant emotion information from valence-dependent differences (all p<0.003), with positive emotions eliciting higher synchrony. Second, DNS correlates more strongly with the first-order derivative of arousal than with raw arousal values, revealing that neural synchrony captures the rate of emotional change rather than static intensity. Third, we provide the first systema

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

Student Perceptions and Preferences Regarding AI-Generated Instructional Videos in Computing Education

arXiv:2607.28203v1 Announce Type: new Abstract: Students differ in how they prefer to engage with learning resources, with some favoring textual materials and others visual or video-based content. Recent advances in generative AI have led CS education research to focus on text-based AI tools for developing learning resources. However, advances in AI video models and the rapid proliferation of AI video generation tools have made it possible for instructors to create high-quality personalized educational videos efficiently and cost-effectively. Understanding students' perceptions of AI-generated videos is thus critical for helping CS instructors know when and how to use them purposefully. To address this gap, we conducted a descriptive post-test survey study in which 170 computing students at two U.S. institutions watched three 3-minute AI videos on the Markdown markup language created with Knowlify. Students then completed a survey about their perceptions of the Markdown videos and thei

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

A Mathematical Framework for Reading the Autopsias' Meta - Compositional System

arXiv:2607.28155v1 Announce Type: new Abstract: Background. New forms of music writing using computers have arisen in the past 20 years. Most of them use the capacities of digital manipulation of data like animation, algorithmic processing, cinematic view, and much more. These scores use dynamic musicography, and all of them share a problem. They have readability problems. We will argue that this problem can be addressed by mathematical tools. Aims. Take the Autopsias [Autopsies] meta-compositional system as a study case for starting the construction of a mathematical framework that can overview the readability for cynetic musicography. The Autopsias system is the process of transforming a musical score dynamically. Methodology. We will start to build a mathematical framework by taking a group of basic topological concepts, and applying them after a bridge between a printed score and a dynamic computational re-appropriation of it has taken place. We will observe the writing and the per

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

Investigating Effective Uncertainty Visualizations for Ordinal Crowdsourced Data of Crowding Conditions

arXiv:2607.28072v1 Announce Type: new Abstract: Commuters often encounter crowding in railway systems, particularly in queues where passenger density varies throughout the day. This introduces uncertainty in crowdedness, making it difficult for individuals to anticipate conditions and plan their trips effectively. Crowdsourcing has been a valuable method for collecting localized user data. But the unpredictability of crowds and the uncertainty of crowdsourced information pose new challenges for decision-making. However, we know little about how to effectively visualize uncertainty in crowdedness to support informed commuting decisions, particularly when using crowdsourced ordinal data. Here, we investigated different uncertainty visualizations and their effectiveness in representing the variability and reliability of crowdsourced crowding data. They were evaluated through an online study, and we found that cluster visualization is best suited to reduce cognitive load while maximizing u

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.HC

VizPilot: Automated Onboarding for SVG-based Composite Visualizations using Multimodal LLMs

arXiv:2607.27938v1 Announce Type: new Abstract: Composite visualizations integrate multiple visualizations to represent complex datasets effectively, but their intrinsic composite designs often impose a high initial cognitive load on novice users. Existing visualization onboarding approaches are typically platform-dependent, require substantial manual authoring effort, and struggle with the structural complexity of composite visualizations, limiting their general applicability. We present VizPilot, an automated visualization onboarding approach that reverse-engineers composite visualization structure to generate interactive onboarding experiences directly from raw visualization artifacts. VizPilot consists of two modules: a Composite Visualization Analyzer and an Onboarding Interface. Leveraging Multimodal Large Language Models (MLLMs), the Analyzer employs a two-stage pipeline that decomposes a visualization into visual components, extracts structured explanations, and maps them to pr

Source ↗
Showing 1301–1350 of 1631 signals
← Prev Page 27 of 33 Next →