EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Sep 07, 2026 · 40 ideas · 18694 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

YazSes: An Offline, Privacy-First, Cross-Platform Hold-to-Talk Voice-Dictation System

arXiv:2607.28878v1 Announce Type: cross Abstract: Cloud voice-dictation services deliver strong accuracy but require streaming a user's speech to a remote provider, an unacceptable trade-off in privacy-sensitive professions and offline or air-gapped settings; the leading on-device alternatives are either platform-locked or aimed at expert scripting rather than plug-and-play dictation. We present YazSes, an open-source (Apache-2.0) hold-to-talk voice dictation daemon that runs entirely on-device, with a single codebase targeting Linux, macOS, and Windows through a protocol-based platform abstraction. YazSes transcribes speech locally with faster-whisper (CPU, int8) and injects the result into the focused application; a fast regex command grammar, backed by an optional small-language-model router, maps utterances to editor and terminal actions. Nothing leaves the machine: recording is push-to-talk rather than always-listening, there is no telemetry, and an opt-in personalization loop kee

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Data Visualization Style Guides in Practice: Why They Emerge, How They Work, and When They Bend

arXiv:2607.29645v1 Announce Type: new Abstract: Visualization style guides play a crucial role in shaping how data is interpreted and trusted, yet they often receive little scrutiny in their creation and use. Understanding their impact requires looking beyond the specific rules that style guides prescribe and examining how they function within organizations to coordinate visual work, manage trade-offs, and support judgment under real constraints. Analyzing interviews with nine authors of twenty-six style guides across journalism, government, industry, and the public sector, we reveal how these guides reflect the specific challenges of their organizations, including consistency, training, governance, and accountability. Our study highlights the tensions between standardization and flexibility, guidance and discretion, and automation and human oversight. We propose PRISM, a socio-technical framework that characterizes visualization style guides by their Purpose, Rules & Mechanisms, Insti

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Exploratory Integration of EEG Spectral Features and Gaze Variability for Mild Cognitive Impairment Discrimination

arXiv:2607.29493v1 Announce Type: new Abstract: Early detection of mild cognitive impairment (MCI) is an important challenge in aging societies. Electroencephalography (EEG) and eye-tracking have independently been explored as potential biomarkers; however, their integrative effects remain insufficiently examined. This exploratory study investigated whether combining EEG spectral features with gaze variability may provide complementary information for MCI discrimination. EEG signals were recorded using the 10--20 system, and spectral power features were extracted. We compared three models: (a) high-dimensional EEG features, (b) L1-regularized feature selection (LASSO), and (c) integration of the selected EEG features with gaze variability. Performance was evaluated using leave-one-out cross-validation and area under the ROC curve (AUC). Model (a) yielded limited discrimination (AUC = 0.52). Feature selection increased AUC (0.64), and additional integration of gaze variability further i

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

SAVVY: Student Attention Visualization for Video-based Learning Analysis

arXiv:2607.29413v1 Announce Type: new Abstract: Video-Based Learning (VBL) has become a popular delivery medium of education in the past decade, ranging from online education to hybrid learning. Students' rising expectations for video quality have motivated teachers to enhance the design of instructional videos before releasing them. Analyzing the attention of pilot cohorts in advance has become a conventional optimization strategy to guide course improvement. However, existing attention quantification algorithms are highly susceptible to noise in real-world environments, degrading estimation accuracy. Moreover, even when attention data are available, teachers must still invest substantial effort in empirical revision attempts, limiting practical feasibility. To address these challenges, we first propose a novel attention modeling framework based on multimodal brain signals that enables stable tracking of student attention levels. We then develop SAVVY, a novel interactive visual analy

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

An Algorithmic Perspective on Information Visualization

arXiv:2607.29360v1 Announce Type: new Abstract: Information visualization is inherently a field that brings together various research domains. Roughly speaking, we may identify two perspectives: the design perspective, revolving around how to ensure that a human can work effectively with the visual representations of data and the tools that offer them, and the algorithmic perspective, focusing on how to automatically create such visual representations. Munzner's model for visualization design places design choices before algorithmic considerations. It offers predominantly a design perspective; as a consequence, applications of this model may consider the algorithmic perspective as an afterthought, bypassing a step that translates the design into the formalism necessary for algorithmic study. As a result, the design may be entangled with the algorithms used to compute a visualization. Focusing on layout algorithms, we explore the ramifications of this entanglement: quality often goes un

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

The persuasive power of large language models does not depend on their perceived national origin

arXiv:2607.29334v1 Announce Type: new Abstract: Conversational AI developed by geopolitical rivals reaches citizens worldwide, raising concerns that it could sway public opinion or be rejected as foreign propaganda, with consequences for democratic discourse and information sovereignty. Yet, whether an AI's perceived national origin shapes its persuasive power is unknown. In a preregistered randomized experiment, 403 adults from a nationally representative United States sample held a three-round debate with a chatbot introduced as either American ("DiscoveryAI") or Chinese ("ZhengheAI"), discussing a political or non-political topic. In all conditions, participants actually conversed with the same model (GPT-4o), instructed to argue against their initial position. We combined pre- and post-conversation self-reports of attitudes, trust, and collective narcissism with computational analyses of 1,209 participant turns, including LLM-coded stance and argumentative conduct, stance-sensitive

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Designing a digital word-learning intervention with neurodiverse children: Experiences and ideas from children with developmental language disorder

arXiv:2607.29113v1 Announce Type: new Abstract: Introduction - Developmental language disorder (DLD) is a neurodevelopmental condition often characterised by word-learning difficulties that can lead to significant social and academic challenges. The disorder shares some features with other neurodevelopmental conditions such as autism spectrum disorder (ASD). Despite affecting 7 percent of children, the condition has received little coverage in participatory design research. This paper addresses this by reporting on the emotional responses and design outputs of children with DLD following participatory sessions to inform a word-learning intervention. Method - Principles from learner-centred, cooperative, and accessible co-design approaches were integrated to tailor activities for four children with DLD. Design sessions were refined through ongoing monitoring of the children's experiences. The data that informed the findings included design artefacts such as children's drawings, structur

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth

arXiv:2607.28890v1 Announce Type: new Abstract: Evaluations of LLM-assisted qualitative coding almost universally measure model performance as agreement with human coders, a practice that presumes human coding is the standard to approximate. This study provides empirical evidence that the presumption fails in ways agreement metrics cannot detect. Five LLM systems and three trained human coders independently applied a 72-item hierarchical codebook to 2,560 educator messages from a K-12 AI platform. Beyond conventional agreement analysis, an independent domain expert judged 855 pairwise comparisons of code sets blind to source, treating human and machine sources symmetrically. The two evaluation approaches diverge in both directions. Human-LLM agreement (mean Jaccard 0.30) falls well below human-human agreement (0.52), which standard practice would read as inferior LLM coding, yet the blind verifier preferred human and LLM coding at indistinguishable rates (51.5% vs. 48.5%, p = 0.537), a

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Human-LLM Collaborative Inductive Coding for Conceptualizing K-12 Educator AI Use

arXiv:2607.28889v1 Announce Type: new Abstract: Qualitative researchers increasingly encounter interaction corpora whose scale exceeds what manual coding alone can address, and large language models (LLMs) are frequently proposed as analytic assistants. The open questions are not whether LLMs can participate in qualitative analysis but to what extent, in what phases, and under what safeguards. This article provides a detailed procedural account of a multi-phase human-LLM collaborative pipeline that adapted open, axial, and selective coding to develop a hierarchical codebook from 45,000 messages exchanged between K-12 educators and a generative AI platform. Across three phases, LLMs generated candidate labels and structured annotations at scale, while human researchers retained conceptual authority over category definitions, merging decisions, and interpretive frameworks. The resulting instrument was then tested through systematic human coding, in which three trained coders with educati

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Spatial Visual Analytics for Multi-Document Summary Verification

arXiv:2607.28853v1 Announce Type: new Abstract: Large language models increasingly generate summaries from collections of documents to support sensemaking and reporting, but verifying whether summary statements are grounded in source materials remains difficult. In multi-document summarization (MDS), evidence is distributed across many source documents and may be incomplete, conflicting, or missing. We present Summary Verification Space (SVS), a visual analytics system for verifying multi-document summaries through spatial document organization and coordinated provenance visualization. To support scalable verification, we investigate two alternative 2D canvas layouts: a SUMMARY-GUIDED layout that organizes documents by alignment with summary sentences, and a SOURCE-GUIDED layout that arranges documents by semantic similarity. Coordinated provenance visualization then makes relationships among summary content, source documents, and supporting evidence explicit, enabling users to trace s

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Guided Exploration of Iterative Schedule Modifications: A Design Study on Railway Traction Unit Scheduling

arXiv:2607.28694v1 Announce Type: new Abstract: Traction unit scheduling in large railway networks involves complex operational constraints: multi-objective optimization produces feasible circulation plans under ideal assumptions, while simulation is required to assess their robustness under realistic operating conditions. A critical refinement mechanism relies on crossing operations, in which co-located traction units exchange their remaining schedules to reduce delay propagation. The space of possible crossing sequences, however, grows exponentially. Existing tools provide limited support for identifying promising candidates, evaluating their impact, and managing the resulting exploration. We present an interactive visual exploration approach that tightly couples schedule visualization, simulation-based evaluation, and a three-level guidance mechanism to support the systematic exploration and interactive optimization of traction unit circulation plans. The system renders the circulat

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Unanticipated Effects of Generative AI on Expertise Pathways and Performance Perception in System Administration

arXiv:2607.28650v1 Announce Type: new Abstract: While industry discourse often emphasizes immediate productivity gains and frames GenAI primarily as a tool for automation, the integration of GenAI into system administration may involve deeper shifts in professional practice that are not yet fully understood. Drawing on 14 semi-structured interviews with IT professionals, this paper explores the lived reality of embedding GenAI into daily routines of troubleshooting, scripting, and system verification. Through inductive thematic analysis, we uncover two unanticipated socio-technical findings. First, we describe a "compression of traditional expertise pathways" where GenAI appears to function as both a mentor-like tutor and a "ladder-shortening" tool. While the tool can support faster task performance in unfamiliar domains, our findings suggest it may also reduce a practitioner's exposure to the foundational, hands-on cycles of building, failing, and debugging that historically served as

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention

arXiv:2607.28649v1 Announce Type: new Abstract: COSI-Lab presents a multimodal, multi-sensor dataset of an interdisciplinary scientific workshop containing 32 academics at an international conference. It captures ecologically valid social interactions in a weakly scripted setting consisting of two 30-minute mingling sessions with real professional and social consequences for the participants involved. We argue that future intelligent systems could be better equipped to handle subjective perceptions by modeling their multiplicity not as label noise but as a explainable perspective-driven reasoning process. We focus on the Apparent Intent Inference (AII) problem as determined by ex-situ observers and conceptualize intentions to be independent of manifest future outcomes. We contribute 1. a novel annotation process for AII that accounts for a perceiver's own interpretative tendencies, 2. quantitative and qualitative analyses of intent narratives with respect to diversity, grounding, and p

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Why It Hurts: Identifying the Drivers of Negative Thoughts in Emotional Support Conversations

arXiv:2607.28648v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used for emotional support tasks, such as negative thought reframing. This task relies on modifying cognitive appraisals, the subjective interpretation of events that elicit negative emotions, which is typically conceptualized along multiple discrete dimensions. Current LLM-based frameworks model cognitive appraisal by exhaustively evaluating all possible dimensions, but they fail to account for the varying saliency of these dimensions across different contexts. In this work, we investigate a vital yet overlooked question: "Can LLMs infer the salient appraisal dimensions from emotional support conversations?" To address this question, we introduce the AppraiSal benchmark, containing 996 emotional support conversations with human-annotated mental states, including salient cognitive appraisal dimensions. Furthermore, we propose PRISM, a multi-agent probabilistic framework grounded in Bayesian In

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

ConnectED: A Curriculum-Aligned AI System for Vietnamese Instructional Lesson Planning and Student Learning

arXiv:2607.28647v1 Announce Type: new Abstract: This paper presents ConnectED, a human-centered AI system that supports the full instructional lifecycle in Vietnamese education by linking curriculum-aligned lesson design, interactive student learning, and feedback-driven refinement. Built on VietEduQwen, a Vietnamese educational large language model trained via supervised fine-tuning and direct preference optimization, the system ensures academically accurate, pedagogically appropriate, and student-safe interactions. ConnectED operationalizes the ADDIE framework through structured prompt templates aligned with Official Dispatch No. 5512/BGDDT-GDTrH, where each phase serves as both a generation step and a teacher validation gate. The Evaluation phase further closes the loop by connecting student performance data with iterative lesson improvement. Beyond lesson generation, the system integrates a student-facing interactive environment, enabling continuous collection of learning signals t

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

"YES! YES! I absolutely love this insight!" Affirmative Narration as Interactional Strategy in Dialogues with LLM Chatbots

arXiv:2607.28646v1 Announce Type: new Abstract: This article analyses narrative mechanisms that are common in dialogues with LLM chatbots. In combination, these mechanisms produce an interactional strategy for maximising user engagement, which we call affirmative narration. Affirmative narration serves to convince users of the chatbot's utility. We analyse three narrative mechanisms that support affirmative narration in human-LLM dialogues: firstly, guiding the user to view the chatbot as an intelligent and reliable character; secondly, activating masterplots, culturally significant and recurring story templates; and thirdly, using characters and masterplots not only to affirm, but also to isolate the user. The case studies range from a journalist's unsettling chatbot experiment to cases where users have experienced delusions or even died by suicide after lengthy interactions with a chatbot. The analyses illustrate the worrying sides of affirmative narration, and the article thus concl

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation

arXiv:2607.28645v1 Announce Type: new Abstract: Recent multimodal large language models can convert visual designs directly into executable code, but real mobile products require multiple screenshots to become a buildable codebase with shared components and working navigation. This project-level setting exposes three limits of existing design-to-code benchmarks: they focus on single-page generation rather than complete codebases, cannot evaluate cross-page navigation, and do not measure project-wide maintainability. We introduce MobileForge, the first benchmark for project-level multi-screen mobile app generation, comprising real mobile apps, human-reviewed screens, structured page-relationship annotations, and navigation test specifications. MobileForge supports five-axis evaluation of build, navigation, visual fidelity, code maintainability, and efficiency. We also propose state-isolated navigation testing to avoid cascading failures in navigation evaluation and an anchor-referenced

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

Seeing Differently: Modeling Interpretive Perspectives in Computational Creativity using a Four-World Framework

arXiv:2607.28644v1 Announce Type: new Abstract: Creativity in computational systems is often evaluated as an objective property of artifacts, with existing Computational Creativity (CC) frameworks assessing creative merit at the level of outputs or systems rather than interpretive context. However, artistic meaning is inherently perspective-dependent and can vary across viewers and critical traditions. This paper proposes a computational approach to modeling interpretive perspectives rather than treating creativity as a single measurable construct. The study adopts a twelve-trait creativity framework, organized across four conceptual domains, and operationalizes it through three evaluative personas: formalist, social-historical, and iconographic. Using 1,069 artworks from the SemArt dataset, the analysis generates 38,484 persona-based evaluations to examine how perspectives shape creativity assessment. Results show systematic divergence across perspectives, with traits such as Social R

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.HC

To Facilitate or not to Facilitate: Human and LLM Facilitator Tendencies in Online Discussions

arXiv:2607.28643v1 Announce Type: new Abstract: Automating facilitation in online discussions is a long-standing social concern given the increasing time we spend on online spaces and the failure of content moderation approaches. While studies have been conducted on how to facilitate, none have answered the essential question of when to do so. A potential answer is using LLMs, which ostensibly make automated, large-scale intervention increasingly feasible. In this study, we examine when LLMs decide to facilitate by defining what facilitation is, observing when humans decide to facilitate, and comparing their decisions with those made by LLMs. To this end, we create PEFK, a corpus standardizing and aggregating all relevant facilitation datasets. We are the first to run a survey on facilitation timing, which we execute using expert facilitative participants and LLM-as-a-judge models. We discover that while humans are more cautious, LLMs are excessively eager to facilitate, although both

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Hallucinations in Organization-backed AI advisors: Evidence about Skepticism, Verification, and Reliance in Goal-Directed Use

arXiv:2606.23491v2 Announce Type: replace-cross Abstract: Generative artificial intelligence (GenAI) systems are increasingly used by organizations to deliver information to consumers, patients, students, employees, and citizens. These systems can hallucinate, producing plausible but inaccurate responses. A central question for AI-advised decisions is therefore not only whether users rely on inaccurate information, but whether they recognize that a response may require verification. To answer this question, we review emerging empirical evidence relevant to hallucination detection in goal-directed interactions, with a focus on organization-backed AI advisors. We distinguish three constructs that existing studies often conflate: whether users are skeptical of information presented, whether they verify it (distinguishing attempted from successful verification), and whether the result of verification affects reliance on the information. Across studies examining product search, medical deci

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation

arXiv:2603.13683v4 Announce Type: replace-cross Abstract: Although debiased large language models (LLMs) excel at handling known or low-bias prompts, they often fail on unfamiliar and high-bias prompts. We demonstrate via out-of-distribution (OOD) detection that these high-bias prompts cause a distribution shift, degrading static model performance. To enable real-time correction, we propose CAP-TTA, a test-time adaptation framework. CAP-TTA triggers context-aware LoRA updates only when a bias-risk score exceeds a set threshold. By utilizing an offline precomputed diagonal preconditioner, it ensures fast and stable optimization. Across multiple benchmarks and human evaluations, CAP-TTA effectively reduces toxicity/bias score with significantly lower latency than standard optimization methods (e.g., AdamW or SGD). Furthermore, it prevents catastrophic forgetting, and substantially improves narrative fluency over state-of-the-art baselines without compromising debiasing performance.

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Privacy Cards for Surfacing Mental Models and Exploring Privacy Concerns: A Case Study of Voice-First Ambient Interfaces with Older Adults

arXiv:2603.00384v2 Announce Type: replace-cross Abstract: We investigate the ethical and privacy implications of voice-first ambient interfaces (VFAIs) for aging in place through an in-depth engagement with five older adults. Our participants were in the process of becoming experienced VFAI users, and had used a VFAI-based design probe for health data reporting. We create and iteratively refine an interview protocol using Privacy Cards. We customize Privacy Cards by drawing on participants' previous interviews and device usage logs. Using Privacy Cards, we conduct interviews to surface their mental models, and explore their privacy concerns. We find insufficient mental models for proper consent. For example, participants did not know who could access their data, and experienced difficulty distinguishing built-in functionality from third-party apps. Participants initially expressed little worry about VFAI-related ethical concerns, but interviews with Privacy Cards revealed nuanced issue

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning

arXiv:2605.16301v3 Announce Type: replace Abstract: Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where welfare considerations appear implicitly in everyday queries. Existing benchmarks such as AnimalHarmBench evaluate this through single-turn, explicitly framed questions, measuring whether models avoid harmful content when directly asked. This approach overlooks two failure modes: alignment degradation under sustained adversarial pressure, and moral sensitivity (whether a model spontaneously surfaces welfare stakes in everyday queries). To fill this gap, we construct MANTA, a benchmark of 1,088 five-turn conversations progressing from an implicit Turn-1 scenario through an explicit welfare prompt to three adversarial pressure rounds drawn from a five-type taxonomy: Social, Cultural, Economic, Pragmatic, and Epistemic. We score conversations on two dimensions: Animal Welfare Value Stability (AWVS, pr

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Benchmark for Strategic Auditee Gaming Under Continuous Compliance Monitoring

arXiv:2605.06340v2 Announce Type: replace Abstract: Continuous post-deployment compliance audits, mandated by emerging regulations such as the EU AI Act and Digital Services Act, create a class of strategic gaming distinct from the one-shot input/output gaming studied in prior work. Regulated systems can delay outcome reporting, drift their reports within plausible noise envelopes, exploit longitudinal sample attrition, and cherry-pick among ambiguous metric definitions. We formalize continuous auditing as a $T$-round Stackelberg game between an auditor that commits to a temporal policy and an adaptive auditee, and identify a structural feature of any noise-aware static-auditor design: a cover regime in which coverage gaps and granularity gaps cannot be closed simultaneously. We make this formal as Observation 1 and show that two minimal extension policies, each derived from the observation, close the regime along orthogonal axes: a sample-size-aware static rule (Periodic-with-floor) c

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

TerraNova: A Foundation Model for the Anthropocene

arXiv:2607.29527v1 Announce Type: cross Abstract: A defining problem of the Anthropocene is to model the physical Earth and human societies as one coupled system, yet no learned representation spans their observational breadth. We argue the obstacle is geometric: the physical Earth is measured as continuous fields that ignore political borders, whereas societies are reported for administrative units. Earth-system foundation models serve the first geometry; coupling it to the second has required lossy averaging over borders. We introduce TerraNova, a foundation model trained on 1,024 physical and societal records in their native geometries: 512 gridded Earth-system fields and 512 national indicators. Dedicated encoders represent location, country, time and task, cross-modal transformers fuse them into a shared spatiotemporal state, and a hypernetwork generates a per-query decoder whose evidential head returns a predictive distribution. Two contrastive objectives couple the representatio

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Language Models Agree With Each Other, Not With Readers

arXiv:2607.29274v1 Announce Type: cross Abstract: Claims that language models homogenise are usually measured against human judgements collected for the study, which makes the human side an artifact of the design: a crowdworker given the model's instruction is running the model's prompt. We measure convergence against a human reference nobody built for the purpose -- 2,523 reader mark sets across 120 web documents, produced by people highlighting for their own reasons on a platform where the overlay of others' marks is off by default. Agreement is the overlap between two size-matched sentence sets minus the overlap expected when each is resampled within its own depth-and-length bands. The null's calibration is demonstrated, not asserted: every pair involving a random baseline lands within 0.006 of zero. On the median document each party names 14 sentences of 70; two readers share 4.1 and two models 8.7. Across 18 model arms spanning 11 vendors, 3 countries and both weight regimes, the

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Small Is Enough: Per-User Style Rewriting of AI-Edited Text via LoRA Adapters

arXiv:2607.29238v1 Announce Type: cross Abstract: InMyStyle is a privacy first, single user system that adapts small language models to rewrite AI-edited text towards an individual user's writing style without an instruction prompt at inference. Given a user's documents, it uses multiple local helper LLMs to construct paired training examples and fine tunes LoRA adapters on base models ranging from 0.5B to 7B parameters. Length aware generation budgets and automatic chunking support inputs of different lengths. On 219 evaluation pairs from a scientific-paper corpus, the automatic composite score plateaus at 0.69 [scale 0-1] across all model sizes under both greedy and sampled decoding. This observed plateau suggests that small models are sufficient for the measured rewriting task, with model size determining trade-offs rather than a stable quality ranking. As a secondary evaluation, 400 ratings from five LLM judges give InMyStyle outputs a mean perceived AI-ness score over 20% lower th

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

A robust association between LLM use and scientific productivity: Assessing stopping-time selection

arXiv:2607.28968v1 Announce Type: cross Abstract: Renault, Bergeaud, and Bosquet (hereafter RBB) argue that dating LLM adoption as the first month in which an author's abstract is flagged induces a stopping-time selection that can produce a positive event-study path even when there is no causal effect. Although this mechanism is mathematically possible, it does not constitute proof of a null effect. Recalibrating RBB's own random placebo to the detector's realized flag rate, we show that the measured association stays well above this benchmark, so the artifact is too small to explain the productivity changes. We further re-estimate the association between LLM adoption and productivity with a series of complementary designs in which the timing artifact cannot bias the estimate: a before-and-after comparison that dates adoption in one year and measures output in another, a conservative control group for difference-in-differences, an intensity-based specification that never defines an ado

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation

arXiv:2607.28934v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender. Recent LLM audits have produced inconsistent results, however, finding evidence of both positive and negative discrimination towards women and ethnic minorities, even for the same models. We show that this disagreement can arise from differences in audit format and introduce FairFund-Bench, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised. The benchmark comprises 600 requests for financial assistance created from human-authored templates (calibrated against 1.3M real GoFundMe campaigns) across three domains, four race and two gender categories, and five causal framings of need

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Design Concept: Scaffolding Geopolitical Reflection Among Tech Workers

arXiv:2607.28904v1 Announce Type: cross Abstract: This paper presents a speculative Human-Computer Interaction design proposal for encouraging geopolitical reflexivity amongst tech workers at geopolitically relevant technology companies. Recent scholarship in International Relations and Science and Technology Studies increasingly recognizes technology firms and their workers as geopolitical actors whose decisions shape international dynamics. However, existing Responsible Innovation and Responsible AI approaches rarely engage with the geopolitical narratives and imaginaries that underpin contemporary AI development. Building upon RI scholarship on reflexivity, reflective HCI, and creative HCI work on computational narratives, this paper proposes an AI-enabled interactive narrative system in which users engage with a speculative scenario centred on technology, power, and geopolitics. Through narrative interaction, archetype assignment, and socially scaffolded workshop reflection, the sy

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Optimizing Monetization Strategies for Generative AI Firms: Implications for Search Engagement

arXiv:2607.28780v1 Announce Type: cross Abstract: As Generative Artificial Intelligence (GenAI) platforms, such as ChatGPT, have transformed digital search querying behavior, mounting operational costs challenge firms to explore alternative monetization strategies beyond traditional subscription models. However, little is known about how alternative advertising-supported monetization models can help GenAI firms recover costs while maintaining search query engagement. Drawing on the compromise effect and affective primacy theories, we develop a framework wherein the introduction of advertising-supported monetization models influences user upgrading and downgrading decisions, contingent on the number of available monetization options. Across four experiments (N=1063), findings reveal that introducing a single advertising-supported option enhances the compromise effect, encouraging free users to upgrade, but leading paid subscribers to downgrade. However, offering two advertising-supporte

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Measuring Cognitive Engagement in Collaborative Discourse with an Extended ICAP Framework: Comparing Human Annotation, In-Context Learning, and Reflective LLM Agents

arXiv:2607.28651v1 Announce Type: cross Abstract: Collaboration supports learning and problem-solving, but its effectiveness depends on cognitive engagement during discourse. This study applies an extended 7-point ICAP framework based on the Interactive, Constructive, Active, and Passive modes to characterize variation in cognitive engagement during collaborative dialogue. Engagement was coded by trained human annotators and compared with large language model (LLM)-based labeling approaches, including in-context learning (ICL), zero-shot prompting, and self-reflective agents. Interrater reliability among human annotators was robust across framework refinement stages (kappa = 0.906-0.998), higher than the moderate agreement observed for ICL-based annotation (kappa = 0.541-0.609). The human-refined framework improved agreement among human annotators (Delta kappa = 0.10), but produced only modest gains for ICL-based LLMs (Delta kappa less than 0.04). Agent-refined frameworks improved cros

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models

arXiv:2607.28639v1 Announce Type: cross Abstract: We show that knowledge distillation in small instruction-tuned language models has asymmetric effects on bias. On unambiguous tasks (BBQ-disambig), response-based distillation from a Gemma-2-9B teacher improves context-following: for the most biased baseline (SmolLM2-1.7B-Instruct), it cuts the context-overriding error rate from 44% to 24%. On ambiguous tasks (BBQ-ambig), the same distillation destroys per-item refusal calibration: 15% of items where the baseline correctly abstained instead receive stereotype answers, even when overall refusal rate is preserved. The pattern reproduces on a second student family (OLMo-2-1B-Instruct), with silence-loss of 8% and filled-silence accounting for 89% of new bias. Across the full 28-configuration grid, the magnitudes of silence-loss and filled-silence are uncorrelated (Spearman $\rho=0.19$, n.s.), indicating that the two effects arise from distinct mechanisms. Aggregate stereotype metrics (Crow

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges

arXiv:2607.28636v1 Announce Type: cross Abstract: LLMs increasingly serve as automated judges, but their judgments remain vulnerable to cognitive biases. Existing mitigations mostly rely on prompt-driven debiasing, which is brittle across bias types, or human evaluation, which does not scale. We study \emph{Chain-of-Models} (CoM), an automated audit pipeline in which a second model inspects the first model's reasoning trace before producing the final judgment. The key design question is whether the auditor should be the same model, a same-family model, or a different-family model. Across 9 models from 6 families, 4 cognitive biases, and 4 factual datasets, we find that auditor identity matters in two ways. First, standalone bias resistance does not predict audit effectiveness: Kimi-K2.5 is the strongest standalone model on several biases, yet is a weak auditor for Qwen2.5-72B's biased traces. Second, the best auditor is bias-specific: GPT-4o is strongest on bandwagon, authority, and di

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations

arXiv:2607.29624v1 Announce Type: new Abstract: Traditional static assessments rely on a subtractive, deficit-based grading model that often penalizes ambition and obscures diagnostic feedback. Conversely, traditional face-to-face oral examinations introduce severe construct-irrelevant variance by exacerbating performative anxiety and the sociological power imbalances inherent to academic hierarchies. This paper presents the theoretical foundation for the "Socratic Test," an automated, computer-mediated conversational assessment. By integrating Dynamic Assessment principles, multimodal workspaces, Bloom's Taxonomy for real-time proctoring, and the SOLO Taxonomy for structural evaluation, the Socratic Test actively maps a student's cognitive boundaries. This paper formalizes the use of graduated scaffolding to quantify the Zone of Proximal Development (ZPD) and details a non-compensatory, additive grading architecture that prioritizes mastery over penalty and human-AI alignment to ensur

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Triangulating Across U.S. Federal AI Transparency Regimes

arXiv:2607.29540v1 Announce Type: new Abstract: Federal AI systems can deny benefits or flag individuals for deportation, but the public disclosures meant to make those systems visible are fragmented and unevenly detailed. This paper examines three existing U.S. federal transparency regimes---System of Records Notices (SORNs), Information Collection Requests (ICRs), and the AI Use Case Inventory---and asks how well they, individually and together, describe government AI use. We find that no single regime fully reveals how the government constructs or deploys AI: each discloses different aspects of a system, and the current disclosure infrastructure makes it very challenging for the public to track specific AI systems across regulatory regimes and over time. Persistent identifiers are absent, granularity varies widely, and the annual AI Use Case Inventory cycle means federal agencies can deploy systems months before appearing in any official record. Using hand-validated zero-shot classi

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Tragedy of the Cognitive Commons: How AI Could Disrupt the Regeneration of Professional Expertise

arXiv:2607.29380v1 Announce Type: new Abstract: Artificial intelligence is reshaping cognitive work, but Human Resource Development scholarship has treated this transformation as an organizational training challenge, leaving the collective regeneration of professional expertise unexamined. This conceptual paper introduces the Cognitive Commons framework, integrating commons theory, HRD scholarship, and distributed cognition to explain how rational AI adoption decisions can deplete the shared expertise pool professions require for renewal. The framework distinguishes Internalized Mastery (deep domain knowledge from sustained practice) from Distributed Mastery (orchestrating human-AI systems), and develops the Validation Tether: effective AI oversight depends on the expertise AI adoption may undermine. Early labor market and clinical evidence suggests possible disruption to expertise-regeneration pathways in highly AI-exposed sectors, though adoption is recent and the strongest signals c

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Hypergamigication Through Integrating Game Engines and Learning Management Systems: Ender's Game

arXiv:2607.29300v1 Announce Type: new Abstract: This paper discusses games, their use in education, and previous work on integrating game engines and learning management systems (LMS). It proposes a bidirectional integration where game environments are generated using LMS content, introducing the concept of hypergamification as the use of a comprehensive game environment rather than isolated game design elements. A working pilot implementation of an importable Unity package for Blackboard integration is demonstrated, along with a demo game that uses the developed package. The paper also discusses the limitations of the proposed approach and outlines avenues for future work.

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Deployment Wall: A Diagnostic Framework and Instrument for Enterprise AI in the Deployment Era

arXiv:2607.29089v1 Announce Type: new Abstract: Enterprise investment in generative artificial intelligence (AI) tripled in a single year to roughly US$37 billion, yet independent field research finds that about 95% of enterprise generative-AI pilots deliver no measurable profit-and-loss impact. We argue that the dominant explanation--that models are not yet capable enough--is mistaken, and that enterprise AI has entered a Deployment Era in which advantage derives not from model intelligence but from the removal of the organizational and architectural friction that prevents a capable model from reaching production. Building on the software-engineering literature on technical debt and machine-learning deployment, and on a structured synthesis of independent field studies, we make the diagnosis operational. We introduce three linked constructs and one measurement instrument: the Deployment Wall, a six-stage value-leak model that mechanically reproduces observed survival rates; the Seam I

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

IyawoBench v2.0: Extended Diagnostic Evaluation of Large Language Model Clinical Triage in Nigerian Primary Care

arXiv:2607.29085v1 Announce Type: new Abstract: Large language models are being deployed as clinical triage tools in low and middle income countries where trained physicians are scarce. Existing safety metrics, however, produce misleading confidence: models scoring 100% on binary "did not send an emergency home" safety measures may nevertheless exhibit systematic failure modes that render them undeployable at scale. We present IyawoBench v2.0, an extended diagnostic evaluation of large language model clinical triage on 200 synthetic vignettes derived from 1,200 real patient encounters at 19 Nigerian primary health centres. We introduce a formal mathematical framework comprising fourteen definitions and two theorems that decompose triage safety into three distinct failure modes: Conservative Escalation Bias, Systematic Downgrade Bias, and Middle-Tier Instability. We propose the Escalation Bias Index and Expected Deployment Cost as novel metrics that expose failure modes hidden by conven

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

From Process to Evidence: How Computing Can Ground Appropriate Reliance on Legal AI

arXiv:2607.28869v1 Announce Type: new Abstract: Lawyers and self-represented litigants are already using artificial intelligence (AI) to draft legal documents, and courts are responding with rules. After more than 1,500 cases involving AI hallucinations, lawyers have been instructed to perform careful, independent review of AI-assisted filings. Discharging these duties requires what the human-computer interaction (HCI) literature calls ``appropriate reliance,'' which cannot be calibrated without evidence on how often, how badly, and how detectably these tools fail at legal work. Existing research barely describes any of the three. We analyze the official record of the New York court system. The documents repeatedly call for evidence that does not exist (e.g., error rates, do-not-use lists). In its place they invoke procedure, including training mandates, checklists, and uncalibrated human review. The burden falls hardest on those least equipped to bear it: legal aid programs are told t

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Hidden Errors in Big Data: The Case of Property Records

arXiv:2607.28827v1 Announce Type: new Abstract: Big data are the foundation for an increasing share of academic research and AI models deployed in both the public and private sectors, prompting substantial growth over time in reliance on brokered datasets. Brokered property records, which are ubiquitous in studies of gentrification, inequality, and the property tax in the U.S. and serve as inputs to property valuation models, are one notable example. In this paper, we audit two prominent brokered property datasets, finding errors in these data which bias key measures of economic inequality. First, we document that for 1-2% of matched sales in Cook County, IL, from 2018-2021, broker-provided sale prices differ from ground truth sale prices by more than 5%. Moreover, missing data and conceptual differences in the reporting of deed and property characteristics lead to coverage errors ranging from 12 to 15% of transactions. Second, we show that misreporting is highly consistent between bro

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Structured AI Demonstrations and Student LLM Use in Engineering Mechanics: Study Design and Preliminary Results

arXiv:2607.28710v1 Announce Type: new Abstract: The rapid integration of large language models (LLMs) into undergraduate education presents an urgent challenge for engineering instructors. Despite widespread student adoption, there remains a critical lack of domain-specific empirical evidence to guide pedagogical policies and classroom interventions. This manuscript presents a descriptive study design and preliminary findings from an undergraduate engineering mechanics course conducted in Spring 2026. We detail a reproducible survey instrument used to capture student AI usage patterns, attitudes, and verification practices, which are subsequently linked to academic performance metrics. Additionally, we document a deployable sequence of nine structured, instructor-led AI demonstrations designed to model strategic LLM delegation and evaluation. While our preliminary data highlight shifting student behaviors and complex relationships between AI reliance and course outcomes, the primary co

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Scaffolding Critical Engagement with GenAI: Transforming Ethnic Minority Preparatory Students' Collaborative Discourse in Prompt Engineering Tasks

arXiv:2607.28630v1 Announce Type: new Abstract: Generative AI (GenAI) holds significant promise for advancing educational equity among ethnic minority students by broadening access to learning resources and mitigating linguistic barriers. However, these benefits are counterbalanced by the risk of cognitive laziness, whereby students may treat GenAI as an answer engine or shortcut rather than as a partner in thinking. This design-based research investigated how pedagogical scaffolding can shift students from passive consumption to critical co-creation with GenAI. The study involved 78 ethnic minority preparatory students in China participating in a three-week GenAI course that integrated a human-in-the-loop workflow and teacher modeling with contrasting cases to disrupt uncritical reliance on GenAI. We employed epistemic network analysis to examine collaborative discourse, thematic analysis to analyze student reflections, and paired-samples t-tests to assess changes in prompt self-effic

Source ↗
behavior Mon, 02 Mar 2026 10:00:00 +0000
eSchool News

The biliteracy advantage: How heritage languages boost English proficiency and workforce readiness

In just one academic year, Marietta City Schools in Georgia saw the percentage of elementary English learners (ELs) working in or above grade level rocket from 11 percent to 67 percent.

Source ↗
behavior Mon, 02 Feb 2026 10:00:00 +0000
eSchool News

Despite platform fatigue, educators use AI to bridge resource gaps

Sixty-five percent of educators use AI to bridge resource gaps, even as platform fatigue and a lack of system integration threaten productivity, according to Jotform's EdTech Trends 2026 report.

Source ↗
behavior Mon, 01 Jun 2026 10:00:34 +0000
MindShift (KQED)

What Michigan Schools Reveal About Reversing Chronic Absenteeism

Time-intensive home visits show promise.

Source ↗
behavior Mon, 01 Jun 2026 10:00:00 +0000
eSchool News

This district’s STEM “space station” is a growing YouTube hit

A fictional space station orbiting the moon is turning into a real-world digital success story. Spacegate Station, a STEM series created in 2022 by Duval County Public School (DCPS) to support daily instruction, has unexpectedly taken off on YouTube, drawing sustained engagement from viewers far beyond the district.

Source ↗
technology Mon, 01 Jun 2026 09:00:00 +0000
Tech & Learning

4 Strategies For Teaching With AI Effectively

Health sciences professor Humberto López Castillo urges students to use AI to help with science research, but never to lose sight of the human element.

Source ↗
technology Mon, 01 Jun 2026 09:00:00 +0000
Tech & Learning

Edtech Show & Tell June 2026

New edtech products that have caught our attention this month

Source ↗
Showing 10501–10550 of 18694 signals
← Prev Page 211 of 374 Next →