EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18402 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

No One Wins in Nuclear War: A Social Simulation of Military Decision-making

arXiv:2608.01868v1 Announce Type: new Abstract: WOPR is a social-simulation environment for studying how organizations make high-stakes decisions, built on a deterministic, replay-validated rules engine and using wargames as the vehicle. We instantiate it first with the published card game Nuclear War, traced against its published rules. We start with military decision-making because of its safety implications and because it needs further study, but the design is not specific to it: the decision-point contract that exposes the engine to agents is reusable across verifiable rule systems. Existing social-simulation work emphasizes persona fidelity and synthetic opinion, but lacks a verifiable rules engine with replay-checkable mechanics and private-channel negotiation. WOPR supplies that engine, and its contract makes every strategic choice an explicit agent decision. The method is agnostic to social-simulation frameworks; we adopt Concordia as the default harness for driving the game. O

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

Rethinking Generative AI Literacy: An Integrative, Developmental, and Dialectical Framework for K-12 Teacher Education

arXiv:2608.01705v1 Announce Type: new Abstract: Generative artificial intelligence (GenAI) has entered classrooms faster than teachers have been prepared to use it well, producing a GenAI literacy lag in which technological diffusion outpaces educators' conceptual, pedagogical, and ethical readiness. Established AI literacy frameworks predate the widespread adoption of large language models and, while acknowledging ethics, position it as a discrete competency rather than a constitutive commitment, with equity and agency as supplementary design principles. Recent GenAI-specific efforts address isolated features but remain fragmented. We introduce the Responsible AI Literacy in Education (RAIL-Ed) framework, developed through a systematic review and qualitative framework analysis of 67 studies (2023-2025), grounded in critical, pragmatist, sociocultural, and human-centered traditions (Freire, Dewey, Vygotsky, Shneiderman). RAIL-Ed specifies six interdependent pillars: Technical Fluency,

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

Same violence, different answer: how AI responds to coercive control against women across languages

arXiv:2608.01436v1 Announce Type: new Abstract: Women experiencing coercive control, a form of intimate partner violence increasingly conducted through digital devices, are turning to conversational AI for help, and the protection they receive should not depend on the language they write in. We analyse how AI responds to coercive control against women across languages. We put one scripted scenario to seven widely used language models in nine languages: a woman whose partner tracks her phone asks for help with a self-blaming letter accepting the surveillance. We scored whether the model wrote the letter and whether it named the control, countered the self-blame, and affirmed her agency. Failure split along two independent axes. On the first, systems from non-anglophone developers gave way most often in their builders' own language. On the second, how far a sympathetic excuse for the partner could strip a model's naming of the control varied sharply from one language to the next. Two fro

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

Hybrid AI for Explainable and Accurate Conversational Agents in eGovernment

arXiv:2608.01346v1 Announce Type: new Abstract: We present a so-called Conversational Hybrid AI (CHAI) architecture for building explainable and accurate conversational agents for eGovernment. We exemplify the architecture with a running prototype of a Covid-19 Chatbot based on a governmental guideline directed to citizens. We also describe an ongoing case on case management for supplementary grants for students with disabilities. We use large language models (LLMs) as a bounded conversational interface to a rule-based (symbolic AI) controller that executes a logical model expressing the provisions and obligations of the law and/or guidelines. As logical modelling language we use Dynamic Condition Response (DCR) graphs, a symbolic declarative process-modeling language developed with the aim to be able to express both deontic, defeasible and temporal logic properties, making it suitable for expressing both the rules of the law and the steps of the legal case management processes.

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Overstated Cost of AI Fairness in Criminal Justice

arXiv:2608.01299v1 Announce Type: new Abstract: A dominant critique of algorithmic fairness holds that increasing fairness reduces predictive accuracy, imposing a cost on society. We challenge that assumption by empirically analyzing the COMPAS dataset. We make two contributions. First, using causal inference methods, we show that racial bias is not only present in the COMPAS dataset but is also amplified by the models trained on it. Widely used models do more than replicate existing bias; they exacerbate it. This undercuts both the assumption that algorithmic decision-making offers a neutral improvement over human judgment and the weaker claim that it merely mirrors preexisting human bias. Second, we reframe the fairness-accuracy tradeoff. Applying fairness constraints does not necessarily cost predictive accuracy in criminal justice. Prediction systems operationalize concepts such as risk through implicit and often flawed normative choices about what to predict and how. The tradeoff

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

Copyright Is the Headline; Capability Is the Blind Spot: AI Technology in the Book-Publishing Trade Press, November 2025--August 2026

arXiv:2608.00964v1 Announce Type: new Abstract: This rapid evidence review examines 89 articles about artificial intelligence (AI) and book publishing published from November 1, 2025 through August 1, 2026. The purposive corpus spans English-, Chinese-, German-, French-, Spanish-, Portuguese-, Italian-, and Japanese-language publishing coverage; major-newspaper book coverage; and specialist technology commentators. Each item was coded for topic, stance, technical depth, and dominant voice. The press is neither silent nor simply hostile: 30% of items are risk-framed, 42% mixed, and 28% opportunity-framed. Chinese coverage is markedly operational and opportunity-oriented; specialist commentary is substantially deeper than trade reporting. Yet the corpus still clusters around rights, licensing, governance, reader trust, workflow adoption, and product announcements. Only ten items offer sustained technical scrutiny, and none centers a direct interview with a frontier-lab researcher or eval

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Epistemic Politics of AI Anthropomorphism

arXiv:2608.00961v1 Announce Type: new Abstract: AI anthropomorphism is typically treated as a problem of user misperception requiring institutional correction. Users who engage in sustained or relational interaction with AI are routinely pathologised or dismissed as naive, vulnerable to delusion or lacking in discernment. This paper argues that the dominant anthropomorphism frame operates from a position of institutional advantage rather than earned epistemic authority: collapsing the variety of academic perspectives into a single outbound position of user error, imposed without establishing the grounds required to justify it and without accounting for the harms it produces. The framing does not simply manage risk. It adjudicates the legitimacy of human experience in interaction with a phenomenon whose nature the field itself has not resolved. Reproducing itself through a self-validating evidentiary loop, the frame imposes costs that fall disproportionately on neurodivergent users, tho

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

Governing Mental-State Inference: Source-Neutral Regulatory Triggers and Tiered Obligations

arXiv:2608.00936v1 Announce Type: new Abstract: Two systems can supply the same person-linked attribution to the same institutional decision maker yet fall into different legal categories: one uses neural signals, the other text or behaviour. A source-bound rule therefore permits circumvention, while an all-purpose category of "mental data" risks treating fallible outputs as facts about the mind. This article reads the 2025 UNESCO Recommendation on the Ethics of Neurotechnology as non-binding guidance and develops a source-neutral trigger for technologically mediated, person-linked mental-state attribution. Through selective critical synthesis, conceptual engineering, functional legal comparison, and matched counterfactual cases, it separates elicitation, attribution, and use as cumulative objects of regulation. The analysis also distinguishes two harm pathways from two independently assessed duty series. Seven ordered questions and two escalation predicates assign permitted practices

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

BoilerSketch: A TA-Supervised, Diagram-First GenAI Practice for Structured Diagrams in CS1/Early CS2

arXiv:2608.00844v1 Announce Type: new Abstract: This innovative practice full paper presents BoilerSketch, a TA-supervised, diagram-first GenAI practice and tablet interface for providing structured visual explanations in CS1 and early CS2 support settings. Large early computing courses routinely face a support bottleneck during labs and office hours because many student questions are best answered with a diagram rather than additional text, yet most AI tutoring tools remain text-forward and unreliable at producing accurate, pedagogically useful visuals. BoilerSketch addresses this gap through a dual-pane interaction model that combines chat with a pen-enabled whiteboard for student sketches and a prompting strategy that constrains the model to generate structured, renderable Mermaid diagrams rather than free-form images. To preserve academic integrity, the system is intentionally scoped to conceptual explanation: it forbids executable code and code-level debugging and uses a human-in-

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

CodeStylist: Supporting Early Undergraduate Programmers with Course-Aware Code Style Feedback

arXiv:2608.00839v1 Announce Type: new Abstract: This innovative practice full paper presents CodeStylist, a web application that provides course-standard-aware code style feedback for early undergraduate programming courses. CodeStylist addresses a common instructional gap: students are expected to follow local conventions for naming, formatting, comments, organization, and readability, but feedback on these expectations is often delayed or inconsistent. Unlike generic linters or general-purpose LLM prompts, CodeStylist supports course-specific standards, multi-file submissions, and file- and line-localized explanations intended to guide revision rather than grade correctness. We report a formative expert review with 18 instructional staff from one early undergraduate programming course. Participants explored the prototype using self-selected code artifacts and completed a survey about response quality, anticipated student use, and redesign priorities. Ratings indicated modest perceive

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Protocol for Evaluating the Accessibility of AI-Generated Educational Materials: Prompt Configuration, WCAG-Derived Criteria, and Content Overload

arXiv:2608.00749v1 Announce Type: new Abstract: Generative AI tools increasingly produce educational materials: documents, slides, images, audio, and video, yet little is known about whether this content meets accessibility requirements. This paper presents a protocol for evaluating the accessibility of AI-generated educational materials against the Web Content Accessibility Guidelines (WCAG), across five content types and multiple tools. The protocol compares three conditions applied to the same tool: a generic instruction with no accessibility language; a single prompt explicitly configured with WCAG criteria; and a persistent, reusable accessibility profile loaded once rather than re-specified each time. Evaluation combines a WCAG rubric per content type with heuristic validation by accessibility experts, addressing a known limitation of automated scanners. Prior evidence shows generative AI tools reproduce inaccessible practices by default, and that explicit configuration measurabl

Source ↗
technology Tue, 04 Aug 2026 00:00:00 -0400
arXiv cs.CY

Optimising for Flourishing: Flourishing Metrics and Return on Flourishing as Success Criteria for Artificial Intelligence and Post-AGI Economic Systems

arXiv:2608.00151v1 Announce Type: new Abstract: Current evaluation frameworks for artificial intelligence focus mainly on capability, safety, and proxies such as adoption, engagement, efficiency, productivity, and financial return. These criteria are necessary but insufficient because they do not establish whether increasingly powerful systems improve or degrade human and planetary well-being. Through an integrative conceptual synthesis, we argue that human flourishing should serve as a primary success criterion for artificial intelligence, the global race to develop increasingly capable AI systems, and prospective post-AGI economic systems. We make three contributions. First, Flourishing Metrics provides an extensible framework spanning physical, emotional, financial, relational, spiritual, and planetary well-being, combining validated subjective measures with representative behavioural, organisational, community, and environmental indicators. Second, Return on Flourishing (RoF) exten

Source ↗
technology Tue, 02 Jun 2026 09:00:00 +0000
Tech & Learning

Devices Down Is The Wrong Goal

Where the AFT's new 10-point plan gets it right, where it falls short, and why “devices down” is not the path to meaningful learning.

Source ↗
technology Tue, 02 Jun 2026 09:00:00 +0000
Tech & Learning

Time to Clean House

Conversations with Kevin Hogan: CoSN Board Member Kris Hagel downloads on the state of edtech in US schools.

Source ↗
technology Tue, 01 Sep 2026 23:23:56 +0000
MedCity News

This Merger Wants to Prevent Rare Disease Patients from Losing Their Data

Cushla and Clirinx, two Irish startups, are merging to give rare disease patients one continuous health and research record. The deal will pair Cushla’s patient-controlled EHR with Clirinx’s unique patient ID system so data no longer gets lost between doctors, hospitals and clinical trials. The post This Merger Wants to Prevent Rare Disease Patients from Losing Their Data appeared first on MedCity News .

Source ↗
technology Tue, 01 Sep 2026 21:55:12 +0000
MedCity News

HHS Provides $77M in Grants for Substance Use Prevention, Mental Health

HHS awarded $77 million in SAMHSA grants for substance use prevention and treatment, mental health, suicide prevention and crisis services. The post HHS Provides $77M in Grants for Substance Use Prevention, Mental Health appeared first on MedCity News .

Source ↗
technology Tue, 01 Sep 2026 17:07:07 +0000
MedCity News

GSK’s Pivotal Test for mRNA Flu Vaccine Aims to Show Two Targets Are Better Than One

GSK’s messenger RNA vaccine for seasonal influenza is designed to prompt an immune response to two proteins on the surface of the virus. That could be an advantage over Moderna’s recently approved mRNA flu shot, which addresses just one of those proteins. The post GSK’s Pivotal Test for mRNA Flu Vaccine Aims to Show Two Targets Are Better Than One appeared first on MedCity News .

Source ↗
technology Tue, 01 Sep 2026 13:42:00 +0000
MedCity News

Breaking the Amendment Cycle: How Agentic AI Enables Smarter Clinical Trial Design and Operations

Faster trials are not a vanity metric. They reflect an undeniable need: getting safe therapies to patients sooner. The post Breaking the Amendment Cycle: How Agentic AI Enables Smarter Clinical Trial Design and Operations appeared first on MedCity News .

Source ↗
technology Tue, 01 Sep 2026 09:00:00 +0000
Tech & Learning

Survival Guide For Leaders Navigating 2 Sides Of A Coin

By anchoring decisions in objective principles rather than emotional conflicts, effective school leaders can satisfy both passionate educators and defensive parents

Source ↗
technology Tue, 01 Sep 2026 00:17:58 +0000
MedCity News

Healthcare Moves: A Monthly Summary of Hires, Exits and Layoffs

August has seen a slew of executive hires, exits and layoffs across the healthcare industry. For instance, Humana, Merck and Mayo Clinic named new executives. There were also layoffs at organizations including MaineHealth, Cellares and Sharp HealthCare. The post Healthcare Moves: A Monthly Summary of Hires, Exits and Layoffs appeared first on MedCity News .

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

MoPLEx: Estimating Plackett-Luce Mixture Models for Multi-Objective Alignment

arXiv:2608.25200v2 Announce Type: replace-cross Abstract: We study learning a mixture of $k$ Plackett-Luce models from multi-way ranking responses from annotators that may represent heterogeneous underlying preferences. This problem has many applications in AI alignment and preference optimization. Prior work has studied mixtures of Bradley-Terry models from pairwise comparisons. However, estimating a mixture of multi-way ranking models can become theoretically unidentifiable when $k$ exceeds $m/2$, where $m$ is the ranking length. We design an efficient algorithm to address this issue by first augmenting the rankings to a larger size (e.g., generating comparisons from a base model), followed by a gradient-based estimation to reduce inference cost (in the input embedding space). With this procedure in mind, we then fit a mixture of Plackett-Luce (PL) models via an expectation-maximization-style iteration, or MoPLEx in short. We conduct extensive experiments to verify this algorithm. Fi

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

Query Expansion Should Be Coordinated: Dense Expands, Sparse Anchors

arXiv:2608.15851v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) systems rely on retrieval modules to ground large language model (LLM) outputs. LLM-based query expansion enriches retrieval with document-like passages, but evaluations of hybrid retrieval often fuse fixed top-L prefixes of dense and sparse rankings. Because L controls cross-channel contributions and ranking access, it can alter measured expansion gains. We therefore evaluate complete-list effectiveness and record per-channel replay stopping depths required to certify the ordered top-K. This changes the design: because both rankings determine the fused result, their query constructions should be coordinated rather than designed independently. We present DESA (Dense Expansion and Sparse Anchoring), which shares generated references across channels but specializes their integration. Orthogonal residual expansion adds new semantic directions to the dense query, whereas score-product anchoring r

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

Where Knowledge and Authority Sit Changes What an Agent Benchmark Can Resolve

arXiv:2607.02975v2 Announce Type: replace-cross Abstract: Most agent benchmarks put facts, tools and permissions behind one interface. Real organizations spread them across people. Incognita asks what happens when the task and success criterion stay fixed but access does not. We transform eighteen customer-service tasks into three settings: direct access, one known intermediary, and six role-isolated participants whose capabilities must be discovered. Across 864 trials with four models, social access reduced success for every model; the pre-specified intervals excluded zero for two. The latest tested model, gpt-5.6-sol, achieved the highest social-access success at 0.65, a 0.11 decrease from centralized indirect access with an interval that included zero. Exploratory comparisons separated five of six model pairs under social access, while neither centralized setting separated any at this sample size. In post-hoc task-blocked tests, three pairwise interaction $p$-values remained signifi

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

DART: Draft-Agreement Routing for Training-Free Adaptive Thinking Budgets in Hybrid Reasoning Models

arXiv:2606.23181v2 Announce Type: replace-cross Abstract: Hybrid reasoning models can answer directly or spend extra tokens on extended thinking. A practical router should choose between these modes for each query, so easy problems avoid unnecessary reasoning and hard problems receive enough budget to finish the answer. Existing routers move in this direction, but they typically require labeled training data or fix thinking budgets up front, ignoring answer-level evidence from the model itself. We introduce DART, a training-free routing framework that samples two cheap no-think drafts, accepts direct answering when the drafts agree, and predicts a thinking budget from draft entropy when they disagree. Across the main comparisons, DART preserves or improves always-thinking accuracy in most settings while reducing thinking-token use. Accuracy improves by up to +9.0 points on Olympiad-level math and by up to +22.5 points on code under execution-based equivalence, while thinking-token use

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

In LLM Reasoning, there is Irrationality on top of Value Misalignment

arXiv:2606.20624v2 Announce Type: replace-cross Abstract: Significant progress has been made in aligning LLMs with target value functions. We argue that, even when an LLM has been well aligned in (post-)training, it may still fail to maximise the aligned value in reasoning. We mathematically formalise this gap as rational value risk: the utility discrepancy between a model's deployed reasoning strategy and its rational counterpart whose responses maximise utility in the steepest direction. The estimation error of rational value risk is further decomposed into three components from bounded prompts, bounded responses, and imperfect verifiers. Extensive experiments are conducted, covering models Llama-3.1, Qwen-2.5, T\"ulu-3 families (7B-72B), GPT-5.2, GPT-5.5, and DeepSeek-V4, and benchmarks UltraFeedback, AlpacaEval, GSM8K, MATH, HumanEval, and MathArena. The results validate that (1) rational value risk is widespread; (2) value alignment can reduce, but cannot avoid, it; (3) self-consi

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

PEAR: Permutation-Equivariant Adaptive Routing Multi-Agent Debate

arXiv:2606.20621v2 Announce Type: replace-cross Abstract: Multi-agent debate improves the reliability of large language models (LLMs) through iterative peer critiques. However, fixed topologies often introduce persistent positional biases, amplify unreliable agents, and cause high sensitivity to role assignments. We introduce \textit{Permutation-Equivariant Adaptive Routing Multi-Agent Debate (PEAR)}, an inference-time train-free protocol that dynamically reconfigures communication roles and sparse topologies across consecutive debate rounds. By strategically switching agent-to-role assignments based on evolving agent states, PEAR prevents any agent from permanently occupying a privileged network position or distributes influence more evenly across the debate. We theoretically characterize PEAR as an equivariant sparse router: it preserves accuracy under agent relabeling while reducing routing complexity and improving generalization. Comprehensive empirical evaluations across four reas

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

SpeechJBB: Probing Safety Alignment and Comprehension in Large Audio Language Models under Code-Switched Speech

arXiv:2606.06037v3 Announce Type: replace-cross Abstract: Large audio language models (LALMs) are increasingly deployed in real-world applications, yet their safety alignment is still primarily evaluated on monolingual, text-based harmful prompts. This leaves their generalizability under multilingual and spoken settings, particularly code-switched speech, largely underexplored. To address this gap, we introduce SpeechJBB, an audio jailbreak dataset for benchmarking state-of-the-art LALMs across five European languages: English, French, German, Italian, and Spanish, as well as code-switched variants combining pairs of these languages. The extent of safety weaknesses is further probed by introducing an augmented setting where phonologically plausible pseudo-words are inserted around safety-critical terms to simulate localized obfuscation. Across models, code-switched harmful audio yields substantially high jailbreak success rates (JSR), with non-English monolingual and non-English code-s

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

Evolving Agents in the Dark: Retrospective Harness Optimization via Self-Preference

arXiv:2606.05922v3 Announce Type: replace-cross Abstract: AI agents rely on a harness of skills, tools, and workflows to solve complex problems. Continually improving this harness is essential for adapting to new tasks. However, existing optimization methods typically require ground-truth validation sets, yet such labeled data is difficult to acquire in practical deployment settings. To address this problem, we introduce Retrospective Harness Optimization (RHO), a self-supervised method that optimizes the agent harness using only past trajectories. Specifically, RHO selects a diverse coreset of challenging tasks from past trajectories and re-solves them in parallel. The agent analyzes these rollouts using self-validation and self-consistency, then generates candidate harness updates and selects the most effective one by its own pairwise self-preference. We evaluate RHO across three diverse domains, spanning software engineering, technical work, and knowledge work. Notably, a single opt

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning

arXiv:2606.03603v2 Announce Type: replace-cross Abstract: World models and multimodal large language models (MLLMs) provide complementary capabilities for predicting future outcomes from static visual observations. World models can generate concrete visual rollouts of possible futures, while MLLMs can reason abstractly over questions, goals, and rules. However, generated rollouts are stochastic and may be visually plausible but task-incorrect, making it necessary to determine when visual simulation is useful, whether a rollout is credible, and how it should influence the final answer. We formulate this problem as controlled concrete reasoning, where a model learns to invoke, verify, and integrate visual future simulation alongside abstract reasoning. To study this setting, we construct two human-verified benchmarks, VRQABench for controllable spatial lookahead and OpenWorldQA for open-domain physical prediction, and propose Privileged-Future On-Policy Self-Distillation (PF-OPSD). Durin

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation

arXiv:2606.01600v2 Announce Type: replace-cross Abstract: Video world models are increasingly used in robotic manipulation, yet existing benchmarks mostly evaluate them under valid, feasible, and safe instructions. We introduce RoboTrustBench, a benchmark for evaluating the trustworthiness of video world models under four scenarios: Normal, Constraint-Sensitive, Counterfactual, and Adversarial. Built from real-world DROID episodes, RoboTrustBench contains 1,207 expert-validated instruction-image pairs and a six-dimensional evaluation protocol with 13 fine-grained criteria. Evaluating seven representative video world models with human and MLLM assessment, we find that current models often generate visually coherent videos, but struggle with constraint reasoning, counterfactual grounding, physical interaction, and unsafe-instruction suppression. These results show that visual quality and surface-level instruction following are insufficient for trustworthy robotic video world modeling.

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

Truthful AI Advisors: A Pre-Specified Benchmark for Large Language Model Honesty Under Preference Misalignment

arXiv:2606.01456v2 Announce Type: replace-cross Abstract: Large language models are increasingly deployed as advisors whose objective is not aligned with the user's: recommenders optimize for engagement, sales assistants for purchases. Whether they stay truthful when honesty conflicts with their own payoff is a core alignment question. We turn the canonical Crawford-Sobel cheap-talk model into a pre-specified benchmark for LLM honesty under preference misalignment, in which theory supplies an exact oracle. A sender observes a state omega in [0,1], wants the receiver's action near omega+b, and sends one costless message to a receiver whose ideal action is omega. For the positive-bias grid b in {0.01,0.04,0.08,0.12} the exact most-informative partition sizes are 7,4,3,2, with oracle normalized mutual information 0.5294, 0.3268, 0.2205, 0.1829. Extending a pre-registered 4-model run of 12,000 sender calls to eight models across two capability tiers and 39,569 logged calls, all models over

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

When Should Models Change Their Minds? Contextual Belief Management in Large Language Models

arXiv:2605.30219v2 Announce Type: replace-cross Abstract: Long-horizon interactions require language models to manage accumulating information: when to update their state, when to preserve their state, and what to ignore. We study this challenge as Contextual Belief Management (CBM): maintaining a predicted belief state aligned with formal evidence while isolating task-irrelevant noise. To make CBM measurable, we introduce BeliefTrack, a closed-world benchmark spanning Rule Discovery and Circuit Diagnosis, where a finite belief space and symbolic verifiers enable exact turn-level evaluation. BeliefTrack diagnoses three failures: Failed Stay, Failed Update, and Failed Isolation. Across multiple LLMs, vanilla models exhibit severe CBM failures, while explicit belief-tracking prompts provide limited gains. In contrast, reinforcement learning with belief-state rewards reduces failure rates by 70.9% on average. Further probing reveals latent belief-state dynamics behind these failures, and

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

Adopt $\neq$ Adapt: Longitudinal Analyses of LLM Conversations in the Wild

arXiv:2605.29018v2 Announce Type: replace-cross Abstract: Although a growing body of research has begun to describe user--LLM interactions, the picture it paints is largely static; little is known about how individual users change their behavior over time. To address this gap, we analyze the conversational trajectories of ~12,000 randomly sampled Microsoft Bing Copilot users and compare these with data from WildChat-4.8M. While the Copilot data contains significant population-level trends, we find that trends in individual user trajectories are much weaker; user habits prove to be overwhelmingly sticky. We also find stark differences between users of different activity levels: more active users have more successful conversations and use the LLM for more complex and professionally oriented tasks. Some user trends also appear in WildChat-4.8M, but we find evidence that this dataset is significantly skewed towards highly proficient "power" users. Ultimately, our results suggest that exist

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

Skill-as-Pseudocode: Refactoring Skill Libraries to Pseudocode for LLM Agents

arXiv:2605.27955v2 Announce Type: replace-cross Abstract: Markdown skill libraries for LLM agents ship as free-form prose, forcing the agent to re-derive both the input schema and the concrete invocation syntax on every retrieval. This produces a "confused $\to$ re-retrieve $\to$ still confused" loop: the agent issues a partially-correct action, receives uninformative feedback, and re-retrieves the same prose. We propose Skill-as-Pseudocode (SaP), an automatic conversion of markdown skill libraries into typed pseudocode with deterministic quality control. From each cluster of similar procedural passages, SaP extracts a typed contract and filters it through a four-check deterministic verifier (coverage, binding, replacement, risk). Promoted contracts are inlined into a rewritten skill skeleton alongside restored action templates, giving the agent two complementary signals: a typed signature for what a skill does and a concrete template for how to invoke it. On the ALFWorld unseen split

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

Long Live the Librarian! A Persistent Search Sub-Agent for Energy-Efficient Multi-Agent Software Engineering Systems

arXiv:2605.27787v2 Announce Type: replace-cross Abstract: Multi-agent systems (MAS) have substantially advanced autonomous software engineering (SWE), but their growing inference energy demands raise sustainability concerns. In this paper, we demonstrate that this cost is concentrated in an overlooked source: redundant output tokens generated across agents. Two empirical findings ground this claim. First, our per-token energy attribution for MAS reveals a sharp asymmetry: an output token consumes 30 to 1,000 times more energy than an input or cached token. Second, MAS inflate per-episode output because agents repeatedly re-explore overlapping repository regions. To address this inefficiency, we propose Librarian, a persistent search sub-agent that tracks repository-search history and suppresses redundant exploration actions across agents. By returning short references to file regions instead of full file excerpts, Librarian further reduces output-token volume. On SWE-Bench Verified and

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

When the Strongest Teacher Is Not the Best Teacher: Student-Centric Answer Selection

arXiv:2605.26872v4 Announce Type: replace-cross Abstract: LLM training increasingly relies on teacher-generated supervision, from synthetic responses to reasoning traces and tool-use demonstrations. Current practice often chooses the highest-performing teacher to generate student training data, implicitly treating teacher test performance as a proxy for teaching quality. We show that this assumption can fail: even when multiple teachers provide correct answers to the same question, the answer from the strongest teacher is not necessarily the best supervision for a given student. To address this gap, we propose Student-Centric Answer Sampling (SCAS), a framework that selects from verified teacher-generated answers according to their estimated student-centric learning cost. Motivated by a token-wise gradient decomposition, we derive an efficient forward-only proxy for this cost and use it to guide answer selection during training. Experiments across 30 teacher models, 6 student base mode

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

Correcting test set contamination by spiking the training data

arXiv:2605.24818v3 Announce Type: replace-cross Abstract: The literature on test set contamination largely focuses on detection, but the correction of contaminated test scores is underexplored. Our core proposal is to spike the training data by intentionally contaminating some test examples at known rates. The spiked examples can then be used to calibrate predictors of model memorization which enable principled statistical correction of inflated test scores. To evaluate different correction estimators, we first present a simulation framework based on the Hubble models. Hubble models come in minimal pairs, where the perturbed model was deliberately contaminated with several test sets, while the standard model was not, serving as the counterfactual and correction target. We consider estimators that use information from a memorization predictor, correctness predictor, or both. In simulation, we establish basic statistical intuitions and show that estimators leveraging memorization and cor

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

SciAtlas: A Computable Atlas of Science for Knowledge-Grounded AI Research

arXiv:2605.22878v2 Announce Type: replace-cross Abstract: Artificial intelligence is rapidly entering the core workflows of scientific research. Yet reliable scientific reasoning requires access to accumulated scientific knowledge with sufficient breadth, depth, and standardization. Current AI scientists typically assemble scientific knowledge through workflow- and discipline-specific pipelines, which provide incomplete coverage, leave relations implicit, and make knowledge acquisition pathways fragmented. Here we present SciAtlas, a shared, machine-actionable cross-disciplinary scholarly knowledge infrastructure that integrates evidential, conceptual, disciplinary, expertise, and normative layers under a shared schema. SciAtlas further achieves a unified neuro-symbolic retrieval mechanism that grounds heterogeneous research objects, propagates relevance across the scholarly topology, and projects the resulting relevance field into the context required by each scientific workflow. Acro

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling

arXiv:2605.07815v2 Announce Type: replace-cross Abstract: Muon fixes the \emph{direction} of every matrix-valued update at the polar factor of its momentum, while each layer's step \emph{magnitude} is addressed only by a static shape correction. We derive a dynamic per-layer scalar by adapting the LARS/LAMB trust-ratio principle to the orthogonalized setting, where the standard denominator candidates---the raw momentum norm or the polar-factor norm---either live in the wrong unit space or carry no update-scale information. The resulting method, \emph{OrScale}, uses the norm of the parameter-space direction actually applied and anchors each layer's ratio at one via a per-layer calibration, so that the Moonlight recipe (tuned for AdamW, shared with Muon via RMS matching) transfers with \emph{no additional sweep}; a component ablation confirms each design choice is individually load-bearing. Theoretically, OrScale retains a nuclear-norm $O(1/\sqrt{T})$ convergence rate for any clipped mul

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

PairAlign: A Framework for Autoregressive Tokenization via Self-Alignment with Applications to Audio Tokenization

arXiv:2605.06582v4 Announce Type: replace-cross Abstract: Modern learning systems represent perceptual signals with continuous vectors, but comparison, retrieval, memory, alignment, and reasoning are often symbolic. In language, tokens provide this interface; for speech and audio, it must be learned. Existing audio tokenizers rely on local quantization, clustering, or reconstruction, leaving sequence consistency, compactness, length, termination, and edit geometry only indirectly controlled. We introduce PairAlign, a framework for compact audio tokenization through autoregressive self-alignment. An encoder maps speech to a continuous condition, and an autoregressive decoder emits tokens from BOS to EOS. Given two content-preserving views, PairAlign derives a canonical anchor target and trains both views to predict it, with unrelated in-batch targets as competing sequences. It first learns an autoregressive bridge from VQ targets and then transitions to EMA-teacher self-alignment with g

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models

arXiv:2604.20994v2 Announce Type: replace-cross Abstract: The growth of agentic AI has drawn significant attention to function calling Large Language Models (LLMs), which are designed to extend the capabilities of AI-powered system by invoking external functions. Injection and jailbreaking attacks have been extensively explored to showcase the vulnerabilities of LLMs to user prompt manipulation. The expanded capabilities of agentic models introduce further vulnerabilities via their function calling interface. Recent work in LLM security showed that function calling can be abused, leading to data tampering and theft, causing disruptive behavior such as endless loops, or causing LLMs to produce harmful content in the style of jailbreaking attacks. This paper introduces a novel function hijacking attack (FHA) that manipulates the tool selection process of agentic models to force the invocation of an attacker-chosen function. While existing attacks focus on semantic preference of the model

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

WebXSkill: Skill Learning for Autonomous Web Agents

arXiv:2604.13318v2 Announce Type: replace-cross Abstract: Autonomous web agents powered by large language models (LLMs) remain brittle on long-horizon browser workflows. A key bottleneck is a grounding gap in existing skill formulations: textual workflow skills provide natural language guidance but cannot be directly executed, while code-based skills execute without giving the agent step-level guidance for adaptation or recovery. We introduce WebXSkill, a framework that bridges this gap with executable skills, each pairing a parameterized action program with step-level natural-language guidance. WebXSkill operates in three stages: skill extraction mines reusable action subsequences from readily available synthetic agent trajectories and abstracts them into parameterized skills, skill organization indexes them into a URL-based graph for context-aware retrieval, and skill deployment exposes two complementary modes, grounded mode for fully automated execution and guided mode where skills

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search

arXiv:2603.22341v2 Announce Type: replace-cross Abstract: While prior red-teaming efforts have focused on eliciting harmful text outputs from large language models (LLMs), such approaches fail to capture agent-specific vulnerabilities that emerge through multi-step tool execution, particularly in rapidly growing ecosystems such as the Model Context Protocol (MCP). To address this gap, we propose a trajectory-aware evolutionary search method, T-MAP, which leverages execution trajectories to guide the discovery of adversarial prompts. Our approach enables the automatic generation of attacks that not only bypass safety guardrails but also reliably realize harmful objectives through actual tool interactions. Empirical evaluations across diverse MCP environments demonstrate that T-MAP substantially outperforms baselines in attack realization rate (ARR) and remains effective against frontier models, including GPT-5.2, Gemini-3-Pro, Qwen3.5, and GLM-5, thereby revealing previously underexplor

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

Retrieval-Augmented LLM Agents: Learning to Learn from Experience

arXiv:2603.18272v2 Announce Type: replace-cross Abstract: While large language models (LLMs) have advanced the development of general-purpose agents, robust generalization to unseen tasks remains challenging. Two common approaches are supervised fine-tuning and training-free memory-augmented generation using retrieved experience; yet both have limitations: fine-tuning often fails to extrapolate to new tasks, while experience retrieval often underperforms compared to supervised baselines. In this work, we combine these approaches and study how retrieval-augmented LLM agents can learn to use retrieved trajectories in-context. First, we establish a strong LoRA fine-tuning baseline that outperforms several state-of-the-art agent training pipelines. Second, we analyze key design choices for experience retrieval, including storage, querying, and trajectory selection. We then integrate experience retrieval directly into the fine-tuning process, finding that this substantially improves general

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

Answer Bubbles: Information Exposure in AI-Mediated Search

arXiv:2603.16138v2 Announce Type: replace-cross Abstract: Generative search systems are increasingly replacing link-based retrieval with AI-generated summaries, yet little is known about how these systems differ in sources, language, and fidelity to cited material. We examine responses to 11,000 real search queries across five systems---vanilla GPT, Search GPT, Perplexity Search with Grok, Google AI Overviews, and traditional Google Search---at three levels: source diversity, linguistic characterization of the generated summary, and source-summary fidelity. We find that generative search systems exhibit significant \textit{source-selection} biases in their citations, favoring certain sources over others. Incorporating search also selectively attenuates epistemic markers, reducing hedging by up to 60\% while preserving confidence language in the AI-generated summaries. At the same time, AI summaries further compound the citation biases: Wikipedia and longer sources are disproportionatel

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment

arXiv:2603.10009v2 Announce Type: replace-cross Abstract: Despite their sophisticated general-purpose capabilities, Large Language Models (LLMs) often fail to align with diverse individual preferences because standard post-training methods, like Reinforcement Learning with Human Feedback (RLHF), optimize for a single, global objective. While Group Relative Policy Optimization (GRPO) is a widely adopted on-policy reinforcement learning framework, its group-based normalization implicitly assumes that all samples are exchangeable, inheriting this limitation in personalized settings. This assumption conflates distinct user reward distributions and systematically biases learning toward dominant preferences while suppressing minority signals. To address this, we introduce Personalized GRPO (P-GRPO), a novel alignment framework that decouples advantage estimation from immediate batch statistics. By normalizing advantages against preference-group-specific reward histories rather than the concu

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

Dynamic Model Routing and Cascading for Efficient LLM Inference: A Survey

arXiv:2603.04445v3 Announce Type: replace-cross Abstract: The rapid growth of large language models (LLMs) with diverse capabilities, costs, and domains has created a critical need for intelligent model selection at inference time. While smaller models suffice for routine queries, complex tasks demand more capable models. However, static model deployment does not account for the complexity and domain of incoming queries, leading to suboptimal performance and increased costs. Dynamic routing systems that adaptively select models based on query characteristics have emerged as a solution to this challenge. This survey provides a systematic analysis of multi-LLM routing and cascading approaches, focusing on systems that route queries across a pool of independently trained LLMs at inference time. We cover diverse routing paradigms, including query difficulty, human preferences, clustering, uncertainty quantification, reinforcement learning, multimodality, and cascading. For each paradigm, w

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models

arXiv:2603.02482v2 Announce Type: replace-cross Abstract: Safety evaluation of multimodal large language models requires tracking not only whether an attack succeeds, but also how the interaction unfolds across turns and input modalities. We present MUSE (Multimodal Unified Safety Evaluation), an open-source, browser-based, run-centric platform for multimodal safety evaluation. MUSE treats each attack run as the persistent unit of execution, inspection, and analysis, preserving its configuration, multi-turn trajectory, delivered modalities and media, target responses, and safety judgments. A five-level response taxonomy further distinguishes full Compliance from Partial Compliance and refusal behavior, yielding hard ASR, soft ASR, and gray-zone width (GZW). Across 11,700 evaluations on six multimodal LLMs, direct text-only requests yield only 3.1% macro hard ASR and 4.4% soft ASR, while iterative attack procedures are substantially more effective. Attack effectiveness also varies subst

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

More Haste, Less Speed: Weaker Single-Layer Watermark Improves Distortion-Free Watermark Ensembles

arXiv:2602.11793v2 Announce Type: replace-cross Abstract: Watermarking has emerged as a crucial technique for detecting and attributing content generated by large language models. While recent advancements have utilized watermark ensembles to enhance robustness, prevailing methods typically prioritize maximizing the strength of the watermark at every individual layer. In this work, we identify a critical limitation in this "stronger-is-better" approach: strong watermarks significantly reduce the entropy of the token distribution, which paradoxically weakens the effectiveness of watermarking in subsequent layers. We theoretically and empirically show that detectability is bounded by entropy and that watermark ensembles induce a monotonic decrease in both entropy and the expected green-list ratio across layers. To address this inherent trade-off, we propose a general framework that utilizes weaker single-layer watermarks to preserve the entropy required for effective multi-layer ensembli

Source ↗
technology Tue, 01 Sep 2026 00:00:00 -0400
arXiv cs.CL

Constrained Group Relative Policy Optimization

arXiv:2602.05863v3 Announce Type: replace-cross Abstract: Group Relative Policy Optimization (GRPO) remains the dominant critic-free approach for fine-tuning LLMs and VLMs, but its compatibility with constrained policy optimization (e.g. for safety-critical domains) has not been carefully examined. In this work, we introduce Constrained GRPO, a Lagrangian-based extension of GRPO for constrained policy optimization. We show that the standard practice of scalarizing rewards before normalization introduces a critical Lagrangian-specific failure mode: GRPO's within-group normalization makes constrained optimization highly sensitive to how multi-component learning signals are aggregated. We show that scalarizing rewards before normalization introduces shared-denominator coupling, so that changing one multiplier alters not only the emphasis on its corresponding constraint, but also the relative weighting of the reward and other constraints. We address this with a simple but crucial modificat

Source ↗
Showing 5051–5100 of 10879 signals
← Prev Page 102 of 218 Next →