EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18402 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

behavior Thu, 30 Jul 2026 13:45:21 +0000
District Admin

To recruit teachers, school districts are building homes

Carolina Sanchez Garcia, a teacher at Rosa Parks Elementary in San Diego, used to wake up at 3 a.m. to leave enough time to drive to work as her children slept in the back seat. The post To recruit teachers, school districts are building homes appeared first on District Administration .

Source ↗
technology Thu, 30 Jul 2026 13:45:00 +0000
MedCity News

The Code Gets You Paid — The Variant Gets You Treated

The between a code that satisfies a payer and the information a physician needs to treat a person, is what nearly every healthcare AI tool is racing past. As we push these systems toward precision medicine and genomics, that gap stops being an annoyance and becomes a patient safety issue. The post The Code Gets You Paid — The Variant Gets You Treated appeared first on MedCity News .

Source ↗
behavior Thu, 30 Jul 2026 13:00:00 +0000
eSchool News

First graders need a chance to explain their math thinking

In first grade, educators shape the trajectories of how students feel about math. Long before they discover algebra or standardized tests, six- and seven-year-olds are forming beliefs about whether they are "math people," if their ideas are worth sharing, and what their mistakes say about them as learners.

Source ↗
regulation Thu, 30 Jul 2026 12:30:00 +0000
The 74

New Mexico Features in National Discussion on Childcare Challenges

A bipartisan coalition of state lawmakers from around the nation on Monday identified lack of facilities and high costs as two leading challenges states are facing in addressing childcare needs. The bipartisan group, which includes a New Mexico lawmaker, presented their findings at the National Conference of State Legislatures in Chicago, along with several proposed […]

Source ↗
technology Thu, 30 Jul 2026 11:45:00 +0000
MedCity News

Moving Beyond Reactive Healthcare: 5 Traits of Successful Healthcare Organizations

[Sponsored] A new whitepaper by League highlights examples from collaborations with Baptist Health. The post Moving Beyond Reactive Healthcare: 5 Traits of Successful Healthcare Organizations appeared first on MedCity News .

Source ↗
regulation Thu, 30 Jul 2026 10:30:00 +0000
The 74

Opinion: How Math Became Relevant Again in a Wisconsin Classroom

It is impossible to imagine a future that does not have data science embedded in our daily lives. The field combines statistics, coding, and critical thinking to extract meaning from data, making it a desirable skill across industries. Yet most American students graduate without any tested knowledge of it. One high school teacher in Wisconsin […]

Source ↗
behavior Thu, 30 Jul 2026 09:53:00 +0000
Getting Smart

Steward Stories: How the PAST Foundation Turned Real-World Problems Into a Learning Ecosystem

What does it look like when a foundation built by a field scientist becomes the connective tissue of an entire regional learning ecosystem? The PAST Foundation's story, told through 25 years of timeline milestones and student voices, offers education leaders a concrete model for weaving together schools, industry, community, and out-of-school time into something greater than any single program. From portable STEM labs to microschool pathways co-launched with universities and industry partners, PAST demonstrates how systemic change begins with a commitment to learning that is rooted in real work and real relationships. This is essential reading for leaders ready to move from coordination to transformation. The post Steward Stories: How the PAST Foundation Turned Real-World Problems Into a Learning Ecosystem appeared first on Getting Smart .

Source ↗
behavior Thu, 30 Jul 2026 09:53:00 +0000
Getting Smart

How the PAST Foundation Turned Real-World Problems Into a Learning Ecosystem

What does it look like when a foundation built by a field scientist becomes the connective tissue of an entire regional learning ecosystem? The PAST Foundation's story, told through 25 years of timeline milestones and student voices, offers education leaders a concrete model for weaving together schools, industry, community, and out-of-school time into something greater than any single program. From portable STEM labs to microschool pathways co-launched with universities and industry partners, PAST demonstrates how systemic change begins with a commitment to learning that is rooted in real work and real relationships. This is essential reading for leaders ready to move from coordination to transformation. The post How the PAST Foundation Turned Real-World Problems Into a Learning Ecosystem appeared first on Getting Smart .

Source ↗
technology Thu, 30 Jul 2026 09:00:00 +0000
Tech & Learning

What is CK-12 and How Can Teachers Use It?

CK-12 is an online education platform designed for K-12 teachers, schools, and families that offers customizable teaching and learning

Source ↗
technology Thu, 30 Jul 2026 07:23:40 -0400
EdTech Mag (K-12)

Security Awareness Training Reduces Phishing Success

Human error is the soft underbelly of cybersecurity. According to IBM, it plays a role in roughly 95% of breaches, a statistic that looms especially large in K–12 education. Schools are uniquely vulnerable, with thousands of users, limited IT resources, and an environment built on openness and trust. “People talk about humans as the weakest link,” says Randy Rose, vice president of security operations and intelligence at the Center for Internet Security. “And the reason they get that rap is because the No. 1 factor in a majority of cyber incidents is social engineering, mostly phishing emails…

Source ↗
audience Thu, 30 Jul 2026 07:00:00 +0000
Inside Higher Ed

Law Schools Are Asking the Wrong Question About AI

Law Schools Are Asking the Wrong Question About AI Elizabeth Redden Thu, 07/30/2026 - 03:00 AM The more interesting one is not about academic integrity. It’s about what it means to assess professional competence in a fast-changing field. Byline(s) James Finkelstein

Source ↗
audience Thu, 30 Jul 2026 07:00:00 +0000
Inside Higher Ed

A First Lawsuit Tests What Universities Are Owed for AI Research

A First Lawsuit Tests What Universities Are Owed for AI Research sara.custer@in… Thu, 07/30/2026 - 03:00 AM The University of Tennessee’s patent suit against Anthropic challenges the technology’s core architecture—and could set a precedent for other institutions. Byline(s) Sara Custer

Source ↗
audience Thu, 30 Jul 2026 07:00:00 +0000
Inside Higher Ed

‘Untenable’ Conditions Shrinking Tenure of Law Deans, Survey Finds

‘Untenable’ Conditions Shrinking Tenure of Law Deans, Survey Finds kathryn.palmer… Thu, 07/30/2026 - 03:00 AM Byline(s) Kathryn Palmer

Source ↗
audience Thu, 30 Jul 2026 07:00:00 +0000
Inside Higher Ed

West Point Won’t Fight Court Ruling on Faculty Gag Order

West Point Won’t Fight Court Ruling on Faculty Gag Order Josh Moody Thu, 07/30/2026 - 03:00 AM Byline(s) Josh Moody

Source ↗
audience Thu, 30 Jul 2026 07:00:00 +0000
Inside Higher Ed

Beyond Articulation Agreements

Beyond Articulation Agreements quintina.barne… Thu, 07/30/2026 - 03:00 AM Recognition as a lever for improving transfer. Byline(s) Lynn Tincher-Ladner

Source ↗
audience Thu, 30 Jul 2026 07:00:00 +0000
Inside Higher Ed

Gates Foundation’s New Goal: 10M ‘Credentials of Value’ by 2045

Gates Foundation’s New Goal: 10M ‘Credentials of Value’ by 2045 jessica.blake@… Thu, 07/30/2026 - 03:00 AM Byline(s) Jessica Blake

Source ↗
audience Thu, 30 Jul 2026 07:00:00 +0000
Inside Higher Ed

Fired Over Speech, Professors Take Their Employers to Court—and Win

Fired Over Speech, Professors Take Their Employers to Court—and Win Emma Whitford Thu, 07/30/2026 - 03:00 AM The recent spate of censorship and academic freedom violations is the worst since the McCarthy era, according to one lawyer. But as faculty rack up legal victories, universities may start to think twice before punishing extramural speech. Byline(s) Emma Whitford

Source ↗
audience Thu, 30 Jul 2026 07:00:00 +0000
Inside Higher Ed

National AAUP Starts Endorsing Political Candidates

National AAUP Starts Endorsing Political Candidates Ryan Quinn Thu, 07/30/2026 - 03:00 AM The group says the move is a first in its over-100-year existence, and part of countering government threats. Critics say AAUP has lost its neutrality. Byline(s) Ryan Quinn

Source ↗
audience Thu, 30 Jul 2026 07:00:00 +0000
Inside Higher Ed

‘Catastrophic’ Tech Issues Postpone Wash. Bar Exam

‘Catastrophic’ Tech Issues Postpone Wash. Bar Exam kathryn.palmer… Thu, 07/30/2026 - 03:00 AM Roughly 645 people will have to take the newly redesigned exam later this year or next. One outraged law dean says those options aren’t good enough for the thwarted test takers. Byline(s) Kathryn Palmer

Source ↗
audience Thu, 30 Jul 2026 05:00:00 -0400
Higher Ed Dive

Texas community college funding model faces ‘growing pains’

Although leaders praised the performance-based formula, they asked lawmakers to provide sufficient money for it.

Source ↗
regulation Thu, 30 Jul 2026 05:00:00 -0400
K-12 Dive

This New York superintendent is embracing a human-first approach to AI in schools

As a “dot-commer” who transitioned to education, Jared Bloom brings a unique perspective to K-12 tech adoption.

Source ↗
regulation Thu, 30 Jul 2026 05:00:00 -0400
K-12 Dive

American support for all-day cellphone bans hits record high, survey finds

Young adults were the least likely to support all-day school bans, but at least half of Americans age 30 or older back them.

Source ↗
need Thu, 30 Jul 2026 05:00:00 +0000
Hechinger Report

As fewer young people choose college, this district wants to ensure they have career options

NIAGARA FALLS, N.Y. — For the last year, Niagara Falls High School senior Ciara Johnson has ripped up tile, torn out walls and built custom cabinets in her high school construction class. To get a taste of each of the trades, the students here have been renovating the former cafeteria kitchen in the old high […] The post As fewer young people choose college, this district wants to ensure they have career options appeared first on The Hechinger Report .

Source ↗
behavior Thu, 30 Jul 2026 00:00:00 GMT
EdSurge

Reality Bites: Students Say They Face Stark Challenges After High School

A new national survey finds a widening gap between classroom preparation and postgraduation life.

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

Learning What Matters: Supervising Global Context Pruning with Causal Evidence Sets

arXiv:2607.21692v2 Announce Type: replace-cross Abstract: Sparse attention prunes a long context to the blocks a model needs, and the usual selector is distilled from a dense teacher's attention. That assumes attention shows which context the answer depends on. We test the assumption on retrieval tasks where the evidence is known exactly, by masking context and measuring whether the answer changes. Attention and causal dependence disagree, and selectors inherit the disagreement. Teachers attend to outdated facts they have learned to ignore, and attend differently across training runs that use the same evidence. In a two-step reference task, attention at the answer position can skip the intermediate step, and how often it skips varies with the training run: selecting one block set per pass, a selector distilled from attention routes at 36% to 98% across teachers, the same selector trained on causal evidence sets reaches 99% or better on every one, and dense accuracy does not say which t

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

VistaHop: Benchmarking Long-Horizon Visual DeepSearch

arXiv:2606.03273v2 Announce Type: replace-cross Abstract: Visual DeepSearch tasks require multimodal large language models (MLLMs) to resolve complex visual queries by repeatedly inspecting image regions, grounding reasoning in visual evidence, and connecting fine-grained clues across multiple steps. However, existing benchmarks primarily evaluate single-step visual understanding or isolated visual-query response generation. They have limited difficulty, limited search horizons, and single-pass image inspection, and thus fail to evaluate models' ability to iteratively revisit visual evidence and reason across multiple steps. In this work, we introduce VistaHop, a benchmark designed specifically to evaluate Visual DeepSearch. It evaluates repeated image inspection, visual-anchor grounding, and long-horizon evidence traversal across different visual regions. VistaHop comprises 600 images, 25 visual search scenarios, and 600 Visual DeepSearch tasks. We also propose VistaArena, a unified e

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents

arXiv:2605.22148v2 Announce Type: replace-cross Abstract: Self-evolving skill libraries, pioneered by Voyager, let frozen LLM agents accumulate reusable knowledge without weight updates, yet recent evaluation shows that LLM-authored skills deliver $+0.0$pp over no-skill baselines while human-curated ones deliver $+16.2$pp: the bottleneck is not skill authoring but lifecycle management. We introduce \textbf{Ratchet}, a single-agent loop in which a frozen LLM writes, retrieves, curates, and retires its own natural-language skills. Ratchet integrates four candidate hygiene mechanisms: outcome-driven retirement, a bounded active-cap, meta-skill authoring guidance, and pattern canonicalisation. On MBPP+ hard-100 with Claude Opus 4.7, Ratchet lifts held-out pass@1 from a $0.258 \pm 0.047$ baseline to a late-window rolling mean of $0.584$ (peak $0.658 \pm 0.042$) across 100 rounds and 3 seeds, a $+0.328 \pm 0.018$ rolling-mean gain where the no-skill control drifts at $+0.002 \pm 0.005$; the

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

LLMAR: A Tuning-Free Recommendation Framework for Sparse and Text-Rich Industrial Domains

arXiv:2604.16379v2 Announce Type: replace-cross Abstract: Industrial B2B applications (e.g., construction site risk prediction, material procurement) face extreme data sparsity yet feature rich textual interactions. In such environments, traditional ID-based collaborative filtering fails lacking co-occurrence signals, while fine-tuning standard Large Language Models (LLMs) incurs high operational costs and struggles with frequent data drift. We propose LLMAR (LLM-Annotated Recommendation), a tuning-free framework. Moving beyond simple embeddings, LLMAR systematically integrates LLM reasoning to capture user "latent motives" without any training process. We introduce three core contributions: (1) Inference-Driven Annotation: uses LLMs to transform behavioral history into structured semantic motives, enabling reasoning-based matching unattainable by ID-based methods; (2) Reflection Loop: a self-correction mechanism that refines generated queries to mitigate hallucinations and resolve "co

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

How memory can affect collective and cooperative behaviors in an LLM-Based Social Particle Swarm

arXiv:2604.12250v2 Announce Type: replace-cross Abstract: This study examines how memory shapes the collective and cooperative dynamics of Large Language Model (LLM) agents in a multi-agent system. To this end, we extend the Social Particle Swarm (SPS) model, in which agents move in a two-dimensional space and play the Prisoner's Dilemma with neighboring agents, by replacing its rule-based agents with LLM agents endowed with Big Five personality scores and varying memory lengths. Using Gemini 2.0 Flash, we find that memory length is a critical parameter governing collective behavior: even a minimal memory drastically suppressed cooperation, transitioning the system from stable cooperative clusters through cyclical formation and collapse of clusters to a state of scattered defection as memory length increased. Big Five personality traits correlated with agent behaviors in partial agreement with findings from experiments with human participants, supporting the validity of the model. This

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

GroupRAG: Cognitively Inspired Group-Aware Retrieval and Reasoning via Knowledge-Driven Problem Structuring

arXiv:2603.26807v2 Announce Type: replace-cross Abstract: The performance of language models is commonly limited by insufficient knowledge and constrained reasoning. Prior approaches such as Retrieval-Augmented Generation (RAG) and Chain-of-Thought (CoT) address these issues by incorporating external knowledge or enforcing linear reasoning chains, but often degrade in real-world settings. Inspired by cognitive science, which characterizes human problem solving as search over structured problem spaces rather than single inference chains, we argue that inadequate awareness of problem structure is a key overlooked limitation. We propose GroupRAG, a cognitively inspired, group-aware retrieval and reasoning framework based on knowledge-driven keypoint grouping. GroupRAG identifies latent structural groups within a problem and performs retrieval and reasoning from multiple conceptual starting points, enabling fine-grained interaction between the two processes. Experiments on MedQA (medical)

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

TANDEM: Temporal-Aware Neural Detection for Multimodal Hate Speech

arXiv:2601.11178v3 Announce Type: replace-cross Abstract: Social media platforms are increasingly dominated by long-form multimodal content, where harmful narratives are constructed through a complex interplay of audio, visual, and textual cues. While automated systems can flag hate speech with high accuracy, they often function as "black boxes" that fail to provide the granular, interpretable evidence, such as precise timestamps and target identities, required for effective human-in-the-loop moderation. In this work, we introduce TANDEM, a unified framework that transforms audio-visual hate detection from a binary classification task into a structured reasoning problem. Our approach employs a novel tandem reinforcement learning strategy where vision-language and audio-language models optimize each other through self-constrained cross-modal context, stabilizing reasoning over extended temporal sequences without requiring dense frame-level supervision. Experiments across three benchmark

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

AdaMARP: An Adaptive Multi-Agent Interaction Framework for General Immersive Role-Playing

arXiv:2601.11007v2 Announce Type: replace-cross Abstract: LLM role-playing aims to portray arbitrary characters in interactive narratives, yet existing systems often suffer from limited immersion and adaptability. They typically under-model dynamic environmental information and assume largely static scenes and casts, offering insufficient support for multi-character orchestration, scene transitions, and on-the-fly character introduction. We propose an adaptive multi-agent role-playing framework, AdaMARP, featuring an immersive message format that interleaves [Thought], (Action), , and Speech, together with an explicit Scene Manager that governs role-playing through discrete actions (init_scene, pick_speaker, switch_scene, add_role, end) accompanied by rationales. To train these capabilities, we construct AdaRPSet for the Actor Model and AdaSMSet for supervising orchestration decisions, and introduce AdaptiveBench for trajectory-level evaluation. Experiments across multiple backbones an

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

DP-MGTD: Privacy-Preserving Machine-Generated Text Detection via Adaptive Differentially Private Entity Sanitization

arXiv:2601.04641v2 Announce Type: replace-cross Abstract: The deployment of Machine-Generated Text (MGT) detection systems necessitates processing sensitive user data, creating a fundamental conflict between authorship verification and privacy preservation. Standard anonymization techniques often disrupt linguistic fluency, while rigorous Differential Privacy (DP) mechanisms typically degrade the statistical signals required for accurate detection. To resolve this dilemma, we propose \textbf{DP-MGTD}, a framework incorporating an Adaptive Differentially Private Entity Sanitization algorithm. Our approach utilizes a two-stage mechanism that performs noisy frequency estimation and dynamically calibrates privacy budgets, applying Laplace and Exponential mechanisms to numerical and textual entities respectively. Crucially, we identify a counter-intuitive phenomenon where the application of DP noise amplifies the distinguishability between human and machine text by exposing distinct sensiti

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM

arXiv:2506.14766v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) frequently hallucinate by over-committing to spurious visual cues. Prior remedies-Visual and Instruction Contrastive Decoding (VCD, ICD)-mitigate this issue, yet the mechanism remains opaque. We first empirically show that their improvements systematically coincide with redistributions of cross-modal attention. Building on this insight, we propose Attention-Steerable Contrastive Decoding (ASCD), which directly steers the attention scores during decoding. ASCD combines (i) positive steering, which amplifies automatically mined text-centric heads-stable within a model and robust across domains-with (ii) negative steering, which dampens on-the-fly identified critical visual tokens. The method incurs negligible runtime and memory overhead and requires no additional training. Across five MLLM backbones and three decoding schemes, ASCD reduces hallucination on POPE, CHAIR, and MMHal-Bench by up

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation

arXiv:2607.07050v3 Announce Type: replace Abstract: Top-K teacher logits make on-policy distillation tractable, but probability mass is not the same as decision support. In a two-teacher tool-use setting, vanilla generalized knowledge distillation raises tool-call recall while also calling on examples that require direct answers. The response teacher's top-32 retains 99.99% of its probability mass yet contains the tool-call behavior-switch token on only 0.4% of 1,500 audited prompts; even top-256 covers only 52.2%. Because omitted logits receive zero direct gradient under the truncated objective, the tool teacher reinforces entry while the response teacher usually cannot oppose it. Frozen replay shows that a wrong entry then amplifies divergence along the generated trajectory. Restoring the tool-call token only at the first response position moves first-token entry but mostly delays eventual calls. Restoring it at every response position closes this gap: across three matched seeds, ful

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

Prompt Framing Distorts Count Based Evaluation of LLM Error Detection: Evidence from Numeric Anchoring

arXiv:2607.01240v2 Announce Type: replace Abstract: Count-based F1 is widely used as a proxy for LLM error-detection quality, but this paper shows that it can rise dramatically without a corresponding improvement in span localization, a gap termed F1 Inflation. The paper introduces ErrorBench, a controlled stress-test protocol for prompt-induced count distortion. ErrorBench evaluates six contemporary LLMs under five prompt conditions over 4,290 responses from 143 CoNLL-2014 passages. Under CoNLL-2014 M2-style scoring, anchored prompts produce up to 0.79 points of F1 Inflation, and up to 0.96 under strict matching. A 100-passage replication using the official ERRANT 3.0.0 pipeline and multi-reference scoring reproduces the pattern: averaged over six models, the Blind-to-Anchored prompt shift raises Count-F1 by +0.21 while raising multi-reference ERRANT F0.5 by only +0.04. The study finds larger count responses in highly instruction-compliant GPT/Claude systems and smaller responses in t

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control

arXiv:2607.01153v3 Announce Type: replace Abstract: Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model followed an instruction, refused appropriately, complied with a policy, or misreported progress in an agentic task. Existing benchmarks compress these into pass/fail labels, obscuring whether failures reflect capability limits, policy ambiguity, instruction conflict, scaffold failure, or unstable evaluator judgments. Adversarial pragmatics is safety-relevant model behaviour under instruction conflict, embedded commands, quotation, scope ambiguity, deixis, and indirect speech acts. It's designed to extend to multi-turn agent transcripts, but the seed set represents that family with a single-turn tool-result contrast. This paper introduces a diagnostic framework, an 18-item seed benchmark, a 54-row pilot, and a six-cell LLM-judge assessment, with a protocol keeping task success, policy compliance, risk, ref

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers

arXiv:2606.21848v2 Announce Type: replace Abstract: Transformer architectures form the foundation of modern natural language processing, making it crucial to address the efficiency and scalability limitations of the standard QKV attention mechanism. The Key-Value (KV) cache is a major bottleneck during long-context inference. We propose Keyless Attention, a novel attention mechanism that introduces a dedicated value-space routing projection to replace the conventional key projection, thereby eliminating key representations from the attention computation. This design yields a Value-Only Cache that reduces KV-cache memory and access overhead by 50% compared with standard attention while improving decode throughput. Experiments across five models and four architectures show that Keyless Attention matches or outperforms standard QKV attention in perplexity on four of five models. Furthermore, it achieves competitive performance on downstream evaluation benchmarks while consistently reducin

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

SkillCAT: Contrastive, Assessment-Augmented and Topology-AwareSkill Self-Evolution for LLM Agents

arXiv:2606.13317v2 Announce Type: replace Abstract: Skill self-evolution methods for LLM agents aim to turn execution trajectories into reusable skill documents. However, current pipelines typically derive skill patches from a single trajectory per task, merge them indiscriminately, and load the entire skill corpus during inference. These choices lead to unreliable evidence extraction, the accumulation of low-quality or even harmful skill edits, and inefficient use of context due to irrelevant or conflicting skill content. We propose SkillCAT, a framework that decomposes this process into three stages. (1) Contrastive Causal Extraction (CCE) samples multiple trajectories per task and contrasts same-task success/failure pairs to find the evidence that explains outcome differences. (2) Assessment-Augmented Evolution (AAE) replays each candidate patch on source-task clones, retains only those that do not damage task outcomes, and then merges the retained patches hierarchically. (3) Topolo

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

StateRAG: Typed State Contracts for Complex Retrieval-Augmented Generation

arXiv:2605.25379v2 Announce Type: replace Abstract: Complex retrieval-augmented generation requires evidence retrieval and control over what to retrieve next, which paths to explore, whether evidence is sufficient, and which intermediate results to retain. Existing RAG paradigms encode these decisions through method-specific model contexts, traversal procedures, verification signals, and memory. We introduce StateRAG, which represents retrieval control as a typed state external to the final reader. The state records the query plan, typed traversal path, candidate evidence, verification verdict, and reusable artifacts, with defined field semantics and designated update sources. Ordered role operators propose field values, and the controller validates and commits accepted proposals. A one-time compact-evidence check may select Bypass. Otherwise, the controller combines the committed verdict with the remaining budget to select Release, Revise, or Fallback. The final reader is invoked only

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents

arXiv:2605.16986v2 Announce Type: replace Abstract: Additional test-time compute can give LLM agents access to more past experience, yet expanding the context or adding rollouts does not necessarily yield greater agent capability. We call this challenge test-time compute-to-capability conversion and propose SkillTTA, which retrieves task-relevant training trajectories and synthesizes a temporary skill conditioned on the visible target context for a solver with fixed parameters. To pursue a higher performance ceiling, SkillTTA further uses meta prompt optimization (MPO) to adapt the policy that writes these skills. MPO evaluates candidate prompts on paired tasks and emphasizes informative transitions. It also confines updates to benchmark-specific atomic slots, reducing the variance caused by observing each edit only indirectly through skill synthesis and solver rollout. Across ALFWorld, SpreadsheetBench, BigCodeBench, and WebShop, SkillTTA outperforms state-of-the-art reuse and optimiz

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators

arXiv:2605.08334v2 Announce Type: replace Abstract: We present CustomerSim, an environment and benchmark to evaluate the extent to which Multimodal Large Language Models (MLLMs) can simulate realistic, persona-driven customer behavior in chat-based retail environments. While prior work treats user simulation as surface-level dialog generation, we focus on a model's ability to seek information and make decisions that adhere to customer specifications in multiturn, agentic simulations. CustomerSim consists of a human-curated set of 360 personas over five product categories, alongside a suite of metrics measuring consistency between a customer simulator's actions and its specifications and conversational quality. We find several behavioral gaps across five open and closed-source state-of-the-art models. First, while models produce fluent conversations, they display significantly lower lexical diversity than human shoppers, and open-source models overdisclose their criteria in the opening

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

LAMUS: A Large-Scale Corpus for Legal Argument Mining from U.S. Caselaw using LLMs

arXiv:2603.08286v2 Announce Type: replace Abstract: Legal argument mining aims to identify and classify the functional components of judicial reasoning, such as facts, issues, rules, analysis, and conclusions. Progress in this area is limited by the lack of large-scale, high-quality annotated datasets for U.S. caselaw, particularly at the state level. This paper introduces LAMUS, a sentence-level legal argument mining corpus constructed from U.S. Supreme Court decisions and Texas criminal appellate opinions. The dataset is created using a data-centric pipeline that combines large-scale case collection, LLM-based automatic annotation, and targeted human-in-the-loop quality refinement. We formulate legal argument mining as a six-class sentence classification task and evaluate multiple general-purpose and legal-domain language models under zero-shot, few-shot, and chain-of-thought prompting strategies, with LegalBERT as a supervised baseline. Results show that chain-of-thought prompting s

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue

arXiv:2603.06194v3 Announce Type: replace Abstract: Reinforcement learning (RL) for large language models (LLMs) has shown strong performance in single-turn tasks, but extending it to multi-turn interaction remains challenging due to sparse rewards and poor per-turn credit assignment. In emotional support dialogues, responses shape future user states, so matched-state step-wise comparison is unavailable, while trajectory-level supervision is insufficient. We propose MICA (Multi-granularity Intertemporal Credit Assignment), a critic-free RL framework for multi-turn emotional support tasks. MICA derives both immediate and delayed credit from a shared potential function over the user's structured support state. Incremental Distance Reward measures the per-turn decrease in residual distance to the target state, while its Monte Carlo return captures delayed effects. After scope-specific normalization, the two signals form a mixed advantage for stable per-turn optimization without matched-st

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

Making Implicit Premises Explicit in Logical Understanding of Enthymemes

arXiv:2603.06114v2 Announce Type: replace Abstract: Real-world arguments in text and dialogues are normally enthymemes (i.e. some of their premises and/or claims are implicit). Natural language processing (NLP) methods for handling enthymemes can potentially identify enthymemes in text but they do not decode their underlying logic, whereas logic-based approaches for handling them assume a knowledgebase with sufficient formulae that can be used to decode them via abduction. There is therefore a lack of a systematic method for translating textual components of an enthymeme into a logical argument and generating the logical formulae required for their decoding, and thereby showing logical entailment. To address this, we propose a pipeline that integrates: (1) a large language model (LLM) to generate intermediate implicit premises based on the explicit premise and claim; (2) another LLM to translate the natural language into logical formulas; and (3) a neuro-symbolic reasoner based on a SA

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

DialectLLM: A Dialect-Aware Dialog[ue] Generation Framework Beyond Standard American English

arXiv:2601.22888v4 Announce Type: replace Abstract: More than 80% of the 1.6B English speakers do not use Standard American English (SAE), yet LLMs often fail to correctly identify non-SAE dialects and generate stereotyped responses for their speakers. We introduce DialectLLM, the first large-scale framework for generating high-quality multi-dialectal conversational data encompassing the three pillars of written dialect -- lexical (vocabulary), orthographic (spelling), and morphosyntactic (grammar) features. DialectLLM produces a dialect-parallel dialog dataset spanning nine English dialects. Partnering with native linguists, we design and validate SAE-to-dialect transformation rules, ensuring authenticity. Our approach challenges the prevailing practice of applying a single morphosyntactic feature set to both user utterances and model responses, showing that models should not reproduce up to 90% of the grammatical features of a dialect. Human evaluation confirms data quality, with ann

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs

arXiv:2601.06599v3 Announce Type: replace Abstract: Large Language Models (LLMs) often encode whether a statement is true as a vector in their residual stream activations. These vectors, also known as truth vectors, have been studied in prior work, however how they change when context is introduced remains unexplored. We study this question by measuring (1) the directional change ($\theta$) between the truth vectors with and without context and (2) the relative magnitude of the truth vectors upon adding context. Across four LLMs and four datasets, we find that (1) truth vectors are roughly orthogonal in early layers, converge in middle layers, and may stabilize or continue increasing in later layers; (2) adding context generally increases the truth vector magnitude, i.e., the separation between true and false representations in the activation space is amplified; (3) larger models distinguish relevant from irrelevant context mainly through directional change ($\theta$), while smaller mo

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

$\texttt{AMEND++}$: Benchmarking Eligibility Criteria Amendments in Clinical Trials

arXiv:2601.06300v2 Announce Type: replace Abstract: Clinical trial amendments frequently introduce delays, increased costs, and administrative burden, with eligibility criteria being the most commonly amended component. We introduce \textit{eligibility criteria amendment prediction}, a novel NLP task that aims to forecast whether the eligibility criteria of an initial trial protocol will undergo future amendments. To support this task, we release $\texttt{AMEND++}$, a benchmark suite comprising two datasets: $\texttt{AMEND}$, which captures eligibility-criteria version histories and amendment labels from public clinical trials, and $\verb|AMEND_LLM|$, a refined subset curated using an LLM-based denoising pipeline to isolate substantive changes. We further propose $\textit{Change-Aware Masked Language Modeling}$ (CAMLM), a revision-aware pretraining strategy that leverages historical edits to learn amendment-sensitive representations. Experiments across diverse baselines show that CAMLM

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

Large Emotional World Model

arXiv:2512.24149v2 Announce Type: replace Abstract: The world is governed by both physical laws and affective dynamics. Physical laws govern state transitions, while affective dynamics shape human actions, decisions, and interactions. A world model that learns only physical laws can approximate the physical world, but not the human world. In this paper, we introduce human emotion as a key state variable in world models, enabling them to capture both future state transitions and their emotional causes. We first construct Emotion-Why-How (EWH), the first world model dataset centered on emotional state transitions, containing 10,850 emotion-aware transition tuples. Each tuple encodes the pre-state, pre-emotion, action, post-emotion, and post-state, supporting reasoning about why actions occur and how emotions reshape future states. Based on EWH, we propose the Large Emotional World Model (LEWM), which factorizes future prediction into two coupled steps: first predicting the future emotion

Source ↗
technology Thu, 30 Jul 2026 00:00:00 -0400
arXiv cs.CL

ARC-Encoder: learning compressed text representations for large language models

arXiv:2510.20535v2 Announce Type: replace Abstract: Recent techniques such as retrieval-augmented generation or chain-of-thought reasoning have led to longer contexts and increased inference costs. Context compression techniques can reduce these costs, but the most effective approaches require fine-tuning the target model or even modifying its architecture. This can degrade its general abilities when not used for this specific purpose. Here we explore an alternative approach: an encoder that compresses the context into continuous representations which replace token embeddings in decoder LLMs. First, we perform a systematic study of training strategies and architecture choices for the encoder. Our findings led to the design of an Adaptable text Representations Compressor, named ARC-Encoder, which outputs $x$-times fewer continuous representations (typically $x\!\in\!\{4,8\}$) than text tokens. We evaluate ARC-Encoder across a variety of LLM usage scenarios, ranging from in-context learn

Source ↗
Showing 6451–6500 of 18402 signals
← Prev Page 130 of 369 Next →