EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Sep 07, 2026 · 40 ideas · 18624 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Thu, 13 Aug 2026 19:04:35 +0000
MedCity News

PTC Therapeutics’ $211M Bid Wins Bankruptcy Auction for Sangamo Gene Therapy

Sangamo Therapeutics filed for Chapter 11 bankruptcy in June after its search for strategic alternatives failed to find a path forward for the company. PTC Therapeutics’ auction win brings a Sangamo gene therapy candidate for Fabry disease, a rare inherited disorder with few treatment options. The post PTC Therapeutics’ $211M Bid Wins Bankruptcy Auction for Sangamo Gene Therapy appeared first on MedCity News .

Source ↗
regulation Thu, 13 Aug 2026 18:30:00 +0000
The 74

Gap to Adequately Fund Schools Widens in Illinois

Chicago Public Schools is further away from having enough money to meet its students’ needs as defined by the state, according to new data released by the Illinois State Board of Education. Under ISBE’s calculations released Friday, CPS is now considered a little over 70% adequately funded, compared with 73% last year. CPS isn’t alone […]

Source ↗
audience Thu, 13 Aug 2026 16:57:03 +0000
Inside Higher Ed

Judge Dismisses Trump’s Antisemitism Lawsuit Against Harvard

Judge Dismisses Trump’s Antisemitism Lawsuit Against Harvard jessica.blake@… Thu, 08/13/2026 - 12:57 PM The decision is a blow for the administration, which has engaged in a yearslong battle with the Ivy League institution. Byline(s) Jessica Blake

Source ↗
regulation Thu, 13 Aug 2026 16:33:00 -0400
K-12 Dive

New York City Public Schools to place ‘school avoidance’ liaisons in every building

The nation's largest school system will task the personnel with addressing a student behavior that affects up to 15% of youth and fuels chronic absenteeism.

Source ↗
regulation Thu, 13 Aug 2026 16:30:00 +0000
The 74

Opinion: Congress Made it Easier to Pay for College. It Must Keep That Promise

Right now, 15.5 million undergraduate students around the country are getting ready to start or return to college. Seven million of them have one thing in common: relying on a Pell Grant to help pay for their higher education. But the grant program is facing a $15 billion funding shortfall that will undermine college access […]

Source ↗
audience Thu, 13 Aug 2026 15:16:33 -0400
Higher Ed Dive

DOJ’s antisemitism lawsuit against Harvard dismissed

Harmeet Dhillon, assistant attorney general of the agency’s civil rights division, said federal officials are "assessing next steps."

Source ↗
regulation Thu, 13 Aug 2026 14:30:00 +0000
The 74

Indiana Literacy Rates Improve for Fifth Straight Year

Indiana’s early literacy rates improved for the fifth consecutive year, with nearly 89% of Hoosier third graders demonstrating proficiency in foundational reading skills on the state’s IREAD assessment. The Indiana Department of Education released statewide IREAD, ILEARN and SAT results for the 2025-26 school year Tuesday at the State Board of Education’s August meeting. The […]

Source ↗
technology Thu, 13 Aug 2026 13:54:40 +0000
MedCity News

Founders Should Stop Worrying If They’re “AI Enough”

In healthcare, enterprises are moving slower than the hype cycle would lead you to believe. That said, they’re not moving slowly – they are saying yes to AI, as long as it can solve for specific workflow needs. The post Founders Should Stop Worrying If They’re “AI Enough” appeared first on MedCity News .

Source ↗
technology Thu, 13 Aug 2026 13:37:38 -0400
EdTech Mag (Higher)

AI and Enrollment Pressures Are Reshaping Higher Education

Higher education is at an inflection point. The 2026 EDUCAUSE Horizon Report identifies the forces most likely to reshape teaching and learning over the next decade — and the picture is complex. As artificial intelligence continues to evolve, it’s changing relationships across campus. Institutions also continue to face mounting enrollment challenges. And emerging risks related to cybersecurity, policy reform and sustainability are compounding institutional strain. AI’s Influence on Teaching and Learning According to the report, AI is redefining instructional design and how faculty teach…

Source ↗
technology Thu, 13 Aug 2026 13:32:08 -0400
EdTech Mag (K-12)

6 Data Governance Best Practices for K–12

K–12 districts are entrusted with a wide range of student data. Attendance records, grades, behavior logs, Individualized Education Programs and health records are gathered and stored in school systems, often for years after children move on to other districts or graduate. And as schools adopt more ed tech tools, they’re responsible for even more data. With that responsibility comes risk: A single mismanaged account can lead to an exposure that puts students’ information at risk. Strong data governance requires building smart policy; it also requires putting into place deliberate practices…

Source ↗
behavior Thu, 13 Aug 2026 13:15:50 +0000
District Admin

Chronic absenteeism remains high six years after pandemic began

Nearly a quarter of U.S. children did not show up for school regularly in 2025, a slight decline from the previous year but still up from about 15 percent before COVID. The post Chronic absenteeism remains high six years after pandemic began appeared first on District Administration .

Source ↗
behavior Thu, 13 Aug 2026 13:10:33 +0000
District Admin

Why transportation is such an important student success factor

Predictable transportation routines can improve students' readiness to participate in classroom learning, particularly for children with special needs. The post Why transportation is such an important student success factor appeared first on District Administration .

Source ↗
behavior Thu, 13 Aug 2026 13:05:15 +0000
District Admin

The McKinney-Vento Challenge: Turning the “Without Delay” Transportation Standard into Reality

Date & Time: Wednesday, September 09, 2026 at 2 p.m. Join HopSkipDrive and Igor Petrovic, Director of Transportation at Adams 12 Five Star Schools in Colorado, for a look at survey data that reveals a major gap between the McKinney-Vento standard and district reality, adding up to missed instructional time and real compliance risk. Attendees will also gain a district-level view of how to weigh cost, speed, and service when transportation needs can’t wait. The post The McKinney-Vento Challenge: Turning the “Without Delay” Transportation Standard into Reality appeared first on District Administration .

Source ↗
regulation Thu, 13 Aug 2026 12:30:00 +0000
The 74

Opinion: Amid the Emerging AI Economy, We Need a Skilled Trades Pipeline in High School

Across the economy, Americans are watching an artificial intelligence investment boom reshape the job market. Some of the same companies spending hundreds of billions to build the AI future are also announcing sweeping job cuts, adding to an already daunting employment landscape for college graduates. In other sectors, opportunity is booming. Homes, roads, bridges, vehicles, […]

Source ↗
regulation Thu, 13 Aug 2026 10:30:00 +0000
The 74

More Early Childhood Programs Are Providing Free Housing to Teaching Staff

This story was co-published with Mother Jones. A few years ago, Eric Gil was living at his uncle’s place in Waterbury, Connecticut, where he shared a bedroom with his brother and cousin. With eight people in the house, it was crowded. He was trying to get out, but rental prices in Waterbury — which currently […]

Source ↗
behavior Thu, 13 Aug 2026 10:00:00 +0000
eSchool News

The foundation for student success is a positive classroom culture

As the school year begins, teachers often devote significant time to planning lessons, reviewing standards, organizing classrooms, and establishing routines. While these are all essential components of effective teaching, they cannot replace the importance of creating a positive classroom culture.

Source ↗
behavior Thu, 13 Aug 2026 09:15:00 +0000
Getting Smart

Voiceless Collegiality: When Getting Along Gets in the Way of School Leadership

When a leadership team meeting ends with everyone offering to help but no one agreeing to lead, something important has gone wrong beneath the surface. In this sharp, diagnostic piece, education leadership professor Andy Szeto names a pattern hiding in plain sight: voiceless collegiality, the organizational silence that masquerades as team harmony. Szeto identifies three distinct types of silence eroding school effectiveness and offers concrete shifts for leadership teams ready to trade comfort for clarity. The post Voiceless Collegiality: When Getting Along Gets in the Way of School Leadership appeared first on Getting Smart .

Source ↗
technology Thu, 13 Aug 2026 09:00:00 +0000
Tech & Learning

Assessment After AI: Designing Student Work to Show Real Thinking

Assessment after AI is not about fear, but rather about intentional and purposeful design

Source ↗
audience Thu, 13 Aug 2026 07:00:00 +0000
Inside Higher Ed

McMahon’s Trick Questions

McMahon’s Trick Questions Elizabeth Redden Thu, 08/13/2026 - 03:00 AM Universities risk walking right into a trap. Byline(s) Richard Amesbury

Source ↗
audience Thu, 13 Aug 2026 07:00:00 +0000
Inside Higher Ed

The NCAA Can’t Outrun Its Antitrust Problem

The NCAA Can’t Outrun Its Antitrust Problem sara.custer@in… Thu, 08/13/2026 - 03:00 AM A stalled bill, a growing pile of lawsuits and no clear path for the athletes at the center of it all. Byline(s) Sara Custer

Source ↗
audience Thu, 13 Aug 2026 07:00:00 +0000
Inside Higher Ed

How 2 Texas Lawsuits Could Tee Up a National Fight Over Academic Freedom

How 2 Texas Lawsuits Could Tee Up a National Fight Over Academic Freedom Emma Whitford Thu, 08/13/2026 - 03:00 AM The AAUP’s suits against Texas A&M and Texas Tech could bring questions about academic freedom to the “unpredictable” Fifth Circuit, which has yet to weigh in on the issue. Byline(s) Emma Whitford

Source ↗
audience Thu, 13 Aug 2026 07:00:00 +0000
Inside Higher Ed

Survey Shows Students Trust College Leaders More Than Politicians

Survey Shows Students Trust College Leaders More Than Politicians Olivia.sanchez Thu, 08/13/2026 - 03:00 AM Students from all political parties reported higher approval of campus-led policies, compared to state and federal policies, according to data from Gallup and Lumina Foundation. Byline(s) Olivia Sanchez

Source ↗
audience Thu, 13 Aug 2026 07:00:00 +0000
Inside Higher Ed

College Wasn’t Built for Student Parents

College Wasn’t Built for Student Parents Joshua.Bay Thu, 08/13/2026 - 03:00 AM In this week’s Voices of Student Success episode, Generation Hope’s Nicole Lynn Lewis explores how colleges can better support student parents and caregivers. Byline(s) Joshua Bay

Source ↗
audience Thu, 13 Aug 2026 07:00:00 +0000
Inside Higher Ed

Senators Prod State Department to Process Student, Scholar Visas

Senators Prod State Department to Process Student, Scholar Visas sara.custer@in… Thu, 08/13/2026 - 03:00 AM Byline(s) Sara Custer

Source ↗
audience Thu, 13 Aug 2026 07:00:00 +0000
Inside Higher Ed

The NSF’s Ph.D.-to-Industry Pipeline Push

The NSF’s Ph.D.-to-Industry Pipeline Push kathryn.palmer… Thu, 08/13/2026 - 03:00 AM The majority of STEM Ph.D.s take jobs outside academia. While some universities help them prepare, the National Science Foundation is now investing millions to boost those efforts. Byline(s) Kathryn Palmer

Source ↗
audience Thu, 13 Aug 2026 07:00:00 +0000
Inside Higher Ed

S.C. Colleges Could Face Financial Penalties Over New Bathroom Law

S.C. Colleges Could Face Financial Penalties Over New Bathroom Law Katherine Knott Thu, 08/13/2026 - 03:00 AM Byline(s) Katherine Knott

Source ↗
audience Thu, 13 Aug 2026 07:00:00 +0000
Inside Higher Ed

ED Says It Plans to Spend Millions in Expiring Research Dollars. Concerns Remain.

ED Says It Plans to Spend Millions in Expiring Research Dollars. Concerns Remain. jessica.blake@… Thu, 08/13/2026 - 03:00 AM Higher education advocates say delays in education research funding have already caused severe damage. Byline(s) Jessica Blake

Source ↗
audience Thu, 13 Aug 2026 05:00:00 -0400
Higher Ed Dive

Unions sue Education Department over ‘professional’ degree rule

Under a new rule, the designation allows students to take out $200,000 in federal student loans versus $100,000 for other graduate programs.

Source ↗
audience Thu, 13 Aug 2026 05:00:00 -0400
Higher Ed Dive

What do college students think of their institutional leaders?

Students at top-ranked colleges were less likely than others to say campus leaders mostly act in their interests, per a Gallup and Lumina Foundation poll.

Source ↗
regulation Thu, 13 Aug 2026 05:00:00 -0400
K-12 Dive

With disparate impact out, what’s next for systemic discrimination cases?

The Trump administration’s move against unintentional discrimination will likely narrow or close investigations, education civil rights experts say.

Source ↗
regulation Thu, 13 Aug 2026 05:00:00 -0400
K-12 Dive

Dyscalculia affects almost as many students as dyslexia. What can districts do?

Most districts don't have the systems in place to screen for the learning disability related to numbers and math, CRPE said in a report.

Source ↗
behavior Thu, 13 Aug 2026 00:00:00 GMT
EdSurge

On AI Policy, Students Have Plenty to Say

At a gathering in Boston, ‘student senators’ proposed a first-of-its-kind national AI policy for K-12 classrooms.

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CL

RedditPersona: A Modular Framework for Community-Conditioned LLM Adaptation from Reddit

arXiv:2606.06027v2 Announce Type: replace-cross Abstract: Community-conditioned language model adaptation needs choices about data collection, community definition, and evaluation that are currently made independently in each study, making it hard to compare assumptions or reuse artifacts. We present RedditPersona, a modular framework that standardizes these choices: it collects Reddit posts and comments, profiles active users, partitions them under five grouping strategies (subreddit-based, graph-structural, semantic, hybrid, and interaction-based), trains a parameter-efficient adapter per strategy via QLoRA, and evaluates them under a shared metric suite spanning fluency, fidelity, distributional alignment, and community identifiability. Applied to 112 subreddits in the urban well-being domain (301,429 user profiles, 16M+ comments), we find that adapters' behavioral identifiability tracks each strategy's agreement with the subreddit baseline, and that a consistent trade-off between i

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CL

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning

arXiv:2606.04923v2 Announce Type: replace-cross Abstract: Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. However, policy models may exploit latent biases in the judge, leading to reward hacking and ineffective or unsafe training outcomes. In real-world rubric-based RL, such hacking behaviors are often subtle and entangled with multiple judge biases, making them difficult to analyze, detect, and mitigate. In this paper, we introduce CHERRL, a Controllable Hacking Environment for Rubric-based RL. By injecting known biases into LaaJ, CHERRL enables stable reproduction of reward hacking, explicit observation of reward divergence, and identification of hacking onset. This provides a clean experimental testbed for studying the mechanisms and mitigations of reward hacking in rubric-based RL. To demonstrate its utility, we analyze different judge biases from the perspectives of discoverability and exploitability, a

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CL

Moxia: A Trust-First Neuro-Symbolic Execution Architecture for Self-Explaining Mathematical Reasoning

arXiv:2606.00671v3 Announce Type: replace-cross Abstract: We present Moxia (formerly AXIOM), a trust-first neuro-symbolic architecture for self-explaining mathematical reasoning over natural-language input. Its language model is strictly a canonicalizer: it rewrites informal problem text into a narrow schema consumed by a deterministic Computer-Algebra-System (CAS) pipeline, which derives and verifies the answer or abstains as a first-class output. Routing follows a 1:1:1 alignment of problem-shape regex, schema-specific prompt, and closed-form CAS handler, with 4,783 routes shipped, 71% of which answer without invoking the language model, and zero LOST_CORRECT regressions as a standing release gate. Because the answer is derived rather than generated, so is its explanation: every handler emits a step trace of the computation it performed, rendered as prose by a layer covering all 4,785 task files that cannot narrate a step the handler did not take. Derivations export to Lean 4 as well

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CL

Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models

arXiv:2605.27101v2 Announce Type: replace-cross Abstract: A key capability for video understanding is reliably linking subjects to events across time, yet whether Video Large Language Models (VideoLLMs) actually achieve this remains unclear. In this work, we introduce DistractionBench to evaluate whether VideoLLMs can robustly link subjects and events in the presence of unrelated video segments. Through controlled interventions, such as inserting short advertisement clips into longer videos, we show that VideoLLMs frequently hallucinate interactions between entities from different segments, incorrectly attributing actions from injected advertisements to subjects in the main video. We characterize this systematic hallucination as bag-of-events (BoE) behavior, where models process videos as collections of events rather than temporally structured sequences. Evaluating 11 popular VideoLLMs, we find that all models exhibit substantial BoE behavior. Our findings suggest that VideoLLMs lack r

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CL

Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models

arXiv:2605.16409v3 Announce Type: replace-cross Abstract: Optical character recognition (OCR) and multilingual scene-text understanding remain challenging for multimodal large language models (MLLMs), particularly in real-world images containing small or degraded text, cluttered layouts, occlusion, handwriting, and complex typography. We present an OCR-aware multilingual post-training framework that improves visual-text grounding in a general-purpose MLLM without requiring an external OCR engine, OCR-extracted text, or text bounding boxes at inference time. The framework combines large-scale multilingual OCR supervision, approximately 5M additional multilingual training samples, controlled synthetic OCR generation and in-image text translation, LoRA-based supervised fine-tuning (SFT), and lightweight OCR-oriented Chain-of-Thought prompting. On a held-out real-world multilingual OCR benchmark, OCR-SFT improves OCR completeness from 71.3 to 84.6, reduces hallucination rate from 18.3\% to

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CL

SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives

arXiv:2509.13450v3 Announce Type: replace-cross Abstract: We introduce SteeringSafety, a benchmark for evaluating representation steering methods across nine safety perspectives spanning 18 datasets. While prior work highlights the general capabilities of representation steering, we focus on safety perspectives including refusal, bias, hallucination, social behaviors, reasoning, epistemic integrity, and normative judgment. SteeringSafety provides modularized building blocks for state-of-the-art steering methods, enabling unified implementation of DIM, ACE, CAA, PCA, and LAT with recent enhancements such as conditional steering. Results on Gemma-2-2B, Llama-3.1-8B, and Qwen-2.5-7B show that strong steering performance depends on the pairing of method, model, and specific perspective. For instance, DIM is consistently effective, yet all methods exhibit substantial entanglement, where improving effectiveness on one safety perspective often significantly changes performance on others. Soci

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CL

ReXrank: A Public Leaderboard for AI-Powered Radiology Report Generation

arXiv:2411.15122v2 Announce Type: replace-cross Abstract: AI-driven models have demonstrated significant potential in automating radiology report generation for chest X-rays. However, there is no standardized benchmark for objectively evaluating their performance. To address this, we present ReXrank, https://rexrank.ai, a public leaderboard and challenge for assessing AI-powered radiology report generation. Our framework incorporates ReXGradient, the largest test dataset consisting of 10,000 studies, and three public datasets (MIMIC-CXR, IU-Xray, CheXpert Plus) for report generation assessment. ReXrank employs 8 evaluation metrics and separately assesses models capable of generating only findings sections and those providing both findings and impressions sections. By providing this standardized evaluation framework, ReXrank enables meaningful comparisons of model performance and offers crucial insights into their robustness across diverse clinical settings. Beyond its current focus on

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CL

Causal Agent based on Large Language Model

arXiv:2408.06849v3 Announce Type: replace-cross Abstract: The large language model (LLM) has achieved significant success across various domains. However, the inherent complexity of causal problems and causal theory poses challenges in accurately describing them in natural language, making it difficult for LLM to comprehend and use them effectively. Causal methods are not easily conveyed through natural language, which hinders LLM's ability to apply them accurately. Additionally, causal datasets are typically tabular, while LLM excels in handling natural language data, creating a structural mismatch that impedes effective reasoning with tabular data. To address these challenges, we have equipped the LLM with causal tools within an agent framework, named the Causal Agent, enabling it to tackle causal problems. The causal agent comprises tools, memory, and reasoning modules. In the tool module, the causal agent calls Python code and uses the encapsulated causal function module to align t

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CL

MMLA: How Memory Lets the Past Shape the Future

arXiv:2606.28876v3 Announce Type: replace Abstract: Proposal. Long context can replay history, but it does not decide which completed observations deserve authority. MMLA formalizes a bounded resident memory between transient context and slow weight updates. A completed local segment is eventized; for each event, a target-conditioned constructor proposes semantic content and a trusted assembler produces a complete versioned row; deployment either commits that row atomically or returns NULL. Realized futures may price actions during training, while deployment remains causal and future-blind. Validated components. Controlled studies establish narrower ingredients. Lifecycle execution is exact on 300/300 held-out records for each of three seeds. Calibrated selection with full-archive fallback improves over a weak budget-matched dense baseline by 5.5--16.6 F1 and over BM25 by 4.0--6.2 F1 on held-out multi-hop QA; the original Llama budget execution is retained as failed, while the correcte

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CL

Self-Harness: Harnesses That Improve Themselves

arXiv:2606.09498v2 Announce Type: replace Abstract: The performance of LLM-based agents is jointly shaped by their base models and the harnesses that mediate their interaction with the environment. Because different models exhibit distinct behaviors, effective harness design is inherently model-specific. Yet agent harnesses are still largely engineered by human experts, a paradigm that scales poorly as modern LLMs become increasingly diverse and rapidly evolving. In this paper, we introduce Self-Harness, a new paradigm in which an LLM-based agent improves its own operating harness, without relying on human engineers or stronger external agents. We operationalize Self-Harness as an iterative loop with three stages: Weakness Mining, which identifies model-specific failure patterns from execution traces; Harness Proposal, which generates diverse yet minimal harness modifications tied to these failures; and Proposal Validation, which accepts candidate edits only after regression testing. W

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CL

CAR: Query-Guided Confidence-Aware Reranking for Retrieval-Augmented Generation

arXiv:2605.04495v2 Announce Type: replace Abstract: Retrieval-augmented generation (RAG) relies on evidence ranking to determine what information is exposed to the generator, yet existing retrieval and reranking methods primarily estimate query--document relevance. Relevance, however, is not equivalent to generator-side usefulness: a relevant passage may introduce ambiguity or distraction, whereas a lower-ranked passage may stabilize the generator's answer. We present CAR (Confidence-Aware Reranking), a training-free rank-correction framework that uses query-only answer stability as a control and measures each candidate by the change it induces in sampled-answer semantic stability. This controlled contrast estimates a document's marginal contribution to generator behavior without treating semantic stability as relevance or calibrated correctness. CAR converts these confidence changes into coarse precedence constraints and returns the feasible ranking with minimum Kendall distance from

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CL

DORA Explorer: Improving the Exploration Ability of LLMs Without Training

arXiv:2604.17244v2 Announce Type: replace Abstract: Large language model (LLM) agents for sequential decision-making struggle to produce diverse outputs. This leads to insufficient exploration, suboptimal solutions, and repeated actions. Actions are generated at the sequence level, but existing sampling strategies, such as temperature scaling, introduce diversity at the token level, not at the sequence level. We introduce DORA EXPLORER (Diversity-Oriented Ranking of Actions), a training-free, inference-time algorithm for improving exploration in LLM agents. DORA generates multiple candidate actions, scores them using sequence-level log-probability statistics, and samples an action via a tunable exploration parameter. We first study exploration in the classic Multi-Armed Bandit setting, where DORA substantially outperforms temperature-based sampling. Our main evaluation is on the Text Adventure Learning Environment Suite (TALES), where prompting strategies fail to explore but DORA deliv

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CL

TEMPER: Testing Emotional Perturbation in Quantitative Reasoning

arXiv:2604.07801v2 Announce Type: replace Abstract: Large language models are trained and evaluated on quantitative reasoning tasks written in clean, emotionally neutral language. However, real-world queries are often wrapped in frustration, urgency or enthusiasm. Does emotional framing alone degrade reasoning when all numerical content is preserved? To investigate this, a controlled emotion translation framework is developed that rewrites problems into emotional variants while preserving all quantities and relationships. Using this framework, Temper-5400 (5,400 semantically verified emotion-neutral pairs) is constructed across GSM8K, MultiArith, and ARC-Challenge, and evaluated on eighteen models (1B to frontier scale). Two core results emerge: First, emotional framing reduces accuracy by 2-10 percentage points even though all numerical content is preserved. Second, neutralizing emotional variants recovers most of the lost performance, showing both that the degradation is tied to emot

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CL

Uncertainty as a Planning Signal: Multi-Turn Decision Making for Goal-Oriented Conversation

arXiv:2604.03924v2 Announce Type: replace Abstract: Goal-oriented conversational systems require making sequential decisions under uncertainty about the user's intent, where the algorithm must balance information acquisition and target commitment over multiple turns. Existing approaches address this challenge from different perspectives: structured methods enable multi-step planning but rely on predefined schemas, while LLM-based approaches support flexible interactions but lack long-horizon decision making, resulting in poor coordination between information acquisition and target commitment. To address this limitation, we formulate goal-oriented conversation as an uncertainty-aware sequential decision problem, where uncertainty serves as a guiding signal for multi-turn decision making. We propose a Conversation Uncertainty-aware Planning framework (CUP) that integrates language models with structured planning: a language model proposes feasible actions, and a planner evaluates their l

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CL

Social Meaning in Large Language Models: Structure, Magnitude, and Pragmatic Prompting

arXiv:2604.02512v2 Announce Type: replace Abstract: Large language models (LLMs) increasingly exhibit human-like patterns of pragmatic and social reasoning. This paper addresses two related questions: do LLMs approximate human social meaning not only qualitatively but also quantitatively, and can prompting strategies informed by pragmatic theory improve this approximation? To address the first, we introduce two calibration-focused metrics distinguishing structural fidelity from magnitude calibration: the Effect Size Ratio (ESR) and the Calibration Deviation Score (CDS). To address the second, we derive prompting conditions from two pragmatic assumptions: that social meaning arises from reasoning over linguistic alternatives, and that listeners infer speaker knowledge states and communicative motives. Applied to a case study on numerical (im)precision across three frontier LLMs, we find that all models reliably reproduce the qualitative structure of human social inferences but differ su

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CL

LLM Router: Rethinking Routing with Prefill Activations

arXiv:2603.20895v3 Announce Type: replace Abstract: Existing routers rely on semantic query features or handcrafted features, which often fail to capture model-specific failures or intrinsic task difficulty. We instead route using internal LLM activations, specifically the residual stream. Our key idea, Encoder-Target Decoupling, separates the model that produces the predictive signal (the Encoder) from the model whose correctness is being estimated (the Target), allowing open-weight encoders to predict the performance of closed-source target models. We evaluate layerwise geometric probes, finding that Fisher Separability ($J$) effectively identifies informative layers, supported by Effective Dimensionality ($d_{\mathrm{eff}}$) diagnostics. We then utilize a SharedTrunkNet, a joint multi-output MLP that predicts simultaneous correctness probabilities across candidate models using concatenated prefill features. In our experiments, SharedTrunkNet consistently outperforms semantic baselin

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CL

Large Language Models Reproduce Racial Stereotypes When Used for Text Annotation

arXiv:2603.13891v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used for automated text annotation in tasks ranging from academic research to content moderation and hiring. Across 19 LLMs and two experiments totaling more than 4 million annotation judgments, we show that subtle identity cues embedded in text systematically bias annotation outcomes in ways that mirror racial stereotypes. In a names-based experiment spanning 39 annotation tasks, texts containing names associated with Black individuals are rated as more aggressive by 18 of 19 models and more gossipy by 18 of 19. Asian names produce a bamboo-ceiling profile: 17 of 19 models rate individuals as more intelligent, while 18 of 19 rate them as less confident and less sociable. Arab names elicit cognitive elevation alongside interpersonal devaluation, and all four minority groups are consistently rated as less self-disciplined. In a matched dialect experiment, the same sentence is judged signifi

Source ↗
technology Thu, 13 Aug 2026 00:00:00 -0400
arXiv cs.CL

LLM-Powered Automatic Translation and Urgency in Crisis Scenarios

arXiv:2602.13452v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly proposed for crisis preparedness and response, particularly for multilingual communication. However, their suitability for high-stakes crisis contexts remains insufficiently evaluated. This work examines the performance of state-of-the-art LLMs and machine translation systems in crisis-domain translation, with a focus on preserving urgency, a critical property for effective crisis communication and triage. Using multilingual crisis data (TICO-19, 30 languages) and a newly introduced urgency-annotated dataset of 100 scenarios translated into 29 languages, we show that dedicated translation models and LLMs exhibit substantial quality degradation, particularly for low-resource languages. Beyond translation quality, we conduct a human annotation study revealing a striking asymmetry: human assessors maintain consistent urgency judgments regardless of prompt language, while LLM-based urgency cla

Source ↗
Showing 7601–7650 of 18624 signals
← Prev Page 153 of 373 Next →