EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18402 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Tue, 12 May 2026 09:00:00 +0000
Tech & Learning

The Southern Surge Proves Science of Reading Works. Why Aren't More Districts Listening?

Conversations with Kevin Hogan: Karl Rectanus brings his edtech evidence background to the nation's original science of reading organization — and is betting on outcomes-based contracting to close the literacy gap.

Source ↗
behavior Tue, 11 Nov 2025 10:01:00 +0000
eSchool News

10 reasons to upgrade to Windows 11 ASAP

K-12 IT leaders are under pressure from all sides--rising cyberattacks, the end of Windows 10 support, and the need for powerful new learning tools.

Source ↗
behavior Tue, 11 Nov 2025 10:00:00 +0000
eSchool News

From explaining to empowering: How tutors can support student thinking

Consider the work of a personal trainer. They can explain and model a workout perfectly, but if the athlete isn’t the one doing the lifting, their muscles won’t grow. The same is true for student learning.

Source ↗
technology Tue, 11 Aug 2026 23:14:49 +0000
MedCity News

NOCD Announces Plans for National Expansion of Intensive Outpatient Therapy

NOCD, a virtual provider for OCD, announced plans to expand its virtual intensive outpatient program for people with severe OCD nationwide. The post NOCD Announces Plans for National Expansion of Intensive Outpatient Therapy appeared first on MedCity News .

Source ↗
technology Tue, 11 Aug 2026 21:26:28 +0000
MedCity News

Infinimmune Secures $75M for a More Human Approach to Developing Better Antibody Drugs

Immunology startup Infinimmune is developing antibody drugs with potential advantages over traditional monoclonal antibodies. The Series A financing will support two lead programs entering the clinic in atopic dermatitis. The post Infinimmune Secures $75M for a More Human Approach to Developing Better Antibody Drugs appeared first on MedCity News .

Source ↗
regulation Tue, 11 Aug 2026 20:23:30 +0000
The 74

Trump Signs Order on Childhood Vaccines Against Medical Groups’ Guidance

WASHINGTON (AP) — President Donald Trump on Monday signed an executive order calling for revamped childhood vaccine recommendations that promote his long-held but discredited theory that childhood shots should be spaced out into separate medical visits. The order advocates separating the measles, mumps and rubella (MMR) vaccine into three different single-disease shots and administering all childhood immunizations at separate appointments whenever […]

Source ↗
regulation Tue, 11 Aug 2026 17:40:00 -0400
K-12 Dive

As districts navigate screen time, CoSN advises against time limits

The association does not believe setting time limits on technology use in schools is helpful in measuring quality screen time, its board chair said.

Source ↗
audience Tue, 11 Aug 2026 17:33:00 -0400
Higher Ed Dive

Antitrust lawsuit against 32 colleges with early decision can proceed, judge rules

However, U.S. District Judge Angel Kelley dismissed the claims against Common App and other noncollege defendants.

Source ↗
regulation Tue, 11 Aug 2026 16:30:00 +0000
The 74

Opinion: As Students Head Back to School, Let’s Rethink What’s Success in Sex Education

As millions of young people head back to school this fall, parents are buying backpacks, teachers are preparing lesson plans, and policymakers are once again debating what belongs in the classroom. One topic will predictably return to the center of those debates: sex education. Too often, the conversation begins and ends with one question: How […]

Source ↗
behavior Tue, 11 Aug 2026 15:13:39 +0000
MindShift (KQED)

The School Year Begins in LA with One Big Change: Limited Screen Time

As educators in the nation's second-largest district prepare for the new limits on classroom screen time, some teachers say it will be a difficult adjustment.

Source ↗
technology Tue, 11 Aug 2026 14:49:46 -0400
EdTech Mag (K-12)

K–12 Schools are Shifting To a Longer-Term Cyber Resilience Strategy

Data from a recent Comparitech report shows a 9% decline in the number of confirmed ransomware attacks on U.S. educational institutions from 2024 to 2025. While incidents may not be rising as fast as they once were, and though average ransom amounts are down 33%, the overall rate of attacks and the severity of breaches is still high. The bottom line is that cyberattacks of any kind aren't going anywhere, and that’s leading schools to shift from emergency response mode toward a longer-term cyber resilience strategy. Brock Boggs, IT director of Cityscape Schools in Dallas, Texas, says that his…

Source ↗
regulation Tue, 11 Aug 2026 14:30:00 +0000
The 74

A Bad Jobs Report for Public Education

A version of this analysis originally appeared in the “Aldeman on Education” Substack. According to the latest jobs report from the Bureau of Labor Statistics, K-12 schools lost 50,000 workers from June to July. The numbers for May were also revised down by another 47,000. These are seasonally adjusted numbers. Each year, public schools “lose” […]

Source ↗
technology Tue, 11 Aug 2026 13:23:00 +0000
MedCity News

From Friction to Fix: Measuring What’s Breaking Payer-Provider Trust

Tolerance for this abrasion is gone. It wastes time, drains staff capacity, drives up costs, and affects the patient experience. The post From Friction to Fix: Measuring What’s Breaking Payer-Provider Trust appeared first on MedCity News .

Source ↗
technology Tue, 11 Aug 2026 12:50:00 +0000
MedCity News

Which Health Tech Companies Are Getting Customer Experience Right?

[Sponsored] A new report from Forrester offers a score card assessing health tech vendors on customer experience for healthcare organizations considering adopting them. The post Which Health Tech Companies Are Getting Customer Experience Right? appeared first on MedCity News .

Source ↗
regulation Tue, 11 Aug 2026 12:30:00 +0000
The 74

NYC Reading Scores Sink, With Third Grade Proficiency Falling 14 points

This story was originally published by Chalkbeat. Sign up for their newsletters at ckbe.at/newsletters Reading scores sank in New York City last school year, with a nearly 14 percentage point decline among third graders, according to preliminary data released Friday. Among students in grades 3-8, about half were proficient in reading, a decline of 6 […]

Source ↗
behavior Tue, 11 Aug 2026 11:50:05 +0000
District Admin

Judge rules Virginia school board discriminated against Black students by restoring Confederate names

The Shenandoah County School Board must rename two schools after a federal court found its 2024 decision violated students' rights. The post Judge rules Virginia school board discriminated against Black students by restoring Confederate names appeared first on District Administration .

Source ↗
behavior Tue, 11 Aug 2026 11:47:04 +0000
District Admin

The school year begins in LA with one big change

This year brings a new challenge for Armaghan Khan and many other teachers in the nation's second-largest school district: Ban screen time for younger students, and significantly limit it for older kids. The post The school year begins in LA with one big change appeared first on District Administration .

Source ↗
regulation Tue, 11 Aug 2026 10:30:00 +0000
The 74

More Dads Are Taking Paid Paternity Leave

Updated A.J. Johnson’s oldest son was born in 2019, the year before residents of his home state of Washington could receive benefits from its newly passed paid family leave program. Johnson, a full-time firefighter, had to save up accrued sick leave to take time off around his son’s birth. He took six weeks off to […]

Source ↗
behavior Tue, 11 Aug 2026 10:00:34 +0000
MindShift (KQED)

The Lasting Legacy of Rita Pierson’s ‘Every Kid Needs a Champion’

More than a decade after Rita Pierson stepped onto the TED stage and declared, “Every child deserves a champion,” her words still echo through schools, teacher prep programs and staff meetings across the country.

Source ↗
behavior Tue, 11 Aug 2026 10:00:00 +0000
eSchool News

How to build a positive classroom in the first 5 days of school

There’s a wise quote that says, “Kids don’t care about your expectations until they know you care about them. Connection is the foundation everything else stands on.”

Source ↗
behavior Tue, 11 Aug 2026 09:15:00 +0000
Getting Smart

Steward Stories: CommunityShare Proves That the Community Is the Curriculum

What if every community member, from a landscape architect to a university researcher to a sitting senator, could become part of a student's education? CommunityShare has been quietly building that reality since 2015, connecting more than 85,000 students and their educators with community partners across 12 states through a model it calls a human library. This Steward Stories profile traces the vision, the timeline, and the hard-won lessons behind one of the most compelling ecosystem intermediaries in the country. Education leaders who are wrestling with relevance, learner agency, and community trust will find both inspiration and a practical framework here. The post Steward Stories: CommunityShare Proves That the Community Is the Curriculum appeared first on Getting Smart .

Source ↗
technology Tue, 11 Aug 2026 09:00:00 +0000
Tech & Learning

View From The Top: How Edtech CEOs Lead Through Market Uncertainty

Tammy Wincup, CEO of Securly, discusses the state of the edtech market and how edtech CEOs can effectively lead in the current climate.

Source ↗
audience Tue, 11 Aug 2026 07:00:00 +0000
Inside Higher Ed

Homer Is Having a Resurgence—but You Wouldn’t Know It in Higher Ed

Homer Is Having a Resurgence—but You Wouldn’t Know It in Higher Ed Elizabeth Redden Tue, 08/11/2026 - 03:00 AM Popular demand for The Odyssey couldn’t be higher, even as classics departments find themselves targeted for cuts. Byline(s) Richard A. Greenwald

Source ↗
audience Tue, 11 Aug 2026 07:00:00 +0000
Inside Higher Ed

3 AI-Related Questions for Notre Dame’s Sonia Howell

3 AI-Related Questions for Notre Dame’s Sonia Howell joshua.m.kim@d… Tue, 08/11/2026 - 03:00 AM An artificial intelligence and digital learning conversation with UND’s director of the Office of Digital Learning. Byline(s) Joshua Kim

Source ↗
audience Tue, 11 Aug 2026 07:00:00 +0000
Inside Higher Ed

Transfer Without the Surprises, Part 2

Transfer Without the Surprises, Part 2 quintina.barne… Tue, 08/11/2026 - 03:00 AM What Michigan students should be able to count on when they transfer. Byline(s) Sarah Szurpicki Amy Reddinger Darryl Gardner Mariah Orzolek

Source ↗
audience Tue, 11 Aug 2026 07:00:00 +0000
Inside Higher Ed

Professor Sues Senator Over Alleged Role in His Firing

Professor Sues Senator Over Alleged Role in His Firing Emma Whitford Tue, 08/11/2026 - 03:00 AM The now-reinstated Austin Peay State University theater professor won a settlement from the university in January. Byline(s) Emma Whitford

Source ↗
audience Tue, 11 Aug 2026 07:00:00 +0000
Inside Higher Ed

Deferred Maintenance Backlog Is ‘Ticking Time Bomb’ for Higher Ed

Deferred Maintenance Backlog Is ‘Ticking Time Bomb’ for Higher Ed Ryan Quinn Tue, 08/11/2026 - 03:00 AM Legislatures have different approaches for addressing—or not addressing—billions of dollars in needed repairs, according to new research. Byline(s) Ryan Quinn

Source ↗
audience Tue, 11 Aug 2026 07:00:00 +0000
Inside Higher Ed

DOJ Sues 3 More States Over In-State Tuition for Undocumented Students

DOJ Sues 3 More States Over In-State Tuition for Undocumented Students Sara Weissman Tue, 08/11/2026 - 03:00 AM Byline(s) Sara Weissman

Source ↗
audience Tue, 11 Aug 2026 07:00:00 +0000
Inside Higher Ed

Students in Need More Likely to Miss Mental Health Support

Students in Need More Likely to Miss Mental Health Support Joshua.Bay Tue, 08/11/2026 - 03:00 AM A new Trellis Strategies analysis finds financially vulnerable students face greater mental health challenges but are less aware of campus supports. Byline(s) Joshua Bay

Source ↗
audience Tue, 11 Aug 2026 07:00:00 +0000
Inside Higher Ed

Native Scholarship Group to Create Endowment With MacKenzie Scott Gift

Native Scholarship Group to Create Endowment With MacKenzie Scott Gift Sara Weissman Tue, 08/11/2026 - 03:00 AM Byline(s) Olivia Sanchez

Source ↗
audience Tue, 11 Aug 2026 07:00:00 +0000
Inside Higher Ed

Trump Admin Invests $180M in Mining Colleges

Trump Admin Invests $180M in Mining Colleges jessica.blake@… Tue, 08/11/2026 - 03:00 AM Byline(s) Jessica Blake

Source ↗
audience Tue, 11 Aug 2026 07:00:00 +0000
Inside Higher Ed

When Funding Gets Scarce, Savvy Scientists Get … OnlyFans?

When Funding Gets Scarce, Savvy Scientists Get … OnlyFans? kathryn.palmer… Tue, 08/11/2026 - 03:00 AM As federal funding cuts threaten scientific research, a group of researchers is following in the footsteps of a fictional single mom who launched an OnlyFans page to make ends meet. Byline(s) Kathryn Palmer

Source ↗
audience Tue, 11 Aug 2026 05:00:00 -0400
Higher Ed Dive

Most 4-year graduates work in same state as their alma mater, report finds

States looking to grow their college-educated workforce should focus on "homegrown talent” over wooing outside workers, according to the analysis.

Source ↗
regulation Tue, 11 Aug 2026 05:00:00 -0400
K-12 Dive

Students miss over 200M more school days annually since COVID

If a student is chronically absent every year, they miss up to a full instructional year by the time they graduate high school, a Bellwether analysis found.

Source ↗
regulation Tue, 11 Aug 2026 05:00:00 -0400
K-12 Dive

Special education state complaints jump nearly 50%

The trend may be due to reduced capacity at the Education Department’s Office for Civil Rights, among other challenges, CEC and NASDSE say.

Source ↗
regulation Tue, 11 Aug 2026 02:30:00 +0000
The 74

‘What Did We Do Wrong Now?’: Rural Iowa Schools Weigh Impact of ESA Expansion

As Iowa’s Education Savings Account program continues to expand to include all K-12 families, rural public school districts are still trying to determine how the scholarships will affect their classrooms, budgets and enrollment. Supporters argue ESAs give families more educational choices, while critics say the program diverts money away from public schools that remain the […]

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CL

Which Decisions Low-Bit Quantization Breaks, and How to Predict Them

arXiv:2608.06564v2 Announce Type: replace-cross Abstract: Quantization is how large language models are actually deployed, and below four bits it hurts. What nobody can say is which decisions change at a given bit-width -- which matters most where a model acts rather than answers, since a tool call it declines to make is a failure no score reports. A compressed agent stops calling its tools, then loses half its safety refusals, while benchmark scores barely move. Prior work assumes the added noise has a roughly fixed size, which would make confident decisions safe. We measure the decision instead: the margin between the option a model picks and its best alternative, before and after quantization, across 16 models, three methods, and 8 down to 2 bits. Kinds of decision do not break together -- at 3 bits the decision to call a tool collapses toward inaction while the choice of which tool is untouched -- and the damage is proportional rather than fixed, the margin multiplied by a factor t

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CL

LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation

arXiv:2608.00267v2 Announce Type: replace-cross Abstract: Coding agent infrastructure is shifting from harness engineering toward loop engineering as coding agents are deployed for sustained long-horizon software development. Existing benchmarks often center on localized tasks or end-state outcomes, offering limited insight into sustained execution. We introduce LOOPSBENCH, a long-horizon benchmark for loop engineering in coding agent evaluation. Each task is a dependency DAG over separately testable development units with source-evidenced prerequisite edges. LOOPSBENCH comprises 112 tasks from authentic sources spanning 8 programming languages and 9 domains. Its flow-aware runtime releases tests along the ready frontier and retains completed nodes as regression obligations. We evaluate frontier coding agents paired with widely used loop implementations. The strongest configuration, Opus-4.7 with Claude Code and outer continuation, resolves 25.00% of tasks. Recorded plans recover only

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CL

Decodable but Not Faithful: Coupling Natural-Language Rationales to Programmatic Verifiers

arXiv:2606.21678v2 Announce Type: replace-cross Abstract: Language models can generate plausible rationales for their predictions, but these explanations may not faithfully represent the model's internal reasoning. We propose verifier-coupled reasoning, a framework that inserts inline claims into reasoning traces and trains an auxiliary consistency head to predict programmatic verifier outputs from rationale-span hidden states. The central finding is a gap between decodability and faithfulness: consistency training reliably makes verifier information decodable from rationale representations, but decodability does not guarantee faithful generation. In LeanCheck (formal theorem proving), rationale-only and proof-only pooling achieve perfect directional separation under counterfactual conflict. In KataGo (Go engine), commentary spans encode 10-way win-rate buckets at 81% accuracy. Yet in a code setting, the model achieves 98.6% coupling while its generated explanations remain unfaithful:

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CL

ATLAS: Agentic Taxonomy of Large-Scale Software Ecosystems

arXiv:2606.21597v2 Announce Type: replace-cross Abstract: The open-source ecosystem on GitHub lacks a systematic hierarchical taxonomy of software repositories. GitHub Topics, the dominant organizational mechanism, is flat, inconsistent, and covers only 67% of projects. We present ATLAS, the first framework that automatically constructs a hierarchical taxonomy for software repositories and classifies projects into it end-to-end. By combining LLM global knowledge with real repository distributions, ATLAS proposes meaningful splitting dimensions and iteratively corrects those that fail to accommodate real projects. A Designer Agent proposes splitting dimensions while a Classifier Agent assigns repositories; a self-corrective refinement loop uses classification failures to drive dimension revision through escalating strategies. We evaluate ATLAS on 54,387 GitHub repositories against six baselines spanning four paradigms, two downstream tasks, and three model families. On a stratified 2,00

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CL

The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary

arXiv:2606.00376v2 Announce Type: replace-cross Abstract: Extended chain-of-thought reasoning can degrade performance on deterministic state-tracking tasks, not solely because of preference biases but, on the evidence we present, because of information-theoretic limits in the capacity of decoder-only attention. We present: (1) an Attention Bottleneck analysis providing evidence that total state-tracking capacity in bits is bounded in terms of head count, head dimension, and context length under stated modeling assumptions, and that total capacity is not the binding constraint; (2) a context-dependent error model with a depth-dependent quadratic term in the error exponent; (3) the State-Space Jaccard metric measuring state drift; and (4) a Deterministic Horizon $d^* \in [19, 31]$ (at $\alpha = 0.5$) marking the depth at which unaided accuracy crosses 50%. Across twelve models and eight task domains (including SWE-Bench, WebArena, and SQL-Multi), tool-integrated reasoning reaches 76-94%

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CL

DeltaPrompts: Escaping the Zero-Delta Trap in Multimodal Distillation

arXiv:2605.15532v3 Announce Type: replace-cross Abstract: Distillation enables compact Vision-Language Models (VLMs) to obtain strong reasoning capabilities, yet the prompts driving this process are typically chosen via simple heuristics or aggregated from off-the-shelf datasets. We reveal a critical inefficiency in this approach: up to 69% of the prompts in standard chart / document reasoning datasets are effectively zero-delta, meaning the teacher and student already induce the exact same answer distribution. Training on these prompts provides minimal learning signal, causing student improvement to rapidly saturate regardless of data scale. To escape the zero-delta trap, we return to first principles: distillation fundamentally minimizes distributional divergence, and thus a prompt is valuable only if it exposes a functional capability gap between the teacher and student. We quantify this gap through answer divergence ($\Delta$), demonstrating that non-zero divergence is critical for

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CL

Barriers to Universal Reasoning With Transformers (And How to Overcome Them)

arXiv:2604.25800v2 Announce Type: replace-cross Abstract: Chain-of-Thought (CoT) has been shown to empirically improve Transformers' performance, and theoretically increase their expressivity to Turing completeness. However, whether Transformers can learn to generalize to CoT traces longer than those seen during training is understudied. We use recent theoretical frameworks for Transformer length generalization and find that -- under standard positional encodings and a finite alphabet -- Transformers with CoT cannot solve problems beyond $TC^0$, i.e. the expressivity benefits do not hold under the stricter requirement of length-generalizable learnability. However, if we allow the vocabulary to grow with problem size, we attain a length-generalizable simulation of Turing machines where the CoT trace length is linear in the simulated runtime up to a constant. Our construction overcomes two core obstacles to reliable length generalization: repeated copying and last-occurrence retrieval. W

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CL

Process Supervision of Confidence Margin for Calibrated LLM Reasoning

arXiv:2604.23333v2 Announce Type: replace-cross Abstract: Scaling test-time computation with reinforcement learning (RL) has emerged as a reliable path to improve large language models (LLM) reasoning ability. Yet, outcome-based reward often incentivizes models to be overconfident, leading to hallucinations, unreliable confidence-based control, and unnecessary compute allocation. We introduce Reinforcement Learning with Confidence Margin (RLCM), a calibration-aware RL framework that jointly optimizes correctness and confidence reliability via a margin-enhanced process reward over intermediate-budget completions. Rather than aligning confidence to correctness likelihoods, RLCM encourages the model to widen the confidence margin between correct and incorrect steps within a single reasoning trajectory. Across mathematical, code, logic and science benchmarks, our method substantially improves calibration while maintaining or improving accuracy. We further show that, with calibrated confide

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CL

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks

arXiv:2604.10015v3 Announce Type: replace-cross Abstract: Recent studies demonstrate that tool-calling capability enables large language models (LLMs) to interact with external environments for long-horizon financial tasks. While existing benchmarks have begun evaluating financial tool calling, they focus on limited scenarios and rely on call-level metrics that fail to capture trajectory-level reasoning quality. To address this gap, we introduce FinTrace, a benchmark comprising 800 expert-annotated trajectories spanning 34 real-world financial task categories across multiple difficulty levels. FinTrace employs a rubric-based evaluation protocol with nine metrics organized along four axes -- action correctness, execution efficiency, process quality, and output quality -- enabling fine-grained assessment of LLM tool-calling behavior. Our evaluation of 13 LLMs reveals that while frontier models achieve strong tool selection, all models struggle with information utilization and final answe

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CL

SkillClaw: Let Skills Evolve Collectively with Agentic Evolver

arXiv:2604.08377v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents such as OpenClaw rely on reusable skills to perform complex tasks, yet these skills remain largely static after deployment. As a result, similar workflows, tool usage patterns, and failure modes are repeatedly rediscovered across users, preventing the system from improving with experience. While interactions from different users provide complementary signals about when a skill works or fails, existing systems lack a mechanism to convert such heterogeneous experiences into reliable skill updates. To address these issues, we present SkillClaw, a framework for collective skill evolution in multi-user agent ecosystems, which treats cross-user and over-time interactions as the primary signal for improving skills. SkillClaw continuously aggregates trajectories generated during use and processes them with an autonomous evolver, which identifies recurring behavioral patterns and translates them into upd

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CL

A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges

arXiv:2604.07650v2 Announce Type: replace-cross Abstract: The rapid growth of the large language model (LLM) ecosystem raises a critical question: are seemingly diverse models truly independent? Shared pretraining data, distillation, and alignment pipelines can induce hidden behavioral dependencies, or latent entanglement, that undermine multi-model systems such as LLM-as-a-judge pipelines and ensemble verification, which implicitly assume independent signals. In practice, this manifests as correlated reasoning patterns and synchronized failures, where apparent agreement reflects shared error modes rather than independent validation. To address this, we develop a statistical framework for auditing behavioral entanglement among black-box LLMs. Our approach introduces a multi-resolution hierarchy that characterizes the joint failure manifold through two information-theoretic metrics: (i) a Difficulty-Weighted Behavioral Entanglement Index (BEI), which amplifies synchronized failures on e

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CL

TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories

arXiv:2604.07223v2 Announce Type: replace-cross Abstract: As large language models (LLMs) evolve from static chatbots into autonomous agents, the primary vulnerability surface shifts from final outputs to intermediate execution traces. While safety guardrails are well-benchmarked for natural language responses, their efficacy remains largely unexplored within multi-step tool-use trajectories. To address this gap, we introduce **TraceSafe-Bench**, the first comprehensive benchmark specifically designed to assess mid-trajectory safety. It encompasses 12 risk categories, ranging from security threats (e.g., prompt injection, privacy leaks) to operational failures (e.g., hallucinations, interface inconsistencies), featuring over 1,000 unique execution instances. Our evaluation of 13 LLM-as-a-guard models and 7 specialized guardrails yields three critical findings: 1) *Structural Bottleneck*: Guardrail efficacy is driven more by structural data competence (e.g., JSON parsing) than semantic

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CL

Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family

arXiv:2604.05971v2 Announce Type: replace-cross Abstract: Recent research has shown that contrastive vision-language models such as CLIP often lack fine-grained understanding of visual content. While a growing body of work has sought to address this limitation, we identify a distinct failure mode in the CLIP family, which we term center bias, that persists even in recent model variants. Specifically, CLIP tends to disproportionately focus on the central region of an image, overlooking important objects located near the boundaries. This limitation is fundamental as failure to recognize relevant objects makes it difficult to perform any sophisticated tasks that depend on those objects. To understand the underlying causes of the limitation, we conduct analyses from both representation and attention perspectives. Using interpretability methods, i.e., embedding decomposition and attention map analysis, we find that relevant concepts especially those associated with off-center objects vanish

Source ↗
technology Tue, 11 Aug 2026 00:00:00 -0400
arXiv cs.CL

Large Language Models Align with the Human Brain during Creative Thinking

arXiv:2604.03480v2 Announce Type: replace-cross Abstract: Creative thinking is a fundamental aspect of human cognition, and divergent thinking-the capacity to generate novel and varied ideas-is widely regarded as its core generative engine. Large language models (LLMs) have recently demonstrated impressive performance on divergent thinking tests and prior work has shown that models with higher task performance tend to be more aligned to human brain activity. However, existing brain-LLM alignment studies have focused on passive, non-creative tasks. Here, we explore brain alignment during creative thinking using fMRI data from 170 participants performing the Alternate Uses Task (AUT). We extract representations from LLMs varying in size (270M-72B) and measure alignment to brain responses via Representational Similarity Analysis (RSA), targeting the creativity-related default mode and frontoparietal networks. We find that brain-LLM alignment is positively associated with model size (defau

Source ↗
Showing 4501–4550 of 18402 signals
← Prev Page 91 of 369 Next →