Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
Conversations with Kevin Hogan: Karl Rectanus brings his edtech evidence background to the nation's original science of reading organization — and is betting on outcomes-based contracting to close the literacy gap.
K-12 IT leaders are under pressure from all sides--rising cyberattacks, the end of Windows 10 support, and the need for powerful new learning tools.
Consider the work of a personal trainer. They can explain and model a workout perfectly, but if the athlete isn’t the one doing the lifting, their muscles won’t grow. The same is true for student learning.
NOCD, a virtual provider for OCD, announced plans to expand its virtual intensive outpatient program for people with severe OCD nationwide. The post NOCD Announces Plans for National Expansion of Intensive Outpatient Therapy appeared first on MedCity News .
Immunology startup Infinimmune is developing antibody drugs with potential advantages over traditional monoclonal antibodies. The Series A financing will support two lead programs entering the clinic in atopic dermatitis. The post Infinimmune Secures $75M for a More Human Approach to Developing Better Antibody Drugs appeared first on MedCity News .
WASHINGTON (AP) — President Donald Trump on Monday signed an executive order calling for revamped childhood vaccine recommendations that promote his long-held but discredited theory that childhood shots should be spaced out into separate medical visits. The order advocates separating the measles, mumps and rubella (MMR) vaccine into three different single-disease shots and administering all childhood immunizations at separate appointments whenever […]
The association does not believe setting time limits on technology use in schools is helpful in measuring quality screen time, its board chair said.
However, U.S. District Judge Angel Kelley dismissed the claims against Common App and other noncollege defendants.
As millions of young people head back to school this fall, parents are buying backpacks, teachers are preparing lesson plans, and policymakers are once again debating what belongs in the classroom. One topic will predictably return to the center of those debates: sex education. Too often, the conversation begins and ends with one question: How […]
As educators in the nation's second-largest district prepare for the new limits on classroom screen time, some teachers say it will be a difficult adjustment.
Data from a recent Comparitech report shows a 9% decline in the number of confirmed ransomware attacks on U.S. educational institutions from 2024 to 2025. While incidents may not be rising as fast as they once were, and though average ransom amounts are down 33%, the overall rate of attacks and the severity of breaches is still high. The bottom line is that cyberattacks of any kind aren't going anywhere, and that’s leading schools to shift from emergency response mode toward a longer-term cyber resilience strategy. Brock Boggs, IT director of Cityscape Schools in Dallas, Texas, says that his…
A version of this analysis originally appeared in the “Aldeman on Education” Substack. According to the latest jobs report from the Bureau of Labor Statistics, K-12 schools lost 50,000 workers from June to July. The numbers for May were also revised down by another 47,000. These are seasonally adjusted numbers. Each year, public schools “lose” […]
Tolerance for this abrasion is gone. It wastes time, drains staff capacity, drives up costs, and affects the patient experience. The post From Friction to Fix: Measuring What’s Breaking Payer-Provider Trust appeared first on MedCity News .
[Sponsored] A new report from Forrester offers a score card assessing health tech vendors on customer experience for healthcare organizations considering adopting them. The post Which Health Tech Companies Are Getting Customer Experience Right? appeared first on MedCity News .
This story was originally published by Chalkbeat. Sign up for their newsletters at ckbe.at/newsletters Reading scores sank in New York City last school year, with a nearly 14 percentage point decline among third graders, according to preliminary data released Friday. Among students in grades 3-8, about half were proficient in reading, a decline of 6 […]
The Shenandoah County School Board must rename two schools after a federal court found its 2024 decision violated students' rights. The post Judge rules Virginia school board discriminated against Black students by restoring Confederate names appeared first on District Administration .
This year brings a new challenge for Armaghan Khan and many other teachers in the nation's second-largest school district: Ban screen time for younger students, and significantly limit it for older kids. The post The school year begins in LA with one big change appeared first on District Administration .
Updated A.J. Johnson’s oldest son was born in 2019, the year before residents of his home state of Washington could receive benefits from its newly passed paid family leave program. Johnson, a full-time firefighter, had to save up accrued sick leave to take time off around his son’s birth. He took six weeks off to […]
More than a decade after Rita Pierson stepped onto the TED stage and declared, “Every child deserves a champion,” her words still echo through schools, teacher prep programs and staff meetings across the country.
There’s a wise quote that says, “Kids don’t care about your expectations until they know you care about them. Connection is the foundation everything else stands on.”
What if every community member, from a landscape architect to a university researcher to a sitting senator, could become part of a student's education? CommunityShare has been quietly building that reality since 2015, connecting more than 85,000 students and their educators with community partners across 12 states through a model it calls a human library. This Steward Stories profile traces the vision, the timeline, and the hard-won lessons behind one of the most compelling ecosystem intermediaries in the country. Education leaders who are wrestling with relevance, learner agency, and community trust will find both inspiration and a practical framework here. The post Steward Stories: CommunityShare Proves That the Community Is the Curriculum appeared first on Getting Smart .
Tammy Wincup, CEO of Securly, discusses the state of the edtech market and how edtech CEOs can effectively lead in the current climate.
Homer Is Having a Resurgence—but You Wouldn’t Know It in Higher Ed Elizabeth Redden Tue, 08/11/2026 - 03:00 AM Popular demand for The Odyssey couldn’t be higher, even as classics departments find themselves targeted for cuts. Byline(s) Richard A. Greenwald
3 AI-Related Questions for Notre Dame’s Sonia Howell joshua.m.kim@d… Tue, 08/11/2026 - 03:00 AM An artificial intelligence and digital learning conversation with UND’s director of the Office of Digital Learning. Byline(s) Joshua Kim
Transfer Without the Surprises, Part 2 quintina.barne… Tue, 08/11/2026 - 03:00 AM What Michigan students should be able to count on when they transfer. Byline(s) Sarah Szurpicki Amy Reddinger Darryl Gardner Mariah Orzolek
Professor Sues Senator Over Alleged Role in His Firing Emma Whitford Tue, 08/11/2026 - 03:00 AM The now-reinstated Austin Peay State University theater professor won a settlement from the university in January. Byline(s) Emma Whitford
Deferred Maintenance Backlog Is ‘Ticking Time Bomb’ for Higher Ed Ryan Quinn Tue, 08/11/2026 - 03:00 AM Legislatures have different approaches for addressing—or not addressing—billions of dollars in needed repairs, according to new research. Byline(s) Ryan Quinn
DOJ Sues 3 More States Over In-State Tuition for Undocumented Students Sara Weissman Tue, 08/11/2026 - 03:00 AM Byline(s) Sara Weissman
Students in Need More Likely to Miss Mental Health Support Joshua.Bay Tue, 08/11/2026 - 03:00 AM A new Trellis Strategies analysis finds financially vulnerable students face greater mental health challenges but are less aware of campus supports. Byline(s) Joshua Bay
Native Scholarship Group to Create Endowment With MacKenzie Scott Gift Sara Weissman Tue, 08/11/2026 - 03:00 AM Byline(s) Olivia Sanchez
Trump Admin Invests $180M in Mining Colleges jessica.blake@… Tue, 08/11/2026 - 03:00 AM Byline(s) Jessica Blake
When Funding Gets Scarce, Savvy Scientists Get … OnlyFans? kathryn.palmer… Tue, 08/11/2026 - 03:00 AM As federal funding cuts threaten scientific research, a group of researchers is following in the footsteps of a fictional single mom who launched an OnlyFans page to make ends meet. Byline(s) Kathryn Palmer
States looking to grow their college-educated workforce should focus on "homegrown talent” over wooing outside workers, according to the analysis.
If a student is chronically absent every year, they miss up to a full instructional year by the time they graduate high school, a Bellwether analysis found.
The trend may be due to reduced capacity at the Education Department’s Office for Civil Rights, among other challenges, CEC and NASDSE say.
As Iowa’s Education Savings Account program continues to expand to include all K-12 families, rural public school districts are still trying to determine how the scholarships will affect their classrooms, budgets and enrollment. Supporters argue ESAs give families more educational choices, while critics say the program diverts money away from public schools that remain the […]
arXiv:2608.06564v2 Announce Type: replace-cross Abstract: Quantization is how large language models are actually deployed, and below four bits it hurts. What nobody can say is which decisions change at a given bit-width -- which matters most where a model acts rather than answers, since a tool call it declines to make is a failure no score reports. A compressed agent stops calling its tools, then loses half its safety refusals, while benchmark scores barely move. Prior work assumes the added noise has a roughly fixed size, which would make confident decisions safe. We measure the decision instead: the margin between the option a model picks and its best alternative, before and after quantization, across 16 models, three methods, and 8 down to 2 bits. Kinds of decision do not break together -- at 3 bits the decision to call a tool collapses toward inaction while the choice of which tool is untouched -- and the damage is proportional rather than fixed, the margin multiplied by a factor t
arXiv:2608.00267v2 Announce Type: replace-cross Abstract: Coding agent infrastructure is shifting from harness engineering toward loop engineering as coding agents are deployed for sustained long-horizon software development. Existing benchmarks often center on localized tasks or end-state outcomes, offering limited insight into sustained execution. We introduce LOOPSBENCH, a long-horizon benchmark for loop engineering in coding agent evaluation. Each task is a dependency DAG over separately testable development units with source-evidenced prerequisite edges. LOOPSBENCH comprises 112 tasks from authentic sources spanning 8 programming languages and 9 domains. Its flow-aware runtime releases tests along the ready frontier and retains completed nodes as regression obligations. We evaluate frontier coding agents paired with widely used loop implementations. The strongest configuration, Opus-4.7 with Claude Code and outer continuation, resolves 25.00% of tasks. Recorded plans recover only
arXiv:2606.21678v2 Announce Type: replace-cross Abstract: Language models can generate plausible rationales for their predictions, but these explanations may not faithfully represent the model's internal reasoning. We propose verifier-coupled reasoning, a framework that inserts inline claims into reasoning traces and trains an auxiliary consistency head to predict programmatic verifier outputs from rationale-span hidden states. The central finding is a gap between decodability and faithfulness: consistency training reliably makes verifier information decodable from rationale representations, but decodability does not guarantee faithful generation. In LeanCheck (formal theorem proving), rationale-only and proof-only pooling achieve perfect directional separation under counterfactual conflict. In KataGo (Go engine), commentary spans encode 10-way win-rate buckets at 81% accuracy. Yet in a code setting, the model achieves 98.6% coupling while its generated explanations remain unfaithful:
arXiv:2606.21597v2 Announce Type: replace-cross Abstract: The open-source ecosystem on GitHub lacks a systematic hierarchical taxonomy of software repositories. GitHub Topics, the dominant organizational mechanism, is flat, inconsistent, and covers only 67% of projects. We present ATLAS, the first framework that automatically constructs a hierarchical taxonomy for software repositories and classifies projects into it end-to-end. By combining LLM global knowledge with real repository distributions, ATLAS proposes meaningful splitting dimensions and iteratively corrects those that fail to accommodate real projects. A Designer Agent proposes splitting dimensions while a Classifier Agent assigns repositories; a self-corrective refinement loop uses classification failures to drive dimension revision through escalating strategies. We evaluate ATLAS on 54,387 GitHub repositories against six baselines spanning four paradigms, two downstream tasks, and three model families. On a stratified 2,00
arXiv:2606.00376v2 Announce Type: replace-cross Abstract: Extended chain-of-thought reasoning can degrade performance on deterministic state-tracking tasks, not solely because of preference biases but, on the evidence we present, because of information-theoretic limits in the capacity of decoder-only attention. We present: (1) an Attention Bottleneck analysis providing evidence that total state-tracking capacity in bits is bounded in terms of head count, head dimension, and context length under stated modeling assumptions, and that total capacity is not the binding constraint; (2) a context-dependent error model with a depth-dependent quadratic term in the error exponent; (3) the State-Space Jaccard metric measuring state drift; and (4) a Deterministic Horizon $d^* \in [19, 31]$ (at $\alpha = 0.5$) marking the depth at which unaided accuracy crosses 50%. Across twelve models and eight task domains (including SWE-Bench, WebArena, and SQL-Multi), tool-integrated reasoning reaches 76-94%
arXiv:2605.15532v3 Announce Type: replace-cross Abstract: Distillation enables compact Vision-Language Models (VLMs) to obtain strong reasoning capabilities, yet the prompts driving this process are typically chosen via simple heuristics or aggregated from off-the-shelf datasets. We reveal a critical inefficiency in this approach: up to 69% of the prompts in standard chart / document reasoning datasets are effectively zero-delta, meaning the teacher and student already induce the exact same answer distribution. Training on these prompts provides minimal learning signal, causing student improvement to rapidly saturate regardless of data scale. To escape the zero-delta trap, we return to first principles: distillation fundamentally minimizes distributional divergence, and thus a prompt is valuable only if it exposes a functional capability gap between the teacher and student. We quantify this gap through answer divergence ($\Delta$), demonstrating that non-zero divergence is critical for
arXiv:2604.25800v2 Announce Type: replace-cross Abstract: Chain-of-Thought (CoT) has been shown to empirically improve Transformers' performance, and theoretically increase their expressivity to Turing completeness. However, whether Transformers can learn to generalize to CoT traces longer than those seen during training is understudied. We use recent theoretical frameworks for Transformer length generalization and find that -- under standard positional encodings and a finite alphabet -- Transformers with CoT cannot solve problems beyond $TC^0$, i.e. the expressivity benefits do not hold under the stricter requirement of length-generalizable learnability. However, if we allow the vocabulary to grow with problem size, we attain a length-generalizable simulation of Turing machines where the CoT trace length is linear in the simulated runtime up to a constant. Our construction overcomes two core obstacles to reliable length generalization: repeated copying and last-occurrence retrieval. W
arXiv:2604.23333v2 Announce Type: replace-cross Abstract: Scaling test-time computation with reinforcement learning (RL) has emerged as a reliable path to improve large language models (LLM) reasoning ability. Yet, outcome-based reward often incentivizes models to be overconfident, leading to hallucinations, unreliable confidence-based control, and unnecessary compute allocation. We introduce Reinforcement Learning with Confidence Margin (RLCM), a calibration-aware RL framework that jointly optimizes correctness and confidence reliability via a margin-enhanced process reward over intermediate-budget completions. Rather than aligning confidence to correctness likelihoods, RLCM encourages the model to widen the confidence margin between correct and incorrect steps within a single reasoning trajectory. Across mathematical, code, logic and science benchmarks, our method substantially improves calibration while maintaining or improving accuracy. We further show that, with calibrated confide
arXiv:2604.10015v3 Announce Type: replace-cross Abstract: Recent studies demonstrate that tool-calling capability enables large language models (LLMs) to interact with external environments for long-horizon financial tasks. While existing benchmarks have begun evaluating financial tool calling, they focus on limited scenarios and rely on call-level metrics that fail to capture trajectory-level reasoning quality. To address this gap, we introduce FinTrace, a benchmark comprising 800 expert-annotated trajectories spanning 34 real-world financial task categories across multiple difficulty levels. FinTrace employs a rubric-based evaluation protocol with nine metrics organized along four axes -- action correctness, execution efficiency, process quality, and output quality -- enabling fine-grained assessment of LLM tool-calling behavior. Our evaluation of 13 LLMs reveals that while frontier models achieve strong tool selection, all models struggle with information utilization and final answe
arXiv:2604.08377v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents such as OpenClaw rely on reusable skills to perform complex tasks, yet these skills remain largely static after deployment. As a result, similar workflows, tool usage patterns, and failure modes are repeatedly rediscovered across users, preventing the system from improving with experience. While interactions from different users provide complementary signals about when a skill works or fails, existing systems lack a mechanism to convert such heterogeneous experiences into reliable skill updates. To address these issues, we present SkillClaw, a framework for collective skill evolution in multi-user agent ecosystems, which treats cross-user and over-time interactions as the primary signal for improving skills. SkillClaw continuously aggregates trajectories generated during use and processes them with an autonomous evolver, which identifies recurring behavioral patterns and translates them into upd
arXiv:2604.07650v2 Announce Type: replace-cross Abstract: The rapid growth of the large language model (LLM) ecosystem raises a critical question: are seemingly diverse models truly independent? Shared pretraining data, distillation, and alignment pipelines can induce hidden behavioral dependencies, or latent entanglement, that undermine multi-model systems such as LLM-as-a-judge pipelines and ensemble verification, which implicitly assume independent signals. In practice, this manifests as correlated reasoning patterns and synchronized failures, where apparent agreement reflects shared error modes rather than independent validation. To address this, we develop a statistical framework for auditing behavioral entanglement among black-box LLMs. Our approach introduces a multi-resolution hierarchy that characterizes the joint failure manifold through two information-theoretic metrics: (i) a Difficulty-Weighted Behavioral Entanglement Index (BEI), which amplifies synchronized failures on e
arXiv:2604.07223v2 Announce Type: replace-cross Abstract: As large language models (LLMs) evolve from static chatbots into autonomous agents, the primary vulnerability surface shifts from final outputs to intermediate execution traces. While safety guardrails are well-benchmarked for natural language responses, their efficacy remains largely unexplored within multi-step tool-use trajectories. To address this gap, we introduce **TraceSafe-Bench**, the first comprehensive benchmark specifically designed to assess mid-trajectory safety. It encompasses 12 risk categories, ranging from security threats (e.g., prompt injection, privacy leaks) to operational failures (e.g., hallucinations, interface inconsistencies), featuring over 1,000 unique execution instances. Our evaluation of 13 LLM-as-a-guard models and 7 specialized guardrails yields three critical findings: 1) *Structural Bottleneck*: Guardrail efficacy is driven more by structural data competence (e.g., JSON parsing) than semantic
arXiv:2604.05971v2 Announce Type: replace-cross Abstract: Recent research has shown that contrastive vision-language models such as CLIP often lack fine-grained understanding of visual content. While a growing body of work has sought to address this limitation, we identify a distinct failure mode in the CLIP family, which we term center bias, that persists even in recent model variants. Specifically, CLIP tends to disproportionately focus on the central region of an image, overlooking important objects located near the boundaries. This limitation is fundamental as failure to recognize relevant objects makes it difficult to perform any sophisticated tasks that depend on those objects. To understand the underlying causes of the limitation, we conduct analyses from both representation and attention perspectives. Using interpretability methods, i.e., embedding decomposition and attention map analysis, we find that relevant concepts especially those associated with off-center objects vanish
arXiv:2604.03480v2 Announce Type: replace-cross Abstract: Creative thinking is a fundamental aspect of human cognition, and divergent thinking-the capacity to generate novel and varied ideas-is widely regarded as its core generative engine. Large language models (LLMs) have recently demonstrated impressive performance on divergent thinking tests and prior work has shown that models with higher task performance tend to be more aligned to human brain activity. However, existing brain-LLM alignment studies have focused on passive, non-creative tasks. Here, we explore brain alignment during creative thinking using fMRI data from 170 participants performing the Alternate Uses Task (AUT). We extract representations from LLMs varying in size (270M-72B) and measure alignment to brain responses via Representational Similarity Analysis (RSA), targeting the creativity-related default mode and frontoparietal networks. We find that brain-LLM alignment is positively associated with model size (defau