Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
Sangamo Therapeutics filed for Chapter 11 bankruptcy in June after its search for strategic alternatives failed to find a path forward for the company. PTC Therapeutics’ auction win brings a Sangamo gene therapy candidate for Fabry disease, a rare inherited disorder with few treatment options. The post PTC Therapeutics’ $211M Bid Wins Bankruptcy Auction for Sangamo Gene Therapy appeared first on MedCity News .
Chicago Public Schools is further away from having enough money to meet its students’ needs as defined by the state, according to new data released by the Illinois State Board of Education. Under ISBE’s calculations released Friday, CPS is now considered a little over 70% adequately funded, compared with 73% last year. CPS isn’t alone […]
Judge Dismisses Trump’s Antisemitism Lawsuit Against Harvard jessica.blake@… Thu, 08/13/2026 - 12:57 PM The decision is a blow for the administration, which has engaged in a yearslong battle with the Ivy League institution. Byline(s) Jessica Blake
The nation's largest school system will task the personnel with addressing a student behavior that affects up to 15% of youth and fuels chronic absenteeism.
Right now, 15.5 million undergraduate students around the country are getting ready to start or return to college. Seven million of them have one thing in common: relying on a Pell Grant to help pay for their higher education. But the grant program is facing a $15 billion funding shortfall that will undermine college access […]
Harmeet Dhillon, assistant attorney general of the agency’s civil rights division, said federal officials are "assessing next steps."
Indiana’s early literacy rates improved for the fifth consecutive year, with nearly 89% of Hoosier third graders demonstrating proficiency in foundational reading skills on the state’s IREAD assessment. The Indiana Department of Education released statewide IREAD, ILEARN and SAT results for the 2025-26 school year Tuesday at the State Board of Education’s August meeting. The […]
In healthcare, enterprises are moving slower than the hype cycle would lead you to believe. That said, they’re not moving slowly – they are saying yes to AI, as long as it can solve for specific workflow needs. The post Founders Should Stop Worrying If They’re “AI Enough” appeared first on MedCity News .
Higher education is at an inflection point. The 2026 EDUCAUSE Horizon Report identifies the forces most likely to reshape teaching and learning over the next decade — and the picture is complex. As artificial intelligence continues to evolve, it’s changing relationships across campus. Institutions also continue to face mounting enrollment challenges. And emerging risks related to cybersecurity, policy reform and sustainability are compounding institutional strain. AI’s Influence on Teaching and Learning According to the report, AI is redefining instructional design and how faculty teach…
K–12 districts are entrusted with a wide range of student data. Attendance records, grades, behavior logs, Individualized Education Programs and health records are gathered and stored in school systems, often for years after children move on to other districts or graduate. And as schools adopt more ed tech tools, they’re responsible for even more data. With that responsibility comes risk: A single mismanaged account can lead to an exposure that puts students’ information at risk. Strong data governance requires building smart policy; it also requires putting into place deliberate practices…
Nearly a quarter of U.S. children did not show up for school regularly in 2025, a slight decline from the previous year but still up from about 15 percent before COVID. The post Chronic absenteeism remains high six years after pandemic began appeared first on District Administration .
Predictable transportation routines can improve students' readiness to participate in classroom learning, particularly for children with special needs. The post Why transportation is such an important student success factor appeared first on District Administration .
Date & Time: Wednesday, September 09, 2026 at 2 p.m. Join HopSkipDrive and Igor Petrovic, Director of Transportation at Adams 12 Five Star Schools in Colorado, for a look at survey data that reveals a major gap between the McKinney-Vento standard and district reality, adding up to missed instructional time and real compliance risk. Attendees will also gain a district-level view of how to weigh cost, speed, and service when transportation needs can’t wait. The post The McKinney-Vento Challenge: Turning the “Without Delay” Transportation Standard into Reality appeared first on District Administration .
Across the economy, Americans are watching an artificial intelligence investment boom reshape the job market. Some of the same companies spending hundreds of billions to build the AI future are also announcing sweeping job cuts, adding to an already daunting employment landscape for college graduates. In other sectors, opportunity is booming. Homes, roads, bridges, vehicles, […]
This story was co-published with Mother Jones. A few years ago, Eric Gil was living at his uncle’s place in Waterbury, Connecticut, where he shared a bedroom with his brother and cousin. With eight people in the house, it was crowded. He was trying to get out, but rental prices in Waterbury — which currently […]
As the school year begins, teachers often devote significant time to planning lessons, reviewing standards, organizing classrooms, and establishing routines. While these are all essential components of effective teaching, they cannot replace the importance of creating a positive classroom culture.
When a leadership team meeting ends with everyone offering to help but no one agreeing to lead, something important has gone wrong beneath the surface. In this sharp, diagnostic piece, education leadership professor Andy Szeto names a pattern hiding in plain sight: voiceless collegiality, the organizational silence that masquerades as team harmony. Szeto identifies three distinct types of silence eroding school effectiveness and offers concrete shifts for leadership teams ready to trade comfort for clarity. The post Voiceless Collegiality: When Getting Along Gets in the Way of School Leadership appeared first on Getting Smart .
Assessment after AI is not about fear, but rather about intentional and purposeful design
McMahon’s Trick Questions Elizabeth Redden Thu, 08/13/2026 - 03:00 AM Universities risk walking right into a trap. Byline(s) Richard Amesbury
The NCAA Can’t Outrun Its Antitrust Problem sara.custer@in… Thu, 08/13/2026 - 03:00 AM A stalled bill, a growing pile of lawsuits and no clear path for the athletes at the center of it all. Byline(s) Sara Custer
How 2 Texas Lawsuits Could Tee Up a National Fight Over Academic Freedom Emma Whitford Thu, 08/13/2026 - 03:00 AM The AAUP’s suits against Texas A&M and Texas Tech could bring questions about academic freedom to the “unpredictable” Fifth Circuit, which has yet to weigh in on the issue. Byline(s) Emma Whitford
Survey Shows Students Trust College Leaders More Than Politicians Olivia.sanchez Thu, 08/13/2026 - 03:00 AM Students from all political parties reported higher approval of campus-led policies, compared to state and federal policies, according to data from Gallup and Lumina Foundation. Byline(s) Olivia Sanchez
College Wasn’t Built for Student Parents Joshua.Bay Thu, 08/13/2026 - 03:00 AM In this week’s Voices of Student Success episode, Generation Hope’s Nicole Lynn Lewis explores how colleges can better support student parents and caregivers. Byline(s) Joshua Bay
Senators Prod State Department to Process Student, Scholar Visas sara.custer@in… Thu, 08/13/2026 - 03:00 AM Byline(s) Sara Custer
The NSF’s Ph.D.-to-Industry Pipeline Push kathryn.palmer… Thu, 08/13/2026 - 03:00 AM The majority of STEM Ph.D.s take jobs outside academia. While some universities help them prepare, the National Science Foundation is now investing millions to boost those efforts. Byline(s) Kathryn Palmer
S.C. Colleges Could Face Financial Penalties Over New Bathroom Law Katherine Knott Thu, 08/13/2026 - 03:00 AM Byline(s) Katherine Knott
ED Says It Plans to Spend Millions in Expiring Research Dollars. Concerns Remain. jessica.blake@… Thu, 08/13/2026 - 03:00 AM Higher education advocates say delays in education research funding have already caused severe damage. Byline(s) Jessica Blake
Under a new rule, the designation allows students to take out $200,000 in federal student loans versus $100,000 for other graduate programs.
Students at top-ranked colleges were less likely than others to say campus leaders mostly act in their interests, per a Gallup and Lumina Foundation poll.
The Trump administration’s move against unintentional discrimination will likely narrow or close investigations, education civil rights experts say.
Most districts don't have the systems in place to screen for the learning disability related to numbers and math, CRPE said in a report.
At a gathering in Boston, ‘student senators’ proposed a first-of-its-kind national AI policy for K-12 classrooms.
arXiv:2606.06027v2 Announce Type: replace-cross Abstract: Community-conditioned language model adaptation needs choices about data collection, community definition, and evaluation that are currently made independently in each study, making it hard to compare assumptions or reuse artifacts. We present RedditPersona, a modular framework that standardizes these choices: it collects Reddit posts and comments, profiles active users, partitions them under five grouping strategies (subreddit-based, graph-structural, semantic, hybrid, and interaction-based), trains a parameter-efficient adapter per strategy via QLoRA, and evaluates them under a shared metric suite spanning fluency, fidelity, distributional alignment, and community identifiability. Applied to 112 subreddits in the urban well-being domain (301,429 user profiles, 16M+ comments), we find that adapters' behavioral identifiability tracks each strategy's agreement with the subreddit baseline, and that a consistent trade-off between i
arXiv:2606.04923v2 Announce Type: replace-cross Abstract: Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. However, policy models may exploit latent biases in the judge, leading to reward hacking and ineffective or unsafe training outcomes. In real-world rubric-based RL, such hacking behaviors are often subtle and entangled with multiple judge biases, making them difficult to analyze, detect, and mitigate. In this paper, we introduce CHERRL, a Controllable Hacking Environment for Rubric-based RL. By injecting known biases into LaaJ, CHERRL enables stable reproduction of reward hacking, explicit observation of reward divergence, and identification of hacking onset. This provides a clean experimental testbed for studying the mechanisms and mitigations of reward hacking in rubric-based RL. To demonstrate its utility, we analyze different judge biases from the perspectives of discoverability and exploitability, a
arXiv:2606.00671v3 Announce Type: replace-cross Abstract: We present Moxia (formerly AXIOM), a trust-first neuro-symbolic architecture for self-explaining mathematical reasoning over natural-language input. Its language model is strictly a canonicalizer: it rewrites informal problem text into a narrow schema consumed by a deterministic Computer-Algebra-System (CAS) pipeline, which derives and verifies the answer or abstains as a first-class output. Routing follows a 1:1:1 alignment of problem-shape regex, schema-specific prompt, and closed-form CAS handler, with 4,783 routes shipped, 71% of which answer without invoking the language model, and zero LOST_CORRECT regressions as a standing release gate. Because the answer is derived rather than generated, so is its explanation: every handler emits a step trace of the computation it performed, rendered as prose by a layer covering all 4,785 task files that cannot narrate a step the handler did not take. Derivations export to Lean 4 as well
arXiv:2605.27101v2 Announce Type: replace-cross Abstract: A key capability for video understanding is reliably linking subjects to events across time, yet whether Video Large Language Models (VideoLLMs) actually achieve this remains unclear. In this work, we introduce DistractionBench to evaluate whether VideoLLMs can robustly link subjects and events in the presence of unrelated video segments. Through controlled interventions, such as inserting short advertisement clips into longer videos, we show that VideoLLMs frequently hallucinate interactions between entities from different segments, incorrectly attributing actions from injected advertisements to subjects in the main video. We characterize this systematic hallucination as bag-of-events (BoE) behavior, where models process videos as collections of events rather than temporally structured sequences. Evaluating 11 popular VideoLLMs, we find that all models exhibit substantial BoE behavior. Our findings suggest that VideoLLMs lack r
arXiv:2605.16409v3 Announce Type: replace-cross Abstract: Optical character recognition (OCR) and multilingual scene-text understanding remain challenging for multimodal large language models (MLLMs), particularly in real-world images containing small or degraded text, cluttered layouts, occlusion, handwriting, and complex typography. We present an OCR-aware multilingual post-training framework that improves visual-text grounding in a general-purpose MLLM without requiring an external OCR engine, OCR-extracted text, or text bounding boxes at inference time. The framework combines large-scale multilingual OCR supervision, approximately 5M additional multilingual training samples, controlled synthetic OCR generation and in-image text translation, LoRA-based supervised fine-tuning (SFT), and lightweight OCR-oriented Chain-of-Thought prompting. On a held-out real-world multilingual OCR benchmark, OCR-SFT improves OCR completeness from 71.3 to 84.6, reduces hallucination rate from 18.3\% to
arXiv:2509.13450v3 Announce Type: replace-cross Abstract: We introduce SteeringSafety, a benchmark for evaluating representation steering methods across nine safety perspectives spanning 18 datasets. While prior work highlights the general capabilities of representation steering, we focus on safety perspectives including refusal, bias, hallucination, social behaviors, reasoning, epistemic integrity, and normative judgment. SteeringSafety provides modularized building blocks for state-of-the-art steering methods, enabling unified implementation of DIM, ACE, CAA, PCA, and LAT with recent enhancements such as conditional steering. Results on Gemma-2-2B, Llama-3.1-8B, and Qwen-2.5-7B show that strong steering performance depends on the pairing of method, model, and specific perspective. For instance, DIM is consistently effective, yet all methods exhibit substantial entanglement, where improving effectiveness on one safety perspective often significantly changes performance on others. Soci
arXiv:2411.15122v2 Announce Type: replace-cross Abstract: AI-driven models have demonstrated significant potential in automating radiology report generation for chest X-rays. However, there is no standardized benchmark for objectively evaluating their performance. To address this, we present ReXrank, https://rexrank.ai, a public leaderboard and challenge for assessing AI-powered radiology report generation. Our framework incorporates ReXGradient, the largest test dataset consisting of 10,000 studies, and three public datasets (MIMIC-CXR, IU-Xray, CheXpert Plus) for report generation assessment. ReXrank employs 8 evaluation metrics and separately assesses models capable of generating only findings sections and those providing both findings and impressions sections. By providing this standardized evaluation framework, ReXrank enables meaningful comparisons of model performance and offers crucial insights into their robustness across diverse clinical settings. Beyond its current focus on
arXiv:2408.06849v3 Announce Type: replace-cross Abstract: The large language model (LLM) has achieved significant success across various domains. However, the inherent complexity of causal problems and causal theory poses challenges in accurately describing them in natural language, making it difficult for LLM to comprehend and use them effectively. Causal methods are not easily conveyed through natural language, which hinders LLM's ability to apply them accurately. Additionally, causal datasets are typically tabular, while LLM excels in handling natural language data, creating a structural mismatch that impedes effective reasoning with tabular data. To address these challenges, we have equipped the LLM with causal tools within an agent framework, named the Causal Agent, enabling it to tackle causal problems. The causal agent comprises tools, memory, and reasoning modules. In the tool module, the causal agent calls Python code and uses the encapsulated causal function module to align t
arXiv:2606.28876v3 Announce Type: replace Abstract: Proposal. Long context can replay history, but it does not decide which completed observations deserve authority. MMLA formalizes a bounded resident memory between transient context and slow weight updates. A completed local segment is eventized; for each event, a target-conditioned constructor proposes semantic content and a trusted assembler produces a complete versioned row; deployment either commits that row atomically or returns NULL. Realized futures may price actions during training, while deployment remains causal and future-blind. Validated components. Controlled studies establish narrower ingredients. Lifecycle execution is exact on 300/300 held-out records for each of three seeds. Calibrated selection with full-archive fallback improves over a weak budget-matched dense baseline by 5.5--16.6 F1 and over BM25 by 4.0--6.2 F1 on held-out multi-hop QA; the original Llama budget execution is retained as failed, while the correcte
arXiv:2606.09498v2 Announce Type: replace Abstract: The performance of LLM-based agents is jointly shaped by their base models and the harnesses that mediate their interaction with the environment. Because different models exhibit distinct behaviors, effective harness design is inherently model-specific. Yet agent harnesses are still largely engineered by human experts, a paradigm that scales poorly as modern LLMs become increasingly diverse and rapidly evolving. In this paper, we introduce Self-Harness, a new paradigm in which an LLM-based agent improves its own operating harness, without relying on human engineers or stronger external agents. We operationalize Self-Harness as an iterative loop with three stages: Weakness Mining, which identifies model-specific failure patterns from execution traces; Harness Proposal, which generates diverse yet minimal harness modifications tied to these failures; and Proposal Validation, which accepts candidate edits only after regression testing. W
arXiv:2605.04495v2 Announce Type: replace Abstract: Retrieval-augmented generation (RAG) relies on evidence ranking to determine what information is exposed to the generator, yet existing retrieval and reranking methods primarily estimate query--document relevance. Relevance, however, is not equivalent to generator-side usefulness: a relevant passage may introduce ambiguity or distraction, whereas a lower-ranked passage may stabilize the generator's answer. We present CAR (Confidence-Aware Reranking), a training-free rank-correction framework that uses query-only answer stability as a control and measures each candidate by the change it induces in sampled-answer semantic stability. This controlled contrast estimates a document's marginal contribution to generator behavior without treating semantic stability as relevance or calibrated correctness. CAR converts these confidence changes into coarse precedence constraints and returns the feasible ranking with minimum Kendall distance from
arXiv:2604.17244v2 Announce Type: replace Abstract: Large language model (LLM) agents for sequential decision-making struggle to produce diverse outputs. This leads to insufficient exploration, suboptimal solutions, and repeated actions. Actions are generated at the sequence level, but existing sampling strategies, such as temperature scaling, introduce diversity at the token level, not at the sequence level. We introduce DORA EXPLORER (Diversity-Oriented Ranking of Actions), a training-free, inference-time algorithm for improving exploration in LLM agents. DORA generates multiple candidate actions, scores them using sequence-level log-probability statistics, and samples an action via a tunable exploration parameter. We first study exploration in the classic Multi-Armed Bandit setting, where DORA substantially outperforms temperature-based sampling. Our main evaluation is on the Text Adventure Learning Environment Suite (TALES), where prompting strategies fail to explore but DORA deliv
arXiv:2604.07801v2 Announce Type: replace Abstract: Large language models are trained and evaluated on quantitative reasoning tasks written in clean, emotionally neutral language. However, real-world queries are often wrapped in frustration, urgency or enthusiasm. Does emotional framing alone degrade reasoning when all numerical content is preserved? To investigate this, a controlled emotion translation framework is developed that rewrites problems into emotional variants while preserving all quantities and relationships. Using this framework, Temper-5400 (5,400 semantically verified emotion-neutral pairs) is constructed across GSM8K, MultiArith, and ARC-Challenge, and evaluated on eighteen models (1B to frontier scale). Two core results emerge: First, emotional framing reduces accuracy by 2-10 percentage points even though all numerical content is preserved. Second, neutralizing emotional variants recovers most of the lost performance, showing both that the degradation is tied to emot
arXiv:2604.03924v2 Announce Type: replace Abstract: Goal-oriented conversational systems require making sequential decisions under uncertainty about the user's intent, where the algorithm must balance information acquisition and target commitment over multiple turns. Existing approaches address this challenge from different perspectives: structured methods enable multi-step planning but rely on predefined schemas, while LLM-based approaches support flexible interactions but lack long-horizon decision making, resulting in poor coordination between information acquisition and target commitment. To address this limitation, we formulate goal-oriented conversation as an uncertainty-aware sequential decision problem, where uncertainty serves as a guiding signal for multi-turn decision making. We propose a Conversation Uncertainty-aware Planning framework (CUP) that integrates language models with structured planning: a language model proposes feasible actions, and a planner evaluates their l
arXiv:2604.02512v2 Announce Type: replace Abstract: Large language models (LLMs) increasingly exhibit human-like patterns of pragmatic and social reasoning. This paper addresses two related questions: do LLMs approximate human social meaning not only qualitatively but also quantitatively, and can prompting strategies informed by pragmatic theory improve this approximation? To address the first, we introduce two calibration-focused metrics distinguishing structural fidelity from magnitude calibration: the Effect Size Ratio (ESR) and the Calibration Deviation Score (CDS). To address the second, we derive prompting conditions from two pragmatic assumptions: that social meaning arises from reasoning over linguistic alternatives, and that listeners infer speaker knowledge states and communicative motives. Applied to a case study on numerical (im)precision across three frontier LLMs, we find that all models reliably reproduce the qualitative structure of human social inferences but differ su
arXiv:2603.20895v3 Announce Type: replace Abstract: Existing routers rely on semantic query features or handcrafted features, which often fail to capture model-specific failures or intrinsic task difficulty. We instead route using internal LLM activations, specifically the residual stream. Our key idea, Encoder-Target Decoupling, separates the model that produces the predictive signal (the Encoder) from the model whose correctness is being estimated (the Target), allowing open-weight encoders to predict the performance of closed-source target models. We evaluate layerwise geometric probes, finding that Fisher Separability ($J$) effectively identifies informative layers, supported by Effective Dimensionality ($d_{\mathrm{eff}}$) diagnostics. We then utilize a SharedTrunkNet, a joint multi-output MLP that predicts simultaneous correctness probabilities across candidate models using concatenated prefill features. In our experiments, SharedTrunkNet consistently outperforms semantic baselin
arXiv:2603.13891v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used for automated text annotation in tasks ranging from academic research to content moderation and hiring. Across 19 LLMs and two experiments totaling more than 4 million annotation judgments, we show that subtle identity cues embedded in text systematically bias annotation outcomes in ways that mirror racial stereotypes. In a names-based experiment spanning 39 annotation tasks, texts containing names associated with Black individuals are rated as more aggressive by 18 of 19 models and more gossipy by 18 of 19. Asian names produce a bamboo-ceiling profile: 17 of 19 models rate individuals as more intelligent, while 18 of 19 rate them as less confident and less sociable. Arab names elicit cognitive elevation alongside interpersonal devaluation, and all four minority groups are consistently rated as less self-disciplined. In a matched dialect experiment, the same sentence is judged signifi
arXiv:2602.13452v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly proposed for crisis preparedness and response, particularly for multilingual communication. However, their suitability for high-stakes crisis contexts remains insufficiently evaluated. This work examines the performance of state-of-the-art LLMs and machine translation systems in crisis-domain translation, with a focus on preserving urgency, a critical property for effective crisis communication and triage. Using multilingual crisis data (TICO-19, 30 languages) and a newly introduced urgency-annotated dataset of 100 scenarios translated into 29 languages, we show that dedicated translation models and LLMs exhibit substantial quality degradation, particularly for low-resource languages. Beyond translation quality, we conduct a human annotation study revealing a striking asymmetry: human assessors maintain consistent urgency judgments regardless of prompt language, while LLM-based urgency cla