Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
Turnover rates among nurses are holding steady, according to new data. Nursing turnover is always a point of contention for CNOs as they try to build a sustainable workforce. In 2026, RN turnover remains sitting at about 17%, which is the same as last year, according to Press Ganey's State of Nursing 2026 report. Gen Z and Millennials have the highest turnover rates, with Gen Z sitting at 22% and Millennials at 21%. Jeff Doucette , CNO at Press Ganey, previously told HealthLeaders the reason Gen Z nurses are leaving the workforce has to do with unmet needs, centering around purpose, support, and alignment with their organizations. "Gen Z clinical nurses generally have less [of a] feeling of psychological safety and they are experiencing significant cognitive overload and administrative burden, but they are less tolerant of the dysfunction that many of us have just learned to live with in healthcare environments," Doucette said. Here are some facts and figures that CNOs should know from
Once upon a time. For generations, those four words were an invitation. Children leaned in because a story was beginning. They would listen closely, follow the characters, and stay with the plot until the end.
A new educational epistemology for today's teachers.
Creating consistency between classrooms and ensuring curriculum alignment school-wide can be challenging, even in the smallest of districts. Every educator teaches--and grades--differently based on their experience and preferences, and too often, they’re forced into a solution that no longer respects their autonomy or acknowledges their strengths.
If an undergraduate program's graduates don't earn more than workers who never went to college, that program could be cut off from federal student loans. But is a degree just about making more money?
Public schools continued to serve more than 1.5 million students across North Carolina in the 2025-26 school year — or about 84% of market share, a percentage that is about the same as previous years. Market share is a term used to describe how many students are served by different sectors of schools, including public […]
In 2020, a 14-inch tall robot named Moxie was introduced to the world as a way to help children build social and emotional skills through conversations and interactive games guided by artificial intelligence. Kids became enthralled and attached to the oblong teal robot, referring to it as their best friend. Four years later, Embodied, the […] The post Don’t let AI raise your kids appeared first on The Hechinger Report .
Collaborations between states and districts can help improve student outcomes, the governors of Maryland and Wyoming say.
Class Disrupted is an education podcast featuring author Michael Horn and Futre’s Diane Tavenner in conversation with educators, school leaders, students and other members of school communities as they investigate the challenges facing the education system in the aftermath of the pandemic — and where we should go from here. Find every episode by bookmarking […]
Bakersfield College agreed to not require Daymon Johnson to use diversity, equity, inclusion and accessibility principles in his teaching or scholarship.
Clemson Picks UGA Provost as President After Guskiewicz Reversal Ryan Quinn Thu, 07/09/2026 - 11:07 AM Byline(s) Ryan Quinn
The last time I interviewed Dr. Dana Suskind, we discussed the three T’s strategy outlined in her book “Parent Nation: Unlocking Every Child’s Potential, Fulfilling Society’s Promise”: Tune in. Talk more. Take turns. “It doesn’t require fancy gadgets,” she told me, “or a specialized degree.” Though it was only a few years ago, the fancy […]
Just 20 states earned a “meets requirements” designation from the Education Department this year for serving students with disabilities ages 3 to 21 under the Individuals with Disabilities Education Act. The post Most states fail to meet IDEA requirements, feds say appeared first on District Administration .
Date & Time: Thursday, August 06, 2026 at 2 p.m. ET In this webinar, learn the key takeaways from a report suggesting that transportation teams are sitting on information that could directly connect to whether students show up to class, whether families choose your schools, and whether your most vulnerable students, such as those experiencing homelessness or with IEPs, are getting the access the law requires. The post Attendance, Enrollment, and Equity Answers Hiding in Your Transportation Data appeared first on District Administration .
This may sound strange coming from the co-founder of an education technology company, but I think paper is a powerful technology in a classroom. Research on the science of learning consistently says so. So do the piles of paper that good teaching produces, the piles that bury the teachers who produced them. The real question […]
Andy Mann’s model is deceptively simple. Informed by years of experience as a technology teacher, he developed a vision to get $10,000 virtual reality headsets and other sought-after devices in the hands of teachers who wanted to try them in their classrooms without the steep investment of buying them. Using strategic funding, Mann meticulously amassed a collection of STEM items that teachers from the Muskegon Area Intermediate School District in Michigan can check out. At his ISTELive 26 session, “Borrow Don’t Buy: Creating a STEM Loaner Library,” Mann shared how he built his STEM loaner…
One of the greatest group project collaborations in K–12 right now has nothing to do with social studies or science projects. The IT leaders who understand the risk tend to speak in technical language, while the administrators and educators who control the budget speak in outcomes. Finding a shared vocabulary helps them fight — and ultimately fund a defense to — the K–12 cybersecurity crisis. The school year ended with a historic cybersecurity attack across K–12 and higher education, punctuating a point that teachers, IT leaders and school districts have been very aware of, with…
Public schools are beginning to imagine ways they can benefit from the new Federal Scholarship Tax Credit, after the Treasury Department clarified last month that district students will be eligible for scholarships. But for microschools, a growing segment of the private school market, the initial guidance from federal officials has left school leaders worried they […]
Our schools face a paradoxical challenge: chronic absenteeism, which is being treated as an attendance problem, actually isn’t.
Our schools face a paradoxical challenge: chronic absenteeism, which is being treated as an attendance problem, actually isn’t.
If you're new to teaching, it particularly helps to remember that you’re not the only one navigating some difficult waters.
How Do Employers View Community College Baccalaureate Degrees? Sara Weissman Thu, 07/09/2026 - 03:00 AM Byline(s) Sara Weissman
North Texas Denies Faculty Conference Funding, Citing Anti-DEI Provisions gianna.jakubowski Thu, 07/09/2026 - 03:00 AM The institution joins other universities that have worked to limit conference funds for professors due to anti-DEI pressures. Byline(s) Gianna Jakubowski
ICE Detains HBCU Baseball MVP gianna.jakubowski Thu, 07/09/2026 - 03:00 AM Byline(s) Gianna Jakubowski
Donor to Move Scholarship From UNC Wilmington Due to DEI Policy Johanna Alonso Thu, 07/09/2026 - 03:00 AM Byline(s) Johanna Alonso
How Far Can One Supreme Court Ruling Stretch? sara.custer@in… Thu, 07/09/2026 - 03:00 AM The Trump administration is using the ban on affirmative action in admissions to crack down on anything related to ethnicity or race in higher ed. Will it try the same tactic with last week’s Supreme Court ruling on transgender athletes? Byline(s) Sara Custer
From Inmate to College Graduate Joshua.Bay Thu, 07/09/2026 - 03:00 AM CUNY’s Prison-to-College Pathways Pipeline celebrated its first commencement while helping incarcerated students earn associate degrees and find belonging. Byline(s) Joshua Bay
Texas Tech Faculty Sue Over Race, Gender Rules Katherine Knott Thu, 07/09/2026 - 03:00 AM Byline(s) Katherine Knott
The Key Podcast: Students Don’t Want Press Releases, They Want Outcomes sara.custer@in… Thu, 07/09/2026 - 03:00 AM Byline(s) IHE Staff
College Associations Say Expanded List of Professional Degrees Is ‘Incomplete’ jessica.blake@… Thu, 07/09/2026 - 03:00 AM Several health-care degrees like advanced nursing, physician assistants and occupational therapists were added, but others, like master’s in social work and education, were not. Byline(s) Jessica Blake
A proposed rule would rescind regulations for the program, which the agency says would allow it to “explore other means” of delivering those services.
At least one state has gone as far as to prohibit artificial intelligence’s use for grading, discipline or other high-stakes decisions.
The students booing artificial intelligence at commencements across the country are not just worried about jobs. They have learned an urgent lesson from the not-so-distant past. They know that the familiar promise of empowerment and creativity will continue to give way to the pathologies of the online surveillance economy: viral slop, commercial manipulation and addictive […] The post OPINION: The days of ‘good guy’ capitalists are over. College students are right to turn against the tech elites appeared first on The Hechinger Report .
What would chemistry look like if students could do more than read about it?
arXiv:2605.09778v2 Announce Type: replace-cross Abstract: Evaluating softmax attention over a fixed long context requires reading every cached key-value pair for each new query token. For a given context (a book, a manual, a legal corpus) the attention output is a deterministic function of the query. We propose Nectar, which fits a compact neural network to this function for queries drawn from a task-relevant distribution. Nectar fits two networks per layer and KV-head: a target network that predicts the attention output and a score network that predicts the log-normalizer. The pair plugs into the standard masked self-attention at inference time, replacing the $O(n)$ attention over the cache with a forward pass whose cost does not depend on $n$. Each module carries on the order of $|\theta|$ parameters per layer and KV-head, typically much smaller than the $2nd$ KV-cache footprint at the same granularity. We report experiments on models from 1.7B to 8B parameters across five long-conte
arXiv:2604.18360v3 Announce Type: replace-cross Abstract: Audio-text retrieval systems based on Contrastive Language-Audio Pretraining (CLAP) achieve strong performance on traditional benchmarks; however, these benchmarks rely on caption-style queries that differ substantially from real-world search behavior, limiting their assessment of practical retrieval robustness. We present Omni-Embed-Audio (OEA), a retrieval-oriented encoder leveraging multimodal LLMs with native audio understanding. To systematically evaluate robustness beyond caption-style queries, we introduce User-Intent Queries (UIQs) - five formulations reflecting natural search behaviors: questions, commands, keyword tags, paraphrases, and exclusion-based negative queries. For negative queries, we develop a hard negative mining pipeline and propose discrimination metrics (HNSR, TFR) assessing models' ability to suppress acoustically similar distractors. Experiments on AudioCaps, Clotho, and MECAT show that OEA achieves co
arXiv:2604.11950v2 Announce Type: replace-cross Abstract: While recent LLM-based agents can identify many candidate bugs in source code, their reports remain static hypotheses that require manual validation, limiting the practicality of automated bug detection. We frame this challenge as a test generation task: given a candidate report, synthesizing an executable proof-of-concept (PoC) - such as a script, command sequence, or crafted input - to trigger the suspected defect. Automated PoC generation can act as a scalable validation oracle, enabling end-to-end autonomous bug detection by providing concrete execution evidence. However, naive LLM agents are unreliable validators: they are biased toward "success" and may reward-hack by producing plausible but non-functional PoCs or even hallucinated traces. To address this, we present ANYPoC, a general multi-agent framework that (1) analyzes and fact-checks a candidate bug report, (2) iteratively synthesizes and executes a PoC while collect
arXiv:2604.07831v2 Announce Type: replace-cross Abstract: Existing red-teaming studies on GUI agents face two fundamental limitations: adversarial perturbations require white-box access unavailable in commercial deployments, while prompt injection is increasingly neutralized by stronger safety alignment. To study robustness under a more practical threat model, we propose Semantic-level UI Element Injection, a black-box red-teaming paradigm that overlays safety-aligned and harmless UI elements onto screenshots to misdirect the agent's visual grounding. Our method couples a modular Editor--Overlapper--Victim pipeline with iterative search that samples multiple candidate edits, keeps the best cumulative overlay, and adapts future prompt strategies based on previous failures. Experiments across 19 victim models spanning 8 model families show that strategic optimization substantially outperforms random injection (3.5-6.9x on the most robust victims) and transfers near-perfectly across archi
arXiv:2603.19742v2 Announce Type: replace-cross Abstract: Understanding the internal mechanisms of transformer-based large language models (LLMs) is crucial for their reliable deployment and effective operation. While recent efforts have yielded a plethora of attribution methods attempting to balance faithfulness and computational efficiency, dense component attribution remains prohibitively expensive. In this work, we introduce Dual Path Attribution (DPA), a novel framework that faithfully traces information flow on the frozen transformer in one forward and one backward pass without requiring counterfactual examples. DPA analytically decomposes and linearizes the computational structure of the SwiGLU Transformers into distinct pathways along which it propagates a targeted unembedding vector to receive the effective representation at each residual position. This target-centric propagation achieves O(1) time complexity with respect to the number of model components, scaling to long inpu
arXiv:2602.02712v2 Announce Type: replace-cross Abstract: A popular approach to post-training control of large language models (LLMs) is the steering of intermediate latent representations. Namely, identify a well-chosen direction depending on the task at hand and perturbs representations along this direction at inference time. While many propositions exist to pick this direction, considerably less is understood about how to choose the magnitude of the move, whereas its importance is clear: too little and the intended behavior does not emerge, too much and the model's performance degrades beyond repair. In this work, we propose the first theoretical analysis of steering strength. We characterize its effect on next token probability, presence of a concept, and cross-entropy, deriving precise qualitative laws governing these quantities. Our analysis reveals surprising behaviors, including non-monotonic effects of steering strength. We validate our theoretical predictions empirically on e
arXiv:2510.09595v3 Announce Type: replace-cross Abstract: Competitive programming problems are increasingly used to evaluate the coding capabilities of large language models (LLMs) due to their complexity and ease of verification. Yet, current coding benchmarks face limitations such as a lack of exceptionally challenging problems, insufficient test case coverage, and reliance on online platform APIs that limit accessibility. To address these issues, we introduce LiveOIBench, a large-scale competitive programming benchmark featuring 403 expert-curated problems, averaging 60 official test cases each, drawn from 72 contests across 14 Informatics Olympiads held between 2023 and 2025. LiveOIBench has four key features: (1) expert-designed tasks with detailed subtask rubrics and extensive test cases; (2) direct comparison to elite human contestants; (3) continuous updates to reduce contamination risk; and (4) a fully offline, reproducible evaluation system. Benchmarking 34 popular general-pu
arXiv:2508.00554v4 Announce Type: replace-cross Abstract: In financial trading, large language model (LLM)-based agents demonstrate significant potential, but their decisions can be sensitive to noisy and non-stationary market information. We propose ContestTrade, a multi-agent trading system with an internal competitive mechanism inspired by institutional investment workflows. The system consists of two specialized teams: (1) a Data Team that processes and condenses massive market data into diversified textual factors optimized for constrained LLM context windows, and (2) a Research Team that produces parallelized multipath trading decisions via tool-augmented deep research. The core design is a "Quantify-Predict-Allocate" contest mechanism within each team: agent outputs are scored only after market outcomes become observable, future utility is predicted from historical scores, and resources are allocated to agents with positive predicted utility. In a post-2024 A-share backtest, Con
arXiv:2607.04581v2 Announce Type: replace Abstract: Text embeddings for Portuguese have no dedicated benchmark: evaluation rests on translated corpora such as English MS MARCO or on thin multilingual coverage, with native tasks scattered and unconsolidated. We introduce MTEB-BR, a benchmark of 22 native Brazilian-Portuguese tasks across seven categories (classification, multilabel classification, pair classification, semantic textual similarity, clustering, retrieval, and reranking), admitting only data created or found in Portuguese and excluding translations by construction. We evaluate 93 models spanning 23M to 27B parameters: 73 open-weight and 20 closed commercial APIs. Alongside the leaderboard we report a statistical layer for every headline comparison: per-task bootstrap confidence intervals, paired-bootstrap significance, a task- and instance-level discrimination analysis (how sharply each task separates models) adapted from Item Response Theory, and a cross-leaderboard correl
arXiv:2606.05894v2 Announce Type: replace Abstract: Long-horizon agents can archive large histories, but future answers still incur retrieval, rereading, and context costs. When retained memory misses answer-relevant evidence, the system must return to larger portions of the raw history. We study budgeted evidence survival: before the query is known, which source evidence should be retained so that it remains recoverable and usable under a fixed retained source-evidence token budget? We instantiate this setting as Budgeted Pre-Query Retention, where memory is written during ingestion and later read without access to the full raw stream. We introduce EMBER, a learned retention policy that constructs a compact, source-backed evidence state. EMBER stores evidence capsules: verbatim source excerpts paired with retrieval keys and update metadata, preserving both grounding and read-time access. Post-query outcome feedback trains the writer to preserve evidence across the ingestion-retrieval-
arXiv:2605.27000v3 Announce Type: replace Abstract: Repeated sampling with a verifier is the standard way to allocate test-time compute for code generation, with pass@$K$ as the canonical metric. Yet the standard policy class draws $K$ independent samples from a single answer distribution, so attempts often collapse onto near-duplicate reasoning paths and waste the budget on redundant rollouts. This failure is costly in competitive programming, where many problems admit multiple distinct algorithmic strategies and pass@$K$ requires only one correct attempt. We propose Coordinated Pass@$K$ Policy Optimization (CPPO), which turns pass@$K$ generation into joint exploration over strategies: a planner emits a tuple of $K{=}4$ alternative high-level methods, and a shared solver attempts one solution per method. CPPO trains this joint policy with a multiplicative planner reward, $R_{\mathrm{plan}} = J_\psi \cdot R_{\mathrm{out}}$, assigning credit only to valid strategy tuples that lead to ve
arXiv:2605.22140v2 Announce Type: replace Abstract: In recent years, large language models have shown substantial potential in psychological support tasks. However, existing psychological counseling data mostly rely on single-turn question answering or short multi-turn dialogues, making it difficult to characterize how college students' psychological distress accumulates, interacts, and gradually evolves over long periods within campus life events. To address this issue, this paper proposes Psy-Chronicle, a structured data-generation framework for synthesizing long-horizon campus psychological counseling dialogues. We generate a semester-spanning temporal stress event graph to model the chronological order and evolutionary dependencies among campus stress events. Through interactive simulation between a student agent and a counselor agent, together with a structured memory integration mechanism, Psy-Chronicle generates long-horizon dialogues with continuity across counseling sessions.
arXiv:2604.25702v3 Announce Type: replace Abstract: Contemporary neural machine translation (NMT) systems are almost exclusively built by training on supervised parallel data. Despite the tremendous progress achieved, these systems still exhibit persistent translation errors. This paper proposes that a post-training paradigm based on reinforcement learning (RL) can effectively rectify such mistakes. We introduce a novel framework that requires only a general text corpus and an expert translator which can be either human or an AI system to provide iterative feedback. In our experiments, we focus specifically on English-to-German translation as a representative high-resource language pair. Crucially, we implement this RL-based post-training using Direct Preference Optimization (DPO). Applying our DPO-driven framework to the gemma3-1b model yields a significant improvement in translation quality, elevating it's COMET score from 0.703 to 0.747 on the English to German task. The results dem
arXiv:2603.21489v2 Announce Type: replace Abstract: AI agents have become increasingly capable at isolated software engineering (SWE) tasks such as resolving issues on Github. Yet long-horizon tasks involving multiple interdependent subtasks still pose challenges both with respect to accuracy, and with respect to timely completion. A natural approach to solving these long-horizon tasks in a timely manner is asynchronous multi-agent collaboration, where multiple agents work on different parts of the task at the same time. But effective application of multi-agent systems has proven surprisingly difficult: concurrent edits by multiple agents interfere with each other, dependencies are difficult to synchronize, and combining partial progress into a coherent whole is challenging. On the other hand, human developers have long relied on mature collaboration infrastructure to manage these challenges in large software projects. Inspired by these collaboration primitives, we introduce Centralize
arXiv:2603.02150v2 Announce Type: replace Abstract: The extraction of critical information from crime-related documents is a crucial task for law enforcement agencies. The extraction of this information can be interpreted as a Named-Entity Recognition (NER) task. However, there is a considerable lack of adequately annotated data on general real-world crime scenarios. To address this issue, we present CrimeNER, a case study of crime-related NER, and a general crime-related Named-Entity Recognition database (CrimeNER-db), consisting of more than 1.5K annotated documents extracted from public reports of terrorist attacks and the US Department of Justice's press notes. We define 4 coarse types of crime entity and 21 fine-grained entity types. We address the quality of the presented database with experiments using fully supervised finetuned general NER models and zero- and few-shot experiments to address the generalization capabilities. The database is available on GitHub.
arXiv:2602.04521v2 Announce Type: replace Abstract: Modern deployments require LLMs to enforce safety policies at scale, yet many controls rely on inference-time interventions that add recurring compute cost and serving complexity. Activation steering is widely used, but it requires runtime hooks and scales cost with the number of generations; conditional variants improve selectivity by gating when steering is applied but still retain an inference-time control path. We ask whether selective refusal can be moved entirely offline: can a mechanistic understanding of category-specific refusal be distilled into a circuit-restricted weight update that deploys as a standard checkpoint? We propose C-{\Delta}{\theta} Circuit Restricted Weight Arithmetic}, which (i) localizes refusal-causal computation as a sparse circuit using EAP-IG and (ii) computes a constrained weight update {\Delta}{\theta}C supported only on that circuit (typically <5% of parameters). Applying {\Delta}{\theta}C yields a d