Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
These three legal regimes are pulling in different directions, and the developers that navigate them well will be the ones that plan for all three from the start. The post The AI Tension in Healthcare: Patent Strategy, FDA Reality, and HIPAA Constraints appeared first on MedCity News .
Rural Health Transformation Program funding gives states broad flexibility and may help address longstanding access gaps. The post For Millions in Rural America, Medicine Is Still Far Away appeared first on MedCity News .
Slowing birth rates drive declines that took state past pandemic enrollment lows. The post NJ public school enrollment falls to lowest level in quarter century appeared first on District Administration .
Limited funds and low enrollment forced districts to close more than 1,000 schools in 2025-26. The post Shrinking student populations cause spike in public school closures across U.S. appeared first on District Administration .
More than 250 million American adults start the new calendar year with a pledge to complete Dry January. By giving up alcohol for a month, some hope for a physical or mental reset, while others hope to give up drinking for good. While more than a quarter of participants won’t make it through the full […]
Princess Moss has been a leader of the National Education Association since 2014, but soon she will implement her own plans for the union as incoming president. One of her top goals when she takes office Sept. 1 is to shift campaigning and organizing from the union’s centralized, national level to local affiliates. Moss was […]
Critics say the White House proposal could curtail studies of achievement gaps and long-term research.
With the growing movement to limit technology in schools, parents and educators are right to question how much time students spend on screens. But, as schools reconsider the role of technology in the classroom, an important distinction is often getting lost in the public conversation: Not all screen time is equal.
The research behind the influential “science of reading” movement was built over decades through federally supported research spanning cognitive science, neuroscience, psychology, linguistics and classroom studies. Education research organizations say a proposed 100-page sweeping overhaul of federal grant rules, which would add a new review by political appointees, could make it harder to produce similar […] The post Education researchers warn political review of federal grants could reshape what gets studied appeared first on The Hechinger Report .
Insights education companies can use to evolve their marketing and sales strategies and build lasting partnerships with schools
What happens when a school system delegates student achievement to an AI system with a single directive: raise test scores? In this provocative new piece, David Ross draws on AI alignment theory, Goodhart's Law, and a decade of stagnant NAEP data to map the subordinate goals an optimization engine might quietly pursue, from narrowing the curriculum to managing teachers as variables. It is essential reading for any leader navigating the intersection of political pressure and emerging technology. The post When AI Chases Test Scores, What Gets Left Behind? appeared first on Getting Smart .
Aqua from Adobe is a free drawing and illustrator tool designed specifically for kids.
As AI adoption accelerates, a countertrend has emerged that reinforces the old saying: What goes around comes around. After spending years looking for ways to remove friction from everyday life, a growing number of people are embracing it by choosing to perform certain tasks without the assistance of AI and other forms of automation. The post What “friction maxxing” teaches us about IT skill development in the age of AI appeared first on eCampus News .
Tell Your Enrollment Story Before Someone Else Does Elizabeth Redden Mon, 07/20/2026 - 03:00 AM If colleges and universities don’t act now to explain their admission decisions, they will be vulnerable to political attacks. Byline(s) Sean Robins
Taking the Cotton Candy Out of Your Learning Objectives Elizabeth Redden Mon, 07/20/2026 - 03:00 AM Rethinking your learning outcomes to build them around transferable skills is the first step toward better, more nutritious course design. Byline(s) Zachary Nowak
Young Men’s Perceptions of Higher Ed Sara Weissman Mon, 07/20/2026 - 03:00 AM Byline(s) Sara Weissman
How Soon Could Colleges Lose Loan Access Under New Accountability Metric? jessica.blake@… Mon, 07/20/2026 - 03:00 AM For most programs, data from the new test on student earnings will be released in 2027 and failing programs could face penalties in 2028. But some have been granted an extension that student advocates say is harmful. Byline(s) Jessica Blake
New Initiative Targets Transfer Credit Loss Joshua.Bay Mon, 07/20/2026 - 03:00 AM Higher education accreditor SACSCOC’s new consortium aims to create common transfer pathways and shorten time to degree. Byline(s) Joshua Bay
Inside the ‘Culture of Fear’ at One American-Chinese University Emma Whitford Mon, 07/20/2026 - 03:00 AM Faculty at Wenzhou-Kean have accused their employer of obfuscatory employment agreements, wrongful termination, discrimination, retaliation and lack of shared governance. University officials dispute the charges. Byline(s) Emma Whitford
Driving More Students to ‘High-Value’ Programs Doug Lederman Mon, 07/20/2026 - 03:00 AM Defining programs solely by graduates’ earnings isn’t ideal, but it’s where we are. A new report shows that it can be done, and how. Byline(s) Doug Lederman
Do You Have to Be Fancy to Get Published in ‘Science’? kathryn.palmer… Mon, 07/20/2026 - 03:00 AM No, not necessarily. But scientists who work for a prestigious university are three times more likely than their peers who don’t to get published in the top-tier journal, according to a new independent analysis of Science ’s manuscript data. Byline(s) Kathryn Palmer
Tag-Teaming New York Sara Brady Mon, 07/20/2026 - 03:00 AM An update for longtime readers. Byline(s) Matt Reed
Texas Colleges Directed to Slash State Budget Requests by 3% Katherine Knott Mon, 07/20/2026 - 03:00 AM Byline(s) Katherine Knott
Some colleges have cut deals, while others have fought back against federal officials. Both come with risks, legal experts say.
Former university presidents Michael Nietzel and Charles Ambrose discuss faculty votes against college leaders and what to do in the face of one.
We’re rounding up last week’s stories, from Republican efforts to downsize the U.S. Department of Education to Nebraska and Texas universities’ funding cuts.
We’re rounding up last week’s news, from a principal’s conference in Florida to pushback on an Education Department interagency agreement.
From transgender student rights to student free speech, here is what's been decided and what's left for future terms.
For years, early childhood leaders have been asked to solve structural problems with predictable answers: Improve quality. Expand access. Support families. Retain educators. Strengthen transitions into kindergarten. All of it matters, but none of it can be fully accomplished with our fragmented system, one in which Head Start, child care, public pre-K, special education and […] The post OPINION: How a small town in rural Maine built a public school that serves a whole community, including infants and toddlers appeared first on The Hechinger Report .
A new PwC report says the women’s oncology market could reach $110 billion by 2030, though funding and care gaps remain. The post Women’s Oncology Market Could Reach $110B in 2030, but Gaps Remain appeared first on MedCity News .
arXiv:2606.20155v2 Announce Type: replace-cross Abstract: Text-to-image (T2I) models generate realistic likenesses of some individuals when prompted with their names, raising privacy concerns. However, distinguishing whether a generated face is memorized or fabricated currently requires ground-truth photos, access to training data, or white-box access to model internals, limiting applicability. We introduce a fully black-box behavioral probe that distinguishes between memorized and unrecognized names, while requiring no reference photos or prior knowledge of training data. To benchmark this task, we present the NAMESAKES dataset of over one thousand names and faces of public figures spanning a wide range of fame levels, along with perturbed, less famous names. Experiments on state-of-the-art T2I models show that our probe substantially predicts identity memorization and separates memorized from unrecognized names, with further insights into differences across model families.
arXiv:2606.07636v2 Announce Type: replace-cross Abstract: Long-form video editing over heterogeneous footage requires agents to coordinate source selection, multimodal analysis, timeline construction, narration and subtitle alignment, rendering, and revision while exposing intermediate state for inspection and repair. We present Crayotter, an open-source multimodal multi-agent demo system for prompt-driven long-form video editing. Crayotter organizes production around coverage-aware material preparation, artifact-grounded editing research, and tool-grounded timeline execution. Across these stages, retrieval reports, video analyses, editing blueprints, scheduler events, tool calls, intermediate renders, and final exports are treated as first-class artifacts rather than hidden transient state. The workbench supports local assets, agent-assisted retrieval, progress monitoring, artifact preview, failure diagnosis, interrupted-job resumption, and resource-aware asynchronous execution for lo
arXiv:2606.02240v3 Announce Type: replace-cross Abstract: Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such as Gmail, Salesforce, or Jira accessed through tool calls) whose response content the user neither writes nor controls. Existing benchmarks under-measure the threat: most cover only a handful of integrations with the same attack payload replayed across runs, and open-source guards are trained on chat-style data rather than tool-response content. We introduce AGENTREDBENCH, a dynamic LLM-driven redteaming benchmark of 215 subtle underspecified-authorization scenarios across 24 enterprise integrations and five attack types. Across an eight-model panel (Anthropic, OpenAI, Google), no-guard attack success rate ranges from 32% to 81%. To keep the scenario set out of training corpora and preserve headline ASR meaning over time, we release the codebase, integration schemas, and AGENTREDGUARD model o
arXiv:2605.04893v3 Announce Type: replace-cross Abstract: Every attention head defines a degree-normalized transport operator, and a growing family of diagnostics reads model behavior (hallucination among them) from its spectrum. We ask what such diagnostics can and cannot infer. The operator splits orthogonally into a symmetric part governing transport \emph{capacity} and an antisymmetric part encoding \emph{orientation}. We prove an identifiability limit: every transpose-invariant spectral diagnostic is \emph{orientation-blind} (unable to distinguish an operator from its transpose, hence blind to the orientation of information flow), with a transpose-stability bound limiting any Lipschitz diagnostic's transpose sensitivity by the asymmetry coefficient $G$. This bounds what spectral diagnostics of the attention operator can resolve (e.g.\ LapEigvals and the attention-spectral branch of LLM-Check). On the surviving axis, a closed-form bipartite-Cheeger landscape shows uniform causal at
arXiv:2604.15010v2 Announce Type: replace-cross Abstract: When do transformers commit to a decision, and what prevents them from correcting it? We introduce prolepsis: a transformer commits early, task-specific attention heads sustain the commitment, and no layer corrects it. Replicating Lindsey et al.'s (2025) planning-site finding on open models (Gemma 2 2B, Llama 3.2 1B), we ask five questions. (Q1) Planning is invisible to six residual-stream methods; among those tested, only CLT-based steering succeeds. (Q2) The single-site spike replicates in shape, at the final prompt token (Anthropic's site is the newline; see the Note added). (Q3) Specific attention heads route the decision to the output, filling a gap flagged as invisible to attribution graphs. (Q4) The evidence is consistent with search within at most 16 layers and commitment beyond, a two-model hypothesis. (Q5) Factual recall shows the same motif at a different network depth, with zero overlap between recurring planning hea
arXiv:2604.04593v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) grounds large language models in external medical knowledge, yet standard retrievers frequently surface hard negatives that are semantically close to the query but describe clinically distinct conditions. While existing query-expansion methods improve query representation to mitigate ambiguity, they typically focus on enriching target-relevant semantics without an explicit mechanism to selectively suppress specific, clinically plausible hard negatives. This leaves the system prone to retrieving plausible mimics that overshadow the actual diagnosis, particularly when such mimics are dominant within the corpus. We propose Contrastive Hypothesis Retrieval (CHR), a framework inspired by the process of clinical differential diagnosis. CHR generates a target hypothesis $H^+$ for the likely correct answer and a mimic hypothesis $H^-$ for the most plausible incorrect alternative, then scores document
arXiv:2604.03873v4 Announce Type: replace-cross Abstract: Black-box knowledge distillation for large language models presents a strict trade-off. Simple off-policy methods (e.g., sequence-level knowledge distillation) struggle to correct the student's inherent errors. Fully on-policy methods (e.g., Generative Adversarial Distillation) solve this via adversarial training but introduce well-known training instability and crippling computational overhead. To address this dilemma, we propose SODA (Semi On-policy Distillation with Alignment), a highly efficient alternative motivated by the inherent capability gap between frontier teachers and much smaller base models. Because a compact student model's natural, zero-shot responses are almost strictly inferior to the powerful teacher's targets, we can construct a highly effective contrastive signal simply by pairing the teacher's optimal response with a one-time static snapshot of the student's outputs. This demonstrates that exposing the sma
arXiv:2602.24009v4 Announce Type: replace-cross Abstract: Jailbreak techniques for large language models (LLMs) evolve faster than benchmarks, making robustness estimates stale and difficult to compare across papers due to drift in datasets, harnesses, and judging protocols. We introduce JAILBREAK FOUNDRY (JBF), a system that addresses this gap via a multi-agent workflow to translate jailbreak papers into executable modules for immediate evaluation within a unified harness. JBF features three core components: (i) JBF-LIB for shared contracts and reusable utilities; (ii) JBF-FORGE for the multi-agent paper-to-module translation; and (iii) JBF-EVAL for standardizing evaluations. Across 30 reproduced attacks, JBF achieves high fidelity with a mean (reproduced-reported) attack success rate (ASR) deviation of +0.26 percentage points. By leveraging shared infrastructure, JBF reduces attack-specific implementation code by more than half relative to original repositories and achieves an 82.5%
arXiv:2602.17229v2 Announce Type: replace-cross Abstract: The black-box nature of Large Language Models necessitates novel evaluation frameworks that transcend surface-level performance metrics. This study investigates the internal neural representations of cognitive complexity using Bloom's Taxonomy as a hierarchical lens. By analyzing high-dimensional activation vectors from different LLMs, we probe whether different cognitive levels, ranging from basic recall (Remember) to abstract synthesis (Create), are linearly separable within the model's residual streams. Our results demonstrate that linear classifiers achieve approximately 95% mean accuracy across all Bloom levels, providing strong evidence that cognitive level is encoded in a linearly accessible subspace of the model's representations. These findings provide evidence that the model resolves the cognitive difficulty of a prompt early in the forward pass, with representations becoming increasingly separable across layers.
arXiv:2607.13347v2 Announce Type: replace Abstract: LLM-as-a-judge is widely used to provide feedback and selection signals in closedloop regeneration, but this use remains insufficiently validated. We study it in table recognition, where deterministic TEDS evaluation provides a controlled testbed, using FinTabNet and OmniDocBench. Three findings emerge. First, judge signals were weak on both datasets: scores frequently tied, rankings were not reproducible, and no tested judge score policy, whether selecting among candidates or accepting revisions under a conservative score margin, improved on the first output on both datasets. Iteration produced better candidates, but the judge recovered them at most partially on one dataset and not at all on the other. Second, severe losses occurred even without specific judge feedback, supporting target-preservation failure under unconstrained regeneration as a proximate mechanism. Third, a structure-preserving instruction reduced the severe-loss ra
arXiv:2606.17522v2 Announce Type: replace Abstract: Deep neural networks are widely believed to derive their expressive power from their ability to form \textbf{hierarchical representations}, capturing progressively more abstract and compositional features across layers. In language modeling, \textbf{transformers} have emerged as the dominant architecture, with early layers capturing local syntactic patterns and later layers encoding more complex clause-level dependencies. While this intuition has shaped model design, there remains a lack of rigorous theoretical work demonstrating \textbf{how} deep transformers represent such hierarchical structures. In this work, we analyze the expressiveness of deep transformer models through the formal lens of bounded-depth, non-recursive context-free grammars. For this class of grammars, we explicitly construct transformers with positional attention whose depth grows linearly with grammar depth, while the neuron count scales with the number of deri
arXiv:2606.16127v2 Announce Type: replace Abstract: The worldwide rise of authoritarianism and the growing role of Large Language Models (LLMs) in users' everyday lives raise the question of whether specific models exhibit or promote authoritarian attitudes. We introduce AuAu, a comprehensive benchmark for assessing the risk of authoritarian tendencies in LLM responses. AuAu combines three evaluation approaches: (i) psychometric questions from 15 human-validated instruments, (ii) vignettes probing intended behavior in concrete situations, and (iii) responses to realistic user prompts. Unlike prior work, AuAu measures not only overall authoritarian alignment but also its established sub-concepts: Authoritarian Aggression, Authoritarian Submission, and Conventionalism. Evaluating 17 models from China, the EU, Russia, and the USA, we find substantial authoritarian response rates on psychometric instruments across all models, though rates drop significantly on more realistic downstream tas
arXiv:2606.10931v3 Announce Type: replace Abstract: Warning: This paper contains several toxic and offensive statements. Modern large language models (LLMs) are typically aligned through large-scale post-training to ensure fair and reliable behavior. In this work, we investigate how easily such guardrails can be broken by Group Relative Policy Optimization (GRPO). We show that one-shot GRPO training on a single biased example is sufficient to induce systematic bias, with stereotype-driven reasoning generalizing across attributes, categories, and benchmarks. We further find that models differ in their susceptibility based on the initial likelihood of producing biased outputs. Our results reveal a critical vulnerability in post-training: alignment can be overridden by a single example.
arXiv:2606.09421v3 Announce Type: replace Abstract: Large language model agents increasingly rely on skills: reusable procedural documents encoding workflows, tool use, implementation patterns, validation checks, and domain rules. Skill rewriting is often treated as prompt compression, but shorter skills can make agents more expensive by removing sparse operational anchors that prevent exploration, debugging, and recovery. We study skill rewriting through this economic lens. Our controlled framework profiles skill structure, rewrites skills using information-preservation strategies, and evaluates the rewrites under fixed task instructions, environments, and verifiers. Experiments on SkillsBench reveal distinct quality--cost trade-offs across strategies: API/code anchoring, workflow guarding, and rule/formula anchoring benefit different task families, with no universally dominant template. In the main held-out evaluation, the learned policy reduces total cost by 7.0% and downstream agen
arXiv:2605.26431v3 Announce Type: replace Abstract: We show that LLMs encode syntactic distinctions not present in the Universal Dependencies (UD) tree distances that structural probes are trained to recover. On English wh-movement stimuli, we measure the probe distance between an embedded subject and its verb, whose UD tree distance is invariant across conditions. That distance is shorter than baseline when the embedded clause is finite and longer when it is infinitival -- a within-clause sign asymmetry present in all 13 models across four families we test, at a majority of layers. No account based on UD distance, linear order, or monotone structural complexity can produce a sign reversal, while Minimalist phase theory can. Holding the matrix verb fixed while varying only the complement type reproduces the same finite-infinitival ordering in every model, ruling out a lexical-semantic explanation. The cross-clause pair separately reproduces the phase-count ordering of earlier work, val
arXiv:2605.15677v2 Announce Type: replace Abstract: Despite the rapid advancements in Vision-Language Models (VLMs), a critical gap remains in their ability to handle structured, controllable diagrammatic tasks essential for professional workflows. Existing methods predominantly rely on pixel-based synthesis, which operates in probabilistic pixel spaces and is inherently limited in editability and fidelity. Instead, we propose a new Diagram-as-Code paradigm with symbolic logic that leverages mxGraph Extensible Markup Language (XML) for precise diagram generation and editing. We present VCG-Bench, a unified benchmark for visual-centric \texttt{mxGraph} tasks. VCG-Bench comprises: (1) a taxonomized dataset of 1,449 diverse diagrams spanning 6 domains and 15 sub-domains, (2) a paradigm definition that integrates Generation (Vision-to-Code) and Editability (Code-to-Code), (3) a Tailored Evaluation Protocol employing multi-dimensional metrics such as \texttt{mxGraph} Execution Success Rate,
arXiv:2605.06006v2 Announce Type: replace Abstract: Fact-checking articles encode rich supporting evidence and reasoning, yet this evidence remains largely inaccessible to automated verification systems due to unstructured presentation. We introduce PrimeFacts, a methodology and resource for extracting fine-grained evidence from full fact-checking articles. We compile 13,106 PolitiFact articles with claims, verdicts, and all referenced sources, and we identify 49,718 in-article hyperlinks as natural anchors to pinpoint key evidence. Our framework leverages large language models (LLMs) to rewrite these anchor sentences into stand-alone, context-independent premises and investigates the extraction of additional implicit evidence. In evaluations on cross-article evidence retrieval and claim verification, the extracted premises substantially improve performance. Decontextualized evidence yields higher retrievability, achieving up to a 30 percent relative gain in Mean Reciprocal Rank over v
arXiv:2605.04539v4 Announce Type: replace Abstract: Direct Preference Optimization (DPO), the efficient alternative to PPO-based RLHF, falls short on knowledge-intensive generation: standard preference signals from human annotators or LLM judges exhibit a systematic verbosity bias that rewards fluency over logical correctness. This blindspot leaves a logical alignment gap -- SFT models reach NLI entailment of only 0.05-0.22 despite producing fluent text. We propose RLearner-LLM with Hybrid-DPO: an automated preference pipeline that fuses a DeBERTa-v3 NLI signal with a verifier LLM score, removing human annotation while overcoming the "alignment tax" of single-signal optimization. Evaluated across five academic domains (Biology, Medicine, Law) with three base architectures (LLaMA-2-13B, Qwen3-8B, Gemma 4 E4B-it), RLearner-LLM yields up to 6x NLI improvement over SFT, with NLI gains in 11 of 15 cells and consistent answer-coverage gains. On Gemma 4 E4B-it (4.5B effective params), Hybrid-
arXiv:2604.16370v3 Announce Type: replace Abstract: Decoding natural language from non-invasive electroencephalography (EEG) remains constrained by low signal-to-noise ratio and limited information bandwidth. This raises a central question: can sentence-level language be reliably recovered from such signals? Under realistic information constraints, this direct-recovery assumption may be too strong. We introduce a semantic compression hypothesis: non-invasive EEG may preserve recoverable semantic anchors rather than the full lexical--syntactic form of a sentence. From this perspective, direct sentence reconstruction is overly fine-grained relative to the recoverable information scale of EEG. To address this mismatch, we propose Brain-CLIPLM, a two-stage framework that decomposes EEG-to-text decoding into semantic-anchor recovery and anchor-guided sentence reconstruction. Stage 1 uses contrastive learning to align word-level EEG evidence with a fixed keyword vocabulary and recover ordere
arXiv:2604.12069v4 Announce Type: replace Abstract: Robust explanations are increasingly required for user trust in enterprise NLP, yet pre-deployment validation is difficult in the common case of black-box deployment (API-only access) where representation-based explainers are infeasible and existing studies provide limited guidance on whether explanations remain stable under real user noise, especially when organizations migrate from encoder classifiers to decoder LLMs. To close this gap, we propose a unified black-box robustness evaluation framework for token-level explanations based on leave-one-out occlusion, and operationalize explanation robustness with top-token flip rate under realistic perturbations (swap, deletion, shuffling, and back-translation) at multiple severity levels. Using this protocol, we conduct a systematic cross-architecture comparison across three benchmark datasets and six models spanning encoder and decoder families (BERT, RoBERTa, Qwen 7B/14B, Llama 8B/70B;