Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
Ultragenyx Pharmaceuticals’ apazunersen did not meet the goals of its Phase 3 test in Angelman syndrome, a rare genetic neurological disorder with no FDA-approved therapies. The disappointing result has some readthrough to Ionis Pharmaceuticals and Oak Hill Bio, each in clinical development with Angelman drugs designed to work in a similar way as Ultragenyx’s failed molecule. The post Ultragenyx Trial Failure Revives Debate About How to Treat a Rare Neuro Disorder appeared first on MedCity News .
The bill seeks to keep all special education, postsecondary, Native American, and elementary and secondary education activities within the department.
The ban is being implemented for the 2026-27 school year as the nation's second-largest school system reviews its artificial intelligence policies.
The for-profit arts and media college's student population fell by 13.1% from 2025 to 2026, according to a spokesperson.
Nursing students in Virginia may soon get paid while completing the clinical training they are already required to do, under a new state grant program. The Virginia Nursing Workforce Center…
Nursing students in Virginia may soon get paid while completing the clinical training they are already required to do, under a new state grant program. The Virginia Nursing Workforce Center…
If enacted, diversity work deemed illegal by President Donald Trump would disqualify colleges from federal tax exemption.
LAWRENCE, Kan. — The U.S. Department of Justice and Department of Education announced a plan for increased enforcement measures against two Kansas school districts for their gender-inclusive policies. In April, the department found four public school districts — Olathe, Shawnee Mission, Topeka and Kansas City, Kansas — were in violation of federal Title IX civil […]
This story was originally published by Grist and is reprinted with permission. This fall, 54 million K-12 students are headed back to the classroom for another year of lessons in all the classic subjects: math, English, history, biology. But there’s another subject that’s been sneaking into school curricula: plastics. A new report from the nonprofit […] The post When plastics companies write the lesson plans appeared first on The Hechinger Report .
This story was originally published by Grist and is reprinted with permission. This fall, 54 million K-12 students are headed back to the classroom for another year of lessons in all the classic subjects: math, English, history, biology. But there’s another subject that’s been sneaking into school curricula: plastics. A new report from the nonprofit […] The post When plastic companies write the lesson plans appeared first on The Hechinger Report .
This shift is not about replacing clinical care with tradition. It is about recognizing that healing in Native communities has always been deeply connected to culture. The post Healing Through Heritage: Why Culturally Grounded Recovery Matters on Tribal Lands appeared first on MedCity News .
Editor's note: The author is Chief Nursing Officer of the American Nurses Enterprise, whose AI training initiative is discussed in this article. Nurse.org received no compensation for this article. Artificial…
Editor's note: The author is Chief Nursing Officer of the American Nurses Enterprise, whose AI training initiative is discussed in this article. Nurse.org received no compensation for this article. Artificial…
New York City is expected to release its school policy for artificial intelligence on Wednesday, with a ban on student-facing AI tools for 2-K through eighth grade, according to four people briefed on the plans and Education Department documents obtained by Chalkbeat. Under the new rules, the youngest children in the nation’s largest school system […]
All too often, students reach their junior year of high school with only a vague notion of what they want to do after they graduate. Waiting until graduation is within reach to begin college and career conversations can leave both the students and counselors feeling rushed and out of time. Districts that are already using data to inform decisions have pivoted to using this information to improve the post-graduation experience for students. “The innovative shift is moving away from a generic ‘here’s a list of colleges’ approach and toward tracking students longitudinally,” says Eric Crespo,…
The data undercuts a major argument for the vouchers—that they would chiefly help Texans hoist their kids out of failing public schools. The post Most Texas voucher recipients were already enrolled in private school appeared first on District Administration .
When a middle schooler sits down in math class, she might be thinking about the argument she had on the way to school, the stress at home she can’t escape or whether she’s going to eat today. She is probably not thinking about algebra. Adults design policies for her future — the classes she takes, […]
AI technology is advancing rapidly in schools, but research is increasingly showing that Large Language Models (LLMs) lack the common sense, morality, and knowledge of individual classroom dynamics needed to be truly equitable.
These are common mistakes for those new to the teaching profession, but are easy to overcome.
Article URL: https://phys.org/news/2026-09-technology-engagement.html Comments URL: https://news.ycombinator.com/item?id=49547003 Points: 2 # Comments: 0
The Number We’re Not Counting quintina.barne… Thu, 09/03/2026 - 03:00 AM An absenteeism-adjusted participation rate. Byline(s) Carolyn Gentle-Genitty
Another Member of Former Regional Accreditors’ Group Leaves Ryan Quinn Thu, 09/03/2026 - 03:00 AM In its departure announcement, the Middle States Commission on Higher Education said it’s “driving broader partnerships.” Byline(s) Ryan Quinn
Miami University Prohibits Faculty From Teaching Outside for Labor Day Action Emma Whitford Thu, 09/03/2026 - 03:00 AM University officials say the planned event violates the collective bargaining agreement, as well as Ohio law. Faculty say this is “untrue and a gross mischaracterization” of the teach-out, which was intended to be celebratory. Byline(s) Emma Whitford
‘Start From Yes’ Initiative Challenges Transfer Credit Skepticism Johanna Alonso Thu, 09/03/2026 - 03:00 AM Byline(s) Johanna Alonso
Set the World on Fire sara.custer@in… Thu, 09/03/2026 - 03:00 AM As AI reshapes higher ed, could teaching students how to be good human beings save it? Byline(s) Sara Custer
Student Support Needs Outpace College Capacity Joshua.Bay Thu, 09/03/2026 - 03:00 AM A new EAB survey finds financial, mental health and academic challenges are straining colleges already facing staffing and budget constraints. Byline(s) Joshua Bay
The Key Podcast: What Happens if We Get AI Right? sara.custer@in… Thu, 09/03/2026 - 03:00 AM Byline(s) IHE Staff
Syracuse Chancellor Says Budget Deficit ‘Far From a Crisis’ Katherine Knott Thu, 09/03/2026 - 03:00 AM Byline(s) Katherine Knott
More Than Half of College Applicants Reported Key Home-Based Challenges Olivia.sanchez Thu, 09/03/2026 - 03:00 AM New data from Common App’s responsibilities and circumstances question—added for all applicants last year—allows students to share the range of extracurricular tasks they perform beyond sports, clubs and volunteer work. Byline(s) Olivia Sanchez
‘Novel’ Partnership Sours Into ‘Hostile Takeover’ kathryn.palmer… Thu, 09/03/2026 - 03:00 AM Three years after Antioch and Otterbein Universities partnered to widen student access to academic programs, Otterbein’s president led an effort to dissolve Antioch’s board and fire its new president. Antioch has taken legal action to preserve its autonomy. Byline(s) Kathryn Palmer
August saw a property proposal dispute culminate in an Arkansas college firing its leader. Plus, a long-serving president revised his exit timeline.
Proving the effectiveness of paths such as apprenticeships can help students know their options, writes an Urban Institute senior fellow.
The findings come as career and technical education has gained traction and are “concerning” amid broader postsecondary shifts, a Rand report said.
Falling enrollment and the need for consistent and stable student resources have led the district to explore options to "rightsize."
A nationwide investigation reveals massive waitlists for government-funded childcare assistance, leaving low-income families stranded without essential ...
arXiv:2608.28945v3 Announce Type: replace-cross Abstract: Automating alignment research may accelerate progress toward aligned AI, but whether it does is hard to measure. Luckily, many alignment failures, such as deception, sycophancy, and jailbreaks, are already measurable by public benchmarks. We study whether automated alignment researchers (AARs) can post-train to mitigate alignment failures by proposing training methods and data to simultaneously optimize multiple safety benchmarks, while largely preserving general capability. Across 10 alignment failures, the strongest AAR methods significantly reduce the targeted alignment failures and generalize to a held-out benchmark, multi-turn behavioral audits, and models up to 4.7x larger than the target model. As a human baseline, 28 experienced researchers receive up to eight hours to develop one-shot methods for the same benchmarks, but their methods underperform the best AAR methods. Using human ideas as the AARs' initial research dir
arXiv:2607.17117v2 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) decompose language model activations into sparse features, yet these models traditionally encode each token independently, failing to expose information that persists across a sequence. We first show that temporal persistence can naturally emerge in standard SAE features: after a feature activates, the hidden state remains aligned with its direction, and past activations help reconstruct later hidden states. How long this lasts varies widely across features. We therefore introduce Persistent Sparse Autoencoders (Persistent SAEs), an extension of standard SAEs that learns a persistence coefficient for each feature, allowing the model to learn feature-specific timescales from reconstruction alone. Our experiments show that Persistent SAEs retain competitive reconstruction quality while learning a spectrum of timescales: short-timescale (fast) features stay locally interpretable, whereas long-timescale (s
arXiv:2607.10103v2 Announce Type: replace-cross Abstract: As large language models (LLMs) become widely deployed, their outputs can be copied, transformed, and redistributed at scale without reliable evidence of origin, creating risks for trust, accountability, intellectual property (IP) protection, and high-stakes decision-making. LLM watermarking addresses this problem by embedding detectable signals into text during or after generation. However, existing methods vary in design assumptions, threat models, and evaluation criteria, while deployment choices such as watermark placement, detection authority, and key management affect reliability, security, and scalability. This paper systematizes LLM watermarking as provenance infrastructure for large-scale data ecosystems. We organize existing approaches along four deployment dimensions: insertion point, verification authority, operational state, and transformation threat model, and relate them to the big data requirements of Volume, Vel
arXiv:2606.20728v2 Announce Type: replace-cross Abstract: Vision foundation tools such as open-vocabulary detectors, segmentation models, and post-processing operators are powerful building blocks for computer vision, but their effectiveness depends heavily on how they are orchestrated: which tools are used, in what order, with what parameters, and under what visual conditions. Existing visual-programming agents typically generate a fixed solution pipeline, making them brittle under dense objects, occlusion, small targets, and domain shift. We introduce VTOS (Vision Tools Orchestration Search), a framework for adaptive visual tool orchestration through joint solution-observer search. VTOS co-searches executable solution programs that compose vision tools such as Grounding DINO, SAM, NMS, and slice-and-detect, together with observer programs that diagnose candidate solutions, identify failure modes, and generate actionable feedback. These observations are accumulated in a shared VisionT
arXiv:2606.04547v3 Announce Type: replace-cross Abstract: Personalizing large language models requires adapting model behavior to individual users while preserving robustness and deployment-scale efficiency. Existing approaches typically personalize LLMs either at the input level, by retrieving user histories or constructing profile prompts, or at the parameter level, by maintaining user-specific parameter-efficient modules. The former makes personalization sensitive to retrieval quality and prompt design, whereas the latter incurs storage and maintenance costs that grow with the user population. To address these limitations, we propose TAP-PER (Temporal Attentive Prefix for PERsonalization), a prefix-based framework that encodes user preferences as learnable representations, avoiding the serialization of user histories into prompts and replacing heavy per-user adapters with lightweight user-state prefix embeddings. Inspired by personalized recommendation systems, TAP-PER decomposes us
arXiv:2605.03096v2 Announce Type: replace-cross Abstract: In classification tasks, models may rely on confounding variables to achieve strong in-distribution performance, capturing spurious features that fail under distribution shift. This shortcut behavior leads to substantial degradation in out-of-distribution settings. Task arithmetic offers a potential solution by removing unwanted signals via subtraction of secondary model updates, but it typically requires full fine-tuning, which is computationally expensive. Prompt tuning provides a parameter-efficient alternative by adapting models through a small set of trainable virtual tokens. Task arithmetic on the resulting prompts presents an appealing alternative to operations on entire models, but the extent to which this approach can limit reliance on spurious features remains to be established. In this work, we study whether composing soft prompts through task arithmetic improves robustness to confounding shifts. We propose Hybrid Pro
arXiv:2604.26841v2 Announce Type: replace-cross Abstract: When do language diffusion models memorize their training data, and how to quantitatively assess their true generative regime? We address these questions by showing that Uniform-based Discrete Diffusion Models (UDDMs) fundamentally behave as Associative Memories (AMs) $\textit{with emergent creative capabilities}$. The core idea of an AM is to reliably recover stored data points as $\textit{memories}$ by establishing distinct basins of attraction around them. Historically, models like Hopfield networks use an explicit energy function to guarantee these stable attractors. We broaden this perspective by leveraging the observation that energy is not strictly necessary, as basins of attraction can also be formed via conditional likelihood maximization. By evaluating token recovery of $\textit{training}$ and $\textit{test}$ examples, we identify in UDDMs a sharp memorization-to-generalization transition governed by the size of the tr
arXiv:2603.03072v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to assist scientists across diverse workflows. A key challenge is generating high-quality figures from textual descriptions, often represented as TikZ programs that can be rendered as scientific images. Prior research has proposed a variety of datasets and modeling approaches for this task. However, existing datasets for Text-to-TikZ are too small and noisy to capture the complexity of TikZ, causing mismatches between text and rendered figures. Moreover, prior approaches rely solely on supervised fine-tuning (SFT), which does not expose the model to the rendered semantics of the figure, often resulting in errors such as looping, irrelevant content, and incorrect spatial relations. To address these issues, we construct DaTikZ-V4, a dataset more than four times larger and substantially higher in quality than DaTikZ-V3, enriched with LLM-generated figure descriptions. Using this da
arXiv:2512.16310v4 Announce Type: replace-cross Abstract: LLM agents can combine individually non-revealing tool returns and disclose a sensitive conclusion, creating Tools Orchestration Privacy Risk (TOP-R). We formalize TOP-R through three conditions: conclusion sensitivity, single-source non-inferability, and compositional inferability. We introduce Library-Grounded Reverse-Inference Seed Expansion (LRSE), a four-library reverse-construction pipeline, and use it to build TOP-Bench, a 1,000-instance benchmark evaluated under a controlled two-stage tool-use protocol. Across six LLM agents, average task completion, leakage, and H-score are 98.0 percent, 88.6 percent, and 20.4. With native reasoning enabled, four models average 81.4 percent final-response leakage and 82.4 percent reasoning-trace leakage. With reasoning disabled, three prompt-only safeguards improve H-score by an average of about 3.4 points on TOP-Bench. We further propose TOP-Align, an SFT+DPO method for learning safer
arXiv:2505.15276v2 Announce Type: replace-cross Abstract: Large reasoning models (LRMs) have achieved remarkable success on complex tasks, yet their tendency to "overthink" leads to inefficiencies. Although "save-thinking" prompts are intended to mitigate this issue, we find that LRMs still frequently enter the "Still-thinking" mode instead of the expected "No-thinking" mode, especially on difficult queries. To analyze this behavioral divergence, we examine LRMs from three perspectives: confidence at the thinking-termination boundary, divergence in internal attention distributions, and attention allocation across prompt segments. We find that high perplexity is associated with later Still-thinking behavior, and that Still-thinking cases allocate more attention to the original question. Based on these observations, we propose an attention intervention method to regulate this behavior. While this intervention suppresses explicit thinking, it also causes a drop in accuracy, suggesting tha
arXiv:2505.00759v3 Announce Type: replace-cross Abstract: The steady improvements of text-to-image (T2I) generative models lead to slow deprecation of automatic evaluation benchmarks that rely on static datasets, motivating researchers to seek alternative ways to evaluate T2I progress. We present Multimodal Text-to-Image Eval (MT2IE), an evaluation framework in which a single multimodal large language model (MLLM) acts as an evaluator agent, iteratively generating the evaluation prompts and scoring the resulting images. We show that MT2IE's image-text consistency scores have higher correlation with human judgment than metrics previously introduced in the literature. MT2IE generates prompts that are efficient at probing T2I model performance: closely recovering the official T2I model rankings of three structurally distinct benchmarks from just 20 generated evaluation prompts, 28-105x fewer than the benchmarks' own prompt sets. When compared to existing evaluation metrics such as CLIPSco
arXiv:2407.14845v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are widely used in decision-making across diverse domains. Ensuring the generation of safe and reliable responses is critical for the effective deployment of LLM-based applications, particularly in high-stakes domains such as healthcare and finance. Most of these applications typically use carefully crafted prompts to guide response generation; however, the relationship between prompts and the reliability of LLM-generated responses is not yet fully understood. To address this gap, we propose a novel prompt-response concept model that explains the relationship between the amount of task-relevant information (informativeness) provided in the prompt and the LLM-generated response uncertainty by identifying four sources of response uncertainty: prompt underspecification, model quality, task variability, and semantic redundancy. We prove that response uncertainty decreases as prompt informativeness or mod
arXiv:2608.03729v3 Announce Type: replace Abstract: Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large language models (LLMs), treating the model itself as the knowledge source. However, LLMs natively possess no representation of entities, leading to duplicate entries as well as conflations. We propose GPTKB 2.0, a methodology for constructing disambiguated KBs directly from LLMs. GPTKB 2.0 incorporates on-the-fly disambiguation of entities, relations and classes, and is meticulously designed to satisfy both scalability and disambiguation accuracy. We analyze the central design decisions and characterize the trade-offs between accuracy, scale, and cost. We execute GPTKB 2.0 at scale, obtaining a materialized KB containing over 1M disambiguated entities and 38.4M triples. This represents the first million-scale LLM-native KB with explicit internal canonicalization of entities, relations, and classes, a
arXiv:2606.23092v2 Announce Type: replace Abstract: Humans possess an innate ability to understand fine-grained interpersonal relationships, which is central to everyday social interactions. Although such reasoning is inherently multimodal, it remains largely unexplored by existing multimodal large language models (MLLMs). To address this gap, we introduce PIVOTS, the first benchmark built from Social-IQ 2.0 and YouTube data to evaluate MLLMs' ability to predict bidirectional interpersonal relationship dimensions grounded in established psychology research. In addition, PIVOTS includes auxiliary tasks that assess models' ability to identify and leverage the critical visual cues underlying such predictions. We evaluate both proprietary and open-source MLLMs and conduct detailed ablation studies to analyze the effects of visual modalities and explicit social role information in conversational utterances. We further examine how joint and pairwise prediction settings benefit MLLMs in scori