Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
Religious schools accepting public funds are required to follow Maine laws that protect against discrimination based on faith, gender identity and sexual orientation, a federal court has ruled. Crosspoint Church, which runs Bangor Christian Schools, and St. Dominic Academy in Auburn filed separate appeals in the United States Court of Appeals for the First Circuit […]
“Yale has the resources, stature, and responsibility to stand firm,” the American Association of University Professors and other groups said.
Local technology companies may have solved some of their biggest problems after hearing from North Carolina high school students during the annual Teamship Showcase in Durham on June 25. Teamship is an internship experience through the College Board where students are paired in groups to solve real-world problems for the businesses they are assigned to. […]
An American Bankers Association executive said more public-private partnerships and hands-on experiences are crucial.
Having students interpret visual art can strengthen analytical and interpretive skills for reading, veteran English educator Carol Jago says.
Generative artificial intelligence is revolutionizing how colleges and universities operate, streamlining workflows, supporting research and enhancing learning. But as adoption grows, so do the risks. Higher education institutions manage sensitive student data, proprietary research and intellectual property. Without the right guardrails in place, AI systems can expose this information or violate compliance standards, putting institutions at risk for reputational damage. For CISOs and CIOs, securing AI environments must be a strategic priority. Here are four key security considerations when…
I have been a high school teacher for almost three decades, spending almost all that time teaching seniors about American civics. My teaching tenure has overlapped with the rise of the very trends now engulfing our educational system: I have watched my students embrace smartphones, social media, online learning and now artificial intelligence. But recently […]
California’s audacious goal of having half of all K-12 students enrolled in bilingual education programs by 2030 has encountered one big stumbling block. The post Short thousands of bilingual teachers, California schools turn to high school students appeared first on District Administration .
The U.S. Department of Education OKs Arkansas’ requests to consolidate federal funding and reduce the amount of paperwork and red tape required of school administrators. The post Feds approve Arkansas’ request for flexibility in spending U.S. education funds appeared first on District Administration .
The science of reading has largely won the policy debate. Over the last decade, state after state has embraced evidence-based reading instruction. Legislatures have passed literacy laws, and teacher preparation programs are (slowly) shifting their coursework. Those changes are paying off: According to national data from the DIBELS early reading screener, 30% of second graders […]
The national faculty group, alongside a local affiliate, asked a judge to block the directives and rule them unconstitutional.
Schools try to block kids from accessing dangerous content and games online but often fall short. Parents and educators say if devices are here to stay, districts need to have more control, transparency.
Summer break has finally arrived, and if you're like most educators, you're probably feeling a lot of conflicting emotions. There's relief that the school year is behind you, pride in everything you've accomplished, but also exhaustion from months of pouring your time, energy, and heart into your students.
As AI accelerates into every corner of schooling, education leaders face a question that is both urgent and deeply human: who holds the gavel? In this provocative new piece, Eric Tucker argues that the real risk is not AI itself but our failure to govern it proportionally, treating a graduation algorithm with the same scrutiny as a spelling hint. This consequence-tiered framework gives leaders, policymakers, and educators a practical architecture for protecting learners while keeping innovation alive. The post A Hint is Not a Diploma: Consequential Educational Decision Making in the Multimodal AI Era appeared first on Getting Smart .
What to be aware of and do if your student is in a toxic or troubling relationship with an AI chatbot and how teaching AI literacy can help
My father, Jake M. Schrum, would take me with him to cattle shows and sales. As an impressionable teenager, I treasured these outings mostly to have one-on-one time with my dad, but also, to learn about Hereford cattle and how to judge them. The post Nobody’s a loser: What genuine education leaders realize appeared first on eCampus News .
AACC’s President on the Future of Community Colleges Sara Weissman Wed, 07/08/2026 - 03:00 AM DeRionne Pollard wants to see community colleges claim the spotlight as engines of innovation. She’s determined to elevate their work on the national stage. Byline(s) Sara Weissman
‘RAISE US’ Is a Rare Positive Development in AI Transformation jdimaggio@upcea.edu Wed, 07/08/2026 - 03:00 AM Good news on the potential impact of AI on the workforce. Byline(s) Ray Schroeder
The Tension Between Access and Quality Doug Lederman Wed, 07/08/2026 - 03:00 AM Improving credit transfer and learning mobility will make it easier for people to game the system. But that risk is outweighed by the far larger number of students it will help. Byline(s) Doug Lederman
Some Colleges Drop Supplemental Essays for 2026–27 Johanna Alonso Wed, 07/08/2026 - 03:00 AM Administrators said the essays weren’t particularly helpful in making admissions decisions. But it’s too soon to know whether the changes are part of a larger trend. Byline(s) Johanna Alonso
Free Summer Childcare Helps Student Parents Joshua.Bay Wed, 07/08/2026 - 03:00 AM For student parents, finding childcare can mean the difference between stopping out and staying on track. LaGuardia Community College is stepping in to help. Byline(s) Joshua Bay
Brown Professor Suspects Majority of His Class Used AI to Cheat Emma Whitford Wed, 07/08/2026 - 03:00 AM Brown University leaders’ response to the alleged cheating incident has been “meek,” the professor said. Byline(s) Emma Whitford
Education Dept. Eyes Changing College Merger, Civil Rights Enforcement Regs jessica.blake@… Wed, 07/08/2026 - 03:00 AM Many of the agenda items have to do with culture war issues like defining sex and cracking down on diversity, equity and inclusion. Byline(s) Jessica Blake
Appeals Court Rules Florida’s Stop WOKE Act Violates First Amendment Katherine Knott Wed, 07/08/2026 - 03:00 AM Byline(s) Katherine Knott
Messages Show Depth of Rift Between Rothman and UW Regents Susan H. Greenberg Wed, 07/08/2026 - 03:00 AM Byline(s) Susan H. Greenberg
The Trump administration unveiled its timeline for a host of regulatory changes, including those related to accreditation, diversity initiatives and Title VI.
If accepted, a ruling could impact other cases challenging diversity efforts in major urban school systems, such as those in New York, Boston and Philadelphia.
The Center for American Progress warns that legislative cuts to safety net programs could prevent Community Eligibility Provision participation.
We’ve gathered past installments of our explainer series in one place to help you stay on top of the must-know information on key topics.
When Jodi Carreon’s son returned to school full time after the pandemic, she expected teachers would roll back the use of the laptops they had relied on while students were home. But soon after her son started second grade, Carreon realized he was still using a Chromebook throughout the day. Then the teacher sent a […] The post Schools try to block kids from accessing dangerous content and games online. Little kids are outsmarting them appeared first on The Hechinger Report .
Taking stock of the buzziest ideas to come out of ISTELive 26
Taking stock of the buzziest ideas to come out of ISTELive 26
Streaming solved the problem of access. Now, we must solve the problem of engagement.
arXiv:2606.20023v2 Announce Type: replace-cross Abstract: As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safety-relevant. However, prior tool-selection studies focus on safety-agnostic metadata preferences, leaving privilege-sensitive choices underexplored. To address this gap, we study over-privileged tool selection, in which an agent selects or escalates to a higher-privilege tool despite a sufficient lower-privilege alternative. We introduce ToolPrivBench to evaluate whether agents choose higher-privilege tools despite sufficient lower-privilege alternatives, measuring both initial selection and escalation after transient tool failures. Across eight domains and five recurring risk patterns, we find that over-privileged tool selection is common among mainstream LLM agents and is further amplified by transient failures. We further find that general safety alignment does not reliably transfer to least-privilege tool choi
arXiv:2605.23826v2 Announce Type: replace-cross Abstract: Keyframe selection is a direct way to provide verifiable visual evidence for long-video question answering (QA). Queries differ in what they require, and finding the right frames depends on knowing what to look for. Existing keyframe selectors either score every frame against a single query, or decompose the query into a fixed schema evaluated by a single visual tool. We propose ToolMerge, a keyframe retrieval method based on decomposition and merging: an Large Language Model (LLM) based planner decomposes the query into tool calls and specifies how their per-tool rankings are merged using boolean operators. To evaluate retrieval directly, we construct Molmo-2 Moments (M2M), a benchmark in which every question is anchored to a specific time interval by construction. Across QA, question retrieval, and caption retrieval, ToolMerge is competitive with prior keyframe selectors, most notably on caption retrieval, outperforming other
arXiv:2604.24222v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have achieved strong performance on general code generation, but their effectiveness drops sharply in enterprise settings where software development relies on internal private libraries absent from public pre-training corpora. Existing Retrieval-Augmented Generation (RAG) methods provide a training-free solution by retrieving static API documentation, but our analysis shows that documentation mainly helps models identify what APIs to use and remains insufficient for teaching how to use them correctly. Even with oracle API-document retrieval, LLMs still make recurring errors at the API, cross-API, and task levels, including API misuse or hallucination, flawed API composition, and incorrect solution strategies. To address this limitation, we propose MEMCoder, a training-free self-evolving memory framework for private-library code generation. MEMCoder augments existing RAG pipelines with a Multi-level E
arXiv:2604.22851v2 Announce Type: replace-cross Abstract: While Vision-Language Models (VLMs) have advanced high-level reasoning in autonomous driving, their ability to ground this reasoning in the underlying physics of ego-motion remains poorly understood. We introduce EgoDyn-Bench [Project page: (https://tum-avs.github.io/EgoDyn-Bench-Website/), Code: (https://github.com/TUM-AVS/EgoDyn-Bench), Dataset: (https://huggingface.co/datasets/fnc1901/EgoDyn-Bench)], a diagnostic benchmark for evaluating the semantic ego-motion understanding of vision-centric foundation models. By mapping continuous vehicle kinematics to discrete motion concepts via a deterministic oracle, we decouple a model's internal physical logic from its visual perception. Our large-scale empirical audit spanning 20$+$ models, including closed-source MLLMs, open-source VLMs across multiple scales, and specialized VLAs, identifies a significant Perception Bottleneck: while models exhibit logical physical concepts, they c
arXiv:2603.15600v2 Announce Type: replace-cross Abstract: Accurate process supervision remains a critical challenge for long-horizon robotic manipulation. A primary bottleneck is that current video MLLMs, trained primarily under a Supervised Fine-Tuning (SFT) paradigm, function as passive "Observers" that recognize ongoing events rather than evaluating the current state relative to the final task goal. In this paper, we introduce PRIMO R1 (Process Reasoning Induced Monitoring), a 7B framework that transforms video MLLMs into active "Critics". We leverage outcome-based Reinforcement Learning to incentivize explicit Chain-of-Thought generation for progress estimation. Furthermore, our architecture constructs a structured temporal input by explicitly anchoring the video sequence between initial and current state images. Supported by the proposed PRIMO Dataset and Benchmark, extensive experiments across diverse in-domain environments and out-of-domain real-world humanoid scenarios demonstr
arXiv:2601.12494v3 Announce Type: replace-cross Abstract: Audio large language models (LLMs) enable unified speech understanding and generation, but adapting them to linguistically complex and dialect-rich settings such as Arabic-English remains challenging. We present a controlled study of multi-task instruction tuning for an Arabic-centric audio LLM across generative tasks, including automatic speech recognition (ASR) and speech and text summarization, as well as discriminative tasks, including dialect identification (DID) and speech emotion recognition (SER), in a resource-constrained setting. To support end-to-end Arabic speech summarization, we introduce AraMega-SSum, the first Arabic speech summarization dataset designed for training and benchmarking Arabic-centric audio LLMs. We compare four training strategies: (i) Uniform Mixing (UM), (ii) Task-Progressive Curriculum (TPC), (iii) Aligner-Based Diverse Sampling (ADS) for training-time batch construction, and (iv) a two-stage TP
arXiv:2601.09173v5 Announce Type: replace-cross Abstract: Representational similarity analysis and related methods compare the internal geometries of neural networks, but they measure only alignment between spaces, leaving a blind spot -- whether a representation's structure is reliably recoverable, not merely similar. We introduce geometric stability, a distinct axis, and \textit{Shesha}, a metric that quantifies it from a single representation by correlating dissimilarity matrices built from complementary random halves of the feature dimensions. Unlike CKA and Procrustes distance, Shesha is provably non-invariant to orthogonal rotations of the feature basis. This is by design: the basis is privileged for learned models, since probes, patching, and steering act on coordinates, and a rotation-invariant metric cannot see whether the targeted structure survives them. A double dissociation isolates the mechanism -- removing the top principal component collapses CKA while Shesha holds, whe
arXiv:2601.06521v2 Announce Type: replace-cross Abstract: While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a crucial fact: state-of-the-art MLLMs consistently fail on basic visual tasks that humans, even 3-year-olds, can solve effortlessly. To systematically investigate this gap, we introduce BabyVision, a benchmark designed to assess core visual abilities independent of linguistic knowledge for MLLMs. BabyVision spans a wide range of tasks, with 388 items divided into 22 subclasses across four key categories. Empirical results and human evaluation reveal that leading MLLMs perform significantly below human baselines. Gemini3-Pro-Preview scores 49.7, lagging behind 6-year-old humans and falling well behind the average adult score of 94.1. These results show despite excelling in knowledge-heavy evaluations, current MLLMs s
arXiv:2512.18542v3 Announce Type: replace-cross Abstract: AI coding assistants produce vulnerable code in 45\% of security-relevant scenarios~\cite{veracode2025}, yet no public training dataset teaches both traditional web security and AI/ML-specific defenses in a format suitable for instruction tuning. We present SecureCode, a production-grade dataset of 2,185 multi-turn security training examples spanning two domains: web application security (1,435 examples covering the OWASP Top 10 2021 across 11 languages and 9 frameworks, 100\% grounded in documented CVEs and security incidents) and AI/ML security (750 examples covering all 10 OWASP LLM Top 10 2025 categories across more than 40 frameworks, including LangChain, OpenAI, and Hugging Face). Every example follows a 4-turn conversational structure -- feature request; vulnerable and secure implementations with attack demonstrations; advanced probing; and defense-in-depth operational guidance -- designed for direct use in instruction tu
arXiv:2510.23636v4 Announce Type: replace-cross Abstract: Flight delay prediction has become a key focus in air traffic management (ATM), as delays reflect inefficiencies in the system. This paper proposes LLM4Delay, a large language model (LLM)-based framework for predicting flight delays from the perspective of air traffic controllers monitoring aircraft after they enter the terminal maneuvering area (TMA). LLM4Delay is designed to integrate textual aeronautical information, including flight data, weather reports, and aerodrome notices, together with multiple trajectories that model airspace conditions, forming a comprehensive delay-relevant context. By jointly leveraging comprehensive textual and trajectory contexts via instance-level projection, an effective cross-modality adaptation strategy that maps multiple instance-level trajectory representations into the language modality, the framework improves delay prediction accuracy. LLM4Delay demonstrates superior performance compared
arXiv:2508.16560v4 Announce Type: replace-cross Abstract: Sparse Autoencoders (SAEs) extract features from LLM internal activations, meant to correspond to interpretable concepts. A core SAE training hyperparameter is L0: how many SAE features should fire per token on average. Existing work compares SAE algorithms using sparsity-reconstruction tradeoff plots, implying L0 is a free parameter with no inherently correct value aside from its effect on reconstruction. In this work we study the effect of L0 on SAEs, and show that if L0 is not set correctly, the SAE fails to disentangle the underlying features of the LLM. If L0 is too low, the SAE will mix correlated features to improve reconstruction. If L0 is too high, the SAE finds degenerate solutions that also mix features. Further, we present a proxy metric that can help guide the search for the correct L0 for an SAE on a given training distribution. We show that our method finds the correct L0 in toy models and coincides with peak spar
arXiv:2505.15516v3 Announce Type: replace-cross Abstract: While eXplainable AI (XAI) has advanced significantly, few methods address interpretability in embedded vector spaces where dimensions represent complex abstractions. We introduce Distance Explainer, a novel method for generating local, post-hoc explanations of embedded spaces in machine learning models. Our approach adapts saliency-based techniques from RISE to explain the distance between two embedded data points by assigning attribution values through selective masking and distance-ranked mask filtering. We evaluate Distance Explainer on cross-modal embeddings (image-image and image-caption pairs) using established XAI metrics including Faithfulness, Sensitivity/Robustness, and Randomization. Experiments with ImageNet and CLIP models demonstrate that our method effectively identifies features contributing to similarity or dissimilarity between embedded data points while maintaining high robustness and consistency. We also exp
arXiv:2409.06067v3 Announce Type: replace-cross Abstract: Previous studies on federated learning (FL) often encounter performance degradation due to data heterogeneity among different clients. In light of the recent advances in multimodal large language models (MLLMs), such as GPT-4v and LLaVA, which demonstrate their exceptional proficiency in multimodal tasks, such as image captioning and multimodal question answering. We introduce a novel federated learning framework, named Multimodal Large Language Model Assisted Federated Learning (MLLM-LLaVA-FL), which employs powerful MLLMs at the server end to address the heterogeneous and long-tailed challenges. Owing to the advanced cross-modality representation capabilities and the extensive open-vocabulary prior knowledge of MLLMs, our framework is adept at harnessing the extensive, yet previously underexploited, open-source data accessible from websites and powerful server-side computational resources. Hence, the MLLM-LLaVA-FL not only enh
arXiv:2607.04107v2 Announce Type: replace Abstract: Word surprisal is a well-established computational predictor of human neural responses during language comprehension, but it remains less clear whether local semantic fit explains neural response variation beyond lexical expectation during naturalistic reading. Using the Dublin EEG-based Reading Experiment Corpus (DERCo), this study examined whether contextual semantic relevance predicts word-locked EEG activity in the N400 and P600 windows. Contextual semantic relevance was computed as an attention-aware measure of how strongly a target word is semantically connected to its recent discourse context, and it was compared with GPT-based word surprisal. Across 22 participants and 32 EEG channels, we tested both predictors using regression-based ERP analyses and generalized additive mixed models while controlling for lexical variables and repeated observations. Both predictors were reliably associated with EEG responses, but they showed p
arXiv:2607.03863v2 Announce Type: replace Abstract: Artificial intelligence has advanced scientific discovery, but most AI4Science systems remain fragmented tools that rely on humans to coordinate problem formulation, literature grounding, model use, simulation, validation, and knowledge reuse. This paper presents \textbf{SCION (Scientific Collaborative Innovation with Agentic Organizational Nexus)}, an agentic scientific operating system that acts as an \textbf{organizational nexus}. Through a Science Agent serving as a \textbf{Meta-Harness}, SCION connects scientific tasks, tools, agents, artifacts, and memory, transforming research into an executable, auditable, and reusable operational process. At its core is the \textbf{Research Execution Plan (REP)}, which compiles high-level scientific intent into staged objectives, dependencies, verification checkpoints, tool requirements, expected artifacts, and fallback conditions. SCION further integrates hierarchical multi-agent execution,
arXiv:2606.22807v2 Announce Type: replace Abstract: As retrieval systems scale, high-quality reranking becomes increasingly important. However, most existing rerankers, whether encoder-based or decoder-based, jointly encode the query and passage, tightly coupling their computation and limiting deployment efficiency as well as flexibility. We present KaLM-Reranker-V1, a fast but not late-interaction (FBNL) reranker that decouples query and passage computation while retaining expressive relevance modeling. Built on an encoder-decoder architecture, KaLM-Reranker-V1 uses the encoder to pre-encode passages with Matryoshka embedding pooling, while the decoder models the system instruction, user instruction, and query intent; cross-attention then captures relevance between the query context and passage representations. This design makes KaLM-Reranker-V1 efficient through decoupled passage encoding, yet not late interaction, by preserving rich relevance modeling through cross-attention. We ins
arXiv:2606.22511v2 Announce Type: replace Abstract: In open-ended generation, LLMs frequently fall into the "likelihood trap", marked by repetitive degeneration and vocabulary dullness, creating a discrepancy between machine-generated and human-written text. While post-hoc tail truncation (e.g., Top-$p$, Min-$p$) avoids sampling from the unreliable tail, it can over-sample from the uncalibrated head and misalign generation with human lexical preferences; fixed scalar repetition penalties likewise ignore variation in logit scale across inference steps, potentially disrupting semantic coherence. To address both limitations, we propose Variance-Calibrated Modulation (VCM), a training-free pre-decoding intervention that reshapes the probability distribution before truncation through two dynamic mechanisms: (1) Contextual Searchlight via PMI, which suppresses global stopwords while elevating context-evoked tokens, and (2) Adaptive Self-Debiasing, which uses real-time logit standard deviatio