EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18402 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

regulation Tue, 25 Aug 2026 14:30:00 +0000
The 74

With a State Cellphone Ban in Place, Michigan Schools Choose a Range of Options

Put it in your locker. Store it in a pouch. Keep it in your backpack. Hand it over to the principal. Leave it at home. With the state of Michigan enacting a cellphone ban this fall that restricts students from using their phones during instructional time, lawmakers have left it up to school districts to […]

Source ↗
technology Tue, 25 Aug 2026 14:08:11 -0400
EdTech Mag (Higher)

How Can AI Agents Support Lean Higher Ed Security Teams?

Growing cyberthreats and limited head count create ongoing challenges for security teams at colleges and universities. According to research from Nile, 94% of higher education IT teams are understaffed amid rising cybersecurity threats. Artificial intelligence could be a way to augment their security strategies to free up staff and resources. “For organizations with lean security teams, there is an exciting opportunity for AI to support some of the heavy lifting,” says Ramya Chitrakar, vice president of engineering at Google Cloud Security. “We are already seeing huge gains for…

Source ↗
technology Tue, 25 Aug 2026 13:08:00 +0000
MedCity News

Before You Sign That AI Contract: 7 Questions Every Healthcare CFO Should Ask

Here’s how finance leaders can evaluate AI investments in revenue cycle management before committing budget, and the risk exposure most vendor pitches leave out. The post Before You Sign That AI Contract: 7 Questions Every Healthcare CFO Should Ask appeared first on MedCity News .

Source ↗
regulation Tue, 25 Aug 2026 12:30:00 +0000
The 74

In South Dakota, a School’s Design Enhances Learning for Younger Kids

Rural schools are asked to solve national problems with local resources. They are expected to close readiness gaps they did not create, respond to poverty they cannot control, and provide opportunities students may not receive anywhere else. Yet these expectations are often placed inside buildings designed for a different era of education. The physical school […]

Source ↗
behavior Tue, 25 Aug 2026 12:16:11 +0000
District Admin

Trump admin touts efforts to shut down Education Department as school year begins

The "wins" the administration is celebrating include highlighting their efforts to shut the department down as well as progress on bolstering school choice. The post Trump admin touts efforts to shut down Education Department as school year begins appeared first on District Administration .

Source ↗
behavior Tue, 25 Aug 2026 12:11:55 +0000
District Admin

When is suspending students discrimination? Education Dept. renews an old fight

The department's Office for Civil Rights (OCR) issued new federal guidance warning school leaders that they should not consider a student's race in discipline matters or they would risk being investigated. The post When is suspending students discrimination? Education Dept. renews an old fight appeared first on District Administration .

Source ↗
regulation Tue, 25 Aug 2026 10:30:00 +0000
The 74

This Superintendent Couldn’t Read as a Kid. Now He Wants to Ensure Every Kid Can

In late July, soon-to-be kindergarteners spent the morning getting their wiggles out at a summer learning program at Robert Bailey Elementary School in Providence. At first, they shyly raised their hands or whispered “that’s me” when their teacher sang an attendance song, but then they jumped up and spent the next hour stretching, dancing and […]

Source ↗
behavior Tue, 25 Aug 2026 10:00:00 +0000
eSchool News

Beyond spotting deepfakes: Teaching multilingual learners to interrogate AI

Long before computers entered our homes and classrooms, photographs helped shape what people believed. Images carried a particular kind of authority: We did not simply look at them; we often treated them as evidence.

Source ↗
technology Tue, 25 Aug 2026 09:00:00 +0000
Tech & Learning

Screen Time Is the Wrong Unit of Measurement

Before restricting screens, schools should ask what students will lose, including access to books, translation, accessibility features, and more.

Source ↗
audience Tue, 25 Aug 2026 07:00:00 +0000
Inside Higher Ed

FIU Violated Student Free Speech Rights, Judge Rules

FIU Violated Student Free Speech Rights, Judge Rules kathryn.palmer… Tue, 08/25/2026 - 03:00 AM Byline(s) Kathryn Palmer

Source ↗
audience Tue, 25 Aug 2026 07:00:00 +0000
Inside Higher Ed

5 Ways Colleges Are Reshaping STEM Education

5 Ways Colleges Are Reshaping STEM Education Joshua.Bay Tue, 08/25/2026 - 03:00 AM Colleges are redesigning gateway courses, expanding hands-on learning and creating new career paths to help more students succeed in STEM. Byline(s) Joshua Bay

Source ↗
audience Tue, 25 Aug 2026 07:00:00 +0000
Inside Higher Ed

The Absurdly Racist Notion That Elite Institutions Routinely Hire A-Minus Black Faculty

The Absurdly Racist Notion That Elite Institutions Routinely Hire A-Minus Black Faculty Sara Brady Tue, 08/25/2026 - 03:00 AM Don’t believe the hype of Black intellectual inferiority. Byline(s) Shaun Harper

Source ↗
audience Tue, 25 Aug 2026 07:00:00 +0000
Inside Higher Ed

A Season of Change for Public Health Accreditor

A Season of Change for Public Health Accreditor Josh Moody Tue, 08/25/2026 - 03:00 AM The Council on Education for Public Health is changing language in its standards on racial disparity, prompting criticism from scholars. It also walked away from federal recognition. Byline(s) Josh Moody

Source ↗
audience Tue, 25 Aug 2026 07:00:00 +0000
Inside Higher Ed

University of Utah to Close Satellite Campus After More Than 25 Years

University of Utah to Close Satellite Campus After More Than 25 Years Olivia.sanchez Tue, 08/25/2026 - 03:00 AM Byline(s) Olivia Sanchez

Source ↗
audience Tue, 25 Aug 2026 07:00:00 +0000
Inside Higher Ed

Court Revives Lawsuit Filed by UK Prof Who Called for War on Israel

Court Revives Lawsuit Filed by UK Prof Who Called for War on Israel kathryn.palmer… Tue, 08/25/2026 - 03:00 AM Byline(s) Kathryn Palmer

Source ↗
audience Tue, 25 Aug 2026 07:00:00 +0000
Inside Higher Ed

Trump Administration Proposes New H-1B Visa Fee

Trump Administration Proposes New H-1B Visa Fee Katherine Knott Tue, 08/25/2026 - 03:00 AM Byline(s) Katherine Knott

Source ↗
audience Tue, 25 Aug 2026 07:00:00 +0000
Inside Higher Ed

Faculty Quit as Texas A&M Victoria Requires Them On Campus 5 Days a Week

Faculty Quit as Texas A&M Victoria Requires Them On Campus 5 Days a Week Emma Whitford Tue, 08/25/2026 - 03:00 AM The new policy, announced in June, is a response to a Texas law restricting remote work for public university faculty. The Texas AAUP says it has prompted chaos and tears. Byline(s) Emma Whitford

Source ↗
audience Tue, 25 Aug 2026 07:00:00 +0000
Inside Higher Ed

In Defense of Nathan Cofnas

In Defense of Nathan Cofnas Sara Brady Tue, 08/25/2026 - 03:00 AM We must not abandon academic freedom and punish anyone for being a harsh critic, even when there are tragic results. Byline(s) John K. Wilson

Source ↗
audience Tue, 25 Aug 2026 05:00:00 -0400
Higher Ed Dive

University of Utah to shutter Sandy location in January

Fall classes will continue as scheduled for the 660 students who have already signed up, the university said.

Source ↗
regulation Tue, 25 Aug 2026 05:00:00 -0400
K-12 Dive

The emerging promise of teacher apprenticeship programs

About four years into the effort to address teacher shortages, two recent studies reveal apprenticeships may be successful in recruiting local educators.

Source ↗
regulation Tue, 25 Aug 2026 05:00:00 -0400
K-12 Dive

3 special education funding strategies states can use to support districts

A K-12 funding expert with the American Institutes for Research says state special education funding systems need to be evidence-based and student-focused.

Source ↗
need Tue, 25 Aug 2026 05:00:00 +0000
Hechinger Report

California encourages undocumented students to go to college. But some are losing hope

FOLSOM, Calif. — At Folsom Lake College’s main administrative building, a flyer on the wall tells undocumented students, “You can still go to college in California even with the current political climate.” “Keep going,” it says. “You are not alone.” Nearby is Alicia Alejo’s office at the Undocu-Falcons Center, where the student support specialist helps undocumented […] The post California encourages undocumented students to go to college. But some are losing hope appeared first on The Hechinger Report .

Source ↗
behavior Tue, 25 Aug 2026 00:00:00 GMT
EdSurge

As Male Teachers Vanish From the Classroom, Schools Look for Ways to Bring Them Back

A new survey of 145 men in the profession points to mentorship, pay and a sense of belonging as the biggest levers for recruitment and retention.

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

Framework for Grounding Healthcare LLMs in a Causal Knowledge Graph: A Cardiovascular Example Pilot

arXiv:2608.15382v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interventions, mechanisms, harms, evidence, and uncertainty. We propose a reproducible, graph-centered evaluation framework for intervention-oriented LLM behavior in healthcare and stress-test it in a cardiovascular pilot. The framework has four components: (i) a domain causal knowledge graph in which assertions are first-class, provenance-preserving nodes with stable identifiers; (ii) a scenario-conditioned subgraph extraction step that, given any clinical scenario, retrieves the relevant reified-assertion subgraph; (iii) four controlled grounding conditions that vary how the retrieved subgraph is composed into the model's context (ungrounded C1, knowledge-graph C2, causal-graph C3, integrated C4); and (iv) an automated scoring pipeline, anchored on asserti

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

Withholding the Completing Chunk: Exact Release-Boundary Equivalence for Production Streaming Guardrails

arXiv:2608.10279v2 Announce Type: replace-cross Abstract: Streaming language-model output creates an enforcement boundary: a control that detects a prohibited pattern after releasing its completing chunk cannot recall it. We study a production policy in which each ordered family is the conjunction of two regular-language predicates. Incremental matching is classical. The problem is exact composition at release time across arbitrary chunk partitions, including end-of-prefix word boundaries that can change on extension. We define an ASCII-explicit policy grammar, compile each predicate to a persistent nondeterministic finite automaton (NFA), distinguish stable from provisional assertion state, apply document-order family priority, and check the decision before releasing each chunk. We show that the resulting monitor is release-boundary equivalent to an absorbing cumulative oracle for every policy in the declared grammar. Production Python and TypeScript implementations were evaluated on

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

The Greatness of Science Cannot Be Planned: Agentic Auto-Research is Fuzz Testing

arXiv:2608.09855v2 Announce Type: replace-cross Abstract: Agentic auto-research is emerging, but most systems treat scientific discovery as goal-oriented optimization against a final benchmark. This paradigm rewards a sparse final verdict and ignores the exploration that precedes it. When agents optimize only the final score, they overfit to the test conditions and sample blindly rather than search. Within a declared research problem, a research agent and a greybox fuzzer for software analysis face the same sparse feedback. A fuzzer rarely finds a bug directly, but coverage makes partial progress observable on every execution. Fuzzers use that dense signal to mutate inputs and allocate effort, rather than merely rank completed runs. Auto-research needs the same two capabilities. First, each experiment must expose a cheap, dense signal of epistemic progress before final scientific validation is available. Second, that signal must determine the next intervention so the agent searches rat

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

A JoLT for the KV cache: Near-lossless KV cache compression via joint Lagrangian allocation of Tucker ranks and a rotated residual for llms

arXiv:2607.12550v3 Announce Type: replace-cross Abstract: The key-value (KV) cache has become the dominant memory cost of transformer inference: it grows with batch size, context length, and depth, and at long context it, rather than the model weights, sets the throughput ceiling. Existing reductions fall into two families. Low-rank methods factor two-dimensional slices of the cache, either per-head matrices or cross-layer feature blocks, and quantization methods lower the bit-width of every entry. Neither exploits the fact that the cache at a layer is naturally a third-order tensor whose three axes, the heads, the tokens, and the features, carry very different amounts of redundancy. We take this tensor view directly. Our method, JoLT (Joint Lagrangian Tucker), applies a partial Tucker decomposition that compresses only the token and feature axes while leaving the head and layer axes intact, then restores the energy that truncation discards with a rotated low-bit residual: a random ort

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG

arXiv:2606.16661v2 Announce Type: replace-cross Abstract: Fixed-length chunking in Retrieval-Augmented Generation (RAG) often leads to boundary fragmentation, where critical evidence is split across segments, degrading retrieval recall. While static windowing and parent retrieval improve recall, they introduce significant token overhead. We propose SCAR (Semantic Continuity-Aware Retrieval), an adaptive retrieval policy that selectively expands neighboring chunks by weighing query-neighbor relevance against a structural continuity penalty. SCAR uses a relative expansion threshold tied to each retrieved chunk's own query-relevance, yielding an approximately scale-invariant decision rule that transfers across embedding models without recalibration. Across four diverse corpora (RFC, GDPR, a 10-K report, and a Merger agreement; N=320 queries; 160 boundary-fragmented), SCAR achieves 92.8% recall on boundary-fragmented queries with only 7.84 chunks, a 22.9% reduction compared to static windo

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

Last But Not Least: Boundary Attention CalibratiON for Multimodal KV Cache Compression

arXiv:2606.14782v3 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) achieve strong vision-language reasoning but incur large KV caches and high decoding latency with long visual contexts. Existing compression methods rely on observation window attention for stable token importance estimation, yet this aggregation can dilute sparse critical evidence and discard answer-relevant tokens under aggressive compression. We identify last query attention as a complementary signal for recovering such evidence, though its irrelevant signals may introduce additional noise. We propose BACON, a plug-and-play method that calibrates observation window attention with last query evidence while suppressing noise through intra-layer coherence and inter-layer persistence. Across diverse benchmarks, models, budgets, and compression methods, BACON improves multimodal KV-cache compression by 7.5% on average under the most aggressive budget, with gains up to 30.9%.

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning

arXiv:2606.11119v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is a promising approach for enhancing reasoning and agentic behavior in large language models. However, rollout-intensive policy optimization is often limited by insufficient reward contrast, arising when overly simple or complex prompts generate low-variance feedback and when outcome-only rewards assign the same terminal assessment to every decision in a multi-turn rollout. Past efforts have focused on allocating available rollout resources to promising prompts, yet they only leverage sample informativeness at the prompt level and neglect variation in prefix-level informativeness across turns within the same rollout. This work targets multi-turn agentic RL by modeling each ReAct-style thought-action-observation turn as a semantically distinct node, allowing budget allocation to extend from prompt roots to turn-level prefixes with further continuations, which naturally forms

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

SocraticPO: Policy Optimization via Interactive Guidance

arXiv:2606.09887v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) for large language models usually supervises reasoning with scalar outcome rewards, such as binary correctness. Such rewards provide an optimization direction but rarely explain how a model should revise its mistaken reasoning, which can encourage shortcut learning and brittle policies. We propose \textbf{SocraticPO} (Socratic Policy Optimization), a policy-optimization framework that augments RL rollouts with Socratic-style natural-language guidance. During rollout, the student first answers independently; if the answer is incorrect, a teacher diagnoses the attempt and provides concise corrective guidance, after which the student continues under the expanded context. Crucially, this guidance is paired with reward decay: correct answers obtained after teacher intervention only receive decayed rewards, preventing the policy from treating teacher help as a free path to reward. Since SocraticPO only modi

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models

arXiv:2606.00566v2 Announce Type: replace-cross Abstract: As language models take on agentic roles that call APIs, read tool outputs, and act on third-party content, their attack surface expands beyond what users type. Whether they treat a malicious instruction the same way regardless of where it arrives has not been studied systematically. We introduce the Safety Asymmetry Score (SAS), measuring how a model's susceptibility to adversarial content shifts depending on whether it arrives in the user message, tool metadata, or tool output, using matched payload pairs that hold the malicious text identical and vary only the channel. Across 10 production LLMs and three attack families, general-purpose models sharply discount instructions arriving as tool metadata relative to identical instructions in the user message, while agent-native models discount them far less. This differential survives an affordance-matched control equalizing tool availability and scoring, and a size-controlled mixe

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

GIM: Evaluating models via tasks that integrate multiple cognitive domains

arXiv:2605.18663v2 Announce Type: replace-cross Abstract: As LLM benchmarks saturate, the evaluation community has pursued two strategies to increase difficulty: escalating knowledge demands (GPQA, HLE) or removing knowledge entirely in favor of abstract reasoning (ARC-AGI). The first conflates memorization with capability; the second divorces reasoning from the practical contexts in which it matters. We take a different approach. The Grounded Integration Measure (GIM) is a benchmark of 820 original problems (615 public, 205 private) where difficulty comes from integration; individual problems require coordinating multiple cognitive operations (constraint satisfaction, state tracking, epistemic vigilance, audience calibration) over broadly accessible knowledge, so that reasoning stays grounded in realistic tasks without being gated on specialized expertise. Each problem is an original expert-authored composition, majority with rubric-decomposed scoring. We calibrate a judge-aware conti

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening

arXiv:2605.17610v2 Announce Type: replace-cross Abstract: The rapid growth of online video platforms and AI-generated content has made reliable video guardrails a key challenge for safety and real-world deployment. While most videos can be screened through fast pattern recognition, a small subset requires deeper reasoning over temporally complex content and nuanced policy constraints. Existing approaches typically rely on large vision-language models applied uniformly across all inputs, resulting in high inference costs and inefficient allocation of computation. We propose SafeLens, a video guardrail framework that introduces a fast-and-slow inference architecture for efficient and accurate content moderation with variable computational cost across inputs. Additionally, we construct a high-quality dataset by applying influence-guided filtering to the SafeWatch Dataset, retaining only 2.4% of the original data. To further address limitations of training-time scaling, we enable test-time

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling

arXiv:2604.25098v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) now exhibit remarkable reasoning capabilities through test-time compute scaling (TTS), with impressive performance across math and coding benchmarks. In parallel, research in model compression has developed pruning methods that seek to remove redundant/detrimental parameters without sacrificing task performance. The intersection of these two research advancements lays the foundation for our work. Specific to reasoning LLMs, prior work has shown that structured pruning (methods which remove entire set of layer blocks), significantly degrades TTS reasoning performance. However, in this work, we revisit this assumption and investigate whether unstructured pruning (methods that carefully remove only certain redundant/detrimental weights) exhibits similar limitations. Surprisingly, our extensive experiments across four reasoning benchmarks on two reasoning LLMs: s1.1-7B and Qwen3-8B, consistently show tha

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

Detecting and Suppressing Reward Hacking with Gradient Fingerprints

arXiv:2604.16242v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) typically optimizes for outcome rewards without imposing constraints on intermediate reasoning. This leaves training susceptible to reward hacking, where models exploit loopholes (e.g., spurious patterns in training data) in the reward function to achieve high scores without solving the intended task. These reward-hacking behaviors are often implicit, as the intermediate chain-of-thought (CoT) may appear plausible on the surface, limiting the effectiveness of purely text-based monitoring. We propose Gradient Fingerprint (GRIFT), a method for detecting reward hacking using models' internal computations. Given a prompt and a model-generated CoT, GRIFT computes gradients of the CoT conditioned on the prompt and compresses them into a compact representation, which is then used to assess whether the CoT reflects reward hacking behavior. Across verifiable reasoning benchmarks spann

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation

arXiv:2604.09746v2 Announce Type: replace-cross Abstract: As large language models (LLMs) are increasingly deployed as autonomous agents, understanding how strategic behavior emerges in multi-agent environments has become an important alignment challenge. We take a neutral empirical stance and construct a controlled environment in which strategic behavior can be directly observed and measured. We introduce a large-scale multi-agent simulation in a simplified model of New York City, where LLM-driven agents interact under opposing incentives. Blue agents aim to reach their destinations efficiently, while Red agents attempt to divert them toward billboard-heavy routes using persuasive language to maximize advertising revenue. Hidden identities make navigation socially mediated, forcing agents to decide when to trust or deceive. We study policy learning through an iterative simulation pipeline that updates agent policies across repeated interaction rounds using Kahneman-Tversky Optimizatio

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

PRAGMA: Revolut Foundation Model

arXiv:2604.08649v2 Announce Type: replace-cross Abstract: Modern financial systems generate vast quantities of transactional and event-level data that encode rich economic signals. This paper presents PRAGMA, a family of foundation models for banking event sequences. Our approach pre-trains a Transformer-based architecture with masked modelling on a large-scale, heterogeneous banking event corpus using a self-supervised objective tailored to the discrete, variable-length nature of financial records. The resulting model supports a wide range of downstream tasks such as credit scoring, fraud detection, and lifetime value prediction: strong performance can be achieved by training a simple linear model on top of the extracted embeddings and can be further improved with lightweight fine-tuning. Through extensive evaluation on downstream tasks, we demonstrate that PRAGMA achieves superior performance across multiple domains directly from raw event sequences, providing a general-purpose repre

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

Safety Training May Persist Through Helpfulness Optimization in LLM Agents

arXiv:2603.02229v2 Announce Type: replace-cross Abstract: Safety post-training has been studied extensively in single-step "chat" settings where safety typically refers to refusing harmful requests. We study an "agentic" (i.e., multi-step, tool-use) setting where safety refers to harmful actions directly taken by the LLM. We investigate the effects of using direct preference optimization (DPO) to optimize safety and/or helpfulness on the ToolEmu agentic benchmark. First, we find that safety training largely persists through subsequent helpfulness training. Second, we find a consistent negative linear correlation ($R^2 = 0.77$) between safety and helpfulness when considering all training configurations together. Even post-training on both metrics simultaneously simply results in another point on the same trend line rather than yielding a "best of both worlds" strategy, despite the presence of such strategies in our dataset. Overall, our findings underscore the need for a better understa

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

Vibe Coding on Trial: Operating Characteristics of Unanimous LLM Juries

arXiv:2602.18492v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are now good enough at coding that developers can describe intent in plain language and let the tool produce the first code draft, a workflow increasingly built into tools like GitHub Copilot, Cursor, and Replit. What is missing is a reliable way to tell which model written queries are safe to accept without sending everything to a human. We study the application of an LLM jury to run this review step. We first benchmark 15 open models on 82 MySQL text to SQL tasks using an execution grounded protocol to get a clean baseline of which models are strong. From the six best models we build unanimous committees of sizes 1 through 6 that see the prompt, schema, and candidate SQL and accept it only when every member says it is correct. This rule matches safety first deployments where false accepts are more costly than false rejects. We measure true positive rate, false positive rate and Youden J and we also

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

You Need Better Attention Priors

arXiv:2601.15380v2 Announce Type: replace-cross Abstract: We generalize the attention mechanism by viewing it through the lens of Entropic Optimal Transport, revealing that standard attention corresponds to a transport problem regularized by an implicit uniform prior. We introduce Generalized Optimal transport Attention with Trainable priors (GOAT), a new attention mechanism that replaces this naive assumption with a learnable, continuous prior. This prior maintains full compatibility with optimized kernels such as FlashAttention. GOAT also provides an EOT-based explanation of attention sinks and materializes a solution for them, avoiding the representational trade-offs of standard attention. Finally, by absorbing spatial information into the core attention computation, GOAT learns an extrapolatable prior that combines the flexibility of learned positional embeddings with the length generalization of fixed encodings.

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

Training Proactive and Personalized LLM Agents

arXiv:2511.02208v2 Announce Type: replace-cross Abstract: Despite rapid progress, current AI agents are primarily optimized for isolated task completion. We argue for a paradigm shift toward training agents as collaborators that communicate and adapt to people. To facilitate this shift in real-world complex applications, we first formalize three dimensions of collaborative AI agents: Productivity, Proactivity, and Personalization (PPP). We introduce UserVille, an interactive environment with configurable LLM-based user simulators and user-centric feedback to evaluate these dimensions, and propose a multi-objective reinforcement learning framework that optimizes them using rewards from task outcomes, question effort, and preference adherence. On two real-world agentic tasks (SWE-Bench and BrowseComp-Plus), PPP-trained agents outperform strong LLM baselines (including GPT-5) by an average of 16.7 points, ask more targeted questions, and generalize to unseen preferences and tasks. A follo

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning

arXiv:2510.17021v2 Announce Type: replace-cross Abstract: Large language model (LLM) unlearning is a key approach for removing undesired data, knowledge, or behaviors from pretrained models while retaining their general utility. Yet, with the rise of open-weight LLMs, we ask: can the unlearning process itself be backdoored, appearing successful under normal conditions yet reverting to pre-unlearned behavior when a hidden trigger is activated? Drawing inspiration from classical backdoor attacks that embed triggers into training data to enforce specific behaviors, we investigate backdooring unlearning, a setting in which models forget as intended in the clean setting but recover forgotten knowledge when the trigger appears. We show that designing such attacks presents unique challenges, hinging on where triggers are placed and how backdoor training is reinforced. We uncover a strong link between the backdoor efficacy and the attention sink phenomenon (i.e., shallow input tokens consisten

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

Beyond Benchmarks: LLM Evaluation with an Anthropomorphic and Lifecycle-oriented Roadmap

arXiv:2508.18646v3 Announce Type: replace-cross Abstract: Despite their rapid advancement, large language models (LLMs) suffer from a critical disconnect between benchmark scores and real-world utility. Current evaluation remains fragmented, prioritizing isolated technical metrics over the holistic, developmental, and societal aspects essential for deployment. Rather than serving merely as a descriptive catalog, this work establishes a diagnostic ontology that causally maps evaluation dimensions to the canonical LLM training pipeline, transforming evaluation from static ranking into a diagnostic tool for root-cause analysis. In this paper, we introduce an anthropomorphic evaluation framework that re-conceptualizes LLM capabilities through a four-dimensional lens: Intelligence Quotient (IQ), Professional Quotient (PQ), Emotional Quotient (EQ), and Value-oriented Quotient (VQ). We operationalize these concepts through a modular evaluation architecture and validate the framework's diagnos

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora

arXiv:2507.09924v2 Announce Type: replace-cross Abstract: Continually updating model-based indexes in generative retrieval with new documents remains challenging, as full retraining is computationally expensive and impractical under resource constraints. We propose MixLoRA-DSI, a novel framework that combines an expandable mixture of Low-Rank Adaptation experts with a layer-wise out-of-distribution (OOD)-driven expansion strategy. Instead of allocating new experts for each new corpus, our proposed expansion strategy enables sublinear parameter growth by selectively introducing new experts only when significant number of OOD documents are detected. Experiments on NQ320k and MS MARCO Passage demonstrate that MixLoRA-DSI outperforms full-model update baselines, with minimal parameter overhead and substantially lower training costs.

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

Scaling Electronic Health Record Foundation Models for Population Health Management

arXiv:2506.00209v3 Announce Type: replace-cross Abstract: Population health management requires scalable methods to identify individuals at risk of chronic diseases such as cardiovascular conditions and cancer, yet existing approaches rely on fragmented data and resource-intensive screening. We present Scaling Electronic Health Record Foundation Models for Population Health Management, an Electronic Health Record Foundation Model that performs large-scale chronic disease prediction using cross-site longitudinal medical records. We pretrain Scaling Electronic Health Record Foundation Models for Population Health Management on billions of medical events from over 5 million patients across Taiwan and the United States, leveraging a unified code alignment framework to address cross-system heterogeneity, and characterize its scaling behavior via IsoFLOP analysis, training compute-optimal models up to 2.4B parameters. Across 11 chronic disease prediction tasks, Scaling Electronic Health Reco

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

Unleashing the Power of LLMs in Dense Retrieval with Query Likelihood Modeling

arXiv:2504.05216v4 Announce Type: replace-cross Abstract: Dense retrieval is a crucial task in Information Retrieval (IR), serving as the basis for downstream tasks such as re-ranking and augmenting generation. Recently, large language models (LLMs) have demonstrated impressive semantic understanding capabilities, making them attractive to researchers focusing on dense retrieval. While LLMs, as decoder-style generative models, excel in language generation, they often fall short in modeling global information due to a lack of attention to subsequent tokens. Drawing inspiration from the classical word-based language modeling approach for IR, specifically the query likelihood (QL) model, we aim to leverage the generative strengths of LLMs through QL maximization. Rather than employing QL estimation for document ranking, we propose an auxiliary task of QL maximization to enhance the backbone for subsequent contrastive learning of the retriever. We introduce our model, LLM-QL, which incorpo

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

Benchmarking and Boosting Multilingual Capabilities of LVLMs via OCR-Centric Reinforcement Learning

arXiv:2503.18484v3 Announce Type: replace-cross Abstract: Evaluating the multilingual capabilities of Large Vision-Language Models (LVLMs) remains challenging because most benchmarks rely on non-parallel corpora, making it unclear whether cross-lingual performance gaps reflect model limitations or dataset inconsistencies. To address this, we introduce PM4Bench, the first multimodal, multilingual, multi-task benchmark built on a strictly parallel 10-language corpus, enabling fair, apples-to-apples cross-lingual comparison of model performance. We further introduce a vision setting that embeds textual inputs directly into images, better approximating deployment scenarios where LVLM-driven agents interact with virtual or physical environments through unified visual observations. Experiments with 10 LVLMs reveal that OCR is a key factor behind cross-lingual disparity when textual content is rendered visually. Motivated by this, we design an OCR-centric GRPO training strategy using fully sy

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

NeST: Neighborhood-aware semantic alignment and temporal modulation for LLM based time series forecasting

arXiv:2412.04806v2 Announce Type: replace-cross Abstract: Adapting Large Language Models (LLMs) trained on discrete text data, to forecast continuous time series signals is challenging. While finetuning the LLMs enables such adaptation, effectively integrating both textual and time series information in the prompt is critical. Current LLM-based time series forecasting methods combine the two modalities through simple concatenation or parameter heavy cross-attention. Moreover, existing methods embed time series data using decomposition techniques that may inadequately capture complex temporal dynamics. To address these limitations, we propose neighborhood-aware semantic alignment and temporal modulation based framework (NEST) to formulate a new text-integrated time series prompt to finetune the LLM. First, we generate neighborhood-aware text prototypes that are optimized to represent local neighborhoods of pretrained word token embeddings of the LLM. Second, we align them with temporal

Source ↗
technology Tue, 25 Aug 2026 00:00:00 -0400
arXiv cs.CL

Time-Series Retrieval for Grounding Multimodal Language Models in Remaining Useful Life Prediction

arXiv:2608.19218v2 Announce Type: replace Abstract: Large language models (LLMs) and agentic AI systems are increasingly being explored for domain-specific maintenance and prognostics tasks, raising the question of whether they can effectively support prognostics and health management (PHM). In this paper, we investigate remaining useful life (RUL) estimation with multimodal large language models (MLLMs) grounded through time-series retrieval. We propose a framework in which historically similar degradation segments are retrieved from the training set and, together with the test trajectory, transformed into a visual comparison artifact that is processed by the MLLM through a structured multimodal prompt. The approach is evaluated on the FD001 partition of the C-MAPSS benchmark under repeated experiments comparing retrieval-based inference against a non-retrieval baseline based on random reference selection. The results show that time-series retrieval consistently improves MLLM-based RU

Source ↗
Showing 3151–3200 of 18402 signals
← Prev Page 64 of 369 Next →