EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18402 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

regulation Thu, 27 Aug 2026 14:30:00 +0000
The 74

Beyond Policy: Why Literacy Reform Is Entering a New Phase

Over the past several years, states across the country have rewritten early literacy policy at an unprecedented pace. Science of reading legislation, teacher training requirements, instructional material reforms, and early intervention initiatives have reshaped the national literacy landscape. Today, nearly every state has taken meaningful steps to strengthen reading instruction. But passing legislation was intended […]

Source ↗
behavior Thu, 27 Aug 2026 13:44:22 +0000
District Admin

Pa. may ban cellphones in schools. Some educators say they’re already seeing results—without pouches

The “bell-to-bell” bans face skepticism from students and parents. After going phone-free, early adopters say the difference is striking. The post Pa. may ban cellphones in schools. Some educators say they’re already seeing results—without pouches appeared first on District Administration .

Source ↗
technology Thu, 27 Aug 2026 13:34:00 +0000
MedCity News

Precision Medicine Transformed Oncology —It’s Time to Do the Same for Sepsis

It’s time we reignite treatment development for sepsis and other immune-mediated conditions that strike when patients are most vulnerable. The post Precision Medicine Transformed Oncology —It’s Time to Do the Same for Sepsis appeared first on MedCity News .

Source ↗
technology Thu, 27 Aug 2026 12:41:10 -0400
EdTech Mag (Higher)

Why Higher Ed CIOs Should Embrace a Federated Data Governance Strategy

In higher education, the standard advice on data governance has been pretty simple for a long time: Build a council, centralize decisions and route everything through IT. For a while, that works. You stand up a committee, you write a charter, you centralize definitions and approvals. And then, quietly, the model stops working. IT becomes the bottleneck for every data question on campus. You’re chasing report definitions, field names and artificial intelligence (AI)-related concerns one request at a time. The council still meets, but governance isn’t really operating day to day. The problem…

Source ↗
regulation Thu, 27 Aug 2026 12:30:00 +0000
The 74

Opinion: What Districts Actually Need to Implement College & Career Pathways at Scale

In the past five years, there has been a surge in states playing a prominent role in improving outcomes for K-12 students. They’ve adopted evidence-based literacy policies, focused attention on math, and invested public dollars to sustain high-dosage tutoring. But gaining perhaps the most momentum across the country recently is state attention on expanding career […]

Source ↗
behavior Thu, 27 Aug 2026 12:19:29 -0700
Nurse.org

APRN Burnout and Workforce Recovery: What NCSBN’s Latest Research Tells Us

In this contributor piece, Brendan Martin, PhD, Director of Research at the National Council of State Boards of Nursing (NCSBN), examines what his organization's 2022 and 2024 National Nursing Workforce…

Source ↗
technology Thu, 27 Aug 2026 11:30:00 +0000
MedCity News

What Do the Latest Iteration of Consumer Drug Platforms Offer?

Consumer drug platforms and how they fit into the rise of the consumer in healthcare will be part of the conversation at INVEST Digital Health, scheduled for October 29 in Dallas. Register today! The post What Do the Latest Iteration of Consumer Drug Platforms Offer? appeared first on MedCity News .

Source ↗
regulation Thu, 27 Aug 2026 10:30:00 +0000
The 74

Veteran Educator Robert Franklin Wins GOP Runoff for Oklahoma Ed Chief

Robert Franklin, a retired Oklahoma educator, was enjoying life as a grandfather and giving lectures on the politics of education when he decided to become a candidate himself. Now he’s likely to be the next state superintendent after defeating teacher and conservative pastor James Taylor in the Republican runoff Tuesday. “I thought to myself, ‘If […]

Source ↗
technology Thu, 27 Aug 2026 10:01:30 -0400
EdTech Mag (K-12)

How To Run a Successful K–12 Ed Tech Pilot Program

Adopting new technology across a school district is a significant investment that’s hard to reverse if the rollout misses the mark. A pilot program can mitigate the risks by first launching a small-scale, time-limited test of a technology tool with a defined group of users before committing to full schoolwide or districtwide adoption. “A pilot helps districts validate that a solution will work in their unique environment and determine if it will integrate with the rest of their ed tech ecosystem and instructional practices,” says Kris Astle, learning and adoption manager at SMART Technologies…

Source ↗
behavior Thu, 27 Aug 2026 10:00:00 +0000
eSchool News

Students won’t remember every lesson, but here’s what’ll stick

For a lot of students, school can feel like a transaction: You do the work, you get the check mark, and then you move on. That model can unintentionally cap what students show us about what they know and can do.

Source ↗
technology Thu, 27 Aug 2026 09:00:00 +0000
Tech & Learning

What My Freshman Daughter Taught Me About AI

Students are already observing how AI is affecting their peers, their classrooms, and their own sense of what counts as authentic work.

Source ↗
behavior Thu, 27 Aug 2026 09:00:00 +0000
Getting Smart

Core Plus Explore: How Embedded Career Relevance Improves Engagement in Foundational Courses

What if the answer to student disengagement lived inside the curriculum itself? A new partnership between Edmentum and Roadtrip Nation embeds short, unscripted career exploration interviews directly into core high school courses, connecting algebra, biology, and English to real professionals and real possibilities. With data showing fewer absences, greater relevance, and students discovering careers they never knew existed, this model reframes a familiar question: school is not just preparation for graduation. It is preparation for possibility. Education leaders will want to see how this approach works, and what it signals about the future of instructional design. The post Core Plus Explore: How Embedded Career Relevance Improves Engagement in Foundational Courses appeared first on Getting Smart .

Source ↗
audience Thu, 27 Aug 2026 07:00:00 +0000
Inside Higher Ed

Postdocs Aren’t Burger Flippers: Why Unionization Is a Bad Idea

Postdocs Aren’t Burger Flippers: Why Unionization Is a Bad Idea Elizabeth Redden Thu, 08/27/2026 - 03:00 AM Unionization disrupts the bespoke mentor-mentee relationships that are central to the advancement of science. Byline(s) Evan D. Morris

Source ↗
audience Thu, 27 Aug 2026 07:00:00 +0000
Inside Higher Ed

AI Tends to Mark Students’ Essays Higher Than Humans, Study Shows

AI Tends to Mark Students’ Essays Higher Than Humans, Study Shows Susan H. Greenberg Thu, 08/27/2026 - 03:00 AM LLMs cannot be relied on to give accurate indication of student performance, researchers find, as universities explore ways to relieve pressure on graders. Byline(s) Juliette Rowsell for Times Higher Education

Source ↗
audience Thu, 27 Aug 2026 07:00:00 +0000
Inside Higher Ed

Wexner Steps Down as Ohio State Medical Center Board Chair Amid Epstein Lawsuits

Wexner Steps Down as Ohio State Medical Center Board Chair Amid Epstein Lawsuits Emma Whitford Thu, 08/27/2026 - 03:00 AM Byline(s) Emma Whitford

Source ↗
audience Thu, 27 Aug 2026 07:00:00 +0000
Inside Higher Ed

DOD Demands 30 American Universities Audit Foreign Partnerships

DOD Demands 30 American Universities Audit Foreign Partnerships jessica.blake@… Thu, 08/27/2026 - 03:00 AM Byline(s) Jessica Blake

Source ↗
audience Thu, 27 Aug 2026 07:00:00 +0000
Inside Higher Ed

Behind the Scenes of the College Presidency

Behind the Scenes of the College Presidency sara.custer@in… Thu, 08/27/2026 - 03:00 AM For three years, The Sandbox has revealed a side of higher education leadership most of us will never see. Byline(s) Sara Custer

Source ↗
audience Thu, 27 Aug 2026 07:00:00 +0000
Inside Higher Ed

Arkansas Baptist College Board Terminates President

Arkansas Baptist College Board Terminates President Susan H. Greenberg Thu, 08/27/2026 - 03:00 AM Byline(s) Susan H. Greenberg

Source ↗
audience Thu, 27 Aug 2026 07:00:00 +0000
Inside Higher Ed

New Paper Takes Aim at ‘Meritocratic’ Admissions Policies

New Paper Takes Aim at ‘Meritocratic’ Admissions Policies kathryn.palmer… Thu, 08/27/2026 - 03:00 AM Wealthy students with high standardized test scores make up the majority of students at the most selective universities, which also spend far more on student instruction than less selective institutions. So where does that leave everyone else? Byline(s) Kathryn Palmer

Source ↗
audience Thu, 27 Aug 2026 07:00:00 +0000
Inside Higher Ed

Commencement Speakers’ Politics Vary by Institution Type

Commencement Speakers’ Politics Vary by Institution Type Johanna Alonso Thu, 08/27/2026 - 03:00 AM A new study from researchers at Georgia State and New York Universities finds that commencement speakers skew liberal. But the type and location of the institution can make a difference in who is invited to speak. Byline(s) Johanna Alonso

Source ↗
audience Thu, 27 Aug 2026 07:00:00 +0000
Inside Higher Ed

Transfer Isn’t Supposed to Be This Hard

Transfer Isn’t Supposed to Be This Hard Joshua.Bay Thu, 08/27/2026 - 03:00 AM In this week’s Voices of Student Success episode, Aspen’s Josh Wyner and CCRC’s John Fink discuss how colleges can help more community college students transfer and earn bachelor’s degrees. Byline(s) Joshua Bay

Source ↗
audience Thu, 27 Aug 2026 07:00:00 +0000
Inside Higher Ed

The Fight for Fair Funding at One Historically Black Land-Grant University

The Fight for Fair Funding at One Historically Black Land-Grant University Olivia.sanchez Thu, 08/27/2026 - 03:00 AM A new book from the former president of Tennessee State University reveals her battle to regain $544 million in documented underfunding. Byline(s) Olivia Sanchez

Source ↗
need Thu, 27 Aug 2026 05:30:00 +0000
Hechinger Report

Who gets held back? Black boys, by a 3-1 ratio compared to white boys

More than a decade of data from schools across the U.S. shows stark divides in who is held back: Boys are more likely than girls to repeat a grade, and Black boys are held back at nearly three times the rate of white boys. Overall, schools are holding back fewer students than they were more […] The post Who gets held back? Black boys, by a 3-1 ratio compared to white boys appeared first on The Hechinger Report .

Source ↗
audience Thu, 27 Aug 2026 05:00:00 -0400
Higher Ed Dive

Anna Maria bankruptcy sale of property faces objection from neighbor

The Roy family says the now-shuttered college promised to gift them a 2.4-acre parcel. The institution is disputing their property claim.

Source ↗
regulation Thu, 27 Aug 2026 05:00:00 -0400
K-12 Dive

Only 2 large school districts are without fiscal ‘red flags,’ analysis finds

The nation’s 100 largest districts averaged 2.4 out of 8 concerning indicators — the worst performance at any government level, the Reason Foundation said.

Source ↗
regulation Thu, 27 Aug 2026 05:00:00 -0400
K-12 Dive

HHS to create national autism ‘elopement’ alert

Advocacy groups and researchers suggest measures to prevent children from wandering away from school or caregivers and into harm's way.

Source ↗
behavior Thu, 27 Aug 2026 00:00:00 GMT
EdSurge

Student Artists Wrestle with AI’s Promise and Peril

Four student artists map out what artificial intelligence means for their communities today — and 50 years from now.

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

A Layered Security Framework Against Prompt Injection in RAG-Based Chatbots

arXiv:2606.19660v2 Announce Type: replace-cross Abstract: Prompt injection is ranked as the most critical vulnerability in large language model (LLM) deployments by the OWASP Top 10 for LLM Applications, yet existing defenses operate at isolated pipeline stages and remain incomplete. Input filters cannot inspect retrieved documents, while output monitors cannot prevent malicious payloads from reaching the model. Consequently, retrieval-augmented generation (RAG) chatbots remain vulnerable to indirect injection, where a poisoned knowledge-base document compromises every user whose query retrieves it. We present a three-layer framework that intercepts both direct and indirect prompt injection throughout the inference pipeline. Layer 1 screens user input using a rule-based pattern library and a fine-tuned semantic anomaly classifier. Layer 2 enforces a provenance-based instruction hierarchy during context assembly, preventing retrieved content from overriding operator policy. Layer 3 audi

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

Amplifying, Not Learning: The Price of Out-of-Distribution Generalization in AI-Text Detection

arXiv:2605.21653v2 Announce Type: replace-cross Abstract: AI-text detectors gate decisions in education, hiring, and publishing, yet they flag the most fluent, formal human writing as machine-generated: they rate the median formal-native human essay as 99.5% likely AI while clearing genuine high-temperature AI at 10.5%. Deployed detectors share it (chatgpt-detector-roberta flags 56% of formal essays at a 1% false-alarm rate). This is not a calibration bug but the signature of one mechanism: a fine-tuned detector does not learn an AI-versus-human boundary, it amplifies an inherited typicality axis (predictability under a language model) that pre-exists fine-tuning, rescaling it rather than constructing one. Decomposing the detector into this inherited reading and a fine-tuned residual, the inherited part carries the bulk of cross-generator transfer and produces the over-flagging of formal humans, while the residual is generator-specific and does not transfer; a frozen-representation pro

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

Reconstruction of Personally Identifiable Information from Proprietary Data in Supervised Fine-Tuned Models

arXiv:2605.12264v2 Announce Type: replace-cross Abstract: Supervised Finetuning (SFT) has become one of the primary methods for adapting a large language model (LLM) with extensive pre-trained knowledge to domain-specific, instruction-following tasks. SFT datasets, composed of instruction-response pairs, often include user-provided information that may contain sensitive data such as personally identifiable information (PII), raising privacy concerns. This paper studies the problem of targeted PII reconstruction from models fine- tuned on proprietary SFT data, in which an adversary attempts to recover PII associated with a specific identity. We construct multi-turn, user-centric Q&A datasets in sensitive domains, specifically medical and legal settings, that incorporate PII to enable realistic evaluation of leakage. We then propose COVA, a coverage-aware decoding algorithm for targeted PII reconstruction under prefix-based attacks. Using COVA, we study how the amount of information avai

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

Cubit: Token Mixer with Kernel Ridge Regression

arXiv:2605.06501v3 Announce Type: replace-cross Abstract: Since its introduction in 2017, the Transformer has become one of the most widely adopted architectures in modern deep learning. Despite extensive efforts to improve positional encoding, attention mechanisms, and feed-forward networks, the core token-mixing mechanism in Transformers remains attention. In this work, we show that the attention module in Transformers can be interpreted as performing Nadaraya-Watson regression, where it computes similarities between tokens and aggregates the corresponding values accordingly. Motivated by this perspective, we propose Cubit, a potential next-generation architecture that leverages Kernel Ridge Regression (KRR), while the vanilla Transformer relies on Nadaraya-Watson regression. Specifically, Cubit modifies the classical attention computation by incorporating the closed-form solution of KRR, combining value aggregation through kernel similarities with normalization via the inverse of th

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

Corpus2Skill: Distilling Enterprise Knowledge into Navigable Agent Skills for QA and RAG

arXiv:2604.14572v4 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) grounds LLM responses in external evidence but treats the model as a passive consumer of search results, with no view of how the corpus is organized or what it has not yet seen. We present Corpus2Skill, a system-level retrieval architecture for bounded, structurally coherent corpora such as enterprise knowledge bases: an offline compiler distills the corpus into a hierarchical skill directory, and at serve time an LLM agent navigates it, drilling from a bird's-eye view through progressively finer summaries down to documents and backtracking when a branch is unproductive. On an enterprise customer-support benchmark, Corpus2Skill improves both answer quality and grounding over single-shot dense, hybrid, hierarchical-retrieval, and agentic RAG baselines at a moderate cost tradeoff, and the lead persists under encoder-matched controls and paired significance tests. An eleven-dataset study shows t

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

MMEmb-R1: Reasoning-Enhanced Multimodal Embedding with Pair-Aware Selection and Adaptive Control

arXiv:2604.06156v2 Announce Type: replace-cross Abstract: MLLMs have been successfully applied to multimodal embedding tasks, yet their generative reasoning capabilities remain underutilized. Directly incorporating chain-of-thought reasoning into embedding learning introduces two fundamental challenges. First, structural misalignment between instance-level reasoning and pairwise contrastive supervision may lead to shortcut behavior, where the model merely learns the superficial format of reasoning. Second, reasoning is not universally beneficial for embedding tasks. Enforcing reasoning for all inputs may introduce unnecessary computation and latency, and can even obscure salient semantic signals for simple cases. To address these issues, we propose MMEmb-R1, an adaptive reasoning-based multimodal embedding framework. We formulate reasoning as a latent variable and introduce pair-aware reasoning selection that employs counterfactual intervention to identify reasoning paths beneficial fo

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

Align Then Adapt: Label-Efficient Adapter Learning for Asymmetric Dense Retrieval

arXiv:2604.03403v2 Announce Type: replace-cross Abstract: Dense retrieval systems increasingly face an asymmetry between complex instruction-like queries and relatively simple, static document collections. While stronger embedders can better understand such queries, re-embedding large corpora or fine-tuning large models is often impractical. We propose Efficient Retrieval Adapter (ERA), a query-side adapter learning framework for re-index-free retrieval adaptation. ERA first aligns the embedding spaces of a strong query embedder and a lightweight document embedder using unlabeled corpus documents, and then adapts the aligned query representation with a small number of labeled query-document pairs. Across 126 MAIR retrieval tasks from six domains, ERA improves average nDCG@10 by up to 8.2 points in symmetric settings and by more than 12 points in asymmetric settings, while using substantially fewer labels than supervised adapter training. These results show that retrieval systems can be

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

Mitigating LLM biases toward spurious social contexts using direct preference optimization

arXiv:2604.02585v3 Announce Type: replace-cross Abstract: LLMs are increasingly used for high-stakes decision-making, yet their sensitivity to spurious context can introduce harmful biases. This is a critical concern when models are deployed for tasks like evaluating teachers' instructional quality, where biased assessment can affect teachers' professional development and career. We investigate model robustness to spurious social contexts about teachers using the largest publicly available dataset of U.S. classroom transcripts (NCTE) paired with expert evaluation scores. Evaluating seven frontier and open-weight models across seven categories of spurious contexts -- including teacher experience, education level, demographic identity, and sycophancy-inducing framings -- we find that irrelevant contexts can shift model-generated ratings by up to 1.48 points on a 7-point scale. Prompt-based mitigations and popular post-training methods, such as Supervised Fine-Tuning (SFT) and Direct Pref

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

OpenSanctions Pairs: Large-Scale Entity Matching with LLMs

arXiv:2603.11051v2 Announce Type: replace-cross Abstract: We release OpenSanctions Pairs, the first large-scale public benchmark for entity matching on sanctions and OSINT data. The dataset includes 755,540 expert-labeled pairs over 1 million entities, aggregated from 293 source datasets across 45 jurisdictions. It captures real-world diversity in compliance data, spanning multiple languages and writing systems (e.g., Latin, Cyrillic, Arabic), inconsistent structure, and time-varying provenance, and is substantially more heterogeneous than prior entity matching benchmarks. As baselines, we evaluate the production rule-based matcher (nomenklatura RegressionV1) alongside open- and closed-source LLMs in both zero- and few-shot settings, each tested with and without MIPROv2 prompt optimization to control for prompt sensitivity. The rule-based baseline reaches 91.3\% F1; GPT-4o achieves the best result at 99.0\% F1, and a locally deployable open-source model (DeepSeek-R1-Distill-Qwen-14B) a

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

Tool Verification for Test-Time Reinforcement Learning

arXiv:2603.02203v2 Announce Type: replace-cross Abstract: Test-time reinforcement learning (TTRL) has emerged as a promising paradigm for Recursive Self-Improving AI (RSI) by adapting Large Reasoning Models (LRMs) on unlabeled test inputs, using self-consensus rewards derived from majority voting over sampled rollouts. However, majority voting can mistake popularity for correctness: a spurious yet high-frequency unverified consensus may become a biased reward signal, causing test-time RL to reinforce frequent but wrong answers and collapse into an incorrect mode. We address this false-popular failure mode with T$^3$RL (Tool-Verification for Test-Time Reinforcement Learning), a verification-aware test-time RL framework. T$^3$RL grounds pseudo-label construction in external tool evidence. Concretely, a verifier utilizes external tool evidence (e.g., from code execution) to upweight verified rollouts during a verification-aware voting, producing more reliable pseudo-labels for training. A

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

ONNX-Net: Towards Universal Representations and Instant Performance Prediction for Neural Architectures

arXiv:2510.04938v2 Announce Type: replace-cross Abstract: Neural architecture search (NAS) automates the design process of high-performing architectures, but remains bottlenecked by expensive performance evaluation. Most existing studies that achieve faster evaluation are mostly tied to cell-based search spaces and graph encodings tailored to those individual search spaces, limiting their flexibility and scalability when applied to more expressive search spaces. In this work, we aim to close the gap of individual search space restrictions and search space dependent network representations. We present ONNX-Bench, a benchmark consisting of a collection of neural networks in a unified format based on ONNX files. ONNX-Bench includes all open-source NAS-bench-based neural networks, resulting in a total size of more than 600k {architecture, accuracy} pairs. This benchmark allows creating a shared neural network representation, ONNX-Net, able to represent any neural architecture using natural

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

Reasoning or Rambling? Exploring the Effect of Thinking on Agent Persuasion

arXiv:2509.21054v2 Announce Type: replace-cross Abstract: Understanding persuasion is critical for the safety and reliability of multi-agent systems built on large language models (LLMs). This paper studies persuasion dynamics by contrasting general LLMs with Large Reasoning Models (LRMs) that employ explicit ``thinking'' processes. Through large-scale experiments on objective (MMLU) and subjective (PersuasionBench and Perspectrum) tasks, we identify Persuasion Duality: reasoning enhances an agent's persuasive power while simultaneously increasing its resistance to persuasion. For LRMs, adding thinking content increases persuasion rates by 21 pp on average, yet reduces susceptibility to incorrect persuasion by up to 10 pp on objective tasks. Despite these gains, we uncover a critical vulnerability: persuasiveness often stems from superficial cues such as response length and repetition rather than logical validity. Non-semantic padding or repeated conclusions can match or exceed the per

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

Recurrence Meets Transformers for Universal Multimodal Retrieval

arXiv:2509.08897v2 Announce Type: replace-cross Abstract: With the rapid advancement of multimodal retrieval and its application in LLMs and multimodal LLMs, increasingly complex retrieval tasks have emerged. Existing methods predominantly rely on task-specific fine-tuning of vision-language models and are limited to single-modality queries or documents. In this paper, we propose ReT-2, a unified retrieval model that supports multimodal queries, composed of both images and text, and searches across multimodal document collections where text and images coexist. ReT-2 leverages multi-layer representations and a recurrent Transformer architecture with LSTM-inspired gating mechanisms to dynamically integrate information across layers and modalities, capturing fine-grained visual and textual details. We evaluate ReT-2 on the challenging M2KR and M-BEIR benchmarks across different retrieval configurations. Results demonstrate that ReT-2 consistently achieves state-of-the-art performance acro

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

Emergent Abilities in Large Language Models: A Survey

arXiv:2503.05788v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are leading a new technological revolution as one of the most promising research streams toward artificial general intelligence. The scaling of these models, accomplished by increasing the number of parameters and the magnitude of the training datasets, has been linked to various so-called emergent abilities that were previously unobserved. These emergent abilities, ranging from advanced reasoning and in-context learning to coding and problem-solving, have sparked an intense scientific debate: Are they truly emergent, or do they simply depend on external factors, such as training dynamics, the type of problems, or the chosen metric? What underlying mechanism causes them? Despite their transformative potential, emergent abilities remain poorly understood, leading to misconceptions about their definition, nature, predictability, and implications. In this work, we shed light on emergent abilities by con

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

Infer Human's Intentions Before Following Natural Language Instructions

arXiv:2409.18073v2 Announce Type: replace-cross Abstract: For AI agents to be helpful to humans, they should be able to follow natural language instructions to complete everyday cooperative tasks in human environments. However, real human instructions inherently possess ambiguity, because the human speakers assume sufficient prior knowledge about their hidden goals and intentions. Standard language grounding and planning methods fail to address such ambiguities because they do not model human internal goals as additional partially observable factors in the environment. We propose a new framework, Follow Instructions with Social and Embodied Reasoning (FISER), aiming for better natural language instruction following in collaborative embodied tasks. Our framework makes explicit inferences about human goals and intentions as intermediate reasoning steps. We implement a set of Transformer-based models and evaluate them over a challenging benchmark, HandMeThat. We empirically demonstrate th

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

Noise Contrastive Estimation-based Matching Framework for Low-Resource Security Attack Pattern Recognition

arXiv:2401.10337v5 Announce Type: replace-cross Abstract: Tactics, Techniques and Procedures (TTPs) represent sophisticated attack patterns in the cybersecurity domain, described encyclopedically in textual knowledge bases. Identifying TTPs in cybersecurity writing, often called *TTP mapping*, is an important and challenging task. Conventional learning approaches often target the problem in the classical multi-class or multi-label classification setting. This setting hinders the learning ability of the model due to a large number of classes (i.e., TTPs), the inevitable skewness of the label distribution and the complex hierarchical structure of the label space. We formulate the problem in a different learning paradigm, where the assignment of a text to a TTP label is decided by the direct semantic similarity between the two, thus reducing the complexity of competing solely over the large labeling space. To that end, we propose a neural matching architecture with a sampling-based learn-

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

Reliable Financial Named Entity Recognition Under Domain Shift: Confidence Estimation and Selective Prediction

arXiv:2608.19558v2 Announce Type: replace Abstract: Financial AI systems often train information extractors on one textual register and deploy them across filings, news, and user-generated content, and standard F1 scores do not indicate which predictions remain safe to automate when that input distribution changes. We study confidence estimation and selective prediction for financial named entity recognition (NER) on a three-tier stress test spanning SEC filings, financial news, and general-topic social media as an extreme out-of-domain condition, evaluating a BERT tagger and LoRA-tuned Qwen2.5-0.5B/1.5B models with five inference-time confidence signals, three training seeds, and bootstrap intervals. Confidence rankings themselves change under shift: whole-output probability is the strongest in-domain error detector but deteriorates out of domain, whereas entity-span probability and self-consistency are more robust; self-consistency is also better calibrated without post-hoc fitting.

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation

arXiv:2608.15062v4 Announce Type: replace Abstract: Scaling transformer language models creates an inherent tension between expressivity and memory efficiency. While unique weights across layers preserve functional specialization---from input-grounding to abstract refinement---they incur a substantial memory footprint. Conversely, standard depth-sharing enforces uniform transformations that collapse representational diversity and degrade modeling quality. We introduce Gated Recurrent Transformer, a recurrent depth transformer where fixed-depth prelude and coda blocks bracket a single shared core iterated R times. Inspired by gated recurrent neural networks, we employ a lightweight projection and an elementwise update gate---conditioned on the hidden state, the fixed prelude output, and noise resampled at every step---to modulate the recurrent update. This allows the model to specialize the input to the same few layers across recurrences, rather than requiring many unique layers to achi

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

SAG: SQL-Retrieval Augmented Generation with Query-Time Dynamic Hyperedges

arXiv:2606.15971v2 Announce Type: replace Abstract: While retrieval-augmented generation (RAG) has proven effective at giving LLMs access to external knowledge, mainstream dense-retrieval implementations remain inherently limited in handling structured constraints and multi-hop reasoning. Graph-based methods address this by constructing knowledge graphs offline, but they often fragment semantics, incur high maintenance, and complicate incremental updates. We propose SAG (SQL-Retrieval Augmented Generation), a structured retrieval architecture that organizes documents into an event-entity index without building a global knowledge graph. SAG represents each chunk as a semantically complete event paired with its entities, forming a latent hyperedge that preserves n-ary relations without decomposing them into triples. At query time, SAG treats shared entities as join keys to connect related chunks. This dynamically yields a query-scoped neighborhood of events, and yet every piece of eviden

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

When Probing Accuracy Saturates, Fragility Resolves: A Complementary Metric for LLM Pre-Training Analysis

arXiv:2606.11375v2 Announce Type: replace Abstract: Standard linear probing declares a property "encoded" when a classifier on hidden states achieves high accuracy. The protocol works well on a snapshot but breaks across pre-training, with probe accuracy saturating within the first few thousand steps, leaving most of training invisible to the instrument. We introduce fragility, a complementary per-layer metric defined as the activation-noise level at which probe accuracy collapses. Fragility is sensitive to both the margin of separability and the redundancy of representation, both of which keep evolving long after accuracy plateaus. Applied to open-checkpoint language models, fragility recovers structure that accuracy alone cannot see. Moralized representations, our interest, emerge in stages, with high-accuracy detection of morally loaded words first and compositional encoding later. Because probe accuracy on its own tracks how lexically separable a dataset is, we establish the compos

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

From Layers to Submodules: Rethinking Granularity in Replacement-Based LLM Compression

arXiv:2606.02559v2 Announce Type: replace Abstract: Post-training compression of Large Language Models (LLMs) removes entire architectural components, either deleting them or replacing them with fitted modules. Existing replacement-based methods share two design constraints: full-layer granularity and contiguous selection. We argue that this is overly restrictive: in fact, redundancy in pretrained transformers is not confined to contiguous regions, nor does it evenly distribute between Attention and FeedForward outputs, implying that different strategies best approximate different submodule types and that removable components need not cluster within contiguous depth ranges. Based on this intuition, we introduce SubFit (Submodule-level Fitted residual replacement), which compresses LLMs at the submodule level: Attention and FeedForward submodules are selected non-contiguously, and each receives its own lightweight fitted residual bypass. SubFit operates post-training and requires only c

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

ResMerge: Residual-based Spectral Merging of Large Language Models

arXiv:2606.02252v2 Announce Type: replace Abstract: Model merging offers a training-free way to combine multiple post-trained expert models, but merging experts obtained through reinforcement learning (RL) remains challenging. Existing spectral merging methods often assume that leading singular directions contain the main task signal, while lower-energy residual components can be compressed, selected, or attenuated to reduce interference. We find that this assumption does not hold for RL task vectors: after decomposing each task vector into a leading spectral head and a residual component, both parts can independently recover substantial behavior knowledge, while exhibiting different merging properties. The head is highly concentrated and informative but more prone to sharp cross-expert conflicts, whereas the residual component is more dispersed and provides a more stable basis for aggregation. Based on this observation, we propose ResMerge, a residual-based spectral merging framework

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CL

Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification

arXiv:2605.28500v2 Announce Type: replace Abstract: Large language models have shown impressive capabilities in code generation, yet they often produce functionally incorrect code. Uncertainty quantification (UQ) methods have emerged as a promising approach for detecting hallucinations in natural language generation, but their effectiveness for code generation tasks remains underexplored. We systematically evaluate how UQ techniques transfer to code generation across three programming languages, five LLMs, and over 1,700 problems. We find that some token-probability-based methods generalize effectively without modification, while sampling-based methods relying on natural language inference (NLI) fail because NLI models cannot distinguish functionally different code, causing most responses to collapse into a single semantic cluster. To address this, we introduce \emph{functional equivalence methods}, a family of code-specific methods that replace NLI-based semantic equivalence with an L

Source ↗
Showing 6651–6700 of 18402 signals
← Prev Page 134 of 369 Next →