EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Sep 07, 2026 · 40 ideas · 18694 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Fri, 21 Aug 2026 16:07:09 -0400
EdTech Mag (K-12)

Why Every K–12 Leader Should Be Paying Attention to E-Rate Funding Right Now

For nearly three decades, E-Rate has made connectivity and network infrastructure accessible for even the smallest and most rural schools. Now, the Federal Communications Commission is taking another look at the program and deciding whether it should continue in its current form. The reason? Leaders have raised questions about whether screen time on E-Rate–funded networks can be tied to educational purposes, especially in light of school board concerns about screen time. Our team at CDW has been talking to K–12 districts about E-Rate for years, and lately we've been hearing that school…

Source ↗
regulation Fri, 21 Aug 2026 15:30:00 +0000
The 74

Hawaiʻi Freshmen Failing: Ninth Grade Repeater Rate Among Nation’s Highest

High school came with more responsibility and less accountability for Josiah Ulukita. Courses were much harder, teachers were less strict about attendance and it was easier to skip school with friends or go on his phone during class, he said. Ulukita went from easily passing middle school to failing freshman English and social studies. After […]

Source ↗
technology Fri, 21 Aug 2026 14:21:09 +0000
MedCity News

The Pulse of Innovation: How 3D Models Can Prepare Surgeons for the Operating Room

Today, 3D modeling and printing technology can make a dramatic impact — even shortening some operations by 30-90 minutes. Looking ahead, that reduction will only grow as the technology is further integrated and becomes more advanced. The post The Pulse of Innovation: How 3D Models Can Prepare Surgeons for the Operating Room appeared first on MedCity News .

Source ↗
technology Fri, 21 Aug 2026 13:56:00 +0000
MedCity News

Modernizing Payment Integrity in an Era of Systemic Fraud

Gaps in reimbursement oversight and recovery will only widen, unless plans modernize how they detect, investigate, and recoup improper payments. The post Modernizing Payment Integrity in an Era of Systemic Fraud appeared first on MedCity News .

Source ↗
regulation Fri, 21 Aug 2026 13:30:00 +0000
The 74

Missouri Education Board Wary of $300 Million School Funding Increase

Missouri education officials are preparing to ask lawmakers for roughly $300 million more for public schools next year, setting up a potentially difficult budget fight as the state faces a projected revenue shortfall. The numbers are yet to be finalized, but the Department of Elementary and Secondary Education’s preliminary report shows that recent changes to […]

Source ↗
behavior Fri, 21 Aug 2026 12:33:32 +0000
District Admin

Nebraska’s largest school district asks police to stop using electric shock gloves on students

The suburban Bellevue police department that patrols two Omaha schools and the large Bellevue district said its officers will continue carrying the gloves. The post Nebraska’s largest school district asks police to stop using electric shock gloves on students appeared first on District Administration .

Source ↗
regulation Fri, 21 Aug 2026 10:30:00 +0000
The 74

Opinion: When Parents Stop Expecting More: The Missing Variable in Academic Decline

American education has spent decades searching for explanations for declining achievement. Poverty, inequality, school funding, curriculum, technology, teacher quality and the effects of the pandemic all matter. Yet one of the most consequential variables receives far less attention: what parents expect from their children and what happens when those expectations weaken. The issue is not […]

Source ↗
behavior Fri, 21 Aug 2026 10:00:00 +0000
eSchool News

Attendance is the scoreboard–connection is the game

If I had been asked a few years ago, I would have told you that chronic absenteeism was a school system's problem to fix. With better data systems, staff trained on outreach, and improved intervention protocols, school districts could move the numbers on their own.

Source ↗
technology Fri, 21 Aug 2026 09:00:00 +0000
eCampus News

Student coaching can help thousands of HBCU students re-enroll and graduate

A new report from InsideTrack, a national student success coaching nonprofit, reveals that a five-year coaching initiative across 39 Historically Black Colleges and Universities (HBCUs) helped stopped-out students re-enroll at 2.5 times the national average while keeping actively enrolled students on track at a fall-to-spring retention rate of 83.8 percent. The post Student coaching can help thousands of HBCU students re-enroll and graduate appeared first on eCampus News .

Source ↗
technology Fri, 21 Aug 2026 08:47:59 +0000
Tech & Learning

Best AI-Powered Tutors for Education

The best AI-powered tutors guide students towards genuine learning

Source ↗
audience Fri, 21 Aug 2026 07:00:00 +0000
Inside Higher Ed

There’s Little Evidence Institutions Care About Writing

There’s Little Evidence Institutions Care About Writing johnw@mcsweeneys.net Fri, 08/21/2026 - 03:00 AM If writing matters, why don’t we put sufficient resources toward teaching it? Byline(s) John Warner

Source ↗
audience Fri, 21 Aug 2026 07:00:00 +0000
Inside Higher Ed

New (School) Year’s Resolutions

New (School) Year’s Resolutions sara.custer@in… Fri, 08/21/2026 - 03:00 AM At the start of every academic year, many of us resolve to work differently or finally take better care of ourselves. But resolutions are often built to fail, and a more compassionate approach can help us make real change that lasts. Byline(s) Jessi Gold

Source ↗
audience Fri, 21 Aug 2026 07:00:00 +0000
Inside Higher Ed

Friday Fragments

Friday Fragments Sara Brady Fri, 08/21/2026 - 03:00 AM F-1 visas, indefensibly expensive in-person sections and a few more aphorisms. Byline(s) Matt Reed

Source ↗
audience Fri, 21 Aug 2026 07:00:00 +0000
Inside Higher Ed

Tenure-Track Ranks Grew Last Year, Survey Finds

Tenure-Track Ranks Grew Last Year, Survey Finds kathryn.palmer… Fri, 08/21/2026 - 03:00 AM Byline(s) Kathryn Palmer

Source ↗
audience Fri, 21 Aug 2026 07:00:00 +0000
Inside Higher Ed

Berkeley Report Says Trusting Campus Climate ‘Vital’ to University’s Mission

Berkeley Report Says Trusting Campus Climate ‘Vital’ to University’s Mission Emma Whitford Fri, 08/21/2026 - 03:00 AM Byline(s) Emma Whitford

Source ↗
audience Fri, 21 Aug 2026 07:00:00 +0000
Inside Higher Ed

Some Students Can Now Venmo Their College Tuition

Some Students Can Now Venmo Their College Tuition Johanna Alonso Fri, 08/21/2026 - 03:00 AM A massive portion of Gen Z uses money-transfer apps like Venmo every day, so allowing students to use them to pay their tuition was a natural step, the company said. Byline(s) Johanna Alonso

Source ↗
audience Fri, 21 Aug 2026 07:00:00 +0000
Inside Higher Ed

University Suspends Academic Who Accused Jason Arday of Plagiarism

University Suspends Academic Who Accused Jason Arday of Plagiarism sara.custer@in… Fri, 08/21/2026 - 03:00 AM Nathan Cofnas says he will “probably be fired” as Ghent begins disciplinary proceedings. Byline(s) Seher Asaf for Times Higher Education

Source ↗
audience Fri, 21 Aug 2026 07:00:00 +0000
Inside Higher Ed

An Earn-While-You-Learn Path Into Teaching

An Earn-While-You-Learn Path Into Teaching Joshua.Bay Fri, 08/21/2026 - 03:00 AM A report from New America finds teacher degree apprenticeships could lower barriers to the profession for working adults and nontraditional students. Byline(s) Joshua Bay

Source ↗
audience Fri, 21 Aug 2026 07:00:00 +0000
Inside Higher Ed

Ivy Tech Receives Approval for Workforce Pell Program

Ivy Tech Receives Approval for Workforce Pell Program Olivia.sanchez Fri, 08/21/2026 - 03:00 AM Byline(s) Olivia Sanchez

Source ↗
audience Fri, 21 Aug 2026 07:00:00 +0000
Inside Higher Ed

Universities Settled Over Allegedly Baseless Probes. They’re Not Expressing Regret.

Universities Settled Over Allegedly Baseless Probes. They’re Not Expressing Regret. Ryan Quinn Fri, 08/21/2026 - 03:00 AM Former Justice Department lawyers say investigations into Brown and Columbia over alleged antisemitism were politically motivated. But it’s unclear whether federal officials will investigate the whistleblowers’ letter. Byline(s) Ryan Quinn

Source ↗
audience Fri, 21 Aug 2026 05:00:00 -0400
Higher Ed Dive

Colleges made deep staff cuts while adding tenure-track faculty last year

The latest survey of higher education employees from CUPA-HR found the biggest spike in tenure-track professors since at least 2016.

Source ↗
regulation Fri, 21 Aug 2026 05:00:00 -0400
K-12 Dive

Test yourself on the past week’s K-12 news

From increased vaccine exemptions in kindergarten to E-rate’s uncertain future under the FCC, what did you learn from our recent stories?

Source ↗
regulation Fri, 21 Aug 2026 05:00:00 -0400
K-12 Dive

Greater Head Start flexibility will empower program leaders

Policy changes under the Trump administration will improve quality and access, write two top federal officials.

Source ↗
behavior Fri, 21 Aug 2026 00:00:00 GMT
EdSurge

EdSurge Welcomes Top Journalists and Educators to Advisory Board

Bringing together award-winning reporters and leading educators, the board will provide strategic guidance on newsroom priorities and emerging issues ...

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

RequestRouter: Request-Boundary Routing for Efficient Single-GPU LLM Inference

arXiv:2605.23057v2 Announce Type: replace-cross Abstract: RequestRouter is a lightweight request-boundary controller for reducing the latency and energy cost of single-GPU large language model inference. Rather than serving all requests with one static configuration, the system uses cheap request-level features to select one fixed inference mode per request, including FP16, quantized inference, speculative decoding, prefix caching, continuous batching, and hybrid modes such as GPTQ plus prefix caching and INT8 plus continuous batching. We evaluate RequestRouter using an 8B instruction-tuned language model served through vLLM on NVIDIA A100 GPUs. Across the full-scale A100 evaluation--26,500 fixed-mode evaluations followed by 3,500 online-controller evaluations, for 30,000 measured inference executions in total--the controller achieves a 2.10x mean latency speedup over FP16 and a 0.48x energy ratio on deployment-style workloads. A smaller matched evaluation with repeated measurements pr

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

Transformer See, Transformer Do: Copying as an Intermediate Step in Learning Analogical Reasoning

arXiv:2604.06501v2 Announce Type: replace-cross Abstract: Analogical reasoning is a hallmark of human intelligence, enabling us to solve new problems by transferring knowledge from one situation to another. Yet, developing artificial intelligence systems capable of robust human-like analogical reasoning has proven difficult. In this work, we train transformers using Meta-Learning for Compositionality (MLC) on an analogical reasoning task (letter-string analogies) and assess their generalization capabilities. We find that letter-string analogies become learnable when guiding the models to attend to the most informative problem elements, induced by including copy tasks in the training data. Furthermore, generalization to new alphabets improves when models are trained with more heterogeneous datasets. For the best training run, our 3-layer encoder-decoder model performs on par with frontier models on our letter-string analogy datasets. The MLC approach also enables some generalization to

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

Disentangling Speaker Traits for Deepfake Source Verification via Chebyshev Polynomial and Riemannian Metric Learning

arXiv:2603.21875v2 Announce Type: replace-cross Abstract: Speech deepfake source verification systems aims to determine whether two synthetic speech utterances originate from the same source generator, often assuming that the resulting source embeddings are independent of speaker traits. However, this assumption remains unverified. In this paper, we first investigate the impact of speaker factors on source verification. We propose a speaker-disentangled metric learning (SDML) framework incorporating two novel loss functions. The first leverages Chebyshev polynomial to mitigate gradient instability during disentanglement optimization. The second projects source and speaker embeddings into hyperbolic space, leveraging Riemannian metric distances to reduce speaker information and learn more discriminative source features. Experimental results on MLAAD benchmark, evaluated under four newly proposed protocols designed for source-speaker disentanglement scenarios, demonstrate the effectivene

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

Towards Audio Token Compression in Large Audio Language Models

arXiv:2511.20973v2 Announce Type: replace-cross Abstract: Large Audio Language Models (LALMs) deliver strong performance across speech and audio tasks, but their audio encoders generate high-rate token sequences (e.g., 25 tokens/s), making attention computation costly and limiting scalability. In this paper, we explore techniques such as unsupervised segmentation, uniform average pooling, etc., to reduce the number of audio tokens before they are consumed by the LLM decoder. To mitigate potential performance degradation, we employ low-rank adapters during finetuning. We evaluate our proposed models on two tasks, automatic speech recognition and speech-to-speech translation tasks, that are dependent on effectively uncovering the underlying lexical content of the input signal, and study the effect of downsampling on these tasks. Experimental results show that compressed LALMs can achieve performance closer to frame-level LALMs while reducing the input audio token count up to three times

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

LODESTAR: Robust Entropy-Based Answer Selection in Retrieval-Augmented Generation for Question Answering -- Directing Frozen-LLM Entropy with a Reinforcement-Learned Prompt Polarizer under Misleading Passages

arXiv:2608.11922v2 Announce Type: replace Abstract: Predictive-distribution entropy is a strong answer-selection rule in retrieval-augmented generation (RAG) for question answering: across five QA benchmarks, selecting the answer a frozen respondent LLM produces with the lowest answer-token entropy lifts mean $F_1$ from 0.4769 to 0.5148 over the retriever's top-ranked passage, without gold answers. Yet this rule, which prior entropy-based selectors adopt, fails: a misleading passage makes the respondent confidently wrong, driving entropy down where the uncertainty signal looks most trustworthy. The failure comes from the passage the respondent reads, and the context it is read in is an input we can intervene on. We introduce LODESTAR: to our knowledge the first method to score a text intervention by the uncertainty it induces in a third-party frozen respondent, compared within one question. LODESTAR uses reinforcement learning (GRPO) to train, once and offline, a polarizer -- a short f

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

SPyCE: Skill-Policy Co-evolution for Multimodal Agents

arXiv:2607.13854v2 Announce Type: replace Abstract: Multimodal agents that think with images iteratively manipulate visual evidence and invoke tools across many steps. Existing reinforcement learning methods reduce trajectories to scalar rewards, forcing the policy to discover reusable tool-use patterns from scratch on every new task; memory-based alternatives retain past experience, yet they rely on test-time retrieval, without updating the policy to absorb reusable patterns from that experience. Our key insight is that multimodal reasoning trajectories should be distilled into reusable skills that co-evolve with the policy during training, rather than being consumed as rewards or retrieved from a static store. To this end, we propose SPyCE (Skill-Policy Co-evolution), a framework that distills trajectories into a hierarchical skill library and updates it throughout reinforcement learning. Execution skills capture local visual operations, while workflow skills encode high-level priors

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

RepSelect: Robust LLM Unlearning via Representation Selectivity

arXiv:2606.17168v3 Announce Type: replace Abstract: When LLM weights are open or fine-tuning is available through an API, suppressing hazardous knowledge and tendencies is not enough: removal has to be deep enough that an adversary cannot restore it. Existing unlearning is shallow by this standard: fine-tuning or a handful of in-context examples brings the behaviour back, and it often degrades general capabilities in the process. We identify a root cause: existing methods edit representations shared with the retain set and lying in the subspace that a fine-tuning attacker recovers, making unlearning simultaneously easy to undo and disruptive. Leveraging this, we propose RepSelect (Representation Selectivity), which isolates forget-set-specific representations by collapsing the top principal components of the weight gradients before each unlearning update, preserving general capabilities while limiting what fine-tuning can recover. Across five unlearning datasets spanning both knowledge

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

Calibrated Triage, Not Autonomy: Confidence Estimation for Medical Vision-Language Models

arXiv:2606.15910v2 Announce Type: replace Abstract: A vision-language model can answer a question about a chest radiograph or a pathology slide fluently and confidently while barely using the image, relying instead on language priors. In medicine this is the failure that matters most: the answer looks trustworthy and is not, and the natural safeguard is a confidence score reliable enough to say when the model should abstain. We ask a deployment question rather than an accuracy one: how much medical imaging work a vision-language model can safely defer on its own, and which confidence signal makes that possible. We evaluate nine confidence estimators, spanning training-free logit baselines, prompt-based self-reports, and trained internal probes, across five open-weight LVLMs and three medical VQA datasets covering broad clinical imaging, radiology, and pathology, every probe trained only on natural images and applied to medicine without adaptation. Recast as bounded selective prediction

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

A Finite-Calibration Regime Map for LLM Judge Panels

arXiv:2606.01034v2 Announce Type: replace Abstract: Deploying an LLM judge panel spends human labels on fitting a calibrator, constructing candidate judge paths, and validating which candidate to deploy. We study when finite labels should support a low-dimensional stacker or reliability model, and when an unrestricted joint output table is worth its cell-count and unseen-pattern cost. We cast this as a finite-calibration regime map and instantiate it as Finite-Calibration Panel Selection (FCPS), a validation selector over judge path, deployed panel size, and aggregator family with support diagnostics. Across RewardBench, LLMBar, SummEval, and Arena100K with a seven-judge pool, scalar/reliability aggregation has lower MSE than unrestricted joint-table calibration in 16 of 20 real dataset--budget cells by point estimate, while paired 95% intervals exclude zero in 11 cells; richer backoff/shrinkage tables narrow some gaps while preserving the finite-support bottleneck. Controlled calibrat

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization

arXiv:2605.24960v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) faithfulness, i.e., whether CoTs genuinely reflect large language models' (LLM) underlying behavior, is typically evaluated with metrics under two disjoint paradigms: contextual faithfulness, measured by perturbing the input or CoT trace, and parametric faithfulness, assessed by intervening on a model's parametric knowledge. Yet prior work compares them only descriptively. We fill this gap by proposing FaithMATE, a unified preference-alignment interface for optimizing models towards either faithfulness paradigm. It enables us to investigate the interplay between the two paradigms, examining whether and to what extent faithfulness gains generalize within and across paradigms. Across three models, two datasets, and six faithfulness metrics, we find that the two paradigms are positively coupled, yet asymmetric: optimizing towards parametric faithfulness yields consistent gains across both paradigms, whereas the con

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

EvoSelect: Data-Efficient LLM Evolution for Targeted Task Adaptation

arXiv:2604.26170v2 Announce Type: replace Abstract: Adapting large language models (LLMs) to a targeted task efficiently and effectively remains a fundamental challenge. Such adaptation often requires iteratively improving the model toward a targeted task, yet collecting high-quality human-labeled data to support this process is costly and difficult to scale. As a result, synthetic data generation has emerged as a flexible and scalable alternative. One straightforward approach is through an iterative generation-training loop, where candidate data are synthesized through an external generator, the model is updated using these data and the process is repeated over iterations. However, generated samples can be noisy, highly redundant, or even misaligned with the targeted task distribution. Training indiscriminately on such data can dilute useful learning signals and even degrade model performance. To address this, we introduce a refined paradigm, namely an iterative generation-selection-t

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

Remask, Don't Replace: Token-to-Mask Refinement in Diffusion Language Models

arXiv:2604.18738v3 Announce Type: replace Abstract: Diffusion language models (dLLMs) generate text through iterative denoising, filling multiple masked positions at each step. Positions filled in the same step are predicted without conditioning on one another's newly filled values and can therefore be mutually inconsistent; once retained, these inconsistencies become context for later predictions. We introduce \emph{Token-to-Mask} (T2M), a training-free inference-time correction method that identifies low-confidence positions using the model's probability of the current token, remasks them, and reconstructs them in later denoising steps. On dLLMs equipped with correction mechanisms, a single T2M configuration transfers across tasks and models without retuning and broadly improves task metrics over each model's native correction mechanism. In controlled experiments, we decompose correction methods into a detector that identifies suspicious tokens and an action that determines how to re

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

Doc-V*:Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA

arXiv:2604.13731v2 Announce Type: replace Abstract: Multi-page Document Visual Question Answering requires reasoning over semantics, layouts, and visual elements in long, visually dense documents. Existing OCR-free methods face a trade-off between capacity and precision: end-to-end models scale poorly with document length, while visual retrieval-based pipelines are brittle and passive. We propose Doc-$V^*$, an \textbf{OCR-free agentic} framework that casts multi-page DocVQA as sequential evidence aggregation. Doc-$V^*$ begins with a thumbnail overview, then actively navigates via semantic retrieval and targeted page fetching, and aggregates evidence in a structured working memory for grounded reasoning. Trained by imitation learning from expert trajectories and further optimized with Group Relative Policy Optimization, Doc-$V^*$ balances answer accuracy with evidence-seeking efficiency. Across five benchmarks, Doc-$V^*$ outperforms open-source baselines and approaches proprietary model

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

Qworld: Question-Specific Evaluation Criteria for LLMs

arXiv:2603.23522v2 Announce Type: replace Abstract: Evaluating large language models (LLMs) on open-ended questions is difficult because response quality depends on the question's context. Binary scores and static rubrics fail to capture these context-dependent requirements. Existing methods define criteria at the dataset level or generate them in a single pass, which limits their ability to explore the evaluation space implied by each question. We introduce One-Question-One-World (Qworld), a method that generates question-specific evaluation criteria using a recursive expansion tree. Given a question, Qworld decomposes it into scenarios, perspectives, and fine-grained binary criteria through hierarchical and horizontal expansion. The resulting criteria specify what a high-quality answer must address for that question. On HealthBench, Qworld covers 89% of expert-authored criteria and generates 79% novel criteria validated by human experts. Experts rate Qworld criteria higher in insight

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

When Contextual Inference Fails: Cancelability in Interactive Instruction Following

arXiv:2603.19997v2 Announce Type: replace Abstract: We investigate the separation of literal interpretation from contextual inference in a collaborative block-building tasks, where an agent must resolve underspecified instructions using context. We adapt an existing two-speaker psycholinguistic paradigm into an interactive benchmark called Build What I Mean (BWIM). This setup contrasts a pragmatically cooperative speaker with one who is only literally reliable. In BWIM, models face underspecified instructions and must choose between making a contextual inference or requesting clarification at a small communication cost. Evaluating several state-of-the-art LLMs, we find a clear dissociation between judgment and action. Although models successfully detect speaker unreliability in explicit confidence ratings, they fail to leverage this awareness when taking action. Instead of deploying efficient clarification strategies, models default to suboptimal behaviors. These include partner-blind

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

DiverValue-Bench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values

arXiv:2509.08022v3 Announce Type: replace Abstract: Aligning large language models (LLMs) with diverse human values is essential for safe and effective deployment, yet existing benchmarks often overlook cultural and demographic variation. We introduce DiverValue-Bench, a population-aware benchmark for evaluating multi-dimensional value alignment across 74 countries/regions. It contains 23,763 quality-controlled instances derived from PRISM user feedback and audited through large-scale human validation, with fine-grained value labels, personalized questions, contrastive reference answers, and rich demographic metadata. Using DiverValue-Bench, we evaluate representative LLMs and reveal substantial geographic and demographic disparities that are masked by aggregate performance. We further show that lightweight preference-based fine-tuning with Low-Rank Adaptation (LoRA) and Direct Preference Optimization (DPO) substantially improves in-domain value alignment while yielding consistent out-

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

HiFi-KPI: A Dataset for Hierarchical KPI Extraction from Earnings Filings

arXiv:2502.15411v5 Announce Type: replace Abstract: Accurate tagging of earnings reports can yield significant short-term returns for stakeholders. The machine-readable inline eXtensible Business Reporting Language (iXBRL) is mandated for public financial filings. Yet, its complex, fine-grained taxonomy limits the cross-company transferability of tagged Key Performance Indicators (KPIs). To address this, we introduce the Hierarchical Financial Key Performance Indicator (HiFi-KPI) dataset, a large-scale corpus of 1.65M paragraphs and 198k unique, hierarchically organized labels linked to iXBRL taxonomies. HiFi-KPI supports multiple tasks and we evaluate three: KPI classification, KPI extraction, and structured KPI extraction. For rapid evaluation, we also release HiFi-KPI-Lite, a manually curated 8K paragraph subset. Baselines on HiFi-KPI-Lite show that encoder-based models achieve over 0.906 macro-F1 on classification, while Large Language Models (LLMs) reach 0.440 F1 on structured ext

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

Explainable Multimodal Depression Recognition in Clinical Interviews via PHQ-Aligned Symptom Summarization

arXiv:2501.16106v2 Announce Type: replace Abstract: Recent advances in multimodal depression recognition for clinical interviews (MDRC) have demonstrated the potential of AI systems by integrating textual, acoustic, and facial cues. However, existing methods pay limited attention to interpretability, thereby constraining reproducibility and clinician review. To address this, we introduce Explain-MDRC, an explainable MDRC framework that mirrors clinical workflows by generating structured symptom summaries from text and integrating them with nonverbal cues for recognition. Specifically, we construct Explain-DAIC, a dataset based on DAIC-WOZ and enriched with PHQ-8-aligned summary annotations, providing a foundation for developing models with built-in interpretability. We further propose PhqCML, a model that combines PHQ-8-aligned symptom summarization with PHQ-aware contrastive learning and summary-informed multimodal fusion. Automated metrics and expert evaluations show that Explain-MDR

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction

arXiv:2608.20320v1 Announce Type: cross Abstract: Travel behavior research increasingly combines digital data collection with predictive modeling, yet these stages are often developed and evaluated separately. This study proposes a three-agent workflow integrating conversational data collection, structured data processing, and behavioral prediction. A chatbot-administered, image-augmented stated-preference survey collected mode choices from student commuters across five predefined weather scenarios, yielding 454 respondent-scenario observations. Weather-related associations were analyzed using a multinomial logit model, while logistic regression and random forest provided machine-learning benchmarks. Nine locally deployed large language models (LLMs), ranging from 2 to 35 billion parameters, were evaluated across four zero-shot prompt-and-context conditions and extended through persona, few-shot, and vision-based configurations. Random forest achieved 69.6% five-class accuracy, while t

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

arXiv:2608.20318v1 Announce Type: cross Abstract: Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems, so that the next system inherits the improvement. That process is the training algorithm: a better objective or update rule improves the compute\mbox{-}capability exchange rate for every subsequent run, including the one that produces the next agent. Whether RSI is feasible therefore turns on whether an agent can design training algorithms. No benchmark isolates that ability: existing suites are won by collecting data or by tuning hyperparameters, and none tells a change to how a run is executed apart from a change to how the model learns. We present AI4AI\mbox{-}Bench, 10 frozen research repositories spanning 10 training algorithm families. In each task, an agent has 4 hours on one B300 to rewrite the training algorithm; its code is then rerun from scratch for up to 12 hours and scored by a fixed evaluator hidden from the agent,

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

Phantom Gains: Auditing Self-Improvement Against a Measured Null

arXiv:2608.20290v1 Announce Type: cross Abstract: Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and loses. Tracking these transitions means differencing two noisy estimates, leaving them vulnerable to measurement artifacts. Auditing three rounds of rank-$32$ LoRA self-training on Qwen3-8B against a frozen control pushed through the identical pipeline, we identify seven measurement failures, each of which inverts a reported finding when its control is absent. Several are standard practice. A ledger built on a single greedy decode manufactures capability changes on an untrained model, largely an artifact of inference batching; the expansion statistic separating acquisition from sharpening assigns that same model a rate of $0.280$. The natural threshold repair does not survive replication: estimated across the frozen comparisons such a design already contains, its null stays non-zero. We replace it with a

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

Break It Down, Pass It On: Cross-Task Skill Transfer in LLM Agents

arXiv:2608.20274v1 Announce Type: cross Abstract: Large language model (LLM) agents can induce skills from completed tasks and reuse them later to grow more capable with experience. In practice, induced skills may transfer unreliably and can even harm the agent that retrieves them. When agent-induced skills transfer reliably across tasks remains an open question. We conduct a comprehensive and controlled study of how the way skills are induced shapes their transfer across tasks. Specifically, we compare task-level with subtask-level skill induction and text with code skill formats, the two axes along which existing methods differ. Task-level skills mostly reduce the agent's performance below its no-memory baseline while subtask-level skills raise it above on average, and text skills transfer better than code skills. To further understand our findings, we examine two complementary properties of the induced skills: specificity, which measures how closely a skill matches real tasks, and a

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference

arXiv:2608.20210v1 Announce Type: cross Abstract: Small language models are usually built like large ones and then squeezed onto a CPU afterwards. We did the opposite: we fixed the target first, one user, one token at a time, 4-bit weights, ordinary CPU, and chose the architecture to suit it. The result keeps full attention in only 6 of its 18 blocks. The other 12 use short convolutions whose memory is two timesteps wide no matter how long the conversation gets, so two thirds of the network never re-reads a growing cache. Trained from scratch on 59.9B tokens, the model scores 47.31 on a five-task benchmark against a bar of 42.20 that was fixed before training began. It beats GPT-2 124M, Pythia-160M, OPT-125M and GPT-neo-125M, all trained on three to six times more data, and exceeds MobileLLM-125M's published score despite that model seeing a trillion tokens. Validation bits-per-byte is 0.8685. To check the architecture rather than the training recipe, we trained a conventional all-atte

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

ContractScrub: A benchmark for final review of legal contracts

arXiv:2608.20204v1 Announce Type: cross Abstract: Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs. Contract ``scrubbing,'' the final review of transactional agreements for errors and inconsistencies, is a particularly suitable task for automation, because it is routine, painstaking work requiring detailed attention to long documents. Scrubbing also seems to align naturally with the general capabilities expected of frontier LLMs around long-context reasoning, consistency checking, and named entity recognition (NER). Despite the economic value and potential for automation, no formal evaluations of LLMs performing contract scrubbing have been conducted. We introduce ContractScrub, the first benchmark designed to evaluate contract scrubbing capabilities, comprising contracts hand-crafted by experienced lawyers over diverse error categories such as misuse of defined terms, incorrect references, a

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

Multi-Agent Orchestration with the Common-Sense Reasoning Capabilities of LLMs for Autonomous Driving

arXiv:2608.20129v1 Announce Type: cross Abstract: Autonomous vehicles require robust perception and decision-making capabilities to operate in diverse and unseen scenarios. While reinforcement learning and rule-based methods can provide effective control and safety mechanisms, their performance may degrade in situations requiring contextual reasoning. Large Language Models (LLMs) have demonstrated strong capabilities in understanding multimodal information and generating contextual reasoning, however, their use for direct vehicle control can introduce latency and hallucination risks. To address these limitations, a hybrid framework is proposed. This system uses an orchestrator to coordinate PPO-trained reinforcement learning and PID control, with LLM common-sense reasoning applied throughout the framework. LLM reasoning is further employed iteratively to refine the RL reward function for dynamic driving environments. The proposed framework is evaluated in highly randomized CARLA scenar

Source ↗
technology Fri, 21 Aug 2026 00:00:00 -0400
arXiv cs.CL

Reward-Guided Autoregressive Graph Generation for Efficient Multi-Agent Communication Topology Design

arXiv:2608.20099v1 Announce Type: cross Abstract: LLM-based Multi-Agent Systems (MAS) achieve strong performance on complex reasoning tasks by coordinating multiple agents, but at the cost of substantial token consumption. Recent work on automatic topology design, ARG-Designer, has reframed this problem as autoregressive graph generation. However, its training objective provides no explicit incentive for the model to generate sparse and efficient topologies. We address this limitation by introducing a Reward-Guided Autoregressive Graph Generation (RGA-Designer) inspired by Reinforcement Learning from Human Feedback (RLHF). We train a reward model that jointly captures task correctness and structural compactness, and then fine-tune the pretrained graph generator using the reward model as feedback. Our method preserves task accuracy at the level of ARG-Designer while reducing token consumption by an average of 20.5%.

Source ↗
Showing 11501–11550 of 18694 signals
← Prev Page 231 of 374 Next →