EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Sep 07, 2026 · 40 ideas · 18624 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

behavior Thu, 07 May 2026 15:06:01 +0000
HN: online learning

Show HN: RVW – A transformer model capable of online continual learning

Article URL: https://zenodo.org/records/20064618 Comments URL: https://news.ycombinator.com/item?id=48050336 Points: 1 # Comments: 0

Source ↗
behavior Thu, 07 May 2026 10:00:00 +0000
eSchool News

Oracy is the missing link for multilingual learners

For multilingual learners, language is not just a subject to be learned--it is the very medium through which they access the curriculum.

Source ↗
technology Thu, 07 May 2026 09:00:00 +0000
Tech & Learning

From "Portrait of a Graduate" to "Portrait of a Learner": Prioritizing Executive Functioning in K-12

Three ways South Fayette Township School District brings their “Portrait of a Learner” to life.

Source ↗
behavior Thu, 07 Aug 2025 14:47:11 +0000
HN: online learning

What would convince you to try a new way of learning online?

Article URL: https://pagezero.ai Comments URL: https://news.ycombinator.com/item?id=44825187 Points: 2 # Comments: 3

Source ↗
behavior Thu, 06 Nov 2025 10:00:00 +0000
eSchool News

More parents are homeschooling–and turning to podcasts for syllabus support

A revolution quietly underway in American education: the rise of homeschooling. In the past decade, there’s been a 61 percent increase in homeschool students across the United States, making it the fastest growing form of education in the country.

Source ↗
technology Thu, 06 Aug 2026 23:11:39 +0000
MedCity News

Moderna Gets FDA Approval for mRNA Flu Vaccine, But CDC Input Still Uncertain

The FDA approved Moderna’s flu vaccine mFLUSIVA, a regulatory decision that comes six months after the agency refused to review the mRNA shot and then quickly reversed its position. But a recommendation from a Centers for Disease Control and Prevention advisory committee is key for payer coverage, and that committee has not met all year due to ongoing litigation. The post Moderna Gets FDA Approval for mRNA Flu Vaccine, But CDC Input Still Uncertain appeared first on MedCity News .

Source ↗
technology Thu, 06 Aug 2026 22:21:39 +0000
MedCity News

Report: What Employers Think About ICHRAs

A new survey found that many employers are considering ICHRAs, though concerns about awareness, affordability and the individual insurance market continue to slow adoption. The post Report: What Employers Think About ICHRAs appeared first on MedCity News .

Source ↗
regulation Thu, 06 Aug 2026 19:51:01 +0000
The 74

Back-to-School Costs Hit Families with Sticker Shock

Source ↗
technology Thu, 06 Aug 2026 18:50:45 -0400
EdTech Mag (Higher)

Higher Ed IT Budgets Face a Dual Threat: Federal Research Cuts and State Funding Pressure

Higher education leadership and IT teams are battling ghost students, establishing AI governance and managing competing priorities as students return to campus. Adding to this strain is a major shift in federal philosophy about National Institutes of Health grants, National Science Foundation funding and broader federal research funding cuts in what has historically been a university-centric system. In previous years, Congress has approved the federal budget, so that agencies such as NIH, NSF and the Department of Energy’s Office of Science could move swiftly to release funding…

Source ↗
regulation Thu, 06 Aug 2026 18:35:33 +0000
The 74

Head Start’s Antipoverty Mission Threatened by Trump Administration Overhaul

Most of Cassie Reed’s life has been marked by instability. During her childhood, she moved between nine different foster homes, often enduring abuse. Violence followed her into her adult life, as she met and fled partners who mistreated her. The turning point came several years ago for Reed, who lives in Spokane, Washington, when she […]

Source ↗
technology Thu, 06 Aug 2026 18:28:03 +0000
MedCity News

Debunked Episode 29: CMS GLP-1 Drug Pilot Program, RPM Strategy Raise Questions

At MedCity News’ inaugural Bullseye conference in Chicago, Debunked Podcast co-hosts Arundhati Parmar and Samir Batra discussed recent CMS actions and Mayo Clinic’s former research director, who has filed a complaint against the medical center over how it’s deploying AI. The post Debunked Episode 29: CMS GLP-1 Drug Pilot Program, RPM Strategy Raise Questions appeared first on MedCity News .

Source ↗
audience Thu, 06 Aug 2026 18:16:00 -0400
Higher Ed Dive

Alum launches as college-town hospitality concept, debuting in Alabama

The move intends to bring luxury condo-hotels and year-round programming near campuses as hospitality players move in on college markets.

Source ↗
need Thu, 06 Aug 2026 17:30:00 +0000
Hechinger Report

Writing skills are getting more attention in school

Each year, we get a pretty good idea of how kids across the country are doing with their math and reading skills. But writing? That’s a different story. Few states test students on how well they can write, and the country hasn’t published national writing test results for elementary-age children since 2002. Regardless, experts say […] The post Writing skills are getting more attention in school appeared first on The Hechinger Report .

Source ↗
regulation Thu, 06 Aug 2026 16:30:00 +0000
The 74

Which Families Are Using ESAs? What We Know About Texas’ First School Vouchers

More than 85,000 families are participating in Texas’ private school voucher program during its first year, a number that could increase as state officials finalize enrollment data, according to a new state report. White families make up the largest share of participants so far, as do those considered low-income, the comptroller’s office noted in the […]

Source ↗
regulation Thu, 06 Aug 2026 14:30:00 +0000
The 74

Teachers Have Some Big Wins in This Year’s California Budget. Here Are Their Victories

California teachers had some big wins in this year’s state budget, including about $700 million for stipends and training, and $218 million to support paid pregnancy leave. The funding for teacher recruitment and retention builds on the more than $1 billion California has spent since 2018 to end a persistent school staffing shortage that has […]

Source ↗
behavior Thu, 06 Aug 2026 13:45:05 +0000
District Admin

When a Texas school police officer got ‘riled up’ over a defiant teen

The incident, one of thousands of times that school officers in Texas have used physical force in recent years, underscores long-held concerns about policing in schools. The post When a Texas school police officer got ‘riled up’ over a defiant teen appeared first on District Administration .

Source ↗
behavior Thu, 06 Aug 2026 13:41:49 +0000
District Admin

Presidential prospect Glenn Youngkin launches group focused on education policy

The former governor of Virginia is targeting several battleground states with an ad campaign promoting President Donald Trump’s federal school choice program. The post Presidential prospect Glenn Youngkin launches group focused on education policy appeared first on District Administration .

Source ↗
technology Thu, 06 Aug 2026 13:40:00 +0000
MedCity News

Connecting the Dots: How Oral Health Can Enable Whole-Person Care

The industry has made real progress in recognizing that oral health belongs at the center of whole-person care, but intent alone isn’t enough. To translate this vision into something tangible, the industry must invest in infrastructure that makes integration not only possible but scalable. The post Connecting the Dots: How Oral Health Can Enable Whole-Person Care appeared first on MedCity News .

Source ↗
technology Thu, 06 Aug 2026 13:30:00 +0000
MedCity News

Should You Centralize Cold Chain Fulfillment? Only If You’ve Solved Distribution First

Cold chain is a logistics problem — not just a throughput problem. Pharmacies must rethink their strategy, not just automate it. The post Should You Centralize Cold Chain Fulfillment? Only If You’ve Solved Distribution First appeared first on MedCity News .

Source ↗
regulation Thu, 06 Aug 2026 10:30:00 +0000
The 74

Opinion: Accountability Works. Outcomes-Based Contracting Makes It Work Better

As John Adams observed, “facts are stubborn things.” The facts, as education policy writer Chad Aldeman recently observed, are that No Child Left Behind coincided with some of the fastest achievement gains on record for Black, Hispanic and low-income students, and when that accountability framework loosened after 2015, scores fell — fastest for the students […]

Source ↗
behavior Thu, 06 Aug 2026 10:00:00 +0000
eSchool News

The grading paradox: Better data is key to understanding real subject mastery

There’s a confusing trend going on in the nation’s classrooms: While national test scores have dropped to levels not seen in decades, average high school GPAs are actually climbing.

Source ↗
regulation Thu, 06 Aug 2026 09:55:00 -0400
K-12 Dive

HHS proposes sweeping Head Start changes

The federal agency aims to give more control to states over the design of their early learning programs, but critics fear a weakening of the landmark program.

Source ↗
regulation Thu, 06 Aug 2026 09:55:00 -0400
K-12 Dive

HHS proposes sweeping Head Start reforms

The federal agency aims to give more control to states over the design of their early learning programs, but critics fear a weakening of the landmark program.

Source ↗
technology Thu, 06 Aug 2026 09:00:00 +0000
Tech & Learning

Empowering The Fall Classroom: A Strategic Guide to 6 Essential AI Tools for Educators and Students

By integrating these AI tools thoughtfully, educators can empower students to become critical editors and auditors of data

Source ↗
audience Thu, 06 Aug 2026 07:00:00 +0000
Inside Higher Ed

Education Department Approves First Workforce Pell Program

Education Department Approves First Workforce Pell Program Sara Weissman Thu, 08/06/2026 - 03:00 AM Byline(s) Sara Weissman

Source ↗
audience Thu, 06 Aug 2026 07:00:00 +0000
Inside Higher Ed

Let’s Call a Truce Between Universities and the Trump Administration

Let’s Call a Truce Between Universities and the Trump Administration Elizabeth Redden Thu, 08/06/2026 - 03:00 AM We’ll tell you what: We’ll publish information about our planned reforms if the Trump administration does the same. Byline(s) Mark Hlavacik Jonathan Zimmerman

Source ↗
audience Thu, 06 Aug 2026 07:00:00 +0000
Inside Higher Ed

Harvard Among Institutions That Cut Personnel in July

Harvard Among Institutions That Cut Personnel in July Josh Moody Thu, 08/06/2026 - 03:00 AM Multiple campuses laid off employees last month to plug budget gaps created by the loss of federal research funds, declining enrollment, shrinking state support and more. Byline(s) Josh Moody

Source ↗
audience Thu, 06 Aug 2026 07:00:00 +0000
Inside Higher Ed

Barriers Block Aid for Student Parents

Barriers Block Aid for Student Parents Joshua.Bay Thu, 08/06/2026 - 03:00 AM A new analysis finds that just 29 percent of California student parents received supplemental aid meant to cover housing, childcare and other basic needs. Byline(s) Joshua Bay

Source ↗
audience Thu, 06 Aug 2026 07:00:00 +0000
Inside Higher Ed

The Key Podcast: How AI Can Help Researchers Fail

The Key Podcast: How AI Can Help Researchers Fail sara.custer@in… Thu, 08/06/2026 - 03:00 AM Byline(s) IHE Staff

Source ↗
audience Thu, 06 Aug 2026 07:00:00 +0000
Inside Higher Ed

U. of Alaska Plans to End Out-of-State Tuition Rates

U. of Alaska Plans to End Out-of-State Tuition Rates Johanna Alonso Thu, 08/06/2026 - 03:00 AM Byline(s) Johanna Alonso

Source ↗
audience Thu, 06 Aug 2026 07:00:00 +0000
Inside Higher Ed

Japan Officially Joins Europe’s Flagship Research Program

Japan Officially Joins Europe’s Flagship Research Program sara.custer@in… Thu, 08/06/2026 - 03:00 AM The Asian giant follows South Korea in associating with Horizon Europe, gaining access to funds for research that addresses societal challenges. Byline(s) Seher Asaf for Times Higher Education

Source ↗
audience Thu, 06 Aug 2026 07:00:00 +0000
Inside Higher Ed

Public Colleges Anchor State Workforce Pipelines

Public Colleges Anchor State Workforce Pipelines kathryn.palmer… Thu, 08/06/2026 - 03:00 AM Graduates who attended a broad access college or university are also more likely to remain in the state for work, according to a new report from the Strada Institute for the Future of Work. Byline(s) Kathryn Palmer

Source ↗
audience Thu, 06 Aug 2026 07:00:00 +0000
Inside Higher Ed

No Carrots and No Sticks?

No Carrots and No Sticks? sara.custer@in… Thu, 08/06/2026 - 03:00 AM The Education Secretary has called on higher ed leaders to declare how they plan to rebuild public trust. But the sector has reason to read between the lines. Byline(s) Sara Custer

Source ↗
audience Thu, 06 Aug 2026 07:00:00 +0000
Inside Higher Ed

Cassidy Eyes Final Push to Pass College Transparency Act

Cassidy Eyes Final Push to Pass College Transparency Act jessica.blake@… Thu, 08/06/2026 - 03:00 AM The Louisiana Republican is trying to use his final months in office to turn it into law. But once again, he faces a steep climb. Byline(s) Jessica Blake

Source ↗
audience Thu, 06 Aug 2026 05:00:00 -0400
Higher Ed Dive

Clemson, UF, UVU and other colleges get new leaders

We’re rounding up July’s leadership transitions, including two public colleges that faced unexpected hurdles when installing new presidents.

Source ↗
regulation Thu, 06 Aug 2026 05:00:00 -0400
K-12 Dive

Where do Ten Commandment laws stand?

As schools head into the 2026-27 school year, some classrooms will be reopening with religious displays despite parents' efforts to stop them.

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CL

When Outputs Disperse, Does Epistemic Revision Follow? A Black-Box Diagnostic for Machine Collectives

arXiv:2608.03722v2 Announce Type: replace-cross Abstract: Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise. In LLM collectives this proxy can break: agents can produce diverse-looking arguments while preserving the same conclusion. We operationalize dispersion-revision coupling: the degree to which an intervention that verifiably increases the dispersion of a collective's outputs in embedding space is accompanied by genuine revision of its epistemic stance rather than premise-preserving reformulation. The diagnostic is black-box: it operates on generated text alone and makes no claims about the internal representations of the generating models. Two channels are measured independently: an output channel, the Coherence Index (CI), verifies that the intervention changed output dispersion; an epistemic channel, per-turn stance annotation, measures whether the collective rev

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CL

Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Research Agents

arXiv:2608.02751v2 Announce Type: replace-cross Abstract: Existing deep-research agents use a Search--Visit workflow that retrieves whole webpages without considering the structure they expose through titles, headings, sections, and metadata. This prevents agents from directly constraining retrieval to parts of a webpage and often carries irrelevant content into their context. We introduce \textsc{Sieve}, a search--inspect--fetch strategy driven by a Boolean Query Language (BQL): it searches webpage fields to filter candidates, uses an interchangeable ranker to order them, presents structure-rich result cards for inspection, and fetches only selected sections. Across three QA collections, \textsc{Sieve} is more accurate than the strongest conventional Search--Visit configuration on each collection while using $20.7$--$50.6\%$ fewer tokens. Boolean filtering improves every tested ranker, and the accuracy--context advantage persists across retriever choices and agent backbones. Our imple

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CL

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy

arXiv:2608.02087v2 Announce Type: replace-cross Abstract: Post-training Large Language Models (LLMs) with Reinforcement Learning (RL) has become an important tool for improving model capabilities, but the LLM action-space structure introduces challenges distinct from classical RL, with implications for inducing exploration. New methods are required that leverage the broad knowledge and flexibility of pre-trained LLMs to deliberately generate diverse experience at training time. We propose Instruction-Conditioned Exploration (ICE), which appends one of a small fixed set of instructions to task prompts during training, using the same set for every problem, increasing the coverage of behaviours attempted. To facilitate ICE, we combine RL on the instruction-conditioned policy with self-distillation of its correct rollouts into the unconditioned test-time policy. ICE with this objective improves Qwen3-1.7B held-out pass@1 performance at 4K response length on mathematical reasoning tasks by

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CL

Theory-Level Autoformalization: From Isolated Statements to Unified Formal Knowledge Bases

arXiv:2607.13292v2 Announce Type: replace-cross Abstract: Autoformalization translates informal natural language into formal, machine-verifiable languages. While most work focuses on individual statements, real formalization efforts are inherently theory-level: they require an entire web of axioms, definitions, and lemmas before target theorems can even be stated. In this position paper, we argue for theory-level autoformalization: formalizing complete theories, including all their inter-dependencies, as structured libraries. We examine the significance of this shift, address alternative views, identify open challenges, and propose three promising paths forward. Our survey of autoformalization is available at https://github.com/marcusm117/Awesome-Autoformalization.

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CL

Not Truly Multilingual: Script Consistency as a Missing Dimension in VLM Evaluation

arXiv:2606.17188v3 Announce Type: replace-cross Abstract: Current multilingual evaluations for Vision-Language Models (VLMs) assume a one-to-one mapping between language and orthography, overlooking billions of users of multi-script languages. We introduce PuMVR (Punjabi Multimodal Visual Reasoning), a benchmark of 1,000 strictly parallel image-text instances across Punjabi's three active scripts: Gurmukhi, Shahmukhi, and Roman. Evaluating 10 state-of-the-art VLMs, we expose a substantial and systematic Script Gap. Models frequently solve visual tasks in one script while failing identical tasks in another, with accuracy deltas reaching 16%. Crucially, visual input boosts absolute performance uniformly yet does not close the orthographic gap. Furthermore, cross-script in-context transfer is highly brittle, exposing script-locked knowledge representation. Supported by McNemar tests across all script pairs, our findings demonstrate that current "multilingual" VLMs are not truly multi-scri

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CL

CogniFold: Always-On Proactive Memory via Cognitive Folding

arXiv:2605.13438v4 Announce Type: replace-cross Abstract: Existing agent memory remains predominantly reactive and retrieval-based, lacking the capacity to autonomously organize experience into persistent cognitive structure. Toward genuinely autonomous agents, we introduce CogniFold, a brain-inspired "always-on" agent memory designed for the next generation of proactive assistants. CogniFold continuously folds fragmented event streams into self-emerging cognitive structures, bootstrapping progressively higher-level cognition from incoming events and accumulated knowledge. We ground this by extending Complementary Learning Systems (CLS) theory from two layers (hippocampus, neocortex) to three, adding a prefrontal intent layer. Emulating the prefrontal cortex as the locus of intentional control and decision-making, CogniFold achieves this through graph-topology self-organization: cognitive structures proactively assemble under the stream, merge when semantically similar, decay when stal

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CL

Leakage-Audited Benchmarking Reveals Limited Evidence for Cross-Subject Auditory-Evoked EEG Vowel Perception Decoding

arXiv:2605.00865v2 Announce Type: replace-cross Abstract: We tested whether auditory-evoked EEG supports subject-independent five-vowel perception decoding when trial identity, model identity, prediction provenance, and participant-level inference are controlled within a single benchmark. We reconstructed Study 2 event tables from OpenNeuro ds006104 version 1.0.1 and analyzed the consonant-vowel pair task. One-to-one marker-stimulus pairing yielded 3,840 independent trials; control-condition selection and artifact rejection retained 1,094 epochs from 16 participants and 61 EEG channels. Thirteen unique implementations were evaluated using leave-one-subject-out testing, with participant metrics reconstructed from 36,102 trial predictions across 33 complete prediction replicas. Random Forest was numerically highest at 21.474% balanced accuracy (95% participant-bootstrap interval, 19.526-23.482%; chance, 20%), but neither its participant-level tests nor any implementation survived correct

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CL

Contextual Agentic Memory is a Memo, Not True Memory

arXiv:2604.27707v2 Announce Type: replace-cross Abstract: Current agentic memory systems (vector stores, retrieval-augmented generation, scratchpads, and context-window management) do not implement memory: they implement lookup. We argue that treating lookup as memory is a category error with provable consequences for agent capability, long-term learning, and security. Retrieval generalizes by similarity to stored cases; weight-based memory generalizes by applying abstract rules to inputs never seen before. Conflating the two produces agents that accumulate notes indefinitely without developing expertise, face a provable generalization ceiling on compositionally novel tasks that no increase in context size or retrieval quality can overcome, and are structurally vulnerable to persistent memory poisoning as injected content propagates across all future sessions. Drawing on Complementary Learning Systems theory from neuroscience, we show that biological intelligence solved this problem by

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CL

Terminal Agents Suffice for Enterprise Automation

arXiv:2604.00073v3 Announce Type: replace-cross Abstract: There has been growing interest in building agents that can interact with digital platforms to execute meaningful enterprise tasks autonomously. Among the approaches explored are tool-augmented agents built on abstractions such as Model Context Protocol (MCP) and web agents that operate through graphical interfaces. Yet, it remains unclear whether such complex agentic systems are necessary given their cost and operational overhead. We argue that a coding agent equipped only with a terminal and a filesystem can solve many enterprise tasks more effectively by interacting directly with platform APIs. We evaluate this hypothesis across diverse real-world systems and show that these low-level terminal agents match or outperform more complex agent architectures at a fraction of the cost. Our findings suggest that simple, flexible programmatic interfaces combined with strong foundation models should be the backbone of enterprise automa

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CL

SleepVLM: A Rule-Grounded Vision-Language Model for Auditable Sleep Staging

arXiv:2603.26738v4 Announce Type: replace-cross Abstract: Sleep staging is essential for sleep assessment and disorder diagnosis. In recent years, automatic sleep staging systems have achieved accuracy approaching that of human experts, but the black-box nature of their predictions hinders clinical adoption. Existing interpretability methods offer partial insight into model behavior, but their outputs still require expert reinterpretation and do not provide a direct basis for auditing individual predictions. To improve trustworthiness, we propose the task of auditable sleep staging. To solve this task, we present SleepVLM, a vision-language model that casts sleep staging as visual reasoning over rendered polysomnography (PSG) waveform images. For each epoch, SleepVLM outputs a stage together with the applicable American Academy of Sleep Medicine (AASM) rules and an auditable rationale. The model is trained using a two-stage framework: Waveform-Perceptual Pre-training followed by Rule-G

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CL

RingSQL: Schema-Independent Synthetic Data Generation for Text-to-SQL Reinforcement Learning

arXiv:2601.05451v2 Announce Type: replace-cross Abstract: Recent advances in text-to-SQL have been driven by larger models, better datasets, and new training methods like RLVR. However, progress remains limited by scarce high-quality training data, a problem RLVR is especially sensitive to since noisy data can produce spurious rewards. Manual data creation is expensive, and existing synthetic methods trade off reliability for scalability: template-based approaches guarantee correct SQL but need schema-specific templates and lack diversity, while LLM-based generation scales easily but lacks quality guarantees. We introduce RingSQL, a hybrid framework for generating question-SQL pairs that combines schema-independent query templates with LLM-based paraphrasing of natural language questions. By grounding question generation in complete template questions, RingSQL preserves question-query correctness across all levels of query complexity, a property purely LLM-based methods fail to maintai

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CL

Unforgettable Generalization in Language Models

arXiv:2409.02228v2 Announce Type: replace-cross Abstract: When language models (LMs) are trained to forget (or "unlearn'') a skill, how precisely does their behavior change? We study the behavior of transformer LMs in which tasks have been forgotten via fine-tuning on randomized labels. Such LMs learn to generate near-random predictions for individual examples in the "training'' set used for forgetting. Across tasks, however, LMs exhibit extreme variability in whether LM predictions change on examples outside the training set. In some tasks (like entailment classification), forgetting generalizes robustly, and causes models to produce uninformative predictions on new task instances; in other tasks (like physical commonsense reasoning and scientific question answering) forgetting affects only the training examples, and models continue to perform the "forgotten'' task accurately even for examples very similar to those that appeared in the training set. Dataset difficulty is not predictiv

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CL

Prompt-Induced Waste in Coding Agents: Reasoning Structure, Tool Behavior, and End-to-End Cost

arXiv:2608.01347v2 Announce Type: replace Abstract: Coding agents do not simply execute instructions; the wording of those instructions changes how much work they perform, what kind of work they perform, and how much that work costs. We present a preregistered study across multiple reasoning models, two real coding-agent harnesses, and controlled software tasks with hidden evaluation. The main finding is that several common prompt habits create substantial extra work without improving success. Asking for multiple approaches causes agents to develop and discard several solution paths before implementing one. Telling them to think deeply mainly produces longer visible reasoning, while demanding maximum certainty encourages repeated checking, extra tests, additional turns, and longer execution. Misleading architectural hints can also push agents toward unsupported lines of investigation. By contrast, prompts that define scope, request the smallest sufficient change, and include a clear st

Source ↗
technology Thu, 06 Aug 2026 00:00:00 -0400
arXiv cs.CL

CardioBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios

arXiv:2607.25186v2 Announce Type: replace Abstract: Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal, multimodal, and safety-critical workflow of cardiovascular care. Objective: To develop CardioBench, a real-world benchmark spanning the cardiovascular care continuum, and assess LLM performance across clinical dimensions and specialist tasks. Methods: CardioBench includes 2,263 items from 13 task-specific datasets derived from de-identified cardiovascular records and examination data. Sixteen cardiology physicians conducted annotation and reference construction, followed by cross-review from two senior cardiologists. Seven LLMs generated 15,841 outputs under standardized zero-shot settings. Open-ended tasks were evaluated using key-point coverage and holistic clinical quality, while CardioEthics was scored by accuracy. Results: GPT-5.4 achieved the highest macro-average (62.55) and item-

Source ↗
Showing 7951–8000 of 18624 signals
← Prev Page 160 of 373 Next →