EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Sep 07, 2026 · 40 ideas · 18694 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

Platform Choice, Trust, and Privacy in the Consumer AI Assistant Market

arXiv:2607.15134v1 Announce Type: new Abstract: We study how a representative sample of United States adult AI-assistant users (n=1,999; June 2026) choose among platforms, allocate tasks across them, evaluate provider trustworthiness, and value data-handling features. Estimates are weighted to the AI-user population using external adoption benchmarks. Four patterns emerge. The market is concentrated but internally differentiated: ChatGPT is the primary assistant for 58% of users and Gemini for 25%, yet smaller platforms hold defensible task niches--Claude captures a third of coding tasks despite a 7% overall share. Task allocation is thus organized by platform far more than by user, and technical use falls steeply with age. Trust is earned through use rather than reputation: Claude is ranked most trustworthy in every head-to-head among users of both platforms, and shows by far the largest gap between how its users and non-users rate it. Finally, privacy concern is near-universal but ac

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

SCITUS: A Multi-Jurisdictional Framework for Adapting NIST AI RMF to the Canadian Regulatory Context

arXiv:2607.15051v1 Announce Type: new Abstract: Canadian organizations deploying artificial intelligence systems face a fragmented regulatory landscape spanning federal requirements (the Treasury Board Directive on Automated Decision-Making) and divergent provincial regulations across Ontario, Quebec, Alberta, Manitoba, and British Columbia. The death of Bill C-27 (Artificial Intelligence and Data Act) in January 2025 - and the federal government's June 2026 confirmation that it will pursue targeted instruments rather than omnibus AI legislation - leaves organizations without unified compliance guidance. Global frameworks such as NIST AI RMF 1.0, the EU AI Act, and ISO/IEC 42001 provide valuable guidance but lack systematic methodologies for adaptation to multi-jurisdictional national contexts. We present SCITUS (Systematic Canadian Integration for Trustworthy and Unified Standards), a comprehensive framework adapting NIST AI RMF 1.0 to Canadian federal and provincial AI regulations si

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

Penny: Transition Network Analysis of Learner-Chatbot Interactions in Scaffolded EFL Writing

arXiv:2607.14575v1 Announce Type: new Abstract: Generative AI chatbots promise to transform English as a Foreign Language (EFL) writing by providing immediate, personalised feedback. However, their pedagogical value depends on how learners engage with them - a process often treated as a "black box." This study uses Transition Network Analysis to model the temporal dynamics of Japanese EFL learners using "Penny," an LLM-powered writing chatbot. Analysis of over 4,500 writing sessions and 21,000 chatbot interactions reveals two dominant behavioural loops: a "Revision Loop," where feedback leads directly to successful error correction, and a "Chat Loop," where learners engage in sustained dialogue with the chatbot following feedback. Crucially, EFL proficiency significantly shapes interaction: high-proficiency learners engage more in open dialogue and negotiation with the chatbot, while low-proficiency learners rely more heavily on repetitive corrective feedback cycles. The findings demon

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

Probabilistic "Copies" in Generative AI Models

arXiv:2607.14532v1 Announce Type: new Abstract: Recent work shows that it is possible to extract verbatim or near-verbatim text of some copyrighted works from some large language models (LLMs or models). That is evidence that the model weights encode the works in some form - that the model has "memorized" those works from its training data. But LLMs don't store information in the same format as familiar databases. Rather, their weights store statistical relationships between tokens that have been learned from the training data, and those relationships inform a generation process that is often probabilistic rather than deterministic. In the case of memorization, those relationships are strong enough that, in many circumstances, the model might generate a copyrighted work from its training data with some probability. Copyright law has not previously had to decide whether storing information that might or might not produce output similar to a copyrighted work is itself a copy of the work.

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation

arXiv:2607.14479v1 Announce Type: new Abstract: As large language models become increasingly capable, concerns about their potential to assist with biological misuse continue to grow. Prioritization of safety differs across the model ecosystem, with some models freely providing high-risk information that could be misused, and others refusing benign scientific content, potentially hindering legitimate research. Both failures stem from a lack of targeted mitigation to distinguish the most dangerous information from broader scientific content. To address this, we introduce BioTIER (Biological Targeted Information for Exclusion and Refusal), a benchmark designed to enable more targeted biological risk mitigation. BioTIER organizes biological content into three risk sets: Catastrophe Avoidance (CA), Biomedical DURC (BD) and Related Biology (RB). These sets represent a spectrum from extremely narrow high-risk topics to a broad range of benign and beneficial biological knowledge. The benchmar

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

Do Generative AI Assistants Respect robots.txt? Tracing Web Access Beyond Visible Answers

arXiv:2607.14447v1 Announce Type: new Abstract: AI assistants increasingly retrieve web content at inference time to provide fresh and grounded answers, yet it remains unclear whether these search-augmented capabilities respect website-owner restrictions expressed through robots.txt. We present a controlled empirical study of ten widely used AI assistants with advertised web-search capabilities. For each assistant, we first identify a configuration that actually produces observable web-browsing behavior and record the user-agent exposed during retrieval. We then evaluate compliance with controlled robots.txt rules across four complementary conditions: allowed for all user-agents, disallowed for all user-agents, allowed only for the assistant-specific user-agent, and disallowed only for that user-agent. Using server-side logs and secret codes embedded in target pages, we distinguish actual page access from user-visible answer correctness across 200 trials. Our results show substantial v

Source ↗
technology Fri, 17 Jul 2026 00:00:00 -0400
arXiv cs.CY

Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI

arXiv:2607.14353v1 Announce Type: new Abstract: As automated decision-making and data-driven technologies pervade society and are used to manage consequential outcomes, understanding the technology's capabilities, limitations, and attendant risks in context requires analysis of full sociotechnical systems. Sociotechnical analysis of risks in highly complex systems provides clear lessons for the design and evaluation of AI systems, transcending a technical focus on reliable or "responsibly designed" components to understand risks at a systems level. Human-made catastrophes have been studied for decades because of the severity of these events: consider Chernobyl, Three Mile Island, Fukushima-Daiichi, Bhopal, the Challenger disaster. A common misconception is that these kinds of events are freak accidents, resulting from the inherently unforeseeable interactions in complex systems. Closer examination reveals that the risks and hazards were well-known beforehand but not acted upon due to s

Source ↗
behavior Fri, 17 Apr 2026 10:00:00 +0000
eSchool News

We can’t wait for another Mississippi Miracle

Recent findings on the negative impacts of AI on learning might be sparking national debate, but they are unsurprising to learning scientists.

Source ↗
behavior Fri, 16 Jan 2026 10:00:00 +0000
eSchool News

Learning the “why” behind the math: How professional learning transformed our teachers

When you walk into a math classroom in Charleston County School District, you can feel the difference. Students aren’t just memorizing steps--they’re reasoning through problems, explaining their thinking, and debating solutions with their peers.

Source ↗
behavior Fri, 15 May 2026 10:00:00 +0000
eSchool News

Building a better bridge: Prioritizing infrastructure in a pre-K expansion

New York is currently standing at a historic crossroads. With a rare alignment of executive leadership in Albany and NYC and a tireless advocacy community, the state is poised to transform the promise of universal early childhood education (ECE) into a reality for tens of thousands of families.

Source ↗
technology Fri, 15 May 2026 09:00:00 +0000
Tech & Learning

When Highlights Are Easy to Fake With AI, Integrity Matters More

AI may make it easier to manipulate athletic performance, but students often underestimate how easily it can be exposed

Source ↗
technology Fri, 15 Aug 2025 18:05:35 +0000
HN: medical education

Using Large Language Models to Simulate History Taking for Medical Education

Article URL: https://www.mdpi.com/2078-2489/16/8/653 Comments URL: https://news.ycombinator.com/item?id=44915580 Points: 2 # Comments: 0

Source ↗
behavior Fri, 14 Nov 2025 10:00:00 +0000
eSchool News

Preserving critical thinking amid AI adoption

AI is now at the center of almost every conversation in education technology. It is reshaping how we create content, build assessments, and support learners. The opportunities are enormous.

Source ↗
technology Fri, 14 Aug 2026 21:20:21 +0000
MedCity News

Patient Group Sues AMA, Says the Org Shouldn’t Be Charging for CPT Codes

PatientRightsAdvocate.org sued the American Medical Association this week, arguing the trade group has no valid copyright over the CPT codes that providers must use to bill for care. It follows a public letter Senator Bill Cassidy sent to the AMA last year accusing it of “abus[ing] [its] government-backed monopoly by charging exorbitant fees to anyone using the CPT code set.” The post Patient Group Sues AMA, Says the Org Shouldn’t Be Charging for CPT Codes appeared first on MedCity News .

Source ↗
audience Fri, 14 Aug 2026 20:25:00 +0000
Inside Higher Ed

Former Cambridge Professor Jason Arday Found Dead

Former Cambridge Professor Jason Arday Found Dead Susan H. Greenberg Fri, 08/14/2026 - 04:25 PM Police say the 41-year-old was found unresponsive at an address in south London and pronounced dead at the scene. Byline(s) Tom Williams for Times Higher Education

Source ↗
technology Fri, 14 Aug 2026 20:05:42 +0000
MedCity News

Aligned Marketplace Secures $20M for Advanced Primary Care Marketplace

Venrock led Aligned Marketplace’s Series A round. In total, the company has raised $31 million to date. The post Aligned Marketplace Secures $20M for Advanced Primary Care Marketplace appeared first on MedCity News .

Source ↗
technology Fri, 14 Aug 2026 19:26:23 +0000
MedCity News

Bristol Myers Squibb Protein Degrader Wins First-in-Class FDA Nod in Multiple Myeloma

Bristol Myers Squibb’s Zenbexus is the first FDA-approved drug in a new class of cancer drugs called CELMoDs. The pharma company is positioning this molecule as a successor to its legacy products for multiple myeloma. The post Bristol Myers Squibb Protein Degrader Wins First-in-Class FDA Nod in Multiple Myeloma appeared first on MedCity News .

Source ↗
regulation Fri, 14 Aug 2026 18:30:00 +0000
The 74

SC Students Improve in Math and Reading, but Less Than Half Meet Math Benchmarks

COLUMBIA, S.C. — Elementary and middle school students posted their best scores yet on end-of-year standardized tests, though more than half still can’t calculate as expected for their age, according to state testing data released Monday. Third- through eighth graders continued to improve in both math and reading, though scores were still generally stronger in […]

Source ↗
audience Fri, 14 Aug 2026 17:28:00 -0400
Higher Ed Dive

EEOC lawsuit alleges Washington University segregated DEI training by race

Under the Trump administration, the agency has vocally cracked down on diversity, equity and inclusion work under the auspices of Title VII.

Source ↗
regulation Fri, 14 Aug 2026 16:30:00 +0000
The 74

Opinion: Head Start Was Built to Include Children With Disabilities. Don’t Weaken It

The federal government last week proposed sweeping changes to Head Start, describing them as a way to expand access for children. For families of children with disabilities, the question is what support will still be there once a child enters the classroom. I’ve been around long enough to know that when it comes to meeting […]

Source ↗
audience Fri, 14 Aug 2026 15:40:28 -0400
Higher Ed Dive

UT Austin plans massive overhaul to core curriculum

A task force called for making the university's general education more narrow and focused on topics such as Western civilization and broad U.S. history.

Source ↗
regulation Fri, 14 Aug 2026 14:57:02 -0400
K-12 Dive

Houston ISD stays at B rating in 2025-26

Three years into Texas’ takeover of the district, one of its most troubled schools has risen from seven consecutive F ratings to its first-ever A.

Source ↗
regulation Fri, 14 Aug 2026 14:30:00 +0000
The 74

School Backpack Overload: Heavy Bags Can Strain Kids’ Bodies

Backpacks can be fashionable and functional, but they can also be too heavy — weighed down by digital devices, musical instruments, sports equipment and more. Some kids carry home a laptop or tablet and textbooks, too. It’s good to be prepared, but kids who walk to school or participate in extracurricular activities may be lugging more […]

Source ↗
behavior Fri, 14 Aug 2026 13:57:29 +0000
District Admin

15% of Texas Schools scored below a ‘C’ on the state accountability scores

The scores grade how well specific campuses and school districts as a whole perform academically and prepare students for college and to enter the workforce. The post 15% of Texas Schools scored below a ‘C’ on the state accountability scores appeared first on District Administration .

Source ↗
technology Fri, 14 Aug 2026 13:27:00 +0000
MedCity News

Healthcare Keeps Buying AI. But Nobody’s Building the Workforce to Run It.

Most younger professionals entering the workforce have little visibility into interoperability, digital health infrastructure, or healthcare data architecture as career pathways. The post Healthcare Keeps Buying AI. But Nobody’s Building the Workforce to Run It. appeared first on MedCity News .

Source ↗
regulation Fri, 14 Aug 2026 12:30:00 +0000
The 74

In an Uncertain Job Market, Top Economists Offer Advice for Incoming NC College Students

It’s move-in week for college students across North Carolina. And once students have gotten settled into their dorms, their attention will turn to getting the classes needed to fulfill their chosen major. But which career paths are a safe bet in 2026? U.S. employers cut 23,000 jobs in July. Job creation for May and June […]

Source ↗
regulation Fri, 14 Aug 2026 10:30:00 +0000
The 74

Opinion: The Science of Reading and Teaching English Learners to Read Are Not Incompatible

For all the noise and attention and — laudable — political energy behind it, the science of reading is not a thing. Or, rather, it is not a thing that schools can implement. The science of reading is a series of facts that, in combination, form a field consensus about how children best learn to […]

Source ↗
behavior Fri, 14 Aug 2026 10:00:00 +0000
eSchool News

Public schools at the starting line: Beginning a new pathway with scholarship granting organizations

For the first time, a federal tax credit is opening a clear path for public school students to access expanded learning opportunities through Scholarship Granting Organizations (SGOs).

Source ↗
technology Fri, 14 Aug 2026 09:00:00 +0000
Tech & Learning

What Are Outcomes-Based Contracts and Why Do They Matter?

The outcome-based contract model can be a good way to increase efficiency and get more bang for the edtech budget buck.

Source ↗
technology Fri, 14 Aug 2026 09:00:00 +0000
eCampus News

Cohort connections matter: Strategies to help graduate students persist and succeed

Graduate education is demanding and isolating, especially for adult learners balancing coursework, research, employment, and family responsibilities. National data show half of doctoral students leave their programs before earning their degree. The post Cohort connections matter: Strategies to help graduate students persist and succeed appeared first on eCampus News .

Source ↗
audience Fri, 14 Aug 2026 07:00:00 +0000
Inside Higher Ed

U of Michigan Attempts to Disrupt the Transactional System

U of Michigan Attempts to Disrupt the Transactional System johnw@mcsweeneys.net Fri, 08/14/2026 - 03:00 AM Making space for student self-regulation is good, actually. Byline(s) John Warner

Source ↗
audience Fri, 14 Aug 2026 07:00:00 +0000
Inside Higher Ed

The Small Section Cut Session

The Small Section Cut Session Sara Brady Fri, 08/14/2026 - 03:00 AM Enrollment shifts aren’t evenly distributed. Byline(s) Matt Reed

Source ↗
audience Fri, 14 Aug 2026 07:00:00 +0000
Inside Higher Ed

Michigan Narrows Gap in Who Can Get Free Community College Tuition

Michigan Narrows Gap in Who Can Get Free Community College Tuition Ryan Quinn Fri, 08/14/2026 - 03:00 AM Byline(s) Ryan Quinn

Source ↗
audience Fri, 14 Aug 2026 07:00:00 +0000
Inside Higher Ed

Accreditor Puts Lane Community College on Warning

Accreditor Puts Lane Community College on Warning kathryn.palmer… Fri, 08/14/2026 - 03:00 AM Byline(s) Kathryn Palmer

Source ↗
audience Fri, 14 Aug 2026 07:00:00 +0000
Inside Higher Ed

Stockton Pulls Genocide Reference After Donor Complains

Stockton Pulls Genocide Reference After Donor Complains Josh Moody Fri, 08/14/2026 - 03:00 AM The public university in New Jersey abruptly deleted a reference to genocide in Gaza from its website. The author, a Jewish scholar of genocide, believes officials caved to donor pressure. Byline(s) Josh Moody

Source ↗
audience Fri, 14 Aug 2026 07:00:00 +0000
Inside Higher Ed

Cambridge to Review Hiring Processes After Arday Scandal

Cambridge to Review Hiring Processes After Arday Scandal sara.custer@in… Fri, 08/14/2026 - 03:00 AM The university says its investigation into sociology professor Jason Arday will feed into a wider inquiry on recruitment of senior academics. Byline(s) Tom Williams for Times Higher Education

Source ↗
audience Fri, 14 Aug 2026 07:00:00 +0000
Inside Higher Ed

Troy Punishes Professor for Writing to Lawmakers on University Letterhead

Troy Punishes Professor for Writing to Lawmakers on University Letterhead Emma Whitford Fri, 08/14/2026 - 03:00 AM Byline(s) Emma Whitford

Source ↗
audience Fri, 14 Aug 2026 07:00:00 +0000
Inside Higher Ed

How 4 Colleges Are Supporting Student Parents

How 4 Colleges Are Supporting Student Parents Joshua.Bay Fri, 08/14/2026 - 03:00 AM From scholarships and free childcare to pre-orientation programs and workforce training, institutions are helping student parents and caregivers persist. Byline(s) Joshua Bay

Source ↗
audience Fri, 14 Aug 2026 07:00:00 +0000
Inside Higher Ed

DOJ Declares 3 Race-Based STEM Programs Unconstitutional

DOJ Declares 3 Race-Based STEM Programs Unconstitutional Olivia.sanchez Fri, 08/14/2026 - 03:00 AM The programs received about $104 million in federal funds this past year. Byline(s) Olivia Sanchez

Source ↗
audience Fri, 14 Aug 2026 05:00:00 -0400
Higher Ed Dive

Education Department: Over 1,900 colleges have overdue data submissions

Institutions have until Jan. 15 to submit two rounds of data required by the Biden-era gainful employment and financial value transparency rules.

Source ↗
audience Fri, 14 Aug 2026 05:00:00 -0400
Higher Ed Dive

University of Nebraska board approves cutting programs with low enrollment

The eliminations come after a bruising budget-saving push to eliminate degrees last year at the system’s Lincoln campus.

Source ↗
regulation Fri, 14 Aug 2026 05:00:00 -0400
K-12 Dive

CISA issues K-12 cybersecurity guidance as schools’ risks persist

The agency’s free guides for district leaders come as the education sector continues to face a perfect storm of cyber vulnerability and limited resources.

Source ↗
regulation Fri, 14 Aug 2026 05:00:00 -0400
K-12 Dive

Test yourself on the past week’s K-12 news

From the latest data on written state special education complaints to a lawsuit against the Education Department, what did you learn from our recent stories?

Source ↗
technology Fri, 14 Aug 2026 00:00:00 -0400
arXiv cs.CL

Subtract, Transport, or Replay? Auditable Deletion from Language-Model Memory

arXiv:2607.27539v2 Announce Type: replace-cross Abstract: Exact deletion from persistent language-model memory depends on whether a record's effect remains addressable after later computation. Native Kimi Delta Attention (KDA) gives a negative result for the tested receipt interface: the corpus-pooled raw recurrent contribution changes by 12-49% with the suffix and remains 8-49% after a decay-ledger correction. Native omission also changes later transition and write terms and other active caches. Frozen-input transport succeeds on its fixed-input control; the changed terms place native omission outside the tested receipt classes. Checkpoint replay supplies the evaluated recomputation path; zero residual on final logits and all 80 audited KDA arrays verifies restoration across the declared checkpoint surface. The complementary result is constructive. We retrofit support-vector memory into frozen Gemma 3 without attention transfer, low-rank recovery, distillation, adapters, or language-m

Source ↗
technology Fri, 14 Aug 2026 00:00:00 -0400
arXiv cs.CL

Do Transformers Need Three Projections? Systematic Study of QKV Variants

arXiv:2606.04032v3 Announce Type: replace-cross Abstract: Transformers have become the standard solution for various AI tasks, with the query, key, and value (QKV) attention formulation playing a central role. However, the individual contribution of these three projections and the impact of omitting some remain poorly understood. We systematically evaluate three projection sharing constraints: a) Q-K=V (shared key-value), b) Q=K-V (shared query-key), and c) Q=K=V (single projection). The last two variants produce symmetric attention maps; to address this, we also explore asymmetric attention via 2D positional encodings. Through experiments spanning synthetic tasks, vision (MNIST, CIFAR, TinyImageNet, anomaly), and language modeling (300M and 1.2B parameter models on 10B tokens), we discovered that our transformers perform on par or occasionally better than the QKV transformer. In language modeling, Q-K=V projection sharing achieves 50% KV cache reduction with only 3.1% perplexity degra

Source ↗
technology Fri, 14 Aug 2026 00:00:00 -0400
arXiv cs.CL

Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking

arXiv:2605.18852v2 Announce Type: replace-cross Abstract: Selecting a final checkpoint for multimodal large language models (MLLMs) is challenging when late-stage candidates are closely matched and downstream evaluation signals are noisy. Small observed differences can be comparable to variability introduced by finite evaluation samples, LLM-based judges, and ambiguous multimodal evidence, while validation loss may not identify the checkpoint preferred by downstream evaluation. We formulate late-stage checkpoint selection as a stability-aware decision problem under evaluation uncertainty and propose a progressive framework combining pointwise filtering, listwise ranking, and pairwise refinement. Repeated evaluation-set subsampling is used to characterize ranking stability, while percentile-based aggregation accounts for lower- and upper-tail behavior. Experiments show that multimodal data evaluability is critical: quality-aware curation of OCR-heavy inputs reduces ranking flip rate fro

Source ↗
technology Fri, 14 Aug 2026 00:00:00 -0400
arXiv cs.CL

From Unstructured Recall to Schema-Grounded Memory: Reliable AI Memory via Iterative, Schema-Aware Extraction

arXiv:2604.27906v3 Announce Type: replace-cross Abstract: Persistent AI memory is often reduced to a retrieval problem: store prior interactions as text, embed them, and ask the model to recover relevant context later. This design is useful for thematic recall, but it is mismatched to the kinds of memory that agents need in production: exact facts, current state, updates and deletions, aggregation, relations, negative queries, and explicit unknowns. These operations require memory to behave less like search and more like a system of record. This paper argues that reliable external AI memory must be schema-grounded. Schemas define what must be remembered, what may be ignored, and which values must never be inferred. We present an iterative, schema-aware write path that decomposes memory ingestion into object detection, field detection, and field-value extraction, with validation gates, local retries, and stateful prompt control. The result shifts interpretation from the read path to the

Source ↗
technology Fri, 14 Aug 2026 00:00:00 -0400
arXiv cs.CL

CangjieBench: Benchmarking LLMs on a Low-Resource General-Purpose Programming Language

arXiv:2603.14501v2 Announce Type: replace-cross Abstract: Large Language Models excel in high-resource programming languages but struggle with low-resource ones. Existing research related to low-resource programming languages primarily focuses on Domain-Specific Languages (DSLs), leaving general-purpose languages that suffer from data scarcity underexplored. To address this gap, we introduce CangjieBench, a contamination-free benchmark for Cangjie, a representative low-resource general-purpose language. The benchmark comprises 248 high-quality samples manually translated from HumanEval and ClassEval, covering both Text-to-Code and Code-to-Code tasks. We conduct a systematic evaluation of diverse LLMs under four settings: Direct Generation, Syntax-Constrained Generation, Retrieval-Augmented Generation (RAG), and Agent. Experiments reveal that Direct Generation performs poorly, whereas Syntax-Constrained Generation offers the best trade-off between accuracy and computational cost. Agent

Source ↗
technology Fri, 14 Aug 2026 00:00:00 -0400
arXiv cs.CL

Learning Latency-Aware Orchestration for Multi-Agent Systems

arXiv:2601.10560v2 Announce Type: replace-cross Abstract: Multi-agent systems (MAS) coordinate multiple LLM-powered agents through structured workflows, gaining reasoning power but incurring high inference latency from multi-step execution and repeated model invocations. Existing orchestration methods primarily optimize task performance and inference cost, leaving latency largely unaddressed. In MAS, end-to-end latency is governed by the \textit{critical execution path}, so reducing total cost alone does not reliably reduce latency. Moreover, optimizing latency while preserving accuracy remains non-trivial: naive latency optimization can misassign operator-level credit and degrade task accuracy. To address this gap, we propose \textbf{L}atency-\textbf{A}ware \textbf{M}ulti-\textbf{a}gent \textbf{S}ystem (\textbf{LAMaS}), a latency-aware orchestration framework for learning-based multi-agent systems. LAMaS addresses this challenge at two levels: at \emph{training time}, it learns latenc

Source ↗
technology Fri, 14 Aug 2026 00:00:00 -0400
arXiv cs.CL

CityRiSE: Reasoning Urban Socio-Economic Status in Large Vision-Language Models via Reinforcement Learning

arXiv:2510.22282v2 Announce Type: replace-cross Abstract: Urban socio-economic sensing plays a vital role in advancing global sustainable development goals. With the advent of Large Vision-Language Models (LVLMs), new opportunities have emerged to address this challenge by framing it as a multi-modal perception and reasoning task. However, recent studies show that LVLMs still struggle to make accurate and interpretable socio-economic predictions from visual data. To overcome these limitations and fully exploit the potential of LVLMs, we propose CityRiSE, a novel framework for Reasoning urban Socio-Economic status in LVLMs via reinforcement learning (RL). With carefully curated multi-modal dataset and verifiable reward design, our approach guides the LVLM to focus on semantically meaningful visual cues, enabling structured and goal-oriented reasoning for generalist socio-economic status prediction. Experiments demonstrate that CityRiSE, equipped with emergent reasoning, significantly ou

Source ↗
Showing 11851–11900 of 18694 signals
← Prev Page 238 of 374 Next →