EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Sep 07, 2026 · 40 ideas · 18624 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

behavior Thu, 11 Jun 2026 10:00:00 +0000
HealthLeaders

Infographic: The Reality of Nursing Turnover in 2026

Turnover rates among nurses are holding steady, according to new data. Nursing turnover is always a point of contention for CNOs as they try to build a sustainable workforce. In 2026, RN turnover remains sitting at about 17%, which is the same as last year, according to Press Ganey's State of Nursing 2026 report. Gen Z and Millennials have the highest turnover rates, with Gen Z sitting at 22% and Millennials at 21%. Jeff Doucette , CNO at Press Ganey, previously told HealthLeaders the reason Gen Z nurses are leaving the workforce has to do with unmet needs, centering around purpose, support, and alignment with their organizations. "Gen Z clinical nurses generally have less [of a] feeling of psychological safety and they are experiencing significant cognitive overload and administrative burden, but they are less tolerant of the dysfunction that many of us have just learned to live with in healthcare environments," Doucette said. Here are some facts and figures that CNOs should know from

Source ↗
behavior Thu, 11 Jun 2026 10:00:00 +0000
eSchool News

The hidden skill many kids are losing

Once upon a time. For generations, those four words were an invitation. Children leaned in because a story was beginning. They would listen closely, follow the characters, and stay with the plot until the end.

Source ↗
behavior Thu, 11 Jun 2026 00:00:00 GMT
EdSurge

What TikTok Is Teaching Future Teachers (That We Aren’t)

A new educational epistemology for today's teachers.

Source ↗
behavior Thu, 11 Dec 2025 10:00:00 +0000
eSchool News

A smarter path to standards-based success: How Superior Public Schools united curriculum and data

Creating consistency between classrooms and ensuring curriculum alignment school-wide can be challenging, even in the smallest of districts. Every educator teaches--and grades--differently based on their experience and preferences, and too often, they’re forced into a solution that no longer respects their autonomy or acknowledges their strengths.

Source ↗
behavior Thu, 09 Jul 2026 19:59:45 +0000
MindShift (KQED)

Under a New Federal Rule, Colleges Must Leave Grads Better Off or Lose Financial Aid

If an undergraduate program's graduates don't earn more than workers who never went to college, that program could be cut off from federal student loans. But is a degree just about making more money?

Source ↗
regulation Thu, 09 Jul 2026 18:30:00 +0000
The 74

Market Share Continues to Hold Steady for NC Public Schools

Public schools continued to serve more than 1.5 million students across North Carolina in the 2025-26 school year — or about 84% of market share, a percentage that is about the same as previous years. Market share is a term used to describe how many students are served by different sectors of schools, including public […]

Source ↗
need Thu, 09 Jul 2026 17:45:00 +0000
Hechinger Report

Don’t let AI raise your kids

In 2020, a 14-inch tall robot named Moxie was introduced to the world as a way to help children build social and emotional skills through conversations and interactive games guided by artificial intelligence. Kids became enthralled and attached to the oblong teal robot, referring to it as their best friend. Four years later, Embodied, the […] The post Don’t let AI raise your kids appeared first on The Hechinger Report .

Source ↗
regulation Thu, 09 Jul 2026 17:00:00 -0400
K-12 Dive

Governors call on states to support locally driven K-12 solutions

Collaborations between states and districts can help improve student outcomes, the governors of Maryland and Wyoming say.

Source ↗
regulation Thu, 09 Jul 2026 16:30:00 +0000
The 74

Finale: Takeaways from a Season of AI in Education

Class Disrupted is an education podcast featuring author Michael Horn and Futre’s Diane Tavenner in conversation with educators, school leaders, students and other members of school communities as they investigate the challenges facing the education system in the aftermath of the pandemic — and where we should go from here. Find every episode by bookmarking […]

Source ↗
audience Thu, 09 Jul 2026 15:35:30 -0400
Higher Ed Dive

California community college settles with professor who sued over DEI policy

Bakersfield College agreed to not require Daymon Johnson to use diversity, equity, inclusion and accessibility principles in his teaching or scholarship.

Source ↗
audience Thu, 09 Jul 2026 15:07:39 +0000
Inside Higher Ed

Clemson Picks UGA Provost as President After Guskiewicz Reversal

Clemson Picks UGA Provost as President After Guskiewicz Reversal Ryan Quinn Thu, 07/09/2026 - 11:07 AM Byline(s) Ryan Quinn

Source ↗
regulation Thu, 09 Jul 2026 14:30:00 +0000
The 74

Dana Suskind on How To Protect Childhood in the Age of AI

The last time I interviewed Dr. Dana Suskind, we discussed the three T’s strategy outlined in her book “Parent Nation: Unlocking Every Child’s Potential, Fulfilling Society’s Promise”: Tune in. Talk more. Take turns. “It doesn’t require fancy gadgets,” she told me, “or a specialized degree.” Though it was only a few years ago, the fancy […]

Source ↗
behavior Thu, 09 Jul 2026 13:46:15 +0000
District Admin

Most states fail to meet IDEA requirements, feds say

Just 20 states earned a “meets requirements” designation from the Education Department this year for serving students with disabilities ages 3 to 21 under the Individuals with Disabilities Education Act. The post Most states fail to meet IDEA requirements, feds say appeared first on District Administration .

Source ↗
behavior Thu, 09 Jul 2026 13:31:47 +0000
District Admin

Attendance, Enrollment, and Equity Answers Hiding in Your Transportation Data

Date & Time: Thursday, August 06, 2026 at 2 p.m. ET In this webinar, learn the key takeaways from a report suggesting that transportation teams are sitting on information that could directly connect to whether students show up to class, whether families choose your schools, and whether your most vulnerable students, such as those experiencing homelessness or with IEPs, are getting the access the law requires. The post Attendance, Enrollment, and Equity Answers Hiding in Your Transportation Data appeared first on District Administration .

Source ↗
regulation Thu, 09 Jul 2026 12:30:00 +0000
The 74

Opinion: In NYC District, Technology Works With Pencil and Paper To Help Kids Learn Math

This may sound strange coming from the co-founder of an education technology company, but I think paper is a powerful technology in a classroom. Research on the science of learning consistently says so. So do the piles of paper that good teaching produces, the piles that bury the teachers who produced them. The real question […]

Source ↗
technology Thu, 09 Jul 2026 12:14:45 -0400
EdTech Mag (K-12)

ISTELive 26: This STEM Loaner Library Checks Out

Andy Mann’s model is deceptively simple. Informed by years of experience as a technology teacher, he developed a vision to get $10,000 virtual reality headsets and other sought-after devices in the hands of teachers who wanted to try them in their classrooms without the steep investment of buying them. Using strategic funding, Mann meticulously amassed a collection of STEM items that teachers from the Muskegon Area Intermediate School District in Michigan can check out. At his ISTELive 26 session, “Borrow Don’t Buy: Creating a STEM Loaner Library,” Mann shared how he built his STEM loaner…

Source ↗
technology Thu, 09 Jul 2026 12:13:42 -0400
EdTech Mag (K-12)

ISTELive 26: Cybersecurity Is Everybody’s Job – How K-12 Districts Are Building a Culture of Shared Responsibility

One of the greatest group project collaborations in K–12 right now has nothing to do with social studies or science projects. The IT leaders who understand the risk tend to speak in technical language, while the administrators and educators who control the budget speak in outcomes. Finding a shared vocabulary helps them fight — and ultimately fund a defense to — the K–12 cybersecurity crisis. The school year ended with a historic cybersecurity attack across K–12 and higher education, punctuating a point that teachers, IT leaders and school districts have been very aware of, with…

Source ↗
regulation Thu, 09 Jul 2026 10:30:00 +0000
The 74

Some Microschools in Limbo While Awaiting New Federal Tax Credit Rules

Public schools are beginning to imagine ways they can benefit from the new Federal Scholarship Tax Credit, after the Treasury Department clarified last month that district students will be eligible for scholarships. But for microschools, a growing segment of the private school market, the initial guidance from federal officials has left school leaders worried they […]

Source ↗
behavior Thu, 09 Jul 2026 10:00:00 +0000
eSchool News

Schools need behavioral assessments to fight chronic absenteeism

Our schools face a paradoxical challenge: chronic absenteeism, which is being treated as an attendance problem, actually isn’t.

Source ↗
behavior Thu, 09 Jul 2026 10:00:00 +0000
eSchool News

Schools need SEL assessments to fight chronic absenteeism

Our schools face a paradoxical challenge: chronic absenteeism, which is being treated as an attendance problem, actually isn’t.

Source ↗
technology Thu, 09 Jul 2026 09:00:00 +0000
Tech & Learning

4 Things Every New Teacher Should Remember

If you're new to teaching, it particularly helps to remember that you’re not the only one navigating some difficult waters.

Source ↗
audience Thu, 09 Jul 2026 07:00:00 +0000
Inside Higher Ed

How Do Employers View Community College Baccalaureate Degrees?

How Do Employers View Community College Baccalaureate Degrees? Sara Weissman Thu, 07/09/2026 - 03:00 AM Byline(s) Sara Weissman

Source ↗
audience Thu, 09 Jul 2026 07:00:00 +0000
Inside Higher Ed

North Texas Denies Faculty Conference Funding, Citing Anti-DEI Provisions

North Texas Denies Faculty Conference Funding, Citing Anti-DEI Provisions gianna.jakubowski Thu, 07/09/2026 - 03:00 AM The institution joins other universities that have worked to limit conference funds for professors due to anti-DEI pressures. Byline(s) Gianna Jakubowski

Source ↗
audience Thu, 09 Jul 2026 07:00:00 +0000
Inside Higher Ed

ICE Detains HBCU Baseball MVP

ICE Detains HBCU Baseball MVP gianna.jakubowski Thu, 07/09/2026 - 03:00 AM Byline(s) Gianna Jakubowski

Source ↗
audience Thu, 09 Jul 2026 07:00:00 +0000
Inside Higher Ed

Donor to Move Scholarship From UNC Wilmington Due to DEI Policy

Donor to Move Scholarship From UNC Wilmington Due to DEI Policy Johanna Alonso Thu, 07/09/2026 - 03:00 AM Byline(s) Johanna Alonso

Source ↗
audience Thu, 09 Jul 2026 07:00:00 +0000
Inside Higher Ed

How Far Can One Supreme Court Ruling Stretch?

How Far Can One Supreme Court Ruling Stretch? sara.custer@in… Thu, 07/09/2026 - 03:00 AM The Trump administration is using the ban on affirmative action in admissions to crack down on anything related to ethnicity or race in higher ed. Will it try the same tactic with last week’s Supreme Court ruling on transgender athletes? Byline(s) Sara Custer

Source ↗
audience Thu, 09 Jul 2026 07:00:00 +0000
Inside Higher Ed

From Inmate to College Graduate

From Inmate to College Graduate Joshua.Bay Thu, 07/09/2026 - 03:00 AM CUNY’s Prison-to-College Pathways Pipeline celebrated its first commencement while helping incarcerated students earn associate degrees and find belonging. Byline(s) Joshua Bay

Source ↗
audience Thu, 09 Jul 2026 07:00:00 +0000
Inside Higher Ed

Texas Tech Faculty Sue Over Race, Gender Rules

Texas Tech Faculty Sue Over Race, Gender Rules Katherine Knott Thu, 07/09/2026 - 03:00 AM Byline(s) Katherine Knott

Source ↗
audience Thu, 09 Jul 2026 07:00:00 +0000
Inside Higher Ed

The Key Podcast: Students Don’t Want Press Releases, They Want Outcomes

The Key Podcast: Students Don’t Want Press Releases, They Want Outcomes sara.custer@in… Thu, 07/09/2026 - 03:00 AM Byline(s) IHE Staff

Source ↗
audience Thu, 09 Jul 2026 07:00:00 +0000
Inside Higher Ed

College Associations Say Expanded List of Professional Degrees Is ‘Incomplete’

College Associations Say Expanded List of Professional Degrees Is ‘Incomplete’ jessica.blake@… Thu, 07/09/2026 - 03:00 AM Several health-care degrees like advanced nursing, physician assistants and occupational therapists were added, but others, like master’s in social work and education, were not. Byline(s) Jessica Blake

Source ↗
regulation Thu, 09 Jul 2026 05:00:00 -0400
K-12 Dive

Education Department targets Equity Assistance Centers again

A proposed rule would rescind regulations for the program, which the agency says would allow it to “explore other means” of delivering those services.

Source ↗
regulation Thu, 09 Jul 2026 05:00:00 -0400
K-12 Dive

4 more states require districts to adopt AI policies

At least one state has gone as far as to prohibit artificial intelligence’s use for grading, discipline or other high-stakes decisions.

Source ↗
need Thu, 09 Jul 2026 05:00:00 +0000
Hechinger Report

OPINION: The days of ‘good guy’ capitalists are over. College students are right to turn against the tech elites

The students booing artificial intelligence at commencements across the country are not just worried about jobs. They have learned an urgent lesson from the not-so-distant past. They know that the familiar promise of empowerment and creativity will continue to give way to the pathologies of the online surveillance economy: viral slop, commercial manipulation and addictive […] The post OPINION: The days of ‘good guy’ capitalists are over. College students are right to turn against the tech elites appeared first on The Hechinger Report .

Source ↗
behavior Thu, 09 Jul 2026 00:00:00 GMT
EdSurge

I Built the Chemistry Platform I Needed in My Own Classroom

What would chemistry look like if students could do more than read about it?

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

Nectar: Neural Estimation of Cached-Token Attention via Regression

arXiv:2605.09778v2 Announce Type: replace-cross Abstract: Evaluating softmax attention over a fixed long context requires reading every cached key-value pair for each new query token. For a given context (a book, a manual, a legal corpus) the attention output is a deterministic function of the query. We propose Nectar, which fits a compact neural network to this function for queries drawn from a task-relevant distribution. Nectar fits two networks per layer and KV-head: a target network that predicts the attention output and a score network that predicts the log-normalizer. The pair plugs into the standard masked self-attention at inference time, replacing the $O(n)$ attention over the cache with a forward pass whose cost does not depend on $n$. Each module carries on the order of $|\theta|$ parameters per layer and KV-head, typically much smaller than the $2nd$ KV-cache footprint at the same granularity. We report experiments on models from 1.7B to 8B parameters across five long-conte

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval

arXiv:2604.18360v3 Announce Type: replace-cross Abstract: Audio-text retrieval systems based on Contrastive Language-Audio Pretraining (CLAP) achieve strong performance on traditional benchmarks; however, these benchmarks rely on caption-style queries that differ substantially from real-world search behavior, limiting their assessment of practical retrieval robustness. We present Omni-Embed-Audio (OEA), a retrieval-oriented encoder leveraging multimodal LLMs with native audio understanding. To systematically evaluate robustness beyond caption-style queries, we introduce User-Intent Queries (UIQs) - five formulations reflecting natural search behaviors: questions, commands, keyword tags, paraphrases, and exclusion-based negative queries. For negative queries, we develop a hard negative mining pipeline and propose discrimination metrics (HNSR, TFR) assessing models' ability to suppress acoustically similar distractors. Experiments on AudioCaps, Clotho, and MECAT show that OEA achieves co

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

AnyPoC: Universal Proof-of-Concept Test Generation for Scalable LLM-Based Bug Detection

arXiv:2604.11950v2 Announce Type: replace-cross Abstract: While recent LLM-based agents can identify many candidate bugs in source code, their reports remain static hypotheses that require manual validation, limiting the practicality of automated bug detection. We frame this challenge as a test generation task: given a candidate report, synthesizing an executable proof-of-concept (PoC) - such as a script, command sequence, or crafted input - to trigger the suspected defect. Automated PoC generation can act as a scalable validation oracle, enabling end-to-end autonomous bug detection by providing concrete execution evidence. However, naive LLM agents are unreliable validators: they are biased toward "success" and may reward-hack by producing plausible but non-functional PoCs or even hallucinated traces. To address this, we present ANYPoC, a general multi-agent framework that (1) analyzes and fact-checks a candidate bug report, (2) iteratively synthesizes and executes a PoC while collect

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection

arXiv:2604.07831v2 Announce Type: replace-cross Abstract: Existing red-teaming studies on GUI agents face two fundamental limitations: adversarial perturbations require white-box access unavailable in commercial deployments, while prompt injection is increasingly neutralized by stronger safety alignment. To study robustness under a more practical threat model, we propose Semantic-level UI Element Injection, a black-box red-teaming paradigm that overlays safety-aligned and harmless UI elements onto screenshots to misdirect the agent's visual grounding. Our method couples a modular Editor--Overlapper--Victim pipeline with iterative search that samples multiple candidate edits, keeps the best cumulative overlay, and adapts future prompt strategies based on previous failures. Experiments across 19 victim models spanning 8 model families show that strategic optimization substantially outperforms random injection (3.5-6.9x on the most robust victims) and transfers near-perfectly across archi

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

Dual Path Attribution: Efficient Attribution for SwiGLU-Transformers through Layer-Wise Target Propagation

arXiv:2603.19742v2 Announce Type: replace-cross Abstract: Understanding the internal mechanisms of transformer-based large language models (LLMs) is crucial for their reliable deployment and effective operation. While recent efforts have yielded a plethora of attribution methods attempting to balance faithfulness and computational efficiency, dense component attribution remains prohibitively expensive. In this work, we introduce Dual Path Attribution (DPA), a novel framework that faithfully traces information flow on the frozen transformer in one forward and one backward pass without requiring counterfactual examples. DPA analytically decomposes and linearizes the computational structure of the SwiGLU Transformers into distinct pathways along which it propagates a targeted unembedding vector to receive the effective representation at each residual position. This target-centric propagation achieves O(1) time complexity with respect to the number of model components, scaling to long inpu

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

Towards Understanding Steering Strength

arXiv:2602.02712v2 Announce Type: replace-cross Abstract: A popular approach to post-training control of large language models (LLMs) is the steering of intermediate latent representations. Namely, identify a well-chosen direction depending on the task at hand and perturbs representations along this direction at inference time. While many propositions exist to pick this direction, considerably less is understood about how to choose the magnitude of the move, whereas its importance is clear: too little and the intended behavior does not emerge, too much and the model's performance degrades beyond repair. In this work, we propose the first theoretical analysis of steering strength. We characterize its effect on next token probability, presence of a concept, and cross-entropy, deriving precise qualitative laws governing these quantities. Our analysis reveals surprising behaviors, including non-monotonic effects of steering strength. We validate our theoretical predictions empirically on e

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?

arXiv:2510.09595v3 Announce Type: replace-cross Abstract: Competitive programming problems are increasingly used to evaluate the coding capabilities of large language models (LLMs) due to their complexity and ease of verification. Yet, current coding benchmarks face limitations such as a lack of exceptionally challenging problems, insufficient test case coverage, and reliance on online platform APIs that limit accessibility. To address these issues, we introduce LiveOIBench, a large-scale competitive programming benchmark featuring 403 expert-curated problems, averaging 60 official test cases each, drawn from 72 contests across 14 Informatics Olympiads held between 2023 and 2025. LiveOIBench has four key features: (1) expert-designed tasks with detailed subtask rubrics and extensive test cases; (2) direct comparison to elite human contestants; (3) continuous updates to reduce contamination risk; and (4) a fully offline, reproducible evaluation system. Benchmarking 34 popular general-pu

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

ContestTrade: A Multi-Agent Trading System Based on Internal Contest Mechanism

arXiv:2508.00554v4 Announce Type: replace-cross Abstract: In financial trading, large language model (LLM)-based agents demonstrate significant potential, but their decisions can be sensitive to noisy and non-stationary market information. We propose ContestTrade, a multi-agent trading system with an internal competitive mechanism inspired by institutional investment workflows. The system consists of two specialized teams: (1) a Data Team that processes and condenses massive market data into diversified textual factors optimized for constrained LLM context windows, and (2) a Research Team that produces parallelized multipath trading decisions via tool-augmented deep research. The core design is a "Quantify-Predict-Allocate" contest mechanism within each team: agent outputs are scored only after market outcomes become observable, future utility is predicted from historical scores, and resources are allocated to agents with positive predicted utility. In a post-2024 A-share backtest, Con

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

MTEB-BR: A Text Embedding Benchmark for Brazilian Portuguese

arXiv:2607.04581v2 Announce Type: replace Abstract: Text embeddings for Portuguese have no dedicated benchmark: evaluation rests on translated corpora such as English MS MARCO or on thin multilingual coverage, with native tasks scattered and unconsolidated. We introduce MTEB-BR, a benchmark of 22 native Brazilian-Portuguese tasks across seven categories (classification, multilabel classification, pair classification, semantic textual similarity, clustering, retrieval, and reranking), admitting only data created or found in Portuguese and excluding translations by construction. We evaluate 93 models spanning 23M to 27B parameters: 73 open-weight and 20 closed commercial APIs. Alongside the leaderboard we report a statistical layer for every headline comparison: per-task bootstrap confidence intervals, paired-bootstrap significance, a task- and instance-level discrimination analysis (how sharply each task separates models) adapted from Item Response Theory, and a cross-leaderboard correl

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

EMBER: Efficient Memory via Budgeted Evidence Retention for Long-Horizon Agents

arXiv:2606.05894v2 Announce Type: replace Abstract: Long-horizon agents can archive large histories, but future answers still incur retrieval, rereading, and context costs. When retained memory misses answer-relevant evidence, the system must return to larger portions of the raw history. We study budgeted evidence survival: before the query is known, which source evidence should be retained so that it remains recoverable and usable under a fixed retained source-evidence token budget? We instantiate this setting as Budgeted Pre-Query Retention, where memory is written during ingestion and later read without access to the full raw stream. We introduce EMBER, a learned retention policy that constructs a compact, source-backed evidence state. EMBER stores evidence capsules: verbatim source excerpts paired with retrieval keys and update metadata, preserving both grounding and read-time access. Post-query outcome feedback trains the writer to preserve evidence across the ingestion-retrieval-

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

Cast a Wider Net: Coordinated Pass@K Policy Optimization for Code Reasoning

arXiv:2605.27000v3 Announce Type: replace Abstract: Repeated sampling with a verifier is the standard way to allocate test-time compute for code generation, with pass@$K$ as the canonical metric. Yet the standard policy class draws $K$ independent samples from a single answer distribution, so attempts often collapse onto near-duplicate reasoning paths and waste the budget on redundant rollouts. This failure is costly in competitive programming, where many problems admit multiple distinct algorithmic strategies and pass@$K$ requires only one correct attempt. We propose Coordinated Pass@$K$ Policy Optimization (CPPO), which turns pass@$K$ generation into joint exploration over strategies: a planner emits a tuple of $K{=}4$ alternative high-level methods, and a shared solver attempts one solution per method. CPPO trains this joint policy with a multiplicative planner reward, $R_{\mathrm{plan}} = J_\psi \cdot R_{\mathrm{out}}$, assigning credit only to valid strategy tuples that lead to ve

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

Psy-Chronicle:A Structured Pipeline for Synthesizing Long-Horizon Campus Psychological Counseling Dialogues

arXiv:2605.22140v2 Announce Type: replace Abstract: In recent years, large language models have shown substantial potential in psychological support tasks. However, existing psychological counseling data mostly rely on single-turn question answering or short multi-turn dialogues, making it difficult to characterize how college students' psychological distress accumulates, interacts, and gradually evolves over long periods within campus life events. To address this issue, this paper proposes Psy-Chronicle, a structured data-generation framework for synthesizing long-horizon campus psychological counseling dialogues. We generate a semester-spanning temporal stress event graph to model the chronological order and evolutionary dependencies among campus stress events. Through interactive simulation between a student agent and a counselor agent, together with a structured memory integration mechanism, Psy-Chronicle generates long-horizon dialogues with continuity across counseling sessions.

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

Backtranslation Augmented Direct Preference Optimization for Neural Machine Translation

arXiv:2604.25702v3 Announce Type: replace Abstract: Contemporary neural machine translation (NMT) systems are almost exclusively built by training on supervised parallel data. Despite the tremendous progress achieved, these systems still exhibit persistent translation errors. This paper proposes that a post-training paradigm based on reinforcement learning (RL) can effectively rectify such mistakes. We introduce a novel framework that requires only a general text corpus and an expert translator which can be either human or an AI system to provide iterative feedback. In our experiments, we focus specifically on English-to-German translation as a representative high-resource language pair. Crucially, we implement this RL-based post-training using Direct Preference Optimization (DPO). Applying our DPO-driven framework to the gemma3-1b model yields a significant improvement in translation quality, elevating it's COMET score from 0.703 to 0.747 on the English to German task. The results dem

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

Effective Strategies for Asynchronous Software Engineering Agents

arXiv:2603.21489v2 Announce Type: replace Abstract: AI agents have become increasingly capable at isolated software engineering (SWE) tasks such as resolving issues on Github. Yet long-horizon tasks involving multiple interdependent subtasks still pose challenges both with respect to accuracy, and with respect to timely completion. A natural approach to solving these long-horizon tasks in a timely manner is asynchronous multi-agent collaboration, where multiple agents work on different parts of the task at the same time. But effective application of multi-agent systems has proven surprisingly difficult: concurrent edits by multiple agents interfere with each other, dependencies are difficult to synchronize, and combining partial progress into a coherent whole is challenging. On the other hand, human developers have long relied on mature collaboration infrastructure to manage these challenges in large software projects. Inspired by these collaboration primitives, we introduce Centralize

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

Named-Entity Recognition in the Crime Domain (CrimeNER): Case Study and Dataset

arXiv:2603.02150v2 Announce Type: replace Abstract: The extraction of critical information from crime-related documents is a crucial task for law enforcement agencies. The extraction of this information can be interpreted as a Named-Entity Recognition (NER) task. However, there is a considerable lack of adequately annotated data on general real-world crime scenarios. To address this issue, we present CrimeNER, a case study of crime-related NER, and a general crime-related Named-Entity Recognition database (CrimeNER-db), consisting of more than 1.5K annotated documents extracted from public reports of terrorist attacks and the US Department of Justice's press notes. We define 4 coarse types of crime entity and 21 fine-grained entity types. We address the quality of the presented database with experiments using fully supervised finetuned general NER models and zero- and few-shot experiments to address the generalization capabilities. The database is available on GitHub.

Source ↗
technology Thu, 09 Jul 2026 00:00:00 -0400
arXiv cs.CL

$C$-$\Delta\Theta$: Circuit-Restricted Weight Arithmetic for Selective Refusal

arXiv:2602.04521v2 Announce Type: replace Abstract: Modern deployments require LLMs to enforce safety policies at scale, yet many controls rely on inference-time interventions that add recurring compute cost and serving complexity. Activation steering is widely used, but it requires runtime hooks and scales cost with the number of generations; conditional variants improve selectivity by gating when steering is applied but still retain an inference-time control path. We ask whether selective refusal can be moved entirely offline: can a mechanistic understanding of category-specific refusal be distilled into a circuit-restricted weight update that deploys as a standard checkpoint? We propose C-{\Delta}{\theta} Circuit Restricted Weight Arithmetic}, which (i) localizes refusal-causal computation as a sparse circuit using EAP-IG and (ii) computes a constrained weight update {\Delta}{\theta}C supported only on that circuit (typically <5% of parameters). Applying {\Delta}{\theta}C yields a d

Source ↗
Showing 7801–7850 of 18624 signals
← Prev Page 157 of 373 Next →