EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Wed, 22 Jul 2026 14:19:00 +0000
MedCity News

Hidden in Plain Sight: The Cost Crisis in Women’s Healthcare

Women have to navigate the fragmented system with very little support, coordinating between providers and often interpreting conflicting guidance. The post Hidden in Plain Sight: The Cost Crisis in Women’s Healthcare appeared first on MedCity News .

Source ↗
technology Wed, 22 Jul 2026 14:13:04 +0000
HN: edtech

A 12-Year-Old on Redesigning EdTech Systems

Article URL: https://micahblachman.com/p/on-redesigning-edtech-systems Comments URL: https://news.ycombinator.com/item?id=49007169 Points: 3 # Comments: 0

Source ↗
regulation Wed, 22 Jul 2026 14:00:00 -0400
K-12 Dive

As Massachusetts codifies science of reading, experts expect results to take a few years

The state’s law takes effect in 2027-28, but one former state education chief doesn’t expect meaningful performance data until three years later.

Source ↗
technology Wed, 22 Jul 2026 13:32:00 -0400
EdTech Mag (Higher)

Review: Jabra Biz 1100 EDU Headset Delivers Where It Counts

Walk into any university computer lab during finals week or remote learning sessions, and audio challenges are immediately apparent. Students on adjacent machines bleed sound into each other’s ears. Spoken instructions compete with the hum of the heating and cooling system. Students toss headsets into bags full of books and personal items, subjecting hardware to the kind of punishment that turns most consumer headsets into trash within weeks. The Jabra Biz 1100 EDU was built in direct response to these conditions, and I recently tested it to ensure it could stand up to the demands of higher…

Source ↗
regulation Wed, 22 Jul 2026 13:21:33 -0400
K-12 Dive

Nearly all districts require civics, but fewer offer hands-on opportunities

Only 35% offer experiential civics opportunities, and just 13% mandate them, according to a survey from Rand Corp. and CRPE.

Source ↗
technology Wed, 22 Jul 2026 13:17:00 +0000
MedCity News

What Healthcare Organizations Miss When It Comes to Reimbursement Accuracy

At the end of the day, the most important question is a simple one: Do you actually know if you’re being reimbursed correctly? The post What Healthcare Organizations Miss When It Comes to Reimbursement Accuracy appeared first on MedCity News .

Source ↗
technology Wed, 22 Jul 2026 13:11:00 +0000
MedCity News

Better Prescribing, Not Less Prescribing, Is What Behavioral Health Needs

The policy conversation needs to shift from just deprescribing alone to building the clinical infrastructure that delivers consistent, high-quality behavioral health care. The post Better Prescribing, Not Less Prescribing, Is What Behavioral Health Needs appeared first on MedCity News .

Source ↗
regulation Wed, 22 Jul 2026 12:30:00 +0000
The 74

Opinion: Unjust Child Welfare System Targeted Buttegieg’s Family — Like Too Many Others

When Child Protective Services and a police officer showed up at the home of former Secretary of Transportation Pete Buttigieg, he assumed, as many people do, that this was an abuse of an otherwise well-meaning and well-functioning system. But what happened to Buttigieg’s family is not an anomaly. Every day, parents and children torn apart […]

Source ↗
behavior Wed, 22 Jul 2026 12:01:23 +0000
District Admin

St Louis-area school districts are losing students. A new task force thinks consolidation could help

The Missouri State Board of Education is expected to vote on the accreditation statuses of over 500 districts across the state next year. The post St Louis-area school districts are losing students. A new task force thinks consolidation could help appeared first on District Administration .

Source ↗
behavior Wed, 22 Jul 2026 11:50:06 +0000
District Admin

A $40,000-per-year AI school with no teachers is opening in Oklahoma this August

An AI-powered, billionaire-backed chain of private schools intends to educate students with two hours of virtual learning a day and entrepreneurship workshops. The post A $40,000-per-year AI school with no teachers is opening in Oklahoma this August appeared first on District Administration .

Source ↗
regulation Wed, 22 Jul 2026 10:30:00 +0000
The 74

America’s Trillion-Dollar School System: 5 Trends Explain Where the Money Goes

In April, the National Center for Education Statistics reported that the 50 states and Washington, D.C., took in $1 trillion in federal, state and local revenue for K-12 public education in the 2023-24 school year. One. Trillion. Dollars. That’s a remarkable milestone. And while it doesn’t tell the whole story — it represents only revenues, […]

Source ↗
behavior Wed, 22 Jul 2026 10:00:00 +0000
eSchool News

Chronic absenteeism isn’t the problem–it’s the signal

Every school leader I know has a version of the same moment: You pull up a student's attendance record expecting to find a discipline problem, and instead you find a family going through a crisis none of us would wish on anyone.

Source ↗
technology Wed, 22 Jul 2026 09:57:11 +0000
Tech & Learning

Tech & Learning Launches “Best for Back to School” Contest

Celebrating Exceptional Products That Support Educators Heading Back To School

Source ↗
technology Wed, 22 Jul 2026 09:26:34 +0000
HN: education

Minuet, a KDE application for music education, calls for testers

Article URL: https://sandroandrade.org/minuet-26-08-call-for-testers/ Comments URL: https://news.ycombinator.com/item?id=49003943 Points: 10 # Comments: 1

Source ↗
technology Wed, 22 Jul 2026 09:00:00 +0000
eCampus News

Escaping the tech debt trap: Why governance matters more in the age of AI

Artificial intelligence is now emerging as a strategic priority across higher education, but many institutions are still determining how to move from experimentation to institution-wide adoption. The post Escaping the tech debt trap: Why governance matters more in the age of AI appeared first on eCampus News .

Source ↗
audience Wed, 22 Jul 2026 07:00:00 +0000
Inside Higher Ed

My Study Abroad

My Study Abroad rachel.toor Wed, 07/22/2026 - 03:00 AM “Foreignness” isn’t what makes study abroad meaningful. Attention is. Not what you do or see, but what affects and changes you. Byline(s) Rachel Toor

Source ↗
audience Wed, 22 Jul 2026 07:00:00 +0000
Inside Higher Ed

That Chart From Brown Isn’t Really About Cheating

That Chart From Brown Isn’t Really About Cheating Elizabeth Redden Wed, 07/22/2026 - 03:00 AM We made grades the point. AI just made them cheap. Byline(s) Leo Schumann

Source ↗
audience Wed, 22 Jul 2026 07:00:00 +0000
Inside Higher Ed

Ohio State Prof Who Tackled Journalist Pleads No Contest

Ohio State Prof Who Tackled Journalist Pleads No Contest kathryn.palmer… Wed, 07/22/2026 - 03:00 AM Byline(s) Kathryn Palmer

Source ↗
audience Wed, 22 Jul 2026 07:00:00 +0000
Inside Higher Ed

Academic Freedom at UT ‘Being Dismantled,’ Says Professor Denied Tenure

Academic Freedom at UT ‘Being Dismantled,’ Says Professor Denied Tenure Emma Whitford Wed, 07/22/2026 - 03:00 AM Border-surveillance researcher Iván Chaar López received a near-unanimous recommendation from faculty and outstanding external support letters. Still, he was denied tenure without explanation. Byline(s) Emma Whitford

Source ↗
audience Wed, 22 Jul 2026 07:00:00 +0000
Inside Higher Ed

Degree Attainment Gaps Persist for Latino Students

Degree Attainment Gaps Persist for Latino Students Sara Weissman Wed, 07/22/2026 - 03:00 AM Byline(s) Sara Weissman

Source ↗
audience Wed, 22 Jul 2026 07:00:00 +0000
Inside Higher Ed

Even Failing Dual Enrollment Students More Likely to Enroll in College

Even Failing Dual Enrollment Students More Likely to Enroll in College gianna.jakubowski Wed, 07/22/2026 - 03:00 AM A new study out of Brown’s Annenberg Institute explores the impact that failing dual-enrollment courses has on high school students’ decision to attend college. Byline(s) Gianna Jakubowski

Source ↗
audience Wed, 22 Jul 2026 07:00:00 +0000
Inside Higher Ed

House Republicans Move to Codify the Definition of Biological Sex

House Republicans Move to Codify the Definition of Biological Sex jessica.blake@… Wed, 07/22/2026 - 03:00 AM Democrats argue the bill will actually strip LGBTQ+ students of civil rights protections, going beyond what the Supreme Court intended. Byline(s) Jessica Blake

Source ↗
audience Wed, 22 Jul 2026 07:00:00 +0000
Inside Higher Ed

Survey: Majority of Americans Say College Not Affordable

Survey: Majority of Americans Say College Not Affordable kathryn.palmer… Wed, 07/22/2026 - 03:00 AM Meanwhile, nearly three-quarters of parents want their child to pursue some type of higher education after high school, with the strongest preference for a four-year degree. Byline(s) Kathryn Palmer

Source ↗
audience Wed, 22 Jul 2026 07:00:00 +0000
Inside Higher Ed

DOJ Takes Aim at UC San Diego Med School’s Use of Hardship in Admissions

DOJ Takes Aim at UC San Diego Med School’s Use of Hardship in Admissions Katherine Knott Wed, 07/22/2026 - 03:00 AM Byline(s) Katherine Knott

Source ↗
audience Wed, 22 Jul 2026 05:00:00 -0400
Higher Ed Dive

Out of reach? Adults grade colleges on their affordability.

A Lumina Foundation and Gallup survey found widespread cost concerns — but most parents still want their children to pursue postsecondary education.

Source ↗
regulation Wed, 22 Jul 2026 05:00:00 -0400
K-12 Dive

AI embraced by more students and educators, Instructure finds

However, about a third of educators and K-12 parents support restrictions on the technology in school.

Source ↗
behavior Wed, 22 Jul 2026 00:00:00 GMT
EdSurge

What Does AI Cost When We Skip the Work?

This Week with EdSurge podcast examines what holds up when artificial intelligence moves fast.

Source ↗
behavior Wed, 22 Jul 2026 00:00:00 GMT
EdSurge

We Must Stop Using AI to ‘Level-Down’ Our Students

How AI deprived my students of the reading struggle they needed to succeed — and five better ways to use it.

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

PARSE: Provenance-Aware Retrieval Sanitization for Professional Domain LLM Agents

arXiv:2606.17467v2 Announce Type: replace-cross Abstract: Prompt injection defenses evaluated on synthetic benchmarks do not generalize to real enterprise documents, which are longer, denser, and interleave legitimate authority language with factual content. We demonstrate this gap with a real-document benchmark of 122 tasks across five professional domains (financial, legal, medical, scientific, DevOps) using actual SEC filings, Federal Register rules, PubMed abstracts, arXiv papers, and GitHub postmortems. Paraphrasing, the strongest defense on synthetic benchmarks, shows no statistically significant attack success rate reduction on real documents (p=0.500) while degrading utility from 91.8% to 82.8%. We introduce PARSE (Provenance-Aware Retrieval Sanitization), a domain-aware, fact-preserving sanitization pipeline that classifies each sentence by injection likelihood, extracts structured facts before rewriting, and verifies fact preservation via a consistency-checking loop. A direct

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

More Edits, More Stable: Understanding the Lifelong Normalization in Sequential Model Editing

arXiv:2605.11836v2 Announce Type: replace-cross Abstract: Lifelong Model Editing aims to continuously update evolving facts in Large Language Models while preserving unrelated knowledge and general capabilities, yet it remains plagued by catastrophic forgetting and model collapse. Empirically, we find that recent editors resilient over long horizons share the same core strategy: Lifelong Normalization (LN), which normalizes value gradients using running statistics. Removing LN causes immediate performance collapse, and we observe a counter-intuitive positive cumulative effect where early edits can promote the success of future edits. Yet the mechanism of LN remains a "black box", leaving its precise role in lifelong stability poorly understood. In this work, we provide the first theoretical account of LN in the lifelong regime. Our analysis reveals a self-reinforcing stability loop and proves that, when combined with ridge-regularized regression, LN yields parameter updates with asympt

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

FinRAG-12B: A Production-Validated Recipe for Grounded Question Answering in Banking

arXiv:2605.05482v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are rapidly being adopted across various domains. However, their adoption in banking industry faces resistance due to demands for high accuracy, regulatory compliance, and the need for verifiable and grounded responses. We present a unified, data-efficient framework for training grounded domain-specific LLMs that optimizes answer quality, citation grounding, and calibrated refusal under real-world deployment constraints. First, we describe a data generation pipeline that combines LLM-as-a-Judge filtering, citation annotation, and curriculum learning with only 143M tokens. The resulting 12B model achieves high answer quality outperforming GPT-4.1 on citation grounding, with a modest citation tradeoff versus the untuned base. Second, we propose a calibrated refusal mechanism: training on 22% unanswerable examples yield a 12% "I don't know" rate, substantially improving over the base model's unsafe 4.3%

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

Robust Reasoning Benchmark

arXiv:2604.08571v3 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) achieve high performance on standard mathematical benchmarks, their problem-solving abilities depend on the context and textual formatting. We introduce the Robust Reasoning Benchmark (RRB), a pipeline of 13 deterministic textual perturbations applied to AIME 2024 and AIME 2025. Evaluating 8 state-of-the-art models, we find that frontier models are largely resilient, with the notable exception of Claude, which categorically refuses many transformed prompts. Open-weights reasoning models exhibit a range of failure modes under structural noise (cognitive thrashing, tokenization breakdown, and reasoning collapse), with up to 54% average accuracy drops across perturbations and up to 100% on some. We further study one of these failure modes in isolation: attention dilution caused by the model's own chain-of-thought. By tasking models with solving multiple independent mathematical problems sequential

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs

arXiv:2603.21014v2 Announce Type: replace-cross Abstract: Mechanistic interpretability seeks to understand how Large Language Models (LLMs) represent and process information. Recent approaches based on dictionary learning and transcoders enable representing model computation in terms of sparse, interpretable features and their interactions, giving rise to feature attribution graphs. However, these graphs are often large and redundant, limiting their interpretability in practice. Cross-Layer Transcoders (CLTs) address this issue by sharing features across layers while preserving layer-specific decoding, yielding more compact representations, but remain difficult to train and analyze at scale. We introduce an open-source library for end-to-end training and interpretability of CLTs. Our framework integrates scalable distributed training with model sharding and compressed activation caching, a unified automated interpretability pipeline for feature analysis and explanation, attribution gra

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

Vectorizing the Trie: Efficient Constrained Decoding for LLM-based Generative Retrieval on Accelerators

arXiv:2602.22647v2 Announce Type: replace-cross Abstract: Generative retrieval has emerged as a powerful paradigm for LLM-based recommendation. However, industrial recommender systems often benefit from restricting the output space to a constrained subset of items based on business logic (e.g. enforcing content freshness or product category), which standard autoregressive decoding cannot natively support. Moreover, existing constrained decoding methods that make use of prefix trees (Tries) incur severe latency penalties on hardware accelerators (TPUs/GPUs). In this work, we introduce STATIC (Sparse Transition Matrix-Accelerated Trie Index for Constrained Decoding), an efficient and scalable constrained decoding technique designed specifically for high-throughput LLM-based generative retrieval on TPUs/GPUs. By flattening the prefix tree into a static Compressed Sparse Row (CSR) matrix, we transform irregular tree traversals into fully vectorized sparse matrix operations, unlocking massi

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

Lower-Resource, Higher Scores: Language Bias in LLM Evaluators

arXiv:2607.14480v2 Announce Type: replace Abstract: LLM evaluators (trained reward models and prompted LLM-as-a-Judge) are routinely validated via pairwise accuracy. In a multilingual setting, this operates under the premise that high pairwise accuracy implies reliable, language-neutral scoring. We show that this assumption does not hold. We conduct experiments with semantically identical instruction-response pairs across 23 languages, and find that multilingual evaluators assign significantly different scores to different evaluation languages. The bias is statistically significant and consistent across eight open-weight evaluators of different architectures and training paradigms, persists in frontier judges, and is strongly correlated with language resource level: lower-resource languages are scored more generously. Meanwhile, these biases are invisible to pairwise accuracy: evaluators achieve above 90% pairwise accuracy, yet have up to 43% difference in acceptance rate across langua

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

Measuring Human Value Expression in Social Media Texts: Calibrated LLM Annotation and Encoder Transfer

arXiv:2606.11018v2 Announce Type: replace Abstract: Measuring subjective constructs in naturally occurring social media text requires annotation procedures that are theoretically grounded, empirically validated, and transferable to an encoder model for scalable prediction. Using non-English social media posts annotated according to Schwartz's theory of basic human values, we investigate how different LLMs, prompts, and instruction languages operationalize the expression of values in text. We argue that although texts may permit multiple plausible interpretations, theory-based value definitions can constrain interpretations and reduce spurious value attributions. Beyond precision, recall, and F1, we evaluate structural alignment between values, error structure, confidence-ambiguity relations, and annotation stability. We show that different LLMs produce different value interpretations. Iterative prompt calibration through error analysis reduces misattributions and improves alignment wit

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling

arXiv:2604.27039v2 Announce Type: replace Abstract: Tokens are the fundamental units of computation in modern autoregressive models, and generation length directly influences both inference cost and reasoning performance. Despite its importance, existing approaches model length primarily at the coarse sequence level. We introduce the Length Value Model (LenVM), a token-level framework that estimates the remaining generation length at every decoding step. By formulating length modeling as a value estimation problem and assigning a constant negative reward to each generated token, LenVM predicts a bounded, discounted return that is a monotone proxy for the remaining generation horizon. This value formulation provides annotation-free, dense, unbiased, and scalable supervision. Experiments on LLMs and VLMs show that LenVM supports exact control, continuous performance--efficiency steering, length prediction, and interpretation. On LIFEBench-token, it raises the exact-length score of Qwen2.

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

Large Language Models Explore by Latent Distilling

arXiv:2604.24927v2 Announce Type: replace Abstract: Generating diverse responses is crucial for test-time scaling of large language models (LLMs), yet standard stochastic sampling mostly yields surface-level lexical variation, limiting semantic exploration. In this paper, we propose Exploratory Sampling (ESamp), a decoding approach that explicitly encourages semantic diversity during generation. ESamp is motivated by the well-known observation that neural networks tend to make lower-error predictions on inputs similar to those encountered before, and incur higher prediction error on novel ones. Building on this property, we train a lightweight Distiller at test time to predict deep-layer hidden representations of the LLM from its shallow-layer representations to model the LLM's depth-wise representation transitions. During decoding, the Distiller continuously adapts to the mappings induced by the current generation context. ESamp uses the prediction error as a novelty signal to reweigh

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

Doctorina MedBench-ICD10: A Dialogue-Based Benchmark and Evaluation Framework for Agent-Based Medical AI

arXiv:2603.25821v2 Announce Type: replace Abstract: We present Doctorina MedBench, a comprehensive evaluation framework for agent-based medical AI based on the simulation of realistic physician-patient interactions. Unlike traditional medical benchmarks that rely on solving standardized test questions, the proposed approach models a multi-step clinical dialogue in which either a physician or an AI system must collect medical history, analyze attached materials (including laboratory reports, images, and medical documents), formulate differential diagnoses, and provide personalized recommendations. System performance is evaluated using the D.O.T.S. metric, which consists of four components: Diagnosis, Observations/Investigations, Treatment, and Step Count, enabling assessment of both clinical correctness and dialogue efficiency. The system also incorporates a multi-level testing and quality monitoring architecture designed to detect model degradation during both development and deploymen

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

Estimating near-verbatim extraction risk in language models with decoding-constrained beam search

arXiv:2603.24917v2 Announce Type: replace Abstract: Recent work shows that standard greedy-decoding extraction methods for quantifying memorization in LLMs miss how extraction risk varies across sequences. Probabilistic extraction -- computing the probability of generating a target suffix given a prefix under a decoding scheme -- addresses this, but is tractable only for verbatim memorization, missing near-verbatim instances that pose similar privacy and copyright risks. Quantifying near-verbatim extraction risk is expensive: the set of near-verbatim suffixes is combinatorially large, and reliable Monte Carlo (MC) estimation can require ~100,000 samples per sequence. To mitigate this cost, we introduce decoding-constrained beam search, which yields deterministic lower bounds on near-verbatim extraction risk at a cost comparable to ~20 MC samples per sequence. Across experiments, our approach surfaces information invisible to verbatim methods: many more extractable sequences, substantia

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

PlotTwist: A Creative Plot Generation Framework with Small Language Models

arXiv:2603.16410v2 Announce Type: replace Abstract: Creative plot generation presents a fundamental challenge for language models: transforming a concise premise into a coherent narrative that sustains global coherence, character development, pacing, tone consistency, and emotional progression. Although recent Large Language Models (LLMs) demonstrate strong fluency on general-purpose tasks, they require preference alignment to perform well on domain-specific tasks such as creative plot generation. However, conducting such alignment at the scale of frontier LLMs is computationally prohibitive, significantly limiting accessibility and practical deployment. To address this, we present PlotTwist, a structured framework that enables Small Language Models (SLMs) with $\leq$3B active parameters to generate high-quality, premise-conditioned plots competitive with frontier systems of vastly greater parameter scale. Our approach decomposes generation into three specialized components: (1) an Asp

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

PyPhonPlan: Simulating phonetic planning with dynamic neural fields and task dynamics

arXiv:2603.16299v2 Announce Type: replace Abstract: We introduce PyPhonPlan, a Python toolkit for implementing dynamical models of phonetic planning using coupled dynamic neural fields and task dynamic simulations. The toolkit provides modular components for defining planning, perception and memory fields, as well as between-field coupling, gestural inputs, and using field activation profiles to solve tract variable trajectories. We illustrate the toolkit's capabilities through an example application: simulating production/perception loops with a coupled memory field, which demonstrates the framework's ability to model interactive speech dynamics using representations that are temporally-principled, neurally-grounded, and phonetically-rich. PyPhonPlan is released as open-source software and contains executable examples to promote reproducibility, extensibility, and cumulative computational development for speech communication research.

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

LinguistAgent Technical Report: A Reflective Multi-Model Platform for Automated Linguistic Annotation

arXiv:2602.05493v2 Announce Type: replace Abstract: Data annotation remains a significant bottleneck in the field of humanities and social sciences, particularly for complex linguistic tasks such as metaphor identification. While Large Language Models (LLMs) show promise, a significant gap remains between the theoretical capability of LLMs and their practical utility for researchers. This paper introduces LinguistAgent, an integrated, user-friendly platform that leverages a reflective multi-model architecture to automate linguistic annotation. The platform comprises an Annotator and an optional Reviewer to simulate a peer-review process. This platform supports comparative experiments across three main paradigms: Prompt Engineering (Zero-shot/Few-shot/Chain-of-thought), Retrieval-Augmented Generation, and Fine-tuning. We demonstrate LinguistAgent's efficacy by replicating the task of metaphor identification from a published study, which provides real-time token-level evaluation (F1 and

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

Tears or Cheers? Benchmarking LLMs via Culturally Elicited Distinct Affective Responses

arXiv:2601.13024v2 Announce Type: replace Abstract: Culture serves as a fundamental determinant of human affective processing and profoundly shapes how individuals perceive and interpret emotional stimuli. Despite this intrinsic link extant evaluations regarding cultural alignment within Large Language Models primarily prioritize declarative knowledge such as geographical facts or established societal customs. These benchmarks remain insufficient to capture the subjective interpretative variance inherent to diverse sociocultural lenses. To address this limitation, we introduce CEDAR, a multimodal benchmark constructed entirely from scenarios capturing Culturally \underline{\textsc{E}}licited \underline{\textsc{D}}istinct \underline{\textsc{A}}ffective \underline{\textsc{R}}esponses. To construct CEDAR, we implement a novel pipeline that leverages LLM-generated provisional labels to isolate instances yielding cross-cultural emotional distinctions, and subsequently derives reliable groun

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

PRISP: Privacy-Safe Few-Shot Personalization via Lightweight Adaptation

arXiv:2601.06471v2 Announce Type: replace Abstract: Large language model (LLM) personalization aims to adapt general-purpose models to individual users. Most existing methods, however, are developed under data-rich and resource-abundant settings, often incurring privacy risks. In contrast, realistic personalization typically occurs after deployment under (i) extremely limited user data, (ii) constrained computational resources, and (iii) strict privacy requirements. We propose PRISP, a lightweight and privacy-safe personalization framework tailored to these constraints. PRISP leverages a Text-to-LoRA hypernetwork to generate task-aware LoRA parameters from task descriptions, and enables efficient user personalization by optimizing a small subset of task-aware LoRA parameters together with minimal additional modules using few-shot user data. Experiments on a few-shot variant of the LaMP benchmark demonstrate that PRISP achieves strong overall performance compared to prior approaches, wh

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

Epistemic Familiarity is Associated With Belief Stability in Large Language Models

arXiv:2511.19166v4 Announce Type: replace Abstract: Large language models (LLMs) are widely used as information sources, yet small changes in semantic assumptions can destabilize their beliefs. We introduce P-StaT (Perturbation Stability of Truth), a framework for evaluating belief stability under matched semantic perturbations in both representational and behavioral settings. Across 21 LLMs and three domains, we compare perturbations involving familiar Fictional statements against synthetically generated and unfamiliar Synthetic statements. Unfamiliar Synthetic perturbations generally induce greater epistemic instability than familiar Fictional perturbations, with behavioral belief retraction rates frequently exceeding 0.5 (50%). Finally, exploratory clustering analyses reveal recurring themes among retracted statements, including ambiguity, technical terminology, and obscure concepts. These results show that epistemic familiarity is systematically associated with stability under sema

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

Breaking the MoE LLM Trilemma: Dynamic Expert Clustering with Structured Compression

arXiv:2510.02345v4 Announce Type: replace Abstract: Mixture-of-Experts (MoE) Large Language Models (LLMs) face a trilemma of load imbalance, parameter redundancy, and communication overhead. We introduce a unified framework based on dynamic expert clustering and structured compression to address these issues cohesively. Our method employs an online clustering procedure that periodically regroups experts using a fused metric of parameter and activation similarity, which stabilizes expert utilization. To our knowledge, this is one of the first frameworks to leverage the semantic embedding capability of the router to dynamically reconfigure the model's architecture during training for substantial efficiency gains. Within each cluster, we decompose expert weights into a shared base matrix and extremely low-rank residual adapters, achieving up to fivefold parameter reduction per group while preserving specialization. This structure enables a two-stage hierarchical routing strategy: tokens a

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

Hyperdimensional Probe: Decoding LLM Representations via Vector Symbolic Architectures

arXiv:2509.25045v3 Announce Type: replace Abstract: Despite their capabilities, Large Language Models (LLMs) remain opaque with limited understanding of their internal representations. Current interpretability methods either focus on input-oriented feature extraction, such as supervised probes and Sparse Autoencoders (SAEs), or on output distribution inspection, such as logit-oriented approaches. A full understanding of LLM vector spaces, however, requires integrating both perspectives, something existing approaches struggle with due to constraints on latent feature definitions. We introduce the Hyperdimensional Probe, a hybrid supervised probe that combines symbolic representations with neural probing. Leveraging Vector Symbolic Architectures (VSAs) and hypervector algebra, it unifies prior methods: the top-down interpretability of supervised probes, SAE's sparsity-driven proxy space, and output-oriented logit investigation. By combining the supervised learning paradigm of traditional

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

Guard Vector: Beyond English LLM Guardrails with Task-Vector Composition and Streaming-Aware Prefix SFT

arXiv:2509.23381v2 Announce Type: replace Abstract: We introduce Guard Vector, a safety task vector computed as the parameter difference between a guardrail model (Guard Model) and a same-architecture pretrained language model. Composing this vector with a target language model yields a Target Guard Model (TGM). We then adapt TGM with a streaming-aware approach that combines prefix-based training and evaluation with a classifier that produces a single-token output. With this composition alone, TGM improves classification quality over established Guard Models across standard safety suites and enables language extensibility to Chinese, Japanese, and Korean, requiring neither additional training nor target language labels for this composition step. It also demonstrates model portability across two widely used public guardrail backbones, Llama and Gemma. With prefix SFT (supervised fine-tuning), TGM preserves classification quality under streaming by aligning the behavior between prefix in

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CL

Evaluating Style-Personalized Text Generation: Challenges and Directions

arXiv:2508.06374v3 Announce Type: replace Abstract: With the surge of large language models (LLMs) and their ability to produce customized output, style-personalized text generation--"write like me"--has become a rapidly growing area of interest. However, style personalization is highly specific, relative to every user, and depends strongly on the pragmatic context, which makes it uniquely challenging. Although prior research has introduced benchmarks and metrics for this area, they tend to be non-standardized and have known limitations (e.g., poor correlation with human subjects). LLMs have been found to not capture author-specific style well, it follows that the metrics themselves must be scrutinized carefully. In this work we critically examine the effectiveness of the most common metrics used in the field, such as BLEU, embeddings, and LLMs-as-judges. We evaluate these metrics using our proposed style discrimination benchmark, which spans eight diverse writing tasks across three ev

Source ↗
Showing 551–600 of 18349 signals
← Prev Page 12 of 367 Next →