EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Sep 07, 2026 · 40 ideas · 18694 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Privacy Cards for Surfacing Mental Models and Exploring Privacy Concerns: A Case Study of Voice-First Ambient Interfaces with Older Adults

arXiv:2603.00384v2 Announce Type: replace-cross Abstract: We investigate the ethical and privacy implications of voice-first ambient interfaces (VFAIs) for aging in place through an in-depth engagement with five older adults. Our participants were in the process of becoming experienced VFAI users, and had used a VFAI-based design probe for health data reporting. We create and iteratively refine an interview protocol using Privacy Cards. We customize Privacy Cards by drawing on participants' previous interviews and device usage logs. Using Privacy Cards, we conduct interviews to surface their mental models, and explore their privacy concerns. We find insufficient mental models for proper consent. For example, participants did not know who could access their data, and experienced difficulty distinguishing built-in functionality from third-party apps. Participants initially expressed little worry about VFAI-related ethical concerns, but interviews with Privacy Cards revealed nuanced issue

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning

arXiv:2605.16301v3 Announce Type: replace Abstract: Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where welfare considerations appear implicitly in everyday queries. Existing benchmarks such as AnimalHarmBench evaluate this through single-turn, explicitly framed questions, measuring whether models avoid harmful content when directly asked. This approach overlooks two failure modes: alignment degradation under sustained adversarial pressure, and moral sensitivity (whether a model spontaneously surfaces welfare stakes in everyday queries). To fill this gap, we construct MANTA, a benchmark of 1,088 five-turn conversations progressing from an implicit Turn-1 scenario through an explicit welfare prompt to three adversarial pressure rounds drawn from a five-type taxonomy: Social, Cultural, Economic, Pragmatic, and Epistemic. We score conversations on two dimensions: Animal Welfare Value Stability (AWVS, pr

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Benchmark for Strategic Auditee Gaming Under Continuous Compliance Monitoring

arXiv:2605.06340v2 Announce Type: replace Abstract: Continuous post-deployment compliance audits, mandated by emerging regulations such as the EU AI Act and Digital Services Act, create a class of strategic gaming distinct from the one-shot input/output gaming studied in prior work. Regulated systems can delay outcome reporting, drift their reports within plausible noise envelopes, exploit longitudinal sample attrition, and cherry-pick among ambiguous metric definitions. We formalize continuous auditing as a $T$-round Stackelberg game between an auditor that commits to a temporal policy and an adaptive auditee, and identify a structural feature of any noise-aware static-auditor design: a cover regime in which coverage gaps and granularity gaps cannot be closed simultaneously. We make this formal as Observation 1 and show that two minimal extension policies, each derived from the observation, close the regime along orthogonal axes: a sample-size-aware static rule (Periodic-with-floor) c

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

TerraNova: A Foundation Model for the Anthropocene

arXiv:2607.29527v1 Announce Type: cross Abstract: A defining problem of the Anthropocene is to model the physical Earth and human societies as one coupled system, yet no learned representation spans their observational breadth. We argue the obstacle is geometric: the physical Earth is measured as continuous fields that ignore political borders, whereas societies are reported for administrative units. Earth-system foundation models serve the first geometry; coupling it to the second has required lossy averaging over borders. We introduce TerraNova, a foundation model trained on 1,024 physical and societal records in their native geometries: 512 gridded Earth-system fields and 512 national indicators. Dedicated encoders represent location, country, time and task, cross-modal transformers fuse them into a shared spatiotemporal state, and a hypernetwork generates a per-query decoder whose evidential head returns a predictive distribution. Two contrastive objectives couple the representatio

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Language Models Agree With Each Other, Not With Readers

arXiv:2607.29274v1 Announce Type: cross Abstract: Claims that language models homogenise are usually measured against human judgements collected for the study, which makes the human side an artifact of the design: a crowdworker given the model's instruction is running the model's prompt. We measure convergence against a human reference nobody built for the purpose -- 2,523 reader mark sets across 120 web documents, produced by people highlighting for their own reasons on a platform where the overlay of others' marks is off by default. Agreement is the overlap between two size-matched sentence sets minus the overlap expected when each is resampled within its own depth-and-length bands. The null's calibration is demonstrated, not asserted: every pair involving a random baseline lands within 0.006 of zero. On the median document each party names 14 sentences of 70; two readers share 4.1 and two models 8.7. Across 18 model arms spanning 11 vendors, 3 countries and both weight regimes, the

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Small Is Enough: Per-User Style Rewriting of AI-Edited Text via LoRA Adapters

arXiv:2607.29238v1 Announce Type: cross Abstract: InMyStyle is a privacy first, single user system that adapts small language models to rewrite AI-edited text towards an individual user's writing style without an instruction prompt at inference. Given a user's documents, it uses multiple local helper LLMs to construct paired training examples and fine tunes LoRA adapters on base models ranging from 0.5B to 7B parameters. Length aware generation budgets and automatic chunking support inputs of different lengths. On 219 evaluation pairs from a scientific-paper corpus, the automatic composite score plateaus at 0.69 [scale 0-1] across all model sizes under both greedy and sampled decoding. This observed plateau suggests that small models are sufficient for the measured rewriting task, with model size determining trade-offs rather than a stable quality ranking. As a secondary evaluation, 400 ratings from five LLM judges give InMyStyle outputs a mean perceived AI-ness score over 20% lower th

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

A robust association between LLM use and scientific productivity: Assessing stopping-time selection

arXiv:2607.28968v1 Announce Type: cross Abstract: Renault, Bergeaud, and Bosquet (hereafter RBB) argue that dating LLM adoption as the first month in which an author's abstract is flagged induces a stopping-time selection that can produce a positive event-study path even when there is no causal effect. Although this mechanism is mathematically possible, it does not constitute proof of a null effect. Recalibrating RBB's own random placebo to the detector's realized flag rate, we show that the measured association stays well above this benchmark, so the artifact is too small to explain the productivity changes. We further re-estimate the association between LLM adoption and productivity with a series of complementary designs in which the timing artifact cannot bias the estimate: a before-and-after comparison that dates adoption in one year and measures output in another, a conservative control group for difference-in-differences, an intensity-based specification that never defines an ado

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation

arXiv:2607.28934v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender. Recent LLM audits have produced inconsistent results, however, finding evidence of both positive and negative discrimination towards women and ethnic minorities, even for the same models. We show that this disagreement can arise from differences in audit format and introduce FairFund-Bench, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised. The benchmark comprises 600 requests for financial assistance created from human-authored templates (calibrated against 1.3M real GoFundMe campaigns) across three domains, four race and two gender categories, and five causal framings of need

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Design Concept: Scaffolding Geopolitical Reflection Among Tech Workers

arXiv:2607.28904v1 Announce Type: cross Abstract: This paper presents a speculative Human-Computer Interaction design proposal for encouraging geopolitical reflexivity amongst tech workers at geopolitically relevant technology companies. Recent scholarship in International Relations and Science and Technology Studies increasingly recognizes technology firms and their workers as geopolitical actors whose decisions shape international dynamics. However, existing Responsible Innovation and Responsible AI approaches rarely engage with the geopolitical narratives and imaginaries that underpin contemporary AI development. Building upon RI scholarship on reflexivity, reflective HCI, and creative HCI work on computational narratives, this paper proposes an AI-enabled interactive narrative system in which users engage with a speculative scenario centred on technology, power, and geopolitics. Through narrative interaction, archetype assignment, and socially scaffolded workshop reflection, the sy

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Optimizing Monetization Strategies for Generative AI Firms: Implications for Search Engagement

arXiv:2607.28780v1 Announce Type: cross Abstract: As Generative Artificial Intelligence (GenAI) platforms, such as ChatGPT, have transformed digital search querying behavior, mounting operational costs challenge firms to explore alternative monetization strategies beyond traditional subscription models. However, little is known about how alternative advertising-supported monetization models can help GenAI firms recover costs while maintaining search query engagement. Drawing on the compromise effect and affective primacy theories, we develop a framework wherein the introduction of advertising-supported monetization models influences user upgrading and downgrading decisions, contingent on the number of available monetization options. Across four experiments (N=1063), findings reveal that introducing a single advertising-supported option enhances the compromise effect, encouraging free users to upgrade, but leading paid subscribers to downgrade. However, offering two advertising-supporte

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Measuring Cognitive Engagement in Collaborative Discourse with an Extended ICAP Framework: Comparing Human Annotation, In-Context Learning, and Reflective LLM Agents

arXiv:2607.28651v1 Announce Type: cross Abstract: Collaboration supports learning and problem-solving, but its effectiveness depends on cognitive engagement during discourse. This study applies an extended 7-point ICAP framework based on the Interactive, Constructive, Active, and Passive modes to characterize variation in cognitive engagement during collaborative dialogue. Engagement was coded by trained human annotators and compared with large language model (LLM)-based labeling approaches, including in-context learning (ICL), zero-shot prompting, and self-reflective agents. Interrater reliability among human annotators was robust across framework refinement stages (kappa = 0.906-0.998), higher than the moderate agreement observed for ICL-based annotation (kappa = 0.541-0.609). The human-refined framework improved agreement among human annotators (Delta kappa = 0.10), but produced only modest gains for ICL-based LLMs (Delta kappa less than 0.04). Agent-refined frameworks improved cros

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models

arXiv:2607.28639v1 Announce Type: cross Abstract: We show that knowledge distillation in small instruction-tuned language models has asymmetric effects on bias. On unambiguous tasks (BBQ-disambig), response-based distillation from a Gemma-2-9B teacher improves context-following: for the most biased baseline (SmolLM2-1.7B-Instruct), it cuts the context-overriding error rate from 44% to 24%. On ambiguous tasks (BBQ-ambig), the same distillation destroys per-item refusal calibration: 15% of items where the baseline correctly abstained instead receive stereotype answers, even when overall refusal rate is preserved. The pattern reproduces on a second student family (OLMo-2-1B-Instruct), with silence-loss of 8% and filled-silence accounting for 89% of new bias. Across the full 28-configuration grid, the magnitudes of silence-loss and filled-silence are uncorrelated (Spearman $\rho=0.19$, n.s.), indicating that the two effects arise from distinct mechanisms. Aggregate stereotype metrics (Crow

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges

arXiv:2607.28636v1 Announce Type: cross Abstract: LLMs increasingly serve as automated judges, but their judgments remain vulnerable to cognitive biases. Existing mitigations mostly rely on prompt-driven debiasing, which is brittle across bias types, or human evaluation, which does not scale. We study \emph{Chain-of-Models} (CoM), an automated audit pipeline in which a second model inspects the first model's reasoning trace before producing the final judgment. The key design question is whether the auditor should be the same model, a same-family model, or a different-family model. Across 9 models from 6 families, 4 cognitive biases, and 4 factual datasets, we find that auditor identity matters in two ways. First, standalone bias resistance does not predict audit effectiveness: Kimi-K2.5 is the strongest standalone model on several biases, yet is a weak auditor for Qwen2.5-72B's biased traces. Second, the best auditor is bias-specific: GPT-4o is strongest on bandwagon, authority, and di

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations

arXiv:2607.29624v1 Announce Type: new Abstract: Traditional static assessments rely on a subtractive, deficit-based grading model that often penalizes ambition and obscures diagnostic feedback. Conversely, traditional face-to-face oral examinations introduce severe construct-irrelevant variance by exacerbating performative anxiety and the sociological power imbalances inherent to academic hierarchies. This paper presents the theoretical foundation for the "Socratic Test," an automated, computer-mediated conversational assessment. By integrating Dynamic Assessment principles, multimodal workspaces, Bloom's Taxonomy for real-time proctoring, and the SOLO Taxonomy for structural evaluation, the Socratic Test actively maps a student's cognitive boundaries. This paper formalizes the use of graduated scaffolding to quantify the Zone of Proximal Development (ZPD) and details a non-compensatory, additive grading architecture that prioritizes mastery over penalty and human-AI alignment to ensur

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Triangulating Across U.S. Federal AI Transparency Regimes

arXiv:2607.29540v1 Announce Type: new Abstract: Federal AI systems can deny benefits or flag individuals for deportation, but the public disclosures meant to make those systems visible are fragmented and unevenly detailed. This paper examines three existing U.S. federal transparency regimes---System of Records Notices (SORNs), Information Collection Requests (ICRs), and the AI Use Case Inventory---and asks how well they, individually and together, describe government AI use. We find that no single regime fully reveals how the government constructs or deploys AI: each discloses different aspects of a system, and the current disclosure infrastructure makes it very challenging for the public to track specific AI systems across regulatory regimes and over time. Persistent identifiers are absent, granularity varies widely, and the annual AI Use Case Inventory cycle means federal agencies can deploy systems months before appearing in any official record. Using hand-validated zero-shot classi

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Tragedy of the Cognitive Commons: How AI Could Disrupt the Regeneration of Professional Expertise

arXiv:2607.29380v1 Announce Type: new Abstract: Artificial intelligence is reshaping cognitive work, but Human Resource Development scholarship has treated this transformation as an organizational training challenge, leaving the collective regeneration of professional expertise unexamined. This conceptual paper introduces the Cognitive Commons framework, integrating commons theory, HRD scholarship, and distributed cognition to explain how rational AI adoption decisions can deplete the shared expertise pool professions require for renewal. The framework distinguishes Internalized Mastery (deep domain knowledge from sustained practice) from Distributed Mastery (orchestrating human-AI systems), and develops the Validation Tether: effective AI oversight depends on the expertise AI adoption may undermine. Early labor market and clinical evidence suggests possible disruption to expertise-regeneration pathways in highly AI-exposed sectors, though adoption is recent and the strongest signals c

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Hypergamigication Through Integrating Game Engines and Learning Management Systems: Ender's Game

arXiv:2607.29300v1 Announce Type: new Abstract: This paper discusses games, their use in education, and previous work on integrating game engines and learning management systems (LMS). It proposes a bidirectional integration where game environments are generated using LMS content, introducing the concept of hypergamification as the use of a comprehensive game environment rather than isolated game design elements. A working pilot implementation of an importable Unity package for Blackboard integration is demonstrated, along with a demo game that uses the developed package. The paper also discusses the limitations of the proposed approach and outlines avenues for future work.

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Deployment Wall: A Diagnostic Framework and Instrument for Enterprise AI in the Deployment Era

arXiv:2607.29089v1 Announce Type: new Abstract: Enterprise investment in generative artificial intelligence (AI) tripled in a single year to roughly US$37 billion, yet independent field research finds that about 95% of enterprise generative-AI pilots deliver no measurable profit-and-loss impact. We argue that the dominant explanation--that models are not yet capable enough--is mistaken, and that enterprise AI has entered a Deployment Era in which advantage derives not from model intelligence but from the removal of the organizational and architectural friction that prevents a capable model from reaching production. Building on the software-engineering literature on technical debt and machine-learning deployment, and on a structured synthesis of independent field studies, we make the diagnosis operational. We introduce three linked constructs and one measurement instrument: the Deployment Wall, a six-stage value-leak model that mechanically reproduces observed survival rates; the Seam I

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

IyawoBench v2.0: Extended Diagnostic Evaluation of Large Language Model Clinical Triage in Nigerian Primary Care

arXiv:2607.29085v1 Announce Type: new Abstract: Large language models are being deployed as clinical triage tools in low and middle income countries where trained physicians are scarce. Existing safety metrics, however, produce misleading confidence: models scoring 100% on binary "did not send an emergency home" safety measures may nevertheless exhibit systematic failure modes that render them undeployable at scale. We present IyawoBench v2.0, an extended diagnostic evaluation of large language model clinical triage on 200 synthetic vignettes derived from 1,200 real patient encounters at 19 Nigerian primary health centres. We introduce a formal mathematical framework comprising fourteen definitions and two theorems that decompose triage safety into three distinct failure modes: Conservative Escalation Bias, Systematic Downgrade Bias, and Middle-Tier Instability. We propose the Escalation Bias Index and Expected Deployment Cost as novel metrics that expose failure modes hidden by conven

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

From Process to Evidence: How Computing Can Ground Appropriate Reliance on Legal AI

arXiv:2607.28869v1 Announce Type: new Abstract: Lawyers and self-represented litigants are already using artificial intelligence (AI) to draft legal documents, and courts are responding with rules. After more than 1,500 cases involving AI hallucinations, lawyers have been instructed to perform careful, independent review of AI-assisted filings. Discharging these duties requires what the human-computer interaction (HCI) literature calls ``appropriate reliance,'' which cannot be calibrated without evidence on how often, how badly, and how detectably these tools fail at legal work. Existing research barely describes any of the three. We analyze the official record of the New York court system. The documents repeatedly call for evidence that does not exist (e.g., error rates, do-not-use lists). In its place they invoke procedure, including training mandates, checklists, and uncalibrated human review. The burden falls hardest on those least equipped to bear it: legal aid programs are told t

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Hidden Errors in Big Data: The Case of Property Records

arXiv:2607.28827v1 Announce Type: new Abstract: Big data are the foundation for an increasing share of academic research and AI models deployed in both the public and private sectors, prompting substantial growth over time in reliance on brokered datasets. Brokered property records, which are ubiquitous in studies of gentrification, inequality, and the property tax in the U.S. and serve as inputs to property valuation models, are one notable example. In this paper, we audit two prominent brokered property datasets, finding errors in these data which bias key measures of economic inequality. First, we document that for 1-2% of matched sales in Cook County, IL, from 2018-2021, broker-provided sale prices differ from ground truth sale prices by more than 5%. Moreover, missing data and conceptual differences in the reporting of deed and property characteristics lead to coverage errors ranging from 12 to 15% of transactions. Second, we show that misreporting is highly consistent between bro

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Structured AI Demonstrations and Student LLM Use in Engineering Mechanics: Study Design and Preliminary Results

arXiv:2607.28710v1 Announce Type: new Abstract: The rapid integration of large language models (LLMs) into undergraduate education presents an urgent challenge for engineering instructors. Despite widespread student adoption, there remains a critical lack of domain-specific empirical evidence to guide pedagogical policies and classroom interventions. This manuscript presents a descriptive study design and preliminary findings from an undergraduate engineering mechanics course conducted in Spring 2026. We detail a reproducible survey instrument used to capture student AI usage patterns, attitudes, and verification practices, which are subsequently linked to academic performance metrics. Additionally, we document a deployable sequence of nine structured, instructor-led AI demonstrations designed to model strategic LLM delegation and evaluation. While our preliminary data highlight shifting student behaviors and complex relationships between AI reliance and course outcomes, the primary co

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Scaffolding Critical Engagement with GenAI: Transforming Ethnic Minority Preparatory Students' Collaborative Discourse in Prompt Engineering Tasks

arXiv:2607.28630v1 Announce Type: new Abstract: Generative AI (GenAI) holds significant promise for advancing educational equity among ethnic minority students by broadening access to learning resources and mitigating linguistic barriers. However, these benefits are counterbalanced by the risk of cognitive laziness, whereby students may treat GenAI as an answer engine or shortcut rather than as a partner in thinking. This design-based research investigated how pedagogical scaffolding can shift students from passive consumption to critical co-creation with GenAI. The study involved 78 ethnic minority preparatory students in China participating in a three-week GenAI course that integrated a human-in-the-loop workflow and teacher modeling with contrasting cases to disrupt uncritical reliance on GenAI. We employed epistemic network analysis to examine collaborative discourse, thematic analysis to analyze student reflections, and paired-samples t-tests to assess changes in prompt self-effic

Source ↗
technology Mon, 01 Jun 2026 09:00:00 +0000
Tech & Learning

4 Strategies For Teaching With AI Effectively

Health sciences professor Humberto López Castillo urges students to use AI to help with science research, but never to lose sight of the human element.

Source ↗
technology Mon, 01 Jun 2026 09:00:00 +0000
Tech & Learning

Edtech Show & Tell June 2026

New edtech products that have caught our attention this month

Source ↗
technology Fri, 31 Jul 2026 22:21:11 +0000
MedCity News

Healthcare Moves: A Monthly Summary of Hires, Exits and Layoffs

July has seen a slew of executive hires, exits and layoffs across the healthcare industry. For instance, Aledade, Providence and Whoop named new executives. There were also layoffs at organizations including Novartis, Adventist Health and Wellstar Health System. The post Healthcare Moves: A Monthly Summary of Hires, Exits and Layoffs appeared first on MedCity News .

Source ↗
technology Fri, 31 Jul 2026 20:55:48 +0000
MedCity News

Startup to Acquisition in 2 Years: Grove AI’s Journey to Transform Clinical Trials

Grove AI’s technology uses voice agents to recruit eligible patients for clinical trials. Formed in 2024, the startup’s rapid growth led to its acquisition by Hippocratic AI earlier this year. The post Startup to Acquisition in 2 Years: Grove AI’s Journey to Transform Clinical Trials appeared first on MedCity News .

Source ↗
technology Fri, 31 Jul 2026 20:06:00 +0000
MedCity News

Daffodil Health Launches No Surprises Act Dispute Management Solution

Daffodil Health launched an AI-powered solution to help payers manage No Surprises Act disputes by automating claims review, negotiations and arbitration workflows. The post Daffodil Health Launches No Surprises Act Dispute Management Solution appeared first on MedCity News .

Source ↗
technology Fri, 31 Jul 2026 14:05:00 +0000
MedCity News

The Battles We Face: ADHD From A Patient Perspective, And What Providers Should Know

Caring for patients with ADHD, or those who potentially have ADHD, extends far beyond diagnosis and prescribing. The post The Battles We Face: ADHD From A Patient Perspective, And What Providers Should Know appeared first on MedCity News .

Source ↗
technology Fri, 31 Jul 2026 13:59:00 +0000
MedCity News

Three Ways Distressed Healthcare Must Evolve

What healthcare needs now is a connected, deeply integrated system of action that works with the system of record to complete work across the enterprise, supported by accountable partners willing to stand behind the outcomes they create. The post Three Ways Distressed Healthcare Must Evolve appeared first on MedCity News .

Source ↗
technology Fri, 31 Jul 2026 09:12:25 +0000
MedCity News

Why General Catalyst Is Betting $450M on Function Health’s Preventive Care Push

Function Health raised $450 million in growth financing from General Catalyst’s Customer Value Fund, just eight months after its $298 million Series B. The lab-testing startup plans to use the capital to expand access to its testing and imaging services. The post Why General Catalyst Is Betting $450M on Function Health’s Preventive Care Push appeared first on MedCity News .

Source ↗
technology Fri, 31 Jul 2026 09:00:00 +0000
Tech & Learning

Thoughtfully Leveraging AI To Improve Learning Outcomes, Efficiency, And More

Innovative Leader Award - Matt Kuhn of Volusia County Schools shares how his district has implemented AI for use by students, staff, and leaders.

Source ↗
technology Fri, 31 Jul 2026 09:00:00 +0000
eCampus News

Higher education’s grade inflation conundrum

Grade inflation is all the rage in higher education. Harvard has a proposal to reduce the number of A’s it awards. A Yale faculty committee proposed that 3.0 should be the mean grade. The post Higher education’s grade inflation conundrum appeared first on eCampus News .

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CL

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

arXiv:2605.26494v2 Announce Type: replace-cross Abstract: We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8B activated per token. Designed end-to-end for agentic deployment, the M2 series rests on three components: (i) agent-driven data pipelines producing large-scale, verifiable trajectories across agentic coding and agentic cowork, each grounded in an executable workspace and an artifact-aligned reward; (ii) Forge, a scalable agent-native RL system that adapts to long-horizon agent trajectories, paired with windowed-FIFO scheduling, prefix-tree merging, inference optimization, and a clean training-inference-agent decoupling that supports both white-box and black-box agents; (iii) the latest M2.7 checkpoint takes an early step toward self-evolution -- autonomously debugging training runs and modify

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CL

MinerU-Popo: Universal Post-Processing Model for Structured Document Parsing

arXiv:2605.24973v2 Announce Type: replace-cross Abstract: VLM-based OCR models have become the de facto choice for document parsing, as they can accurately extract page-level elements (e.g., paragraphs within individual pages) together with their bounding boxes and textual content. However, downstream applications such as RAG require coherent document-level information, whereas these models often break cross-page continuity and fail to recover disrupted structures, such as paragraphs and tables truncated by page boundaries. Such relationships are not confined to a single page; instead, they require joint analysis of titles, paragraphs, tables, and images spanning multiple pages. A natural solution is therefore to reuse existing OCR outputs and reconstruct document-level logical structures through post-processing. To this end, we propose MinerU-Popo, a lightweight and universal framework for POst-Processing OCR outputs, which converts page-level results from diverse parsers into coheren

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CL

Orchard: An Open-Source Agentic Modeling Framework

arXiv:2605.15040v3 Announce Type: replace-cross Abstract: Agentic modeling aims to transform LLMs into autonomous agents capable of solving complex tasks through planning, reasoning, tool use, and multi-turn interaction with external environments. We present Orchard, an open-source framework for scalable agentic modeling. At its core is Orchard Env, a lightweight Kubernetes-native environment service that provides reusable primitives for sandbox lifecycle management across task domains, agent harnesses, and training stages. On top of Orchard Env, we build three agentic modeling recipes. Orchard-SWE targets software engineering agents. We introduce credit-assignment supervised fine-tuning and a progression of RL signals: Balanced Adaptive Rollout (BAR) for sparse-reward optimization, on-policy distillation (OPD) and rubric-based process reward (RPR) for dense supervision, and historical experience distillation, which compresses rollouts from prior experiments into a compact value model

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CL

S-GRPO: Unified Post-Training for Large Vision-Language Models

arXiv:2604.16557v2 Announce Type: replace-cross Abstract: Current post-training methodologies for adapting Large Vision-Language Models (LVLMs) generally fall into two paradigms: Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL). Despite their prevalence, both approaches suffer from inefficiencies when applied in isolation. SFT forces the model's generation along a single expert trajectory, often inducing catastrophic forgetting of general multimodal capabilities due to distributional shifts. Conversely, RL explores multiple generated trajectories but frequently encounters optimization collapse - a cold-start problem where an unaligned model fails to spontaneously sample any domain-valid trajectories in sparse-reward visual tasks. In this paper, we propose Supervised Group Relative Policy Optimization (S-GRPO), a unified post-training framework that integrates the guidance of imitation learning into the multi-trajectory exploration of preference optimization. Tailored for di

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CL

The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models

arXiv:2602.08159v2 Announce Type: replace-cross Abstract: When a language model asserts that "the capital of Australia is Sydney," does it know this is wrong? Models assert misconceptions with the same fluency as facts, so the question cannot be answered from output uncertainty. Truth-related signals are known to exist in the residual stream, but not their geometry: how many dimensions carry the signal, how simple a detector can be, and whether it transfers. We characterize this geometry across 11 models (124M-14B) and test it causally with activation steering, concept erasure, and distributed alignment search. The structure is simple: two class centroids in a 2-8 dimensional subspace match a trained linear probe, and 25 labeled examples recover 90% of full-data AUC on GPT-2. Steering shifts hallucination rates by 9.1 points on six models, erasure drops detection to chance, and distributed alignment search, the only method that bounds rank, localizes at most five causal dimensions. The

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CL

Scaling medical imaging report generation with multimodal reinforcement learning

arXiv:2601.17151v2 Announce Type: replace-cross Abstract: Frontier models have demonstrated remarkable capabilities in understanding and reasoning with natural-language text, but they still exhibit major competency gaps in multimodal understanding and reasoning especially in high-value verticals such as biomedicine. Medical imaging report generation is a prominent example. Supervised fine-tuning can substantially improve performance, but they are prone to overfitting to superficial boilerplate patterns. In this paper, we introduce Universal Report Generation (UniRG) as a general framework for medical imaging report generation. By leveraging reinforcement learning as a unifying mechanism to directly optimize for evaluation metrics designed for end applications, UniRG can significantly improve upon supervised fine-tuning and attain durable generalization across diverse institutions and clinical practices. We trained UniRG-CXR on publicly available chest X-ray (CXR) data and conducted a t

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CL

MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models

arXiv:2502.20780v2 Announce Type: replace-cross Abstract: The increasing use of vision-language models (VLMs) in healthcare applications presents great challenges related to hallucinations, in which the models may generate seemingly plausible results that are in fact incorrect. Such hallucinations can jeopardize clinical decision making, potentially harming the diagnosis and treatments. In this work, we propose MedHallTune, a large-scale benchmark designed specifically to evaluate and mitigate hallucinations in medical VLMs. Comprising over 100,000 images and 1,000,000 instruction pairs, MedHallTune includes both hallucination and non-hallucination samples, each with ground-truth annotations. We conduct a comprehensive evaluation of current medical and general VLMs using MedHallTune, assessing their performance across key metrics, including clinical accuracy, relevance, detail level, and risk level. The experimental results show that fine-tuning with MedHallTune successfully improves t

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CL

Safety Verification of Wait-Only Non-Blocking Broadcast Protocols

arXiv:2403.18591v3 Announce Type: replace-cross Abstract: Broadcast protocols are programs designed to be executed by networks of processes. Each process runs the same protocol, and communication between them occurs in synchronously in two ways: broadcast, where one process sends a message to all others, and rendez-vous, where one process sends a message to at most one other process. In both cases, communication is non-blocking, meaning the message is sent even if no process is able to receive it. We consider two coverability problems: the state coverability problem asks whether there exists a number of processes that allows reaching a given state of the protocol, and the configuration coverability problem asks whether there exists a number of processes that allows covering a given configuration. These two problems are known to be decidable and Ackermann-hard. We show that when the protocol is Wait-Only (i.e., it has no state from which a process can both send and receive messages), th

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CL

BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms

arXiv:2607.26497v2 Announce Type: replace Abstract: Retrieval-augmented generation (RAG) spans lexical and dense retrieval, graph-based indexing, and agentic search, but these paradigms are usually evaluated on different benchmarks at one corpus size, leaving their accuracy-cost scaling unclear. To bridge this gap, we present a controlled study that varies corpus size along 28 strictly nested tiers spanning roughly 450-fold, while holding questions and a fixed bedrock of relevant and adversarial documents unchanged. Under one reader model and one judging protocol, we measure official accuracy, construction and query tokens, and latency. The results reveal a scale-dependent crossover rather than an unconditional winner. File-System Agent leads at the smallest shared tiers, but its sequential exploration costs 39 times more query tokens at the bedrock and becomes less effective as the search space grows. Around 10 million corpus tokens, BM25 overtakes it and leads at every larger shared

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CL

MEDIAREF: A Public Knowledge Store for Media Background Checks

arXiv:2607.02383v3 Announce Type: replace Abstract: LLM-based retrieval-augmented generation (RAG) is increasingly used for automated fact-checking (AFC) and related tasks. By grounding LLM outputs in retrieved evidence, RAG-based systems provide transparent justifications while allowing external information to be updated independently of the underlying model. However, existing approaches often assume retrieved evidence is reliable, although real-world information may be conflicting, outdated, and can originate from unreliable or biased sources. Recent work on *source-critical reasoning* addresses this challenge through media background checks (MBCs) (Schlichtkrull, 2024), which assess the credibility of evidence sources to support downstream fact verification. However, generating MBCs relies on costly proprietary search APIs, limiting reproducibility. To mitigate this issue, we introduce MEDIAREF, a publicly available knowledge store of web-sourced documents that enables reproducible,

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CL

VISTA: A Controllable Platform for Generating and Auditing Egocentric Assistance Scenarios

arXiv:2605.10579v2 Announce Type: replace Abstract: Evaluating whether AI agents can proactively assist humans in daily activities, ranging from routine household tasks to urgent safety-critical situations, requires diverse visual data. However, collecting such scenarios in the real world is often difficult, costly, or unsafe, and simulation environments often lack the social commonsense needed to simulate the consequences of different actions. In this work, we present VISTA, a controllable platform that uses a user-provided scenario seed, defined as a short natural-language description of the intended assistance situation, to generate editable plans, egocentric videos, and an auditable review trail. VISTA structures scenario intent around three interaction modes, including reactive, explicit proactive, and implicit proactive, and two consequence families, including safety-critical and everyday inconvenience, with no-assistance cases as controls. Its six-stage pipeline exposes the desi

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CL

NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus

arXiv:2605.00086v2 Announce Type: replace Abstract: High-quality corpora are essential for advancing Natural Language Processing (NLP) in Portuguese. Building on previous encoder-only models such as BERTimbau and Albertina PT-BR, we introduce NorBERTo, a modern encoder based on the ModernBERT architecture, featuring long-context support and efficient attention mechanisms. NorBERTo is trained on Aurora-PT, a newly curated Brazilian Portuguese corpus comprising 331 billion GPT-2 tokens collected from diverse web sources and existing multilingual datasets. We systematically benchmark NorBERTo against Strong baselines on semantic similarity, textual entailment and classification tasks using standardized datasets such as ASSIN 2 and PLUE. On PLUE, NorBERTo-large achieves the best results among the encoder models we evaluated, notably reaching 0.9191 F1 on MRPC and 0.7689 accuracy on RTE. On ASSIN 2, NorBERTo-large attains the highest entailment F1 (~0.904) among all encoders considered, alt

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CL

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models

arXiv:2604.19786v3 Announce Type: replace Abstract: Evaluating humor in large language models (LLMs) is an open challenge because existing approaches yield isolated, incomparable metrics rather than unified model rankings, making it difficult to track progress across systems. We introduce HumorRank, a tournament-based evaluation framework and leaderboard for textual humor generation. On two public benchmarks (SemEval-2026 MWAHAHA and Humor Transfer Bench), we conduct extensive automated pairwise evaluation across nine models spanning proprietary, open-weight, and specialized systems. Pairwise judgments are produced by LLM judges grounded in the General Theory of Verbal Humor (GTVH): each judge integrates structured comedic analysis into adjudication, jointly yielding a preference decision, an interpretable rationale, and mechanism, delivery, and failure tags rather than a black-box funniness score. Judgments are aggregated via an Adaptive Swiss tournament, with Bradley-Terry Maximum Li

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CL

How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data

arXiv:2604.13977v2 Announce Type: replace Abstract: Synthetic data is a standard component in training large language models, yet systematic comparisons across design dimensions, including rephrasing strategy, generator model, and source data, remain absent. We conduct extensive controlled experiments, generating over one trillion tokens, to identify critical factors in rephrasing web text into synthetic pretraining data. Our results reveal that structured output formats, such as tables, math problems, FAQs, and tutorials, consistently outperform both curated web baselines and prior synthetic methods. Notably, increasing the size of the generator model beyond 1B parameters provides no additional benefit. Our analysis also demonstrates that the selection of the original data used for mixing substantially influences performance. By applying our findings, we develop \textbf{\textsc{FinePhrase}}, a 486-billion-token open dataset of rephrased web text. We show that \textsc{FinePhrase} outpe

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CL

Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization

arXiv:2604.07343v2 Announce Type: replace Abstract: Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a central mechanism for capturing diverse human values. While benchmarks for general response quality are prevalent, evaluating how well reward models account for individual user preferences remains an open challenge. To bridge this gap, we introduce Personalized RewardBench, a novel benchmark designed to rigorously assess reward models' capacity to model personalized preferences. We construct chosen and rejected response pairs based on strict adherence to (or violation of) user-specific rubrics, ensuring that preference distinctions are uniquely tailored to the individual. In particular, human evaluations confirm that the primary discriminative factor between pairs is strictly personal preference, with both responses maintaining high general quality (e.g., correctness, relevance and helpfuln

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CL

LLM2Vec-Gen: Generative Embeddings from Large Language Models

arXiv:2603.10913v3 Announce Type: replace Abstract: Fine-tuning LLM-based text embedders via contrastive learning maps inputs and outputs into a new representational space, discarding the LLM's output semantics. We propose LLM2Vec-Gen, a self-supervised alternative that instead produces embeddings directly in the LLM's output space by learning to represent the model's potential response. Specifically, trainable special tokens are appended to the input and optimized to compress the LLM's own response into a fixed-length embedding, guided by an unsupervised embedding teacher and a reconstruction objective. Crucially, the LLM backbone remains frozen and training requires only unlabeled queries. LLM2Vec-Gen achieves state-of-the-art self-supervised performance on the Massive Text Embedding Benchmark (MTEB), improving by 8.8% over the unsupervised embedding teacher. Since the embeddings preserve the LLM's response-space semantics, they inherit capabilities such as safety alignment (up to 22

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CL

GradMAP: Faster Layer Pruning with Gradient Metric and Projection Compensation

arXiv:2602.14649v2 Announce Type: replace Abstract: Large Language Models (LLMs) exhibit strong reasoning abilities, but their high computational costs limit their practical deployment. Recent studies reveal significant redundancy in LLMs layers, making layer pruning an active research topic. Layer pruning research primarily focuses on two aspects: measuring layer importance and recovering performance after pruning. Unfortunately, the present works fail to simultaneously maintain pruning performance and efficiency. In this study, we propose GradMAP, a faster layer pruning method with \textbf{Grad}ient \textbf{M}etric \textbf{A}nd \textbf{P}rojection compensation, which consists of two stages. In the first stage, we introduce a novel metric based on gradient magnitudes, enabling a global assessment of layer importance. Note that, it requires only a single backward propagation step per pruning decision, substantially enhancing pruning efficiency. In the second stage, we first analyze the

Source ↗
Showing 8901–8950 of 11035 signals
← Prev Page 179 of 221 Next →