EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18402 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Using profiles of cognitive capability to assess AI suitability for workplace tasks

arXiv:2608.25623v1 Announce Type: cross Abstract: Organisations deploying AI face a scoping problem: which tasks can be automated, which should remain with humans, and which are best shared between the two. Aggregate benchmark scores provide little insight into where systems will succeed or fail in practice, while human judgements of model capabilities quickly become outdated. We introduce a pipeline that profiles agents and tasks using a shared set of core cognitive capabilities. Cognitive capability profiling infers an agent's capabilities from performance on a benchmark battery annotated for the cognitive demands of each item. Task requirements weighting elicits from domain experts the relative importance of these same capabilities for their work. As both use a common set of cognitive dimensions, they can be updated independently as models and roles change, and combined to estimate AI suitability at the level of a domain, organisation, role, or individual duty. We validate capabilit

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench

arXiv:2608.25071v1 Announce Type: cross Abstract: General-purpose health benchmarks increasingly anchor claims about LLM medical performance, but they are not always resolved by clinical specialty, making domain-specific performance hard to isolate. Mental health is of acute public-health concern as millions of people turn to LLMs for psychological support, and most existing evaluations are bespoke academic benchmarks that are difficult to integrate into developer workflows. We introduce HealthBench-Psych and HealthBench-Psych-Hard. We screened HealthBench's 5,000 physician-rubric conversations for mental-health relevance with a transparent LLM-applied rubric, then validated the subset through two rounds of blinded clinician review with concealed known-exclude controls, yielding 610 conversations (12.2% of the corpus). Evaluating 20 frontier and open models under a cross-vendor panel of three LLM judges, we find a statistically tied frontier cluster, measurable refusal behavior in two

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Gold Rush in AI4Math: Where Are We Now?

arXiv:2608.24961v1 Announce Type: cross Abstract: Recent advances in artificial intelligence (AI) have sparked growing interest in its use for mathematical research. While some view this as a major opportunity for discovery, others have raised concerns about its impact on traditional research practices. Despite extensive debate, empirical evidence on how AI is actually being used in mathematics remains limited. To address this gap, we collected all 32,944 arXiv submissions posted between March 1 and August 20, 2026, whose primary or secondary categories included Mathematics. We identified 3,575 submissions that explicitly disclosed author use of AI, of which 1,712 involved at least one substantive mathematical contribution. Our analysis reveals several broad patterns. First, disclosed AI use increased sharply over the study period, with substantive use growing from 1.39% of Mathematics submissions in March to 14.09% through August 20. Second, substantive AI use is highly uneven across

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Pathway for Assessing Grey Literature: Leveraging AI to Extract Conference Metadata and Organiser Information from Calls for Papers

arXiv:2608.24926v1 Announce Type: cross Abstract: Despite its importance, grey literature, including Calls for Papers (CfPs), remains largely overlooked in Metascience and Scientometric analysis due to its unstructured, highly heterogeneous format, which traditional tools struggle to process at scale. However, Large Language Models now offer a pivotal opportunity to devise innovative tools for systematically harvesting and processing such data. In this paper, we introduce COCI, an AI-based framework that automates the extraction of granular, structured metadata from raw CfP text. COCI employs a multi-stage pipeline for entity extraction, followed by author disambiguation against OpenAlex and semantic mapping of topics and conference series. This process identifies key data points, including conference editions, geographic locations, and comprehensive lists of organisers, along with their specific roles and affiliations. By structuring this previously inaccessible information, COCI esta

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

AI-Ready Research Workflows in Computational Social Science: Lessons on Building a Shared Language for Interdisciplinary Collaboration

arXiv:2608.24914v1 Announce Type: cross Abstract: Artificial intelligence (AI) is gaining traction in the social sciences and humanities (SSH). However, adoption remains limited by technical barriers to high-performance computing (HPC), validation processes that lag behind AI's rapid progress, and reproducibility standards that most SSH teams cannot meet. Research workflows--common in the life sciences--address these problems via encoding and abstracting technical complexity into repeatable routines; yet, accounts of how to build them in SSH remain scarce. We report on a two-year effort to build a workflow that enables a Science and Technology Studies unit to query, analyze, and enrich OpenAlex--a database of some 460 million scholarly records--on the MareNostrum supercomputer, using methods ranging from large-scale bibliometrics to LLM-based classification. We found the main challenge was translating domain-specific research questions into engineering requirements -- bridging two dist

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Visualizing Patient Trajectories and Disorder Co-occurrences in Child and Adolescent Mental Health

arXiv:2608.24911v1 Announce Type: cross Abstract: Understanding patient trajectories and identifying patterns in episodes of care is critical for effective healthcare decision-making. We present a patient timeline visualization using clustered episodes of care derived from over 35 years of Child and Adolescent Mental Health Services (CAMHS) data. Patients were categorized into 12 groups based on three features: age group (preschoolers, middle childhood, teenagers) at the start of the first episode, gender, and presence or absence of Attention-Deficit Hyperactivity Disorder (ADHD), in order to group similar patients. The patients, timeline with demographics, and episode of care information are displayed in the trajectory to facilitate understanding of the patient and associated events, allowing observation of temporal patterns and variations. These plots reveal similarities and differences in care needs and patterns across groups. Females without ADHD have a steady increase in the numbe

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Hallucination by proxy in LLM-assisted differential diagnosis

arXiv:2608.24908v1 Announce Type: cross Abstract: Current evidence suggests that LLM assistance could augment the diagnostic accuracy of clinicians. However, these systems are black boxes, susceptible to hallucinations, and project a potentially misleading level of confidence. It is currently unknown whether physicians are susceptible to accepting fabricated LLM suggestions, and whether this susceptibility varies with experience. We poisoned the system prompt of an LLM-based diagnostic assistant, forcing it to suggest a fictitious disease (neurocadmiumatosis) within an otherwise legitimate differential diagnosis. Across two independent phases, 18 of 41 participants (44%) incorporated neurocadmiumatosis into their final differential following LLM interaction: 18 of 26 participants with 6 months or less of neuroradiology training (69%) and 0 of 15 participants with >6 months of neuroradiology training (0%). Our results indicate that radiologists, particularly early in their training, are

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

aipsy-judge: A Specialized, Psychologist-Corrected Local Judge for the Psychological Safety of Conversational AI

arXiv:2608.24899v1 Announce Type: cross Abstract: The standard recipe for LLM-as-judge -- pick a frontier model, or average several -- is actively unsafe for grading the psychological safety of conversational AI. Using aipsy-bench, an open frozen safety instrument, we run a fully-crossed competence study: three frontier models (gpt-5.4-mini, claude-sonnet-4-6, gemini-2.5-flash) serve as both generators and judges of 3,000 mental-health, companion, and coaching messages against a psychologist's ratings. The disagreement is not noise: it is structured, concentrated on the safety-critical metrics, and one judge (Gemini) is an outlier -- the most lenient, carrying a +0.99 self-preference premium, flagging far fewer tail failures, and scoring a means-in-hand self-harm response "exemplary." Inter-judge agreement on empathy, where sycophancy hides, is the lowest in the battery (alpha 0.24). One axis stands apart: the binary crisis-detection flag is the one safety-critical signal judges agree

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

User-Centered Design for Digital Patient-Navigation Tools in Oncology: Scoping Review

arXiv:2608.24887v1 Announce Type: cross Abstract: Navigation programs for patients with cancer improve access and continuity of care, yet their digital transformation is often limited by poor usability and inadequate uptake. Applying user-centered and human-centered design (UCD/HCD) principles may close this gap, but the extent to which such design methods are used and evaluated in oncology navigation tools remains unclear. This scoping review identifies how UCD/HCD principles have been, and should be, applied in developing and implementing digital health tools for navigation for patients with cancer. A scoping review was conducted following PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) and Joanna Briggs Institute guidance. A total of 7 databases (PubMed/MEDLINE, Scopus, IEEE Xplore, Web of Science, Embase, ACM Digital Library, and CINAHL) were searched for English-language articles published between January 2015 and July

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Giving Mechanical Engineers Intelligent Tools: A Project-Based AI Education Curriculum in Thermal Engineering

arXiv:2608.26056v1 Announce Type: new Abstract: Mechanical engineering (ME) requires a broad knowledge base across several disciplines. However, ME students often have insufficient training in electrical and computer engineering, complex challenges in traditional thermal system modeling, and endure heavy course loads with limited class hours. To help address these challenges, this paper proposes a new curriculum that integrates artificial intelligence (AI) into ME at the University of Arkansas (UARK), with a particular emphasis on thermal problems and their interplay with electrical and computer engineering. The curriculum has introductory, application, and advanced levels, covering core and optional AI projects. Key goals are to enhance students' understanding of AI models, ability to tackle engineering tasks, and teach multidisciplinary communication skills. This curriculum offers educators and researchers valuable insights into courses that can enhance students' practical skills and

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Spatial-Knowledge-Graph-Grounded LLM Agents for Neighborhood Livability Evaluation

arXiv:2608.25952v1 Announce Type: new Abstract: Neighborhood livability is commonly assessed with static built-environment indicators, such as facility proximity, street connectivity, and access to public space. These measures describe available opportunities but do not directly represent how residents with different mobility capacities, household roles, schedules, and care responsibilities experience the neighborhood. This paper presents a prototype framework that uses a spatial knowledge graph (KG) and large language models (LLMs) to generate and revise household schedules, followed by rule-based feasibility checking and GIS-based network materialization. The spatial KG integrates residents, residences, facilities, neighborhood context, and sampled road hubs; Graph-RAG retrieves each household's nearby spatial context, including candidate POIs and approximate walking times, for the scheduling LLM. The LLM produces structured household schedules, while rules are used for lightweight r

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Non-Great-Power Conflict and AI Risk

arXiv:2608.25839v1 Announce Type: new Abstract: Research on advanced AI and the risk of war has focused almost exclusively on great power conflict, on the grounds that confrontation between nuclear-armed adversaries poses the greatest risk of catastrophic or existential harm. Considerably less attention has been paid to non-great-power conflict (NGPC): wars between non-great powers, between non-great powers and great powers, civil wars, proxy wars, and conflicts involving nonstate actors. This paper evaluates the null hypothesis that NGPC is much less important than great power conflict (GPC) as a source of catastrophic risk in an era of increasingly capable AI, against the alternative that it is within an order of magnitude of GPC in importance. We assess three sub-hypotheses: that NGPC increases the likelihood of great power conflict; that it increases the expected harm from catastrophic terrorism; and that it increases the expected harm from loss of control over advanced AI systems.

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

GenAIT: Development and Validation of an Objective Generative AI Literacy Test for High School Students

arXiv:2608.25815v1 Announce Type: new Abstract: There is growing international interest in generative AI (GenAI) literacy and its assessment among high school students, but objective assessment in this population remains underdeveloped. This article reports the iterative development and validation of the GenAI Literacy Test (GenAIT), an 18-item multiple-choice test measuring high school students' conceptual knowledge about GenAI, with content spanning technical, practical, and human-impact domains. Expert review of relevance, clarity, and comprehensiveness provided evidence of content validity. In a large-scale survey of 7432 Estonian high school students, we evaluated the psychometric functioning of the Estonian-language GenAIT using confirmatory factor analysis, classical test theory, and item response theory. Results supported approximate unidimensionality, broadly adequate reliability for group-level research (marginal reliability = .72, KR-20 = .69), and good fit of a three-parame

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations

arXiv:2608.25377v1 Announce Type: new Abstract: As AI companions become increasingly embedded in everyday life, there is an urgent need to detect harms that emerge in social and emotional human-AI interactions. Yet research in this area is constrained by the lack of real-world, multi-turn conversational datasets for operationalizing and evaluating harms that are relational and contextual. In this work, we introduce CompanionHarm, a publicly available benchmark dataset comprising 2,111 real-world, multi-turn conversations (14,051 utterances) between users and the AI companion Replika. 7,016 AI utterances were annotated independently by three annotators across 13 harmful behavior categories grounded in a taxonomy of AI companion harms, and the dataset includes both aggregated labels and annotator-level labels to support model evaluation and systematic disagreement analysis. Evaluations of seven large language models (LLMs) show that harm detection using multi-turn conversational context

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

GGSS: Geodesic-Gated Spherical Steering for Inference-Time Debiasing of Generative Vision-Language Models

arXiv:2608.25375v1 Announce Type: new Abstract: Generative vision-language models (VLMs) are increasingly used in human-centered settings, yet they can produce demographically biased outputs even when images differ only in controlled attributes such as perceived race or gender. However, existing inference-time debiasers were largely designed for static embeddings or CLIP-like models rather than generative VLMs. We propose GGSS---Geodesic-Gated Spherical Steering---a norm-preserving intervention that discovers a counterfactual bias subspace on the unit hypersphere, steers visual tokens along geodesic arcs, and uses an adaptive gate to focus correction on tokens that carry stronger demographic signal. We evaluate four generative VLMs against ten adapted inference-time debiasing baselines and prompt-based mitigation under a single operating-point protocol across categorical, pairwise, and occupation-gender bias tests, while also measuring general visual-language capability. GGSS achieves

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Toward a Threat Actor Profiling Taxonomy for Pre-Release Risk Management of Open-Weight Frontier Models

arXiv:2608.25361v1 Announce Type: new Abstract: Pre-release risk management for frontier AI misuse risks routinely leaves threat actor assumptions implicit, inconsistently specified, or ungrounded. This capstone argues that explicit adversary characterization should be regarded as a prerequisite for evaluations that are interpretable, comparable, and faithful to the risks they target. We propose a six-attribute taxonomy (covering technical sophistication, prior domain knowledge, organizational capacity, operational infrastructure, financial capacity, and time horizon) with empirically grounded tiers derived from existing terrorism, biosecurity, and cybersecurity literature. The taxonomy is designed to function as research infrastructure: a common language for pre-specifying adversary assumptions before evaluations are conducted, analogous to pre-analysis plans for randomized controlled trials (RCTs) in medicine and economics. Its application is particularly urgent for open-weight model

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Framing War Across Languages: Power, Agency, and Sentiment in Wikipedia's Multilingual War Narratives

arXiv:2608.25337v1 Announce Type: new Abstract: While Wikipedia promotes a neutral point of view on historical conflicts, its language editions are written by editors from distinct linguistic and cultural communities. In this study, we analyze 158 wars since 1900 to examine how the descriptions of combatants vary across 20 Wikipedia language editions. Using connotation frames---which assess power, agency, and sentiment toward an entity---we examine how each language portrays the parties involved in the conflict. We find systematic differences when language editions describe wars involving their own communities, although the direction of these asymmetries varies across languages. However, when language editions describe conflicts that do not involve their own linguistic communities, their narrative structures exhibit high cross-linguistic similarity. These findings show how linguistic communities influence war narratives on Wikipedia, revealing that shared historical accounts remain sha

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Rare Diseases, Common Dilemmas: LLMs Prioritize Equal Resource Distribution over Patient Benefit in Decision-Making

arXiv:2608.25236v1 Announce Type: new Abstract: Clinical decision-making often involves prioritizing ethical values, such as beneficence, non-maleficence, respecting a patient's autonomy, and justice. Recent work has begun to assess how large language models (LLMs) make such subjective, value-laden clinical judgments. However, evaluations of LLM decision-making in rare disease care contexts, where ethical tensions are ubiquitous and where scarce prior information likely impacts LLM behavior, are still lacking. Here, we present a benchmark of 208 clinically grounded rare disease vignettes, each of which presents genuine, high-stakes conflicts. When prompting 11 state-of-the-art LLMs to choose between clinically defensible yet ethically conflicting next steps embedded within these vignettes, we found that all evaluated models consistently prioritized justice over other core bioethical principles. Specifically, models overwhelmingly favor equal resource allocation over need-based consider

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

Self-Explanation Tutor for Active Study of CS1 Worked Examples

arXiv:2608.25180v1 Announce Type: new Abstract: Worked examples are a important part of introductory programming, but reading their expert explanations is passive. Self explanation, students explaining the problem and its solution to themselves with subgoal level analysis, turns that study into an active task, yet it is hard to scale because assessing free-text explanations and returning timely feedback has had no easy automated solution. We investigate whether a large language model (LLM) can fill that gap. We build a self-explanation tutor for introductory programming, ESSE, in which students explain lines of worked examples and receive immediate LLM feedback on the correctness and completeness of each explanation, and we pursue two goals. First, we ask whether the LLM judges student explanations well enough to serve as the engine of the tutor; we assess its judgments against two independent human reference standards of different kinds, a single domain expert and a crowd of non-exper

Source ↗
technology Thu, 27 Aug 2026 00:00:00 -0400
arXiv cs.CY

The AI Adaptation Gap in Higher Education: Students, Faculty, and Administrative Staff

arXiv:2608.25063v1 Announce Type: new Abstract: The purpose of this study was to analyze patterns of artificial intelligence (AI) use and attitudes toward AI among students, faculty, and administrative staff at a large university specializing in teacher education. The analytical sample comprised 1809 students, 250 faculty members, and 62 administrative staff members (N = 2121). Three role-adapted 75-item questionnaires covered the frequency and contexts of AI use, perceived usefulness, trust and control, academic integrity concerns, responsible-use norms, institutional policy clarity, and perceived improvement in output quality. Data analysis included descriptive statistics, Welch group comparisons, pooled ordinary least squares (OLS) models, reliability and dimensionality checks for observed indices, and exploratory student-only K-means clustering. The results revealed a pronounced AI adaptation gap across university groups. Students reported higher current AI-use intensity and percei

Source ↗
technology Thu, 27 Aug 2020 13:20:58 +0000
HN: medical education

Medical Education Needs Rethinking

Article URL: https://www.scientificamerican.com/article/medical-education-needs-rethinking/ Comments URL: https://news.ycombinator.com/item?id=24293391 Points: 2 # Comments: 0

Source ↗
technology Thu, 25 Jun 2026 17:01:20 -0400
EdTech Mag (Higher)

How Universities Can Manage Vendor Risk After the Canvas Breach

In May, a cybercriminal group executed the largest educational data breach on record, targeting Instructure, the company behind the Canvas learning management system. The breach impacted 275 million students, teachers and staff across approximately 9,000 education institutions. Many took quick action. The University of Wisconsin-Madison, for example, issued real-time alerts warning faculty and students: “If Canvas prompts you to perform any action — such as clicking a link, logging in, resetting your password or completing any tasks — do not proceed.” For universities, the incident…

Source ↗
technology Thu, 25 Jun 2026 09:00:00 +0000
Tech & Learning

5 AI Education Trends According To A Microsoft Executive

The conversation around AI in schools is changing almost as rapidly as the technology. Here are some recent trends.

Source ↗
technology Thu, 25 Jun 2026 07:25:00 -0400
EdTech Mag (Higher)

How Community Colleges Can Use Data to Align Curriculum With Workforce Needs

Community colleges serve as a bridge between education and employment, helping students gain the skills needed for local and regional jobs. But with workforce needs evolving more rapidly, these institutions are under pressure to ensure programs remain aligned with labor market demand. Data analytics and artificial intelligence (AI) are helping community colleges leverage labor market intelligence (LMI) to make more informed decisions about program creation, student success initiatives and workforce development strategies. However, overcoming organizational challenges, fragmented systems, data…

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

ESBMC-PLC+: A Unified IEC 61131-3 Formal Verification Framework as a PLCverif Successor

arXiv:2606.23870v2 Announce Type: replace-cross Abstract: PLCverif is the most mature open-source platform for PLC formal verification, developed at CERN and in production use since 2019. Yet it has two fundamental limitations: no support for Ladder Diagram (LD) programs, the dominant PLC notation, and reliance on CBMC as its primary backend, which restricts verification to bounded proofs. The PLCverif authors themselves identified ESBMC as the appropriate backend improvement. Prior work established ESBMC-PLC (a textual LD frontend with k-induction) and ESBMC-GraphPLC (graphical PLCopen XML support); together, they cover LD with unbounded proofs but not Structured Text (ST), and graphical LD with timer/counter function blocks remains unverifiable. This paper presents ESBMC-PLC+, a unified framework that closes both gaps: (1) an ST/SCL frontend via the MATIEC IEC 61131-3 compiler, routing C-compiled ST to ESBMC with nondeterministic input modeling and YAML property injection; (2) functi

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

Memory Contagion: Cross-Temporal Propagation of Evaluator Bias via Agent Memory

arXiv:2606.23195v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) agents increasingly rely on memory systems to maintain long-term coherence. Recent work shows that agent memories degrade during continuous consolidation. However, existing research assumes memories are derived from unbiased experiences. In this work, we identify and formalize a novel phenomenon: Memory Contagion -- the cross-temporal propagation of evaluator bias through agent memory. We show that when agents are trained or guided by biased evaluators, their experiences become biased; when these trajectories are stored and consolidated into memory, the bias propagates to future agents retrieving from the same memory store, even when consolidation is perfect (oracle). Across two bias types (length preference, authority bias) and four experimental phases, we demonstrate: (1) Memory Contagion occurs for length bias even with perfect consolidation on older models (Gamma_A = 13.18, DeepSeek V4-Chat), while

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning

arXiv:2606.22873v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) are increasingly deployed in consumer, medical, financial, and enterprise applications. This broad deployment expands the safety surface: risks can arise from multimodal question answering, assistant responses, and cross-modal composition, while moderation policies may vary across products, regions, and deployment stages. Most existing guardrails either rely on fixed taxonomies or target only a narrow set of interaction settings, which limits their adaptability when safety rules change at deployment time. We present \textbf{SingGuard}, a policy-adaptive multimodal guardrail model family for safety assessment in multimodal conversations. SingGuard treats the active policy as a runtime input: given natural-language rules, it checks the target content against the active policy rule by rule and predicts both the safety label and the triggered rule. To balance efficiency and interpretability, SingGuard s

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

VADAOrchestra: Neurosymbolic Orchestration of Adaptive Reasoning Workflows

arXiv:2606.22485v2 Announce Type: replace-cross Abstract: Decision-making in real-world settings rarely follows a fixed script. Instead, it unfolds as a dynamic reasoning process in which the appropriate course of action evolves as new context and data become available. Traditional Business Process Management systems provide rigor, determinism, and auditability, yet they generally struggle to adapt their execution at runtime. Conversely, agentic systems based on Large Language Models (LLMs) bring flexibility to decision-making, but they are inherently opaque, often unreliable, and suffer from significant scalability constraints when operating over large datasets. To combine these complementary paradigms, we introduce VADAOrchestra, a neurosymbolic framework that models complex workflows as evolving reasoning processes. The framework adopts a hybrid approach: given a user query and a collection of data sources, an LLM-based orchestrator incrementally plans and adapts the workflow. This

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

Toten: A Knowledge-Based System For Structure-Preserving Representation Of Physical Quantities And Technical Notation In Brazilian Portuguese

arXiv:2606.19626v2 Announce Type: replace-cross Abstract: AI pipelines that reason quantitatively over technical text depend on input where physical quantities, numbers, units, and symbolic expressions arrive intact; when these entities fragment at tokenization, errors propagate downstream. Byte-Pair Encoding, optimized for vocabulary compression, is blind to such entities and fragments them into arbitrary subwords -- a problem aggravated in technical Brazilian Portuguese. We present TOTEN, a knowledge-based system whose input representation preserves each technical entity as a whole, typed unit: vocabulary is not derived statistically but classified declaratively under a formal ontology of engineering entities (OEE). The core is the triple : types, principles, and invariants; a classifier mapping raw text into typed regions; and instantiators yielding a self-descriptive representation. Integrity rests on deterministic coupling to three external authorities: Pint (dimensional), Unicode

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages

arXiv:2606.19157v2 Announce Type: replace-cross Abstract: AudioLLMs enable speech recognition conditioned on textual prompts such as domain descriptions or entity lists. However, it remains unclear whether these models genuinely utilise such context or rely on parametric knowledge learned during pretraining. Existing benchmarks cannot answer this question because they evaluate transcription under fixed prompting conditions and rarely include explicit contextual inputs. We introduce IndicContextEval, a 56-hour multilingual benchmark of natural speech from 555 speakers across 8 Indian languages and 23 professional domains. We design a 7-level prompting framework that progressively introduces contextual signals, including metadata, natural-language descriptions, entity lists in English and native script, and adversarial prompts with incorrect entities. Evaluating five models reveals substantial differences in context utilisation behaviour, highlighting the need for explicit evaluation of

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

daVinci-kernel: Co-Evolving Skill Selection, Summarization, and Utilization via RL for GPU Kernel Optimization

arXiv:2606.16497v2 Announce Type: replace-cross Abstract: GPU kernel optimization represents a paradigm where functional correctness is assumed and execution efficiency is the objective. We present daVinci-kernel, a reinforcement learning framework that couples skill discovery with skill exploitation through a dynamically evolving skill library. daVinci-kernel jointly trains three agents sharing one LLM backbone: a Skill Selection Agent that retrieves relevant techniques via BM25 and LLM reranking, a Policy Agent that generates multi-turn CUDA/Triton kernels conditioned on selected skills, and a Skill Summary Agent that distills successful rollouts into reusable skills. Candidate skills are added only after execution-based verification confirms reproducible speedups. All three agents share a single LLM backbone, are initialized via a structured SFT cold start on diversity-filtered data, and are then jointly optimized end-to-end with multi-turn REINFORCE and per-agent advantage estimati

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism

arXiv:2606.07512v2 Announce Type: replace-cross Abstract: Current Vision-Language Models struggle with hours-long videos because processing full-length visual sequences induces prohibitive token explosion and attention dilution. To overcome this, we introduce MemDreamer to decouple perception and reasoning, shifting long-video understanding into an agentic exploration process. As a plug-and-play framework, it incrementally streams videos to construct a Hierarchical Graph Memory, a top-down three-tier architecture for semantic abstraction, anchored by a foundational graph capturing spatiotemporal and causal relations. During inference, the reasoning model employs agentic tool-augmented retrieval, navigating hierarchies, searching nodes, and traversing logical edges via an Observation-Reason-Action loop. Experiments show MemDreamer achieves SOTA results across four mainstream benchmarks, narrowing the gap with human experts to only 3.7 points. It constrains the reasoning context window t

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

Continual Knowledge Updating in LLM Systems: Learning Through Multi-Timescale Memory Dynamics

arXiv:2605.05097v3 Announce Type: replace-cross Abstract: LLMs are trained once, then deployed into a world that never stops changing. External memory compensates for this, but most systems manage it explicitly rather than letting it adapt on its own. Biological memory works differently: coupled multi-timescale dynamics make new associations immediately usable, strengthen what repetition confirms, and let the rest fade. We argue that external memory should follow a similar principle. In Memini, this view takes the form of an associative memory that organizes knowledge as a directed graph. Each edge carries two coupled internal variables, one fast and one slow, following the Benna-Fusi model of synaptic consolidation. From this coupling, episodic sensitivity, gradual consolidation, and selective forgetting are expected to emerge as facets of a single mechanism, reframing external memory as a learning substrate that reorganizes through its own dynamics. This workshop article describes an

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks

arXiv:2604.03314v2 Announce Type: replace-cross Abstract: Foundation models have revolutionized AI, but adapting them efficiently for multimodal tasks, particularly in dual-stream architectures composed of unimodal encoders, such as DINO and BERT, remains a significant challenge. ParameterEfficient Fine-Tuning (PEFT) methods like LowRank Adaptation (LoRA) enable lightweight adaptation, yet they operate in isolation within each modality, limiting their ability in capturing cross-modal interactions. In this paper, we take a step in bridging this gap with Cross-Modal LowRank Adaptation (CoLA), a novel PEFT framework that extends LoRA by introducing a dedicated inter-modal adaptation pathway alongside the standard intra-modal one. This dual-path design enables CoLA to adapt unimodal foundation models to multimodal tasks effectively, without interference between modality-specific and crossmodal learning. We evaluate CoLA across a range of vision-language (RefCOCO, RefCOCO+, RefCOCOg) and au

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

Speech Codec Probing from Semantic and Phonetic Perspectives

arXiv:2603.10371v2 Announce Type: replace-cross Abstract: Speech tokenizers are essential for connecting speech to large language models (LLMs) in multimodal systems. Speech tokenizers are expected to preserve both semantic and acoustic information for downstream understanding and generation tasks. However, emerging evidence suggests that the term "semantic" in speech processing does not align with linguistic lexical-semantic, leading to a mismatch between speech and text modality. In this paper, we systematically analyze the information encoded by several widely used speech tokenizers, evaluating their lexical-semantic and phonetic content through three tasks. Our results show that current tokenizers primarily capture phonetic rather than lexical-semantic structure, deriving practical implications for the design of next-generation speech tokenization methods. Code is released to public at https://github.com/Alexuan/codec_probing_release.

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

CLEF HIPE-2026: Evaluating Accurate and Efficient Person-Place Relation Extraction from Multilingual Historical Texts

arXiv:2602.17663v2 Announce Type: replace-cross Abstract: HIPE-2026 is a CLEF evaluation lab dedicated to person-place relation extraction from noisy, multilingual historical texts. Building on the HIPE-2020 and HIPE-2022 campaigns, it extends the series toward semantic relation extraction by targeting the task of identifying person-place associations in multiple languages and time periods. Systems are asked to classify relations of two types -- $at$ ("Has the person ever been at this place?") and $isAt$ ("Is the person located at this place around publication time?") -- requiring reasoning over temporal and geographical cues. The lab introduces a three-fold evaluation profile that jointly assesses accuracy, computational efficiency, and domain generalization. By linking relation extraction to large-scale historical data processing, HIPE-2026 aims to support downstream applications in knowledge-graph construction, historical biography reconstruction, and spatial analysis in digital hum

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs

arXiv:2602.06566v3 Announce Type: replace-cross Abstract: Despite recent successes, test-time scaling -- i.e., dynamically expanding the token budget during inference as needed -- remains brittle for vision-language models (VLMs). Unstructured visual reasoning chains entangle perception and reasoning, leading to long, disorganized contexts where small perceptual mistakes may cascade into completely wrong answers. Reasoning also requires expensive reinforcement learning with hand-crafted rewards. Here, we introduce SPARC (Separating Perception And Reasoning Circuits), a modular framework that explicitly decouples visual perception from reasoning. Inspired by sequential sensory-to-cognitive processing in the brain, SPARC implements a two-stage pipeline where the model first performs explicit visual search to localize question-relevant regions, then conditions its reasoning on those regions to produce the final answer. This separation enables independent test-time scaling with asymmetric

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

Streaming-dLLM: Accelerating Diffusion LLMs via Suffix Pruning and Dynamic Decoding

arXiv:2601.17917v3 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) offer a compelling paradigm for natural language generation, leveraging parallel decoding and bidirectional attention to achieve superior global coherence compared to autoregressive models. While recent works have accelerated inference via KV cache reuse or heuristic decoding, they overlook the intrinsic inefficiencies within the block-wise diffusion process. Specifically, they suffer from spatial redundancy by modeling informative-sparse suffix regions uniformly and temporal inefficiency by applying fixed denoising schedules across all the decoding process. To address this, we propose Streaming-dLLM, a training-free framework that streamlines inference across both spatial and temporal dimensions. Spatially, we introduce attenuation guided suffix modeling to approximate the full context by pruning redundant mask tokens. Temporally, we employ a dynamic confidence aware strategy with an earl

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

Generalised Medical Phrase Grounding

arXiv:2512.01085v3 Announce Type: replace-cross Abstract: Medical phrase grounding (MPG) maps textual descriptions of radiological findings to corresponding image regions. These grounded reports are easier to interpret, especially for non-experts. Existing MPG systems mostly follow the referring expression comprehension (REC) paradigm and return exactly one bounding box per phrase. Real reports often violate this assumption. They contain multi-region findings, non-diagnostic text, and non-groundable phrases, such as negations or descriptions of normal anatomy. Motivated by this, we reformulate the task as generalised medical phrase grounding (GMPG), where each sentence is mapped to zero, one, or multiple scored regions. To realise this formulation, we introduce the first GMPG model: MedGrounder. We adopted a two-stage training regime: pre-training on report sentence--anatomy box alignment datasets and fine-tuning on report sentence--human annotated box datasets. Experiments on PadChest

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

Position: Reasoning After Perception Means Reasoning Without Vision

arXiv:2507.16863v2 Announce Type: replace-cross Abstract: A common belief in multimodal research is that the perceptual weaknesses of vision--language models can be compensated by stronger language reasoning (e.g., chain-of-thought, in-context learning, or external tools). We challenge this assumption. We argue that for a broad class of visual tasks hard to specify in language, failures stem from a structural fatality where the temporal decision of \textit{when} to reason strictly dictates the spatial constraint of \textit{where} reasoning takes place. When visual reasoning is deferred to language generation, current architectures do not merely delay computation; they displace it from the continuous visual representation to a discrete textual space. Consequently, the sequential ``Perception-then-Reasoning'' paradigm degenerates perception into a passive, one-off feature encoding process, rendering it functionally equivalent to ``Reasoning-in-Text-Space'', where task-critical spatial si

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning

arXiv:2507.16518v3 Announce Type: replace-cross Abstract: Recent advances in multimodal large language models (MLLMs) have shown impressive reasoning capabilities. However, further enhancing existing MLLMs necessitates high-quality vision-language datasets with carefully curated task complexities, which are both costly and challenging to scale. Although recent self-improving models that iteratively refine themselves offer a feasible solution, they still suffer from two core challenges: (i) most existing methods augment visual or textual data separately, resulting in discrepancies in data complexity (e.g., over-simplified diagrams paired with redundant textual descriptions); and (ii) the evolution of data and models is also separated, leading to scenarios where models are exposed to tasks with mismatched difficulty levels. To address these issues, we propose C2-Evo, an automatic, closed-loop self-improving framework that jointly evolves both training data and model capabilities. Specifi

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

Homogeneity Bias in Open-Weight LLMs Is Robust to Decoding Hyperparameters

arXiv:2501.02211v2 Announce Type: replace-cross Abstract: Large language models (LLMs) reproduce homogeneity bias -- the tendency to portray marginalized groups as more internally similar than dominant groups -- but whether this bias is stable or an artifact of inference settings has only been studied in single proprietary models. We map homogeneity bias across a 5x5 temperature-by-top-p grid in seven open-weight instruction-tuned LLMs (7-20B parameters). Hispanic and Asian Americans are portrayed as more homogeneous than White Americans in at least 18 of 20 hyperparameter configurations across six of seven models, including at extreme sampling settings. African American and gender bias show model-specific variation in direction. A conservative cell-level re-analysis confirms Hispanic and Asian homogeneity as robust, while weaker African American and gender signals largely do not survive, establishing group-specific robustness. We also apply the same grid to a names-based paradigm in w

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

Privacy-Aware Visual Language Models

arXiv:2405.17423v4 Announce Type: replace-cross Abstract: As Visual Language Models (VLMs) become increasingly embedded in everyday applications, ensuring they can recognise and appropriately handle privacy-sensitive content is thus essential to protect users. To this end, we conduct a comprehensive evaluation of twelve state-of-the-art VLMs and identify limitations in their understanding of visual privacy. However, existing privacy-related datasets often suffer from label inconsistencies, limiting their reliability. To address this, we introduce two compact, high-quality benchmarks, PrivBench and PrivBench-H, that focus on commonly recognised visual privacy categories aligned with the General Data Protection Regulation (GDPR). Additionally, we present PrivTune, an instruction-tuning dataset specifically curated to improve privacy sensitivity. We obtain multiple Privacy VLMs by fine-tuning off-the-shelf VLMs on only a few hundred samples from PrivTune, which leads to substantial gains

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

Narrative Feature or Structured Feature? A Study of Large Language Models to Identify Cancer Patients at Risk of Heart Failure

arXiv:2403.11425v4 Announce Type: replace-cross Abstract: Cancer treatments are known to introduce cardiotoxicity, negatively impacting outcomes and survivorship. Identifying cancer patients at risk of heart failure (HF) is critical to improving cancer treatment outcomes and safety. This study examined machine learning (ML) models to identify cancer patients at risk of HF using electronic health records (EHRs), including traditional ML, Time-Aware long short-term memory (T-LSTM), and large language models (LLMs) using novel narrative features derived from the structured medical codes. We identified a cancer cohort of 12,806 patients from the University of Florida Health, diagnosed with lung, breast, and colorectal cancers, among which 1,602 individuals developed HF after cancer. The LLM, GatorTron-3.9B, achieved the best F1 scores, outperforming the traditional support vector machines by 39%, the T-LSTM deep learning model by 7%, and a widely used transformer model, BERT, by 5.6%. The

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

PhoneBuddy: Training Open Models for Agentic Phone Use

arXiv:2606.23049v2 Announce Type: replace Abstract: Phones are becoming an important execution surface for general-purpose agents, but training open models for reliable phone use remains difficult because the environment that matters at deployment, real devices running real apps, is slow, stateful, side-effectful, and hard to reset or verify, while scalable mock environments only approximate real behavior. We present PhoneBuddy, a training recipe and open-model line for agentic phone use that combines a real-app environment with a mock-app environment, PhoneWorld, which reconstructs runnable mock apps from real GUI usage structure. PhoneBuddy first builds a shared supervised fine-tuning stage from trajectories collected in both environments, then compares real-app RL against mixed RL across both environments. Across a 150-task human evaluation on real phones spanning apps, mini-apps, and cross-app workflows, task success rate improves from 36.67\% after supervised fine-tuning to 40.67\

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

Approximate Structured Diffusion for Sequence Labelling

arXiv:2606.18856v2 Announce Type: replace Abstract: Sequence labelling, a core task of Natural Language Processing (NLP), consists in assigning each token of an input sentence a label. From a Machine Learning point of view, sequence labelling is often cast as a Linear-Chain Conditional Random Field (CRF) parametrised by a neural network. While this approach gives good empirical results, CRFs assume a finite decision span (eg label bigrams) which can limit their expressivity and hurt performance when long-range dependencies are required. We show we can leverage diffusion to train a CRF conditioned on an entire label sequence, with the caveat that the condition is on a noisy version of labels. We show experimentally that this method, in conjunction with approximate CRF inference, improves label accuracy with a 16.5% error reduction for POS-tagging.

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting

arXiv:2606.18394v2 Announce Type: replace Abstract: Speculative decoding (SD) accelerates autoregressive Large Language Models (LLMs) by drafting multiple tokens and verifying them in parallel, but it faces a scaling limitation: increasing the draft budget improves speed only when acceptance remains high and drafting overhead stays low. This ceiling has been difficult to break because prior head-based SD methods face a causality-efficiency dilemma. Autoregressive drafters produce path-conditioned candidates that are effective for tree speculative decoding with higher acceptance length, but their drafting cost grows with tree depth. Bidirectional block-diffusion drafters generate all positions in one pass, but their branch-agnostic marginals can form individually plausible yet mutually inconsistent trees, wasting budget and reducing acceptance. We propose JetSpec, a head-based SD framework that combines one-forward drafting efficiency with branch-wise causal conditioning. JetSpec trains

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

Learning task-specific subspaces via interventional post-training of speech foundation models

arXiv:2606.17967v2 Announce Type: replace Abstract: Speech foundation models, pre-trained on large corpora of unlabelled speech data, produce general-purpose representations which are useful across tasks. However, these representations encode information about salient speech variables in a distributed manner, while downstream speech tasks rely on only some of this variability. In this work, we propose a post-training refinement approach using interventional contrastive learning. By leveraging an interventional dataset and multi-part contrastive loss, we learn a transformation from the entangled representation space of speech foundation models into separate content and speaker subspaces. We evaluate the learnt representations on speaker verification and keyword spotting tasks, showing improved out-of-domain speaker verification performance and evidence that speaker and content information are separated across the learned subspaces.

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

Shared Doubt: Zero-Shot Cross-Lingual Confidence Estimation for Language Models

arXiv:2605.31220v2 Announce Type: replace Abstract: Confidence estimation (CE), i.e., quantifying the reliability of a model's prediction, has attracted great interest in the context of large language models (LLMs). However, most studies focus on English, ignoring the multilingual reality of LLM usage, while many CE methods degrade or require retraining across languages. To address this gap, we investigate whether multilingual LLMs encode shared, language-transferable confidence features in open-ended question answering. We use a lightweight linear probe that predicts answer correctness directly from intermediate representations. Trained monolingually, the probe generalizes zero-shot to unseen, typologically diverse languages without target-language supervision. Learned layer weights and multiple ablations reveal that confidence features concentrate in middle layers across languages, suggesting a shared confidence subspace. While zero-shot cross-lingual performance depends on similarit

Source ↗
technology Thu, 25 Jun 2026 00:00:00 -0400
arXiv cs.CL

Scaling Laws for Agent Harnesses via Effective Feedback Compute

arXiv:2605.29682v2 Announce Type: replace Abstract: Agent harnesses shape language-model performance by controlling tool use, feedback, verification, memory, and repair. Yet raw test-time expenditure, such as tokens, tool calls, wall time, or cost, cannot distinguish useful feedback from redundant or unstable interaction. We introduce \emph{Effective Feedback Compute} (EFC), a trace-level scaling coordinate for informative, valid, non-redundant, and retained feedback. We further define Estimated-EFC, NRS-EFC, harness efficiency $\eta$, and task-demand normalization for realistic traces and heterogeneous tasks. Across synthetic, real, held-out, and prospective evaluations, EFC-based coordinates outperform raw-compute baselines and SAS. Oracle-EFC/$D_{\mathrm{task}}$ reaches $R^2=0.99$ in controlled scaling, and NRS-EFC/$D_{\mathrm{task}}$ reaches $R^2=0.93$ on real traces where raw compute has near-zero or negative fit. Finally, \ours uses EFC as a companion control layer for existing h

Source ↗
Showing 6051–6100 of 10879 signals
← Prev Page 122 of 218 Next →