EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Sep 07, 2026 · 40 ideas · 18624 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Thu, 03 Sep 2026 17:57:53 +0000
MedCity News

Ultragenyx Trial Failure Revives Debate About How to Treat a Rare Neuro Disorder

Ultragenyx Pharmaceuticals’ apazunersen did not meet the goals of its Phase 3 test in Angelman syndrome, a rare genetic neurological disorder with no FDA-approved therapies. The disappointing result has some readthrough to Ionis Pharmaceuticals and Oak Hill Bio, each in clinical development with Angelman drugs designed to work in a similar way as Ultragenyx’s failed molecule. The post Ultragenyx Trial Failure Revives Debate About How to Treat a Rare Neuro Disorder appeared first on MedCity News .

Source ↗
regulation Thu, 03 Sep 2026 17:00:00 -0400
K-12 Dive

Bipartisan House bill aims to stop outsourcing of some Education Department offices

The bill seeks to keep all special education, postsecondary, Native American, and elementary and secondary education activities within the department.

Source ↗
regulation Thu, 03 Sep 2026 17:00:00 -0400
K-12 Dive

LAUSD restricts all students from using AI tools

The ban is being implemented for the 2026-27 school year as the nation's second-largest school system reviews its artificial intelligence policies.

Source ↗
regulation Thu, 03 Sep 2026 16:35:29 +0000
The 74

Is Your School a Bright Spot? Explore The 74’s Reading Map, Lite-Brite Style

Source ↗
audience Thu, 03 Sep 2026 16:25:22 -0400
Higher Ed Dive

Full Sail University cuts 180 jobs

The for-profit arts and media college's student population fell by 13.1% from 2025 to 2026, according to a spokesperson.

Source ↗
behavior Thu, 03 Sep 2026 15:30:05 -0700
Nurse.org

Virginia Nursing Students Could Get Paid for Clinical Work Under New $4.8M Grants

Nursing students in Virginia may soon get paid while completing the clinical training they are already required to do, under a new state grant program. The Virginia Nursing Workforce Center…

Source ↗
behavior Thu, 03 Sep 2026 15:30:05 -0700
Nurse.org

Virginia Nursing Students Could Get Paid for Clinical Work Under New $4.8M Grants

Nursing students in Virginia may soon get paid while completing the clinical training they are already required to do, under a new state grant program. The Virginia Nursing Workforce Center…

Source ↗
audience Thu, 03 Sep 2026 14:54:04 -0400
Higher Ed Dive

Private colleges could lose tax-exempt status under new IRS proposal

If enacted, diversity work deemed illegal by President Donald Trump would disqualify colleges from federal tax exemption.

Source ↗
regulation Thu, 03 Sep 2026 14:30:00 +0000
The 74

Federal Agencies Threaten Kansas School Districts Over Gender-Inclusive Policies

LAWRENCE, Kan. — The U.S. Department of Justice and Department of Education announced a plan for increased enforcement measures against two Kansas school districts for their gender-inclusive policies. In April, the department found four public school districts — Olathe, Shawnee Mission, Topeka and Kansas City, Kansas — were in violation of federal Title IX civil […]

Source ↗
need Thu, 03 Sep 2026 13:30:00 +0000
Hechinger Report

When plastics companies write the lesson plans

This story was originally published by Grist and is reprinted with permission. This fall, 54 million K-12 students are headed back to the classroom for another year of lessons in all the classic subjects: math, English, history, biology. But there’s another subject that’s been sneaking into school curricula: plastics. A new report from the nonprofit […] The post When plastics companies write the lesson plans appeared first on The Hechinger Report .

Source ↗
need Thu, 03 Sep 2026 13:30:00 +0000
Hechinger Report

When plastic companies write the lesson plans

This story was originally published by Grist and is reprinted with permission. This fall, 54 million K-12 students are headed back to the classroom for another year of lessons in all the classic subjects: math, English, history, biology. But there’s another subject that’s been sneaking into school curricula: plastics. A new report from the nonprofit […] The post When plastic companies write the lesson plans appeared first on The Hechinger Report .

Source ↗
technology Thu, 03 Sep 2026 13:03:00 +0000
MedCity News

Healing Through Heritage: Why Culturally Grounded Recovery Matters on Tribal Lands

This shift is not about replacing clinical care with tradition. It is about recognizing that healing in Native communities has always been deeply connected to culture. The post Healing Through Heritage: Why Culturally Grounded Recovery Matters on Tribal Lands appeared first on MedCity News .

Source ↗
behavior Thu, 03 Sep 2026 12:57:37 -0700
Nurse.org

AI Should Work for Nurses, Not the Other Way Around

Editor's note: The author is Chief Nursing Officer of the American Nurses Enterprise, whose AI training initiative is discussed in this article. Nurse.org received no compensation for this article. Artificial…

Source ↗
behavior Thu, 03 Sep 2026 12:57:37 -0700
Nurse.org

AI Should Work for Nurses, Not the Other Way Around

Editor's note: The author is Chief Nursing Officer of the American Nurses Enterprise, whose AI training initiative is discussed in this article. Nurse.org received no compensation for this article. Artificial…

Source ↗
regulation Thu, 03 Sep 2026 12:30:00 +0000
The 74

NYC to Ban Student AI Tools in 2-K Through 8th Grade, Limit Classroom Screen Time

New York City is expected to release its school policy for artificial intelligence on Wednesday, with a ban on student-facing AI tools for 2-K through eighth grade, according to four people briefed on the plans and Education Department documents obtained by Chalkbeat. Under the new rules, the youngest children in the nation’s largest school system […]

Source ↗
technology Thu, 03 Sep 2026 12:17:24 -0400
EdTech Mag (K-12)

Beyond Graduation: How School Data Can Prepare Students for Careers

All too often, students reach their junior year of high school with only a vague notion of what they want to do after they graduate. Waiting until graduation is within reach to begin college and career conversations can leave both the students and counselors feeling rushed and out of time. Districts that are already using data to inform decisions have pivoted to using this information to improve the post-graduation experience for students. “The innovative shift is moving away from a generic ‘here’s a list of colleges’ approach and toward tracking students longitudinally,” says Eric Crespo,…

Source ↗
behavior Thu, 03 Sep 2026 11:32:58 +0000
District Admin

Most Texas voucher recipients were already enrolled in private school

The data undercuts a major argument for the vouchers—that they would chiefly help Texans hoist their kids out of failing public schools. The post Most Texas voucher recipients were already enrolled in private school appeared first on District Administration .

Source ↗
regulation Thu, 03 Sep 2026 10:30:00 +0000
The 74

Opinion: Survey: 1 in 4 Kids Who Needed Mental Health Help Didn’t Get It

When a middle schooler sits down in math class, she might be thinking about the argument she had on the way to school, the stress at home she can’t escape or whether she’s going to eat today. She is probably not thinking about algebra. Adults design policies for her future — the classes she takes, […]

Source ↗
behavior Thu, 03 Sep 2026 10:00:00 +0000
eSchool News

Building an everyday equitable practice for AI in schools

AI technology is advancing rapidly in schools, but research is increasingly showing that Large Language Models (LLMs) lack the common sense, morality, and knowledge of individual classroom dynamics needed to be truly equitable.

Source ↗
technology Thu, 03 Sep 2026 09:00:00 +0000
Tech & Learning

4 Mistakes New Teachers Make and How To Overcome Them

These are common mistakes for those new to the teaching profession, but are easy to overcome.

Source ↗
technology Thu, 03 Sep 2026 07:34:45 +0000
HN: education

With education technology, engagement is not the same as learning

Article URL: https://phys.org/news/2026-09-technology-engagement.html Comments URL: https://news.ycombinator.com/item?id=49547003 Points: 2 # Comments: 0

Source ↗
audience Thu, 03 Sep 2026 07:00:00 +0000
Inside Higher Ed

The Number We’re Not Counting

The Number We’re Not Counting quintina.barne… Thu, 09/03/2026 - 03:00 AM An absenteeism-adjusted participation rate. Byline(s) Carolyn Gentle-Genitty

Source ↗
audience Thu, 03 Sep 2026 07:00:00 +0000
Inside Higher Ed

Another Member of Former Regional Accreditors’ Group Leaves

Another Member of Former Regional Accreditors’ Group Leaves Ryan Quinn Thu, 09/03/2026 - 03:00 AM In its departure announcement, the Middle States Commission on Higher Education said it’s “driving broader partnerships.” Byline(s) Ryan Quinn

Source ↗
audience Thu, 03 Sep 2026 07:00:00 +0000
Inside Higher Ed

Miami University Prohibits Faculty From Teaching Outside for Labor Day Action

Miami University Prohibits Faculty From Teaching Outside for Labor Day Action Emma Whitford Thu, 09/03/2026 - 03:00 AM University officials say the planned event violates the collective bargaining agreement, as well as Ohio law. Faculty say this is “untrue and a gross mischaracterization” of the teach-out, which was intended to be celebratory. Byline(s) Emma Whitford

Source ↗
audience Thu, 03 Sep 2026 07:00:00 +0000
Inside Higher Ed

‘Start From Yes’ Initiative Challenges Transfer Credit Skepticism

‘Start From Yes’ Initiative Challenges Transfer Credit Skepticism Johanna Alonso Thu, 09/03/2026 - 03:00 AM Byline(s) Johanna Alonso

Source ↗
audience Thu, 03 Sep 2026 07:00:00 +0000
Inside Higher Ed

Set the World on Fire

Set the World on Fire sara.custer@in… Thu, 09/03/2026 - 03:00 AM As AI reshapes higher ed, could teaching students how to be good human beings save it? Byline(s) Sara Custer

Source ↗
audience Thu, 03 Sep 2026 07:00:00 +0000
Inside Higher Ed

Student Support Needs Outpace College Capacity

Student Support Needs Outpace College Capacity Joshua.Bay Thu, 09/03/2026 - 03:00 AM A new EAB survey finds financial, mental health and academic challenges are straining colleges already facing staffing and budget constraints. Byline(s) Joshua Bay

Source ↗
audience Thu, 03 Sep 2026 07:00:00 +0000
Inside Higher Ed

The Key Podcast: What Happens if We Get AI Right?

The Key Podcast: What Happens if We Get AI Right? sara.custer@in… Thu, 09/03/2026 - 03:00 AM Byline(s) IHE Staff

Source ↗
audience Thu, 03 Sep 2026 07:00:00 +0000
Inside Higher Ed

Syracuse Chancellor Says Budget Deficit ‘Far From a Crisis’

Syracuse Chancellor Says Budget Deficit ‘Far From a Crisis’ Katherine Knott Thu, 09/03/2026 - 03:00 AM Byline(s) Katherine Knott

Source ↗
audience Thu, 03 Sep 2026 07:00:00 +0000
Inside Higher Ed

More Than Half of College Applicants Reported Key Home-Based Challenges

More Than Half of College Applicants Reported Key Home-Based Challenges Olivia.sanchez Thu, 09/03/2026 - 03:00 AM New data from Common App’s responsibilities and circumstances question—added for all applicants last year—allows students to share the range of extracurricular tasks they perform beyond sports, clubs and volunteer work. Byline(s) Olivia Sanchez

Source ↗
audience Thu, 03 Sep 2026 07:00:00 +0000
Inside Higher Ed

‘Novel’ Partnership Sours Into ‘Hostile Takeover’

‘Novel’ Partnership Sours Into ‘Hostile Takeover’ kathryn.palmer… Thu, 09/03/2026 - 03:00 AM Three years after Antioch and Otterbein Universities partnered to widen student access to academic programs, Otterbein’s president led an effort to dissolve Antioch’s board and fire its new president. Antioch has taken legal action to preserve its autonomy. Byline(s) Kathryn Palmer

Source ↗
audience Thu, 03 Sep 2026 05:00:00 -0400
Higher Ed Dive

Brown, UT-Arlington presidents announce departures

August saw a property proposal dispute culminate in an Arkansas college firing its leader. Plus, a long-serving president revised his exit timeline.

Source ↗
audience Thu, 03 Sep 2026 05:00:00 -0400
Higher Ed Dive

Not enough students know about work-based learning opportunities

Proving the effectiveness of paths such as apprenticeships can help students know their options, writes an Urban Institute senior fellow.

Source ↗
regulation Thu, 03 Sep 2026 05:00:00 -0400
K-12 Dive

High schoolers say school counselors emphasize college pathways — and not CTE

The findings come as career and technical education has gained traction and are “concerning” amid broader postsecondary shifts, a Rand report said.

Source ↗
regulation Thu, 03 Sep 2026 05:00:00 -0400
K-12 Dive

Portland Public Schools ponders up to 18 school closures

Falling enrollment and the need for consistent and stable student resources have led the district to explore options to "rightsize."

Source ↗
behavior Thu, 03 Sep 2026 00:00:00 GMT
EdSurge

Hundreds of Thousands of Eligible Kids Are Waiting for Childcare Assistance

A nationwide investigation reveals massive waitlists for government-funded childcare assistance, leaving low-income families stranded without essential ...

Source ↗
technology Thu, 03 Sep 2026 00:00:00 -0400
arXiv cs.CL

Automated Researchers Can Mitigate Well-characterized Alignment Failures

arXiv:2608.28945v3 Announce Type: replace-cross Abstract: Automating alignment research may accelerate progress toward aligned AI, but whether it does is hard to measure. Luckily, many alignment failures, such as deception, sycophancy, and jailbreaks, are already measurable by public benchmarks. We study whether automated alignment researchers (AARs) can post-train to mitigate alignment failures by proposing training methods and data to simultaneously optimize multiple safety benchmarks, while largely preserving general capability. Across 10 alignment failures, the strongest AAR methods significantly reduce the targeted alignment failures and generalize to a held-out benchmark, multi-turn behavioral audits, and models up to 4.7x larger than the target model. As a human baseline, 28 experienced researchers receive up to eight hours to develop one-shot methods for the same benchmarks, but their methods underperform the best AAR methods. Using human ideas as the AARs' initial research dir

Source ↗
technology Thu, 03 Sep 2026 00:00:00 -0400
arXiv cs.CL

Persistent Sparse Autoencoders: Learning Feature-Specific Timescales in Language Model Representations

arXiv:2607.17117v2 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) decompose language model activations into sparse features, yet these models traditionally encode each token independently, failing to expose information that persists across a sequence. We first show that temporal persistence can naturally emerge in standard SAE features: after a feature activates, the hidden state remains aligned with its direction, and past activations help reconstruct later hidden states. How long this lasts varies widely across features. We therefore introduce Persistent Sparse Autoencoders (Persistent SAEs), an extension of standard SAEs that learns a persistence coefficient for each feature, allowing the model to learn feature-specific timescales from reconstruction alone. Our experiments show that Persistent SAEs retain competitive reconstruction quality while learning a spectrum of timescales: short-timescale (fast) features stay locally interpretable, whereas long-timescale (s

Source ↗
technology Thu, 03 Sep 2026 00:00:00 -0400
arXiv cs.CL

LLM Watermarking as Big Data Provenance: A Deployment-Oriented Systematization

arXiv:2607.10103v2 Announce Type: replace-cross Abstract: As large language models (LLMs) become widely deployed, their outputs can be copied, transformed, and redistributed at scale without reliable evidence of origin, creating risks for trust, accountability, intellectual property (IP) protection, and high-stakes decision-making. LLM watermarking addresses this problem by embedding detectable signals into text during or after generation. However, existing methods vary in design assumptions, threat models, and evaluation criteria, while deployment choices such as watermark placement, detection authority, and key management affect reliability, security, and scalability. This paper systematizes LLM watermarking as provenance infrastructure for large-scale data ecosystems. We organize existing approaches along four deployment dimensions: insertion point, verification authority, operational state, and transformation threat model, and relate them to the big data requirements of Volume, Vel

Source ↗
technology Thu, 03 Sep 2026 00:00:00 -0400
arXiv cs.CL

VTOS: Learning to Orchestrate Vision Tools by Co-Searching Solutions and Observers

arXiv:2606.20728v2 Announce Type: replace-cross Abstract: Vision foundation tools such as open-vocabulary detectors, segmentation models, and post-processing operators are powerful building blocks for computer vision, but their effectiveness depends heavily on how they are orchestrated: which tools are used, in what order, with what parameters, and under what visual conditions. Existing visual-programming agents typically generate a fixed solution pipeline, making them brittle under dense objects, occlusion, small targets, and domain shift. We introduce VTOS (Vision Tools Orchestration Search), a framework for adaptive visual tool orchestration through joint solution-observer search. VTOS co-searches executable solution programs that compose vision tools such as Grounding DINO, SAM, NMS, and slice-and-detect, together with observer programs that diagnose candidate solutions, identify failure modes, and generate actionable feedback. These observations are accumulated in a shared VisionT

Source ↗
technology Thu, 03 Sep 2026 00:00:00 -0400
arXiv cs.CL

Beyond Retrieval: Learning Compact User Representations for Scalable LLM Personalization

arXiv:2606.04547v3 Announce Type: replace-cross Abstract: Personalizing large language models requires adapting model behavior to individual users while preserving robustness and deployment-scale efficiency. Existing approaches typically personalize LLMs either at the input level, by retrieving user histories or constructing profile prompts, or at the parameter level, by maintaining user-specific parameter-efficient modules. The former makes personalization sensitive to retrieval quality and prompt design, whereas the latter incurs storage and maintenance costs that grow with the user population. To address these limitations, we propose TAP-PER (Temporal Attentive Prefix for PERsonalization), a prefix-based framework that encodes user preferences as learnable representations, avoiding the serialization of user histories into prompts and replacing heavy per-user adapters with lightweight user-state prefix embeddings. Inspired by personalized recommendation systems, TAP-PER decomposes us

Source ↗
technology Thu, 03 Sep 2026 00:00:00 -0400
arXiv cs.CL

When Prompts Interact: Assessing Prompt Arithmetic for Deconfounding under Distribution Shift

arXiv:2605.03096v2 Announce Type: replace-cross Abstract: In classification tasks, models may rely on confounding variables to achieve strong in-distribution performance, capturing spurious features that fail under distribution shift. This shortcut behavior leads to substantial degradation in out-of-distribution settings. Task arithmetic offers a potential solution by removing unwanted signals via subtraction of secondary model updates, but it typically requires full fine-tuning, which is computationally expensive. Prompt tuning provides a parameter-efficient alternative by adapting models through a small set of trainable virtual tokens. Task arithmetic on the resulting prompts presents an appealing alternative to operations on entire models, but the extent to which this approach can limit reliance on spurious features remains to be established. In this work, we study whether composing soft prompts through task arithmetic improves robustness to confounding shifts. We propose Hybrid Pro

Source ↗
technology Thu, 03 Sep 2026 00:00:00 -0400
arXiv cs.CL

Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data

arXiv:2604.26841v2 Announce Type: replace-cross Abstract: When do language diffusion models memorize their training data, and how to quantitatively assess their true generative regime? We address these questions by showing that Uniform-based Discrete Diffusion Models (UDDMs) fundamentally behave as Associative Memories (AMs) $\textit{with emergent creative capabilities}$. The core idea of an AM is to reliably recover stored data points as $\textit{memories}$ by establishing distinct basins of attraction around them. Historically, models like Hopfield networks use an explicit energy function to guarantee these stable attractors. We broaden this perspective by leveraging the observation that energy is not strictly necessary, as basins of attraction can also be formed via conditional likelihood maximization. By evaluating token recovery of $\textit{training}$ and $\textit{test}$ examples, we identify in UDDMs a sharp memorization-to-generalization transition governed by the size of the tr

Source ↗
technology Thu, 03 Sep 2026 00:00:00 -0400
arXiv cs.CL

TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning

arXiv:2603.03072v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to assist scientists across diverse workflows. A key challenge is generating high-quality figures from textual descriptions, often represented as TikZ programs that can be rendered as scientific images. Prior research has proposed a variety of datasets and modeling approaches for this task. However, existing datasets for Text-to-TikZ are too small and noisy to capture the complexity of TikZ, causing mismatches between text and rendered figures. Moreover, prior approaches rely solely on supervised fine-tuning (SFT), which does not expose the model to the rendered semantics of the figure, often resulting in errors such as looping, irrelevant content, and incorrect spatial relations. To address these issues, we construct DaTikZ-V4, a dataset more than four times larger and substantially higher in quality than DaTikZ-V3, enriched with LLM-generated figure descriptions. Using this da

Source ↗
technology Thu, 03 Sep 2026 00:00:00 -0400
arXiv cs.CL

Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation

arXiv:2512.16310v4 Announce Type: replace-cross Abstract: LLM agents can combine individually non-revealing tool returns and disclose a sensitive conclusion, creating Tools Orchestration Privacy Risk (TOP-R). We formalize TOP-R through three conditions: conclusion sensitivity, single-source non-inferability, and compositional inferability. We introduce Library-Grounded Reverse-Inference Seed Expansion (LRSE), a four-library reverse-construction pipeline, and use it to build TOP-Bench, a 1,000-instance benchmark evaluated under a controlled two-stage tool-use protocol. Across six LLM agents, average task completion, leakage, and H-score are 98.0 percent, 88.6 percent, and 20.4. With native reasoning enabled, four models average 81.4 percent final-response leakage and 82.4 percent reasoning-trace leakage. With reasoning disabled, three prompt-only safeguards improve H-score by an average of about 3.4 points on TOP-Bench. We further propose TOP-Align, an SFT+DPO method for learning safer

Source ↗
technology Thu, 03 Sep 2026 00:00:00 -0400
arXiv cs.CL

When Can Large Reasoning Models Save Thinking? Mechanistic Analysis of Behavioral Divergence in Reasoning

arXiv:2505.15276v2 Announce Type: replace-cross Abstract: Large reasoning models (LRMs) have achieved remarkable success on complex tasks, yet their tendency to "overthink" leads to inefficiencies. Although "save-thinking" prompts are intended to mitigate this issue, we find that LRMs still frequently enter the "Still-thinking" mode instead of the expected "No-thinking" mode, especially on difficult queries. To analyze this behavioral divergence, we examine LRMs from three perspectives: confidence at the thinking-termination boundary, divergence in internal attention distributions, and attention allocation across prompt segments. We find that high perplexity is associated with later Still-thinking behavior, and that Still-thinking cases allocate more attention to the original question. Based on these observations, we propose an attention intervention method to regulate this behavior. While this intervention suppresses explicit thinking, it also causes a drop in accuracy, suggesting tha

Source ↗
technology Thu, 03 Sep 2026 00:00:00 -0400
arXiv cs.CL

Multimodal Language Models as Text-to-Image Model Evaluators

arXiv:2505.00759v3 Announce Type: replace-cross Abstract: The steady improvements of text-to-image (T2I) generative models lead to slow deprecation of automatic evaluation benchmarks that rely on static datasets, motivating researchers to seek alternative ways to evaluate T2I progress. We present Multimodal Text-to-Image Eval (MT2IE), an evaluation framework in which a single multimodal large language model (MLLM) acts as an evaluator agent, iteratively generating the evaluation prompts and scoring the resulting images. We show that MT2IE's image-text consistency scores have higher correlation with human judgment than metrics previously introduced in the literature. MT2IE generates prompts that are efficient at probing T2I model performance: closely recovering the official T2I model rankings of three structurally distinct benchmarks from just 20 generated evaluation prompts, 28-105x fewer than the benchmarks' own prompt sets. When compared to existing evaluation metrics such as CLIPSco

Source ↗
technology Thu, 03 Sep 2026 00:00:00 -0400
arXiv cs.CL

Prompting the Unknown: Understanding Response Uncertainty in Large Language Models

arXiv:2407.14845v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are widely used in decision-making across diverse domains. Ensuring the generation of safe and reliable responses is critical for the effective deployment of LLM-based applications, particularly in high-stakes domains such as healthcare and finance. Most of these applications typically use carefully crafted prompts to guide response generation; however, the relationship between prompts and the reliability of LLM-generated responses is not yet fully understood. To address this gap, we propose a novel prompt-response concept model that explains the relationship between the amount of task-relevant information (informativeness) provided in the prompt and the LLM-generated response uncertainty by identifying four sources of response uncertainty: prompt underspecification, model quality, task variability, and semantic redundancy. We prove that response uncertainty decreases as prompt informativeness or mod

Source ↗
technology Thu, 03 Sep 2026 00:00:00 -0400
arXiv cs.CL

Direct Construction of Disambiguated Knowledge Bases from Large Language Models

arXiv:2608.03729v3 Announce Type: replace Abstract: Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large language models (LLMs), treating the model itself as the knowledge source. However, LLMs natively possess no representation of entities, leading to duplicate entries as well as conflations. We propose GPTKB 2.0, a methodology for constructing disambiguated KBs directly from LLMs. GPTKB 2.0 incorporates on-the-fly disambiguation of entities, relations and classes, and is meticulously designed to satisfy both scalability and disambiguation accuracy. We analyze the central design decisions and characterize the trade-offs between accuracy, scale, and cost. We execute GPTKB 2.0 at scale, obtaining a materialized KB containing over 1M disambiguated entities and 38.4M triples. This represents the first million-scale LLM-native KB with explicit internal canonicalization of entities, relations, and classes, a

Source ↗
technology Thu, 03 Sep 2026 00:00:00 -0400
arXiv cs.CL

PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models

arXiv:2606.23092v2 Announce Type: replace Abstract: Humans possess an innate ability to understand fine-grained interpersonal relationships, which is central to everyday social interactions. Although such reasoning is inherently multimodal, it remains largely unexplored by existing multimodal large language models (MLLMs). To address this gap, we introduce PIVOTS, the first benchmark built from Social-IQ 2.0 and YouTube data to evaluate MLLMs' ability to predict bidirectional interpersonal relationship dimensions grounded in established psychology research. In addition, PIVOTS includes auxiliary tasks that assess models' ability to identify and leverage the critical visual cues underlying such predictions. We evaluate both proprietary and open-source MLLMs and conduct detailed ablation studies to analyze the effects of visual modalities and explicit social role information in conversational utterances. We further examine how joint and pairwise prediction settings benefit MLLMs in scori

Source ↗
Showing 8201–8250 of 18624 signals
← Prev Page 165 of 373 Next →