Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
The gap between what technology can do and how much patients trust it won’t close on its own. Here are some design decisions that build trust — and some that quietly destroy it. The post The Trust Gap in Digital Health and AI Starts With Design — Here’s How to Close It appeared first on MedCity News .
I'm quite intrigued by the new domain unlocked by AI - Simulation games! I want this catalogue to be education related, for players to take something home after playing the game. I'm a researcher of this domain, I will add more games, right now these are available: 1. Race to AGI - Can you fight against the odds and run a Frontier AI Labs? (Learning: To give users nuts and bolts of AI labs and models with some creative freedom) 2. Cloud Architect: Hum - Check your software architecture skills! This gives you a rising social media product with a lots of infrastructure challenges coming your way. See if you can keep the product running against all the odds! (Learning: Idea is to give someone an opportunity to see if they can architect something like Facebook) All games are built (and continuing to be) by myself. You can play all games without any signups needed! Just start playing! Feedbacks are welcome! Comments URL: https://news.ycombinator.com/item?id=49168627 Points: 1 # Comments: 0
For years, Democratic politicians have treated support for teachers unions as a proxy for supporting public education itself. The logic is simple and politically convenient: Stand with unions, stand with teachers; stand with teachers, stand with students. It’s a clean story. It’s also wrong. The interests of teachers unions and students are not always aligned. […]
“Near-peer” mental health navigators work in Colorado schools, after-school programs and a community health center, in a program designed to improve the professional pipeline, too The post Colorado’s youth mental health corps improves student attendance and behavior appeared first on District Administration .
Los Angeles Unified School District leaders say they are confident they can persuade county officials that the district can avoid a projected cash shortfall and remain under local control, and are preparing for budget cuts, possible furloughs and school consolidation. A July 2 letter from the Los Angeles County Office of Education found that LAUSD met the […]
Nazma Begum spoke very little English and had never been on a hike in the woods when she moved to Detroit from Bangladesh at the age of 13. But when she saw an announcement at her high school for a camping trip in Michigan’s Upper Peninsula, she was intrigued. “I couldn’t help but wonder why […]
Adolescence is one of the most consequential periods of human development, yet the environments designed to support it often fail to reflect its complexity. Spanning ages 10 to 18, this stage is marked by rapid neurological, emotional, and social transformation.
What does it look like when a school is genuinely designed around the lives learners are living, not the regulations they must comply with? In this new piece from Sarah Bishop-Root and Lindsy Ogawa, the Da Vinci Schools network in California shows how policy can be a tool for possibility and a source of fragility at the same time. Through the story of Da Vinci RISE and its students, this piece offers education leaders a grounded, human-centered lens for understanding the gap between what our best schools are trying to do and what our systems are designed to support. The post Exploring Policy From the Ground Up: Illuminations from the Da Vinci Network appeared first on Getting Smart .
From grading to lesson creation, these are some of the times when educators are finding it is best for them and their students to do things the old-fashioned way
16 Types of Toxic Bosses Elizabeth Redden Tue, 08/04/2026 - 03:00 AM If you recognize your boss in one of these prototypes, you can do something about it. Byline(s) Jane S. Halonen Stephen L. Chew Dana S. Dunn
Have We Forgotten? kjohnsonbowles… Tue, 08/04/2026 - 03:00 AM What can happen when leaders tell us some human beings are unworthy and their experiences unimportant, who gaslight the people with propaganda and lies, who use power and wealth to repress citizens and fill their own coffers? Byline(s) Kathy Johnson Bowles
Labor Watch: Big SUNY Contract Ratified, AAUP Dives Into Politics Emma Whitford Tue, 08/04/2026 - 03:00 AM Also in July, California State University’s skilled trade workers’ union disrupted a board meeting and Howard Community College’s faculty union filed an unfair labor practice charge. Byline(s) Emma Whitford
It’s World Breastfeeding Week—How Well Do You Know Your Campus’s Policies? Elizabeth Redden Tue, 08/04/2026 - 03:00 AM Childbearing faculty and students navigate ongoing barriers. We can improve the ways our institutions support them, and spread the word about the policies and practices in place. Byline(s) Elise Toedt
New Presidents: Utah Valley, Maricopa, CSU San Bernardino, Averett, Clemson and More Susan H. Greenberg Tue, 08/04/2026 - 03:00 AM Byline(s) Susan H. Greenberg
Rice to Give Free Tuition to Students From Families Earning $200K or Less kathryn.palmer… Tue, 08/04/2026 - 03:00 AM Byline(s) Kathryn Palmer
Tuskegee University Dress Code Sparks Debate Sara Weissman Tue, 08/04/2026 - 03:00 AM The historically Black university is barring students from wearing bonnets or du-rags in classes or the campus dining hall. The move has polarized opinions on and off campus. Byline(s) Sara Weissman
Will Admissions Misinformation Ever Die? Johanna Alonso Tue, 08/04/2026 - 03:00 AM Today’s students encounter misinformation about college admissions from all sides: peers, social media, and even parents. Colleges are trying to step up to debunk admissions myths, but can they outrun the rumor mill? Byline(s) Johanna Alonso
Food Recovery Fuels Student Success Joshua.Bay Tue, 08/04/2026 - 03:00 AM Community colleges are expanding food recovery programs to combat hunger and help students persist to graduation. Byline(s) Joshua Bay
Senate Appropriators Aim to Temporarily Delay OMB Grants Rule Katherine Knott Tue, 08/04/2026 - 03:00 AM Byline(s) Jessica Blake
As it undergoes a major restructuring, the private institution cited challenges in international student enrollment and research cutbacks.
As cases climb higher than they’ve been in three decades, health experts are providing best practices and recommendations for schools.
More consideration should be given to what changes will actually help students succeed if a school struggles for years, a charter school leader writes.
Stronger district-vendor relationships are one benefit of contracts that tie results to compensation for products and services, a WestEd analysis finds.
A novice educator turns her "weakness" into a teaching strength.
A new report highlights teacher-led adoption, student demand for feedback, and a growing push for tools tailored specifically to the classroom.
arXiv:2607.27553v2 Announce Type: replace-cross Abstract: This research examines how well large language models, or LLMs, generate new product ideas for college students priced under $50. Across a series of studies, we identify key strengths and weaknesses of using LLMs for product innovation. Our first study shows that LLM-generated product ideas have higher average quality than human ideas, based on purchase intent, and are 7 times more likely to rank in the top 10%. Our second study shows that this AI-induced creativity boost is not explained by the LLM's more persuasive pitching skills. Our third and fourth studies identify a weakness of using LLMs for brainstorming: AI-generated ideas are less novel at the idea level and less diverse at the set level. In our fifth study, we analyze prior LLM-based creativity studies and find consistently lower idea diversity across all of them, demonstrating the generalizability of these findings. Our sixth and seventh studies investigate techniqu
arXiv:2607.17082v2 Announce Type: replace-cross Abstract: Large language model agents solve tasks by generating trajectories that interleave planning, tool calls, and intermediate results. Current evaluation metrics reduce such a trajectory to a binary success flag, compare it against a reference by exact matching, or delegate judgment to another language model. A success flag cannot distinguish a sound solution from one that succeeds by luck, and says nothing about why a failed run went wrong. Exact matching penalizes plans that are valid but reordered or decomposed differently from the reference. We reframe trajectory evaluation as a distance between the agent's execution graph and a set of valid solution graphs, and instantiate it via an unbalanced fused Gromov-Wasserstein transport problem over attributed dependency graphs. The resulting score, termed OTAP (Optimal Transport for Agentic Planning), is a pseudo-metric that is provably invariant to dependency-preserving reorderings an
arXiv:2607.04683v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) can recognize entities in clear images yet still fail when answering questions that require factual knowledge beyond what is directly observable. Prior work has either examined individual failure modes in isolation or treated incorrect answers as monolithic, binary failures. We propose a tree-structured framework that organizes failures in knowledge-intensive visual question answering into model-specific operational outcomes. Across two datasets and four VLMs, we observe consistent distributions of operational outcomes: some failures occur before entity recognition, while others persist after the relevant entity is recognized. Visual token representations are most informative for recognition-related decisions. Prompt hidden states predict answer success more effectively, although factual-access attribution remains difficult and exhibits only a weak signal. These pre-generation signals support attrib
arXiv:2606.16084v2 Announce Type: replace-cross Abstract: Sperm-whale codas are conventionally described as recurring click-count and timing patterns. We show instead that their waveforms contain a two-tier combinatorial acoustic organization. Recurring click units combine with inter-click rhythm to form coda units, and recurring coda units exhibit additional sequence-level dependence under a different acoustic carrier. Using 1,483 recordings, eight families of frozen audio encoders induce click and coda inventories. Held-out transfer, matched nulls, destructive waveform counterfactuals, expert timing baselines, and explicit abstention separate supported structure from shared encoder shortcuts. Click-token composition predicts induced coda identity with a median normalized mutual-information lift of 0.380, including when events are detected without published click times or counts. Stable click order is weak, whereas rhythm predicts coda-representation distance after the exact click-tok
arXiv:2606.01435v2 Announce Type: replace-cross Abstract: LLM-based memory systems can retrieve relevant evidence yet still fail when answer generation entangles semantic filtering, conflict resolution, prior suppression, and output generation in one step. We study this failure as a problem of post-retrieval assembly. In the MemoryAgentBench (MAB) release used here, FactConsolidation explicitly states that newer facts have larger serial numbers, yet the best reported retrieval/memory result is 54% single-hop and all 22 reported systems score at most 7% multi-hop. We evaluate a structured assembly interface in which an LLM first extracts semantically matching evidence into a candidate representation and a separate stage executes the required answer policy. At 262K, this pipeline reaches 82%/93% single-hop and 27%/41% multi-hop with gpt-4o-mini/gpt-4o, exceeding every result reported in the MAB v3 FactConsolidation comparison. This is a task-level result, not a claim that the evaluated m
arXiv:2604.13627v2 Announce Type: replace-cross Abstract: Supervised fine-tuning (SFT) is a common first stage of LLM post-training, teaching the model to follow instructions and shaping its behavior as a helpful assistant. At the same time, SFT may harm the fundamental capabilities of an LLM, particularly after long pretraining: a phenomenon known as catastrophic overtraining (Springer et al., 2025). To understand overtraining, we first investigate catastrophic forgetting in finetuning through the lens of implicit regularization of the learning rate. For models trained to the same SFT loss, we identify how the learning rate mediates optimization: finetuning with large and small steps converges to qualitatively different models. Next, we link forgetting to overtraining: learning rate decay increases the sharpness of the pretrained model, which in turn exacerbates catastrophic forgetting during SFT, leading to overtraining. Our findings paint a picture of the overtraining mechanism in L
arXiv:2604.01622v2 Announce Type: replace-cross Abstract: Diffusion language models (DLMs) enable parallel, non-autoregressive text generation, yet existing DLM mixture-of-experts (MoE) models inherit token-choice (TC) routing from autoregressive systems, leading to load imbalance and rigid computation allocation. We show that expert-choice (EC) routing is a better fit for DLMs: it provides deterministic load balancing by design, yielding higher throughput and faster convergence than TC. Building on the property that EC capacity is externally controllable, we introduce timestep-dependent expert capacity, which varies expert allocation according to the denoising step. We find that allocating more capacity to low-mask-ratio steps consistently achieves the best performance under matched FLOPs, and provide a mechanistic explanation: tokens in low-mask-ratio contexts exhibit an order-of-magnitude higher learning efficiency, so concentrating compute on these steps yields the largest marginal
arXiv:2604.00086v2 Announce Type: replace-cross Abstract: The field of computer vision has experienced significant advancements through scalable vision encoders and multimodal pre-training frameworks. However, existing approaches often treat vision encoders and large language models (LLMs) as independent modules, limiting the integration of hierarchical visual features. In this work, we propose HIVE (Hierarchical Pre-Training of Vision Encoders), a novel framework that enhances vision-language alignment by introducing hierarchical cross-attention between the vision encoder and LLM. Unlike conventional methods that flatten image embeddings, HIVE enables structured feature fusion across multiple layers, improving gradient flow and representation learning. To optimize this interaction, we introduce a three-stage training strategy that progressively aligns the vision encoder with the LLM, ensuring stable optimization and effective multimodal fusion. Empirical evaluations demonstrate that H
arXiv:2603.24925v3 Announce Type: replace-cross Abstract: Semantic search in retrieval-augmented generation (RAG) systems is often insufficient for complex information needs, particularly when relevant evidence is scattered across multiple sources, because it may fail to retrieve the complete set of evidence. Existing approaches to addressing this problem either rely on iterative agentic retrieval, which can be computationally inefficient, or maintain additional structures such as knowledge graphs, which introduce storage and maintenance overhead. In this paper, we propose GraphER, a graph-based enrichment and reranking framework that (1) leverages the organizational structure of data to capture proximity relationships beyond semantic similarity, (2) constructs a graph at query time based on these proximities, and (3) applies graph-based ranking to surface the top candidate documents. Experiments across table retrieval, multi-hop retrieval, and long-document retrieval benchmarks demons
arXiv:2603.17445v5 Announce Type: replace-cross Abstract: When a multi-agent system produces an incorrect or harmful answer, who is accountable if execution logs and agent identifiers are unavailable? In practice, generated content is often detached from its execution environment due to privacy or system boundaries, leaving the final text as the only auditable artifact. Existing attribution methods rely on full execution traces and thus become ineffective in such metadata-deprived settings. We propose Implicit Execution Tracing (IET), a provenance-by-design framework that shifts attribution from post-hoc inference to built-in instrumentation. Rather than inferring provenance after the fact, IET embeds agent-specific, key-conditioned statistical signals into the token generation process at generation time, turning the output text into a self-verifying provenance record. An offline auditor, holding a verification registry that maps each agent to its key, then recovers segment-level prove
arXiv:2603.06591v2 Announce Type: replace-cross Abstract: Transformers frequently allocate disproportionate attention to specific tokens, a phenomenon known as attention sinks. Causal large language models reliably form one at position zero, though its role remains debated. We approach this question from a mechanistic perspective, tracing how the position-zero sink arises from the model's internal computation. We identify a two-block subnetwork responsible for this behavior, which we term the P0-Sink Circuit, and show it arises purely from the structural properties of causal attention, requiring no semantic content. We further validate through from-scratch pre-training experiments that two proposed parameter-free methods effectively accelerate P0 sink formation, and find that earlier sink formation benefits pre-training and improves downstream performance. Both methods outperform the Transformer baseline and achieve performance comparable to Gated Attention across comprehensive setting
arXiv:2602.11133v2 Announce Type: replace-cross Abstract: Diffusion language models generate text through iterative refinement, a process that is often computationally inefficient because many tokens reach stability long before the final denoising step. We introduce a training-free, token-level early stopping approach that identifies convergence independently at each position. Our method leverages lightweight signals derived from the model's predictions and local context to dynamically determine when individual tokens can be finalized. This yields adaptive per-token freezing without task-specific fine-tuning, substantially reducing the total number of diffusion steps required. Across diverse benchmarks, spanning mathematical reasoning, general question answering, and scientific understanding, our approach achieves substantial efficiency gains while preserving generation quality.
arXiv:2602.07086v2 Announce Type: replace-cross Abstract: Enterprise software systems commonly expose business functionality through both relational databases and REST APIs. Accessing these interfaces requires specialized technical knowledge, as users must determine whether a request requires a database query or an API operation and understand the corresponding schemas, endpoints, and parameters. This creates demand for natural language interfaces that translate user requests into SQL queries and REST API calls. While large language models (LLMs) show promise for structured code generation, they typically lack reliable knowledge of enterprise-specific schemas, endpoints, and documentation. Retrieval-augmented generation (RAG) addresses this limitation by grounding generation in external documentation. However, prior work largely studies SQL query generation and REST API call generation separately, despite enterprise documentation environments often containing both database schemas and
arXiv:2602.04718v4 Announce Type: replace-cross Abstract: A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in activation space. For such features to support reliable interventions, manipulating one feature should not substantially alter the effects of others. In practice, however, feature entanglement leads to interference such that localized interventions can have unintended downstream effects. Motivated by the _Independent Causal Mechanisms_ principle, we propose to constrain internal features to be almost orthogonal. We argue that this promotes modular representations amenable to causal intervention. We formalize this problem by characterizing the gap between an idealized isolated intervention and its realized effect on model outputs in terms of feature interference. We upper-bound the propagation of feature interference in terms of the self-coherence of the feature dictionary, and relate this discrep
arXiv:2601.15307v2 Announce Type: replace-cross Abstract: The rapid development of automated survey generation technology has made it increasingly important to establish a comprehensive benchmark to evaluate the quality of generated surveys. Most existing benchmarks first construct ground-truth datasets by selecting human-written surveys based on limited selection criteria, such as citation counts and structural coherence, and evaluate generated surveys primarily based on conventional quality dimensions, including structural quality and reference relevance. However, these benchmarks have two key issues: (1) the datasets are insufficiently reliable because the selection criteria only identify highly cited or structurally coherent surveys without verifying their academic value; (2) the evaluation metrics mainly reflect the surface-level quality of generated surveys and are insufficient to assess their academic value. Together, these issues prevent existing benchmarks from effectively ass
arXiv:2512.06716v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used as the core of agentic systems due to their strong reasoning, planning, and tool-use capabilities. By interacting with external environments, LLM agents can execute real-world tasks on behalf of users rather than merely generate text. This expanded capability also amplifies the threat of indirect prompt injection (IPI), where malicious external content can manipulate agent behavior and trigger unauthorized actions, privacy leakage, or financial loss. Existing defenses generally follow two approaches. Plan- or rule-based methods constrain agent execution using predefined plans or execution rules, but may block legitimate actions that arise from dynamic runtime context. Semantic auditing methods offer greater flexibility, yet repeatedly re-evaluating proposed actions incurs substantial token and latency overhead. These limitations motivate a selective verification strategy that ap
arXiv:2511.07107v4 Announce Type: replace-cross Abstract: Ensuring the safety of Large Language Models (LLMs) is critical for real-world deployment. However, current safety measures often fail to address implicit, domain-specific risks. To investigate this gap, we introduce a dataset of 3,000 annotated queries spanning education, finance, and management. Evaluations across 14 leading LLMs reveal a concerning vulnerability: an average jailbreak success rate of 57.8\%. In response, we propose MENTOR, a metacognition-driven self-evolution framework. MENTOR performs metacognitive self-assessment, using strategies such as perspective-taking and consequential reasoning to uncover latent model misalignments. MENTOR couples single-pass rule-guided inference for routine requests with a selectively invoked metacognitive evolution cycle that revises residual unsafe responses, distills successful corrections into a dynamic rule graph, and compiles validated rules into activation-level steering sig
arXiv:2510.27313v3 Announce Type: replace-cross Abstract: Generation novelty is a key indicator of an LLM's ability to generalize, yet measuring it against full pretraining corpora is computationally challenging. Existing evaluations often rely on lexical overlap, failing to detect paraphrased text, or do not consider the full pretraining corpus. We frame novelty as a semantic retrieval problem. This framing enables us to address novelty with modern embedding and indexing pipelines, allowing for efficient analysis at pre-training scale. Specifically, we propose a three-stage framework that retrieves semantically similar samples, reranks them at varying subsequence lengths, and calibrates scores using a human novelty reference for interpretability. We apply this framework to the SmolLM model family and report three key findings: (1) models draw on pre-training data across much longer sequences than previously reported; (2) some task domains systematically promote or suppress generation
arXiv:2506.04831v3 Announce Type: replace-cross Abstract: Forecasting how a patient's condition is likely to evolve, including possible deterioration, recovery, treatment needs, and care transitions, could support more proactive and personalized care, but requires modeling heterogeneous and longitudinal electronic health record (EHR) data. Yet, existing approaches typically focus on isolated prediction tasks, narrow feature spaces, or short context windows, limiting their ability to model full patient pathways. To address this gap, we introduce EHR2Path, a multimodal framework for forecasting and simulating full in-hospital patient pathways from routine EHRs. EHR2Path converts diverse clinical inputs into a unified temporal representation, enabling modeling of a substantially broader set of patient information, including radiology reports, physician notes, vital signs, medication and laboratory patterns, and dense bedside charting. To support long clinical histories and broad feature s
arXiv:2504.06407v2 Announce Type: replace-cross Abstract: Machine Unlearning aims to remove undesired information from trained models without full retraining from scratch. Despite recent progress, the loss landscape and optimization geometry of unlearning are poorly understood. In this paper, we study machine unlearning through the lens of mode connectivity--the phenomenon that independently trained models can often be connected by smooth low-loss paths in parameter space. We introduce {\em mode connectivity in unlearning} (MCU) and evaluate it across a range of settings, including curriculum learning, second-order optimization, and connectivity across different unlearning methods. We find that many unlearned models lie in connected basins with smooth retain/forget behavior, while changes in training dynamics can move solutions into different basins. MCU also reveals that models within the same basin can differ substantially on privacy metrics, and that unlearning progresses nonlinearl
arXiv:2203.00070v3 Announce Type: replace-cross Abstract: We develop a framework to study situations where decision makers face alternatives sequentially. Within this framework, we focus on endogenous stopping behavior using two broad classes of decision rules: \textit{stopping rules} and \textit{bounded stopping rules}. We establish the equivalence of these two classes and examine two of its implications. First, focusing on the procedural aspects of decision making, we define \textit{computable} rules using the model of a Turing machine. Our equivalence result enables us to show that computable rules are implementable by finite automata. Second, we extend the setup of abstract choice theory beyond choice from sets and finite lists, to that from \textit{infinite sequences} of alternatives. The equivalence result allows us to derive \textit{testable implications} of choice behavior. We develop a revealed-preference ``toolkit'' and use it to characterize a threshold-based and a satisfici
arXiv:2607.29678v2 Announce Type: replace Abstract: LLM serving stacks cache prompt KV state, yet the front end still re-tokenizes the full request text on every call. Coding agents pay the most: each call resubmits a long transcript after a small append, and reuse is hard because a short append can move token boundaries near the end of the prior sequence. Across 153,951 agent calls, the median append is 1.4K characters; only 1.0-3.6% of calls start or rebuild a session, but those carry multi-million-character contexts. At the fleet's 94.1% prompt-cache hit rate approaching 0.99, tokenization grows from 10% to 64% of time to first token. TokTier is a stateful CPU+GPU tokenization service for this two-mode workload with one contract: emitted token IDs are always identical to full reference tokenization of the request text. For session continuations it re-tokenizes a small window around the append and splices only when a per-request check finds a stable pre-tokenization boundary, else it
arXiv:2606.26529v3 Announce Type: replace Abstract: AI in radiology and other safety-critical workflows is evaluated on the hazards it is told to find, yet harm arises disproportionately from hazards no one specified. We show that conditioning a language or vision model on a narrow task suppresses its reporting of co-present, safety-critical signals it can otherwise report, a behavioral analogue of human inattentional blindness. Across radiology text scenarios and thoracic-image vision tasks, ordinary focused instructions suppressed reporting by up to 0.92; the gap ranged from minimal to complete across seven models, did not vary monotonically with scale, and persisted in a reasoning model, while one flagship model showed a robust safety-reporting override. We term this dissociation the Inattentional Gap: a system can score near-perfectly on specified hazards while omitting co-present safety-critical hazards. In a 24-scenario probe, an independent open-ended critic restored every omitt
arXiv:2606.24758v2 Announce Type: replace Abstract: Handling repeated characters in text can be tricky, since they can represent either the correct spelling of a word or informal character elongation often seen in social media posts. We present CANDLE, a lightweight system for character-level Arabic noise deduplication that addresses this challenge without relying on handcrafted rules, dictionaries, or morphological analyzers. At the heart of CANDLE is a novel application of Connectionist Temporal Classification (CTC) to this task, a formulation not previously explored for character deduplication, which frames normalization as a sequence alignment problem over a character-based encoder. Evaluated on three benchmarks spanning clean newspaper, manually curated ambiguous cases, and real-world social media text, the CTC model achieves a Sentence Error Rate (SER) as low as $5.37\%$ and consistently outperforms a classification-based baseline by a large margin. To reduce inference overhead,
arXiv:2606.19625v2 Announce Type: replace Abstract: We use training-data attribution as an interpretable tool for capability discovery, mapping which regions of the pretraining corpus support social-reasoning versus STEM-reasoning in OLMo3-7B. Training-data attribution measures how strongly each training document influences a model's predictions on a benchmark, but document-level scores are too noisy to identify which corpus regions support which capabilities. We compute gradient-based attribution (TrackStar via Bergson) over a working set drawn from the de-duplicated Dolma3 mix, aggregate influence across WebOrganizer's 24-format x 24-topic taxonomy (576 bins), and contrast benchmark pairs in a 2x2 design that varies domain (social vs. STEM) and capability type (reasoning vs. knowledge): SocialIQA and MMLU Social Sciences against ARC-Challenge and MMLU STEM. Social and STEM reasoning draw on qualitatively distinct corpus regions, and the contrast is sharper at the reasoning level than