EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Design Concept: Scaffolding Geopolitical Reflection Among Tech Workers

arXiv:2607.28904v1 Announce Type: cross Abstract: This paper presents a speculative Human-Computer Interaction design proposal for encouraging geopolitical reflexivity amongst tech workers at geopolitically relevant technology companies. Recent scholarship in International Relations and Science and Technology Studies increasingly recognizes technology firms and their workers as geopolitical actors whose decisions shape international dynamics. However, existing Responsible Innovation and Responsible AI approaches rarely engage with the geopolitical narratives and imaginaries that underpin contemporary AI development. Building upon RI scholarship on reflexivity, reflective HCI, and creative HCI work on computational narratives, this paper proposes an AI-enabled interactive narrative system in which users engage with a speculative scenario centred on technology, power, and geopolitics. Through narrative interaction, archetype assignment, and socially scaffolded workshop reflection, the sy

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Optimizing Monetization Strategies for Generative AI Firms: Implications for Search Engagement

arXiv:2607.28780v1 Announce Type: cross Abstract: As Generative Artificial Intelligence (GenAI) platforms, such as ChatGPT, have transformed digital search querying behavior, mounting operational costs challenge firms to explore alternative monetization strategies beyond traditional subscription models. However, little is known about how alternative advertising-supported monetization models can help GenAI firms recover costs while maintaining search query engagement. Drawing on the compromise effect and affective primacy theories, we develop a framework wherein the introduction of advertising-supported monetization models influences user upgrading and downgrading decisions, contingent on the number of available monetization options. Across four experiments (N=1063), findings reveal that introducing a single advertising-supported option enhances the compromise effect, encouraging free users to upgrade, but leading paid subscribers to downgrade. However, offering two advertising-supporte

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Measuring Cognitive Engagement in Collaborative Discourse with an Extended ICAP Framework: Comparing Human Annotation, In-Context Learning, and Reflective LLM Agents

arXiv:2607.28651v1 Announce Type: cross Abstract: Collaboration supports learning and problem-solving, but its effectiveness depends on cognitive engagement during discourse. This study applies an extended 7-point ICAP framework based on the Interactive, Constructive, Active, and Passive modes to characterize variation in cognitive engagement during collaborative dialogue. Engagement was coded by trained human annotators and compared with large language model (LLM)-based labeling approaches, including in-context learning (ICL), zero-shot prompting, and self-reflective agents. Interrater reliability among human annotators was robust across framework refinement stages (kappa = 0.906-0.998), higher than the moderate agreement observed for ICL-based annotation (kappa = 0.541-0.609). The human-refined framework improved agreement among human annotators (Delta kappa = 0.10), but produced only modest gains for ICL-based LLMs (Delta kappa less than 0.04). Agent-refined frameworks improved cros

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models

arXiv:2607.28639v1 Announce Type: cross Abstract: We show that knowledge distillation in small instruction-tuned language models has asymmetric effects on bias. On unambiguous tasks (BBQ-disambig), response-based distillation from a Gemma-2-9B teacher improves context-following: for the most biased baseline (SmolLM2-1.7B-Instruct), it cuts the context-overriding error rate from 44% to 24%. On ambiguous tasks (BBQ-ambig), the same distillation destroys per-item refusal calibration: 15% of items where the baseline correctly abstained instead receive stereotype answers, even when overall refusal rate is preserved. The pattern reproduces on a second student family (OLMo-2-1B-Instruct), with silence-loss of 8% and filled-silence accounting for 89% of new bias. Across the full 28-configuration grid, the magnitudes of silence-loss and filled-silence are uncorrelated (Spearman $\rho=0.19$, n.s.), indicating that the two effects arise from distinct mechanisms. Aggregate stereotype metrics (Crow

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges

arXiv:2607.28636v1 Announce Type: cross Abstract: LLMs increasingly serve as automated judges, but their judgments remain vulnerable to cognitive biases. Existing mitigations mostly rely on prompt-driven debiasing, which is brittle across bias types, or human evaluation, which does not scale. We study \emph{Chain-of-Models} (CoM), an automated audit pipeline in which a second model inspects the first model's reasoning trace before producing the final judgment. The key design question is whether the auditor should be the same model, a same-family model, or a different-family model. Across 9 models from 6 families, 4 cognitive biases, and 4 factual datasets, we find that auditor identity matters in two ways. First, standalone bias resistance does not predict audit effectiveness: Kimi-K2.5 is the strongest standalone model on several biases, yet is a weak auditor for Qwen2.5-72B's biased traces. Second, the best auditor is bias-specific: GPT-4o is strongest on bandwagon, authority, and di

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations

arXiv:2607.29624v1 Announce Type: new Abstract: Traditional static assessments rely on a subtractive, deficit-based grading model that often penalizes ambition and obscures diagnostic feedback. Conversely, traditional face-to-face oral examinations introduce severe construct-irrelevant variance by exacerbating performative anxiety and the sociological power imbalances inherent to academic hierarchies. This paper presents the theoretical foundation for the "Socratic Test," an automated, computer-mediated conversational assessment. By integrating Dynamic Assessment principles, multimodal workspaces, Bloom's Taxonomy for real-time proctoring, and the SOLO Taxonomy for structural evaluation, the Socratic Test actively maps a student's cognitive boundaries. This paper formalizes the use of graduated scaffolding to quantify the Zone of Proximal Development (ZPD) and details a non-compensatory, additive grading architecture that prioritizes mastery over penalty and human-AI alignment to ensur

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Triangulating Across U.S. Federal AI Transparency Regimes

arXiv:2607.29540v1 Announce Type: new Abstract: Federal AI systems can deny benefits or flag individuals for deportation, but the public disclosures meant to make those systems visible are fragmented and unevenly detailed. This paper examines three existing U.S. federal transparency regimes---System of Records Notices (SORNs), Information Collection Requests (ICRs), and the AI Use Case Inventory---and asks how well they, individually and together, describe government AI use. We find that no single regime fully reveals how the government constructs or deploys AI: each discloses different aspects of a system, and the current disclosure infrastructure makes it very challenging for the public to track specific AI systems across regulatory regimes and over time. Persistent identifiers are absent, granularity varies widely, and the annual AI Use Case Inventory cycle means federal agencies can deploy systems months before appearing in any official record. Using hand-validated zero-shot classi

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Tragedy of the Cognitive Commons: How AI Could Disrupt the Regeneration of Professional Expertise

arXiv:2607.29380v1 Announce Type: new Abstract: Artificial intelligence is reshaping cognitive work, but Human Resource Development scholarship has treated this transformation as an organizational training challenge, leaving the collective regeneration of professional expertise unexamined. This conceptual paper introduces the Cognitive Commons framework, integrating commons theory, HRD scholarship, and distributed cognition to explain how rational AI adoption decisions can deplete the shared expertise pool professions require for renewal. The framework distinguishes Internalized Mastery (deep domain knowledge from sustained practice) from Distributed Mastery (orchestrating human-AI systems), and develops the Validation Tether: effective AI oversight depends on the expertise AI adoption may undermine. Early labor market and clinical evidence suggests possible disruption to expertise-regeneration pathways in highly AI-exposed sectors, though adoption is recent and the strongest signals c

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Hypergamigication Through Integrating Game Engines and Learning Management Systems: Ender's Game

arXiv:2607.29300v1 Announce Type: new Abstract: This paper discusses games, their use in education, and previous work on integrating game engines and learning management systems (LMS). It proposes a bidirectional integration where game environments are generated using LMS content, introducing the concept of hypergamification as the use of a comprehensive game environment rather than isolated game design elements. A working pilot implementation of an importable Unity package for Blackboard integration is demonstrated, along with a demo game that uses the developed package. The paper also discusses the limitations of the proposed approach and outlines avenues for future work.

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Deployment Wall: A Diagnostic Framework and Instrument for Enterprise AI in the Deployment Era

arXiv:2607.29089v1 Announce Type: new Abstract: Enterprise investment in generative artificial intelligence (AI) tripled in a single year to roughly US$37 billion, yet independent field research finds that about 95% of enterprise generative-AI pilots deliver no measurable profit-and-loss impact. We argue that the dominant explanation--that models are not yet capable enough--is mistaken, and that enterprise AI has entered a Deployment Era in which advantage derives not from model intelligence but from the removal of the organizational and architectural friction that prevents a capable model from reaching production. Building on the software-engineering literature on technical debt and machine-learning deployment, and on a structured synthesis of independent field studies, we make the diagnosis operational. We introduce three linked constructs and one measurement instrument: the Deployment Wall, a six-stage value-leak model that mechanically reproduces observed survival rates; the Seam I

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

IyawoBench v2.0: Extended Diagnostic Evaluation of Large Language Model Clinical Triage in Nigerian Primary Care

arXiv:2607.29085v1 Announce Type: new Abstract: Large language models are being deployed as clinical triage tools in low and middle income countries where trained physicians are scarce. Existing safety metrics, however, produce misleading confidence: models scoring 100% on binary "did not send an emergency home" safety measures may nevertheless exhibit systematic failure modes that render them undeployable at scale. We present IyawoBench v2.0, an extended diagnostic evaluation of large language model clinical triage on 200 synthetic vignettes derived from 1,200 real patient encounters at 19 Nigerian primary health centres. We introduce a formal mathematical framework comprising fourteen definitions and two theorems that decompose triage safety into three distinct failure modes: Conservative Escalation Bias, Systematic Downgrade Bias, and Middle-Tier Instability. We propose the Escalation Bias Index and Expected Deployment Cost as novel metrics that expose failure modes hidden by conven

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

From Process to Evidence: How Computing Can Ground Appropriate Reliance on Legal AI

arXiv:2607.28869v1 Announce Type: new Abstract: Lawyers and self-represented litigants are already using artificial intelligence (AI) to draft legal documents, and courts are responding with rules. After more than 1,500 cases involving AI hallucinations, lawyers have been instructed to perform careful, independent review of AI-assisted filings. Discharging these duties requires what the human-computer interaction (HCI) literature calls ``appropriate reliance,'' which cannot be calibrated without evidence on how often, how badly, and how detectably these tools fail at legal work. Existing research barely describes any of the three. We analyze the official record of the New York court system. The documents repeatedly call for evidence that does not exist (e.g., error rates, do-not-use lists). In its place they invoke procedure, including training mandates, checklists, and uncalibrated human review. The burden falls hardest on those least equipped to bear it: legal aid programs are told t

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Hidden Errors in Big Data: The Case of Property Records

arXiv:2607.28827v1 Announce Type: new Abstract: Big data are the foundation for an increasing share of academic research and AI models deployed in both the public and private sectors, prompting substantial growth over time in reliance on brokered datasets. Brokered property records, which are ubiquitous in studies of gentrification, inequality, and the property tax in the U.S. and serve as inputs to property valuation models, are one notable example. In this paper, we audit two prominent brokered property datasets, finding errors in these data which bias key measures of economic inequality. First, we document that for 1-2% of matched sales in Cook County, IL, from 2018-2021, broker-provided sale prices differ from ground truth sale prices by more than 5%. Moreover, missing data and conceptual differences in the reporting of deed and property characteristics lead to coverage errors ranging from 12 to 15% of transactions. Second, we show that misreporting is highly consistent between bro

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Structured AI Demonstrations and Student LLM Use in Engineering Mechanics: Study Design and Preliminary Results

arXiv:2607.28710v1 Announce Type: new Abstract: The rapid integration of large language models (LLMs) into undergraduate education presents an urgent challenge for engineering instructors. Despite widespread student adoption, there remains a critical lack of domain-specific empirical evidence to guide pedagogical policies and classroom interventions. This manuscript presents a descriptive study design and preliminary findings from an undergraduate engineering mechanics course conducted in Spring 2026. We detail a reproducible survey instrument used to capture student AI usage patterns, attitudes, and verification practices, which are subsequently linked to academic performance metrics. Additionally, we document a deployable sequence of nine structured, instructor-led AI demonstrations designed to model strategic LLM delegation and evaluation. While our preliminary data highlight shifting student behaviors and complex relationships between AI reliance and course outcomes, the primary co

Source ↗
technology Mon, 03 Aug 2026 00:00:00 -0400
arXiv cs.CY

Scaffolding Critical Engagement with GenAI: Transforming Ethnic Minority Preparatory Students' Collaborative Discourse in Prompt Engineering Tasks

arXiv:2607.28630v1 Announce Type: new Abstract: Generative AI (GenAI) holds significant promise for advancing educational equity among ethnic minority students by broadening access to learning resources and mitigating linguistic barriers. However, these benefits are counterbalanced by the risk of cognitive laziness, whereby students may treat GenAI as an answer engine or shortcut rather than as a partner in thinking. This design-based research investigated how pedagogical scaffolding can shift students from passive consumption to critical co-creation with GenAI. The study involved 78 ethnic minority preparatory students in China participating in a three-week GenAI course that integrated a human-in-the-loop workflow and teacher modeling with contrasting cases to disrupt uncritical reliance on GenAI. We employed epistemic network analysis to examine collaborative discourse, thematic analysis to analyze student reflections, and paired-samples t-tests to assess changes in prompt self-effic

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CY

Fill-Side Behavioral Concentration on Polymarket: Identification Limits under Record-Level Attribution

arXiv:2605.11640v2 Announce Type: replace-cross Abstract: This paper studies behavioral concentration in Polymarket's public executed-fill record and formalizes what that record can and cannot identify. A pre-publication reconciliation corrects the empirical scope: the archived extraction covers the legacy CTF Exchange over Polygon blocks 86,008,447-86,107,178, approximately 25 April 2026 17:09 UTC through 28 April 2026 00:00 UTC, rather than the full 21-27 April week stated previously. It contains 13,356,931 OrderFilled records, 77,204 addresses with at least five attributed records, and 43,116 token identifiers; negative-risk markets are absent. The archived feature construction credits both maker and taker addresses on each record. This convention is not invariant to match fragmentation, and mint/burn executions do not admit a universal buyer/seller interpretation. The reported one-cluster result is therefore retained only as a null under the original record-level representation, no

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CY

Epistemic diversity across language models mitigates knowledge collapse

arXiv:2512.15011v3 Announce Type: replace-cross Abstract: Artificial intelligence (AI) increasingly generates the very content used to train future AI systems. This feedback loop can degrade model quality, reduce informational diversity, and ultimately drive knowledge collapse, i.e. a degradation to a narrow and inaccurate set of ideas. We ask: to mitigate collapse, is it better to concentrate the internet's knowledge into a handful of dominant models (referred to as an AI monoculture), or to distribute it across a diverse ecosystem of models? To study the effect of diversity on model performance, we randomly segment the fixed training data across an increasing number of language models and evaluate the resulting ecosystems of models over ten self-training iterations. Our results show that diversity improves long-term performance of models, while monoculture accelerates collapse. Specifically, we observe that the optimal diversity level (i.e., the level that maximizes performance) incr

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CY

Secure human oversight of AI: Threat modeling in a socio-technical context

arXiv:2509.12290v3 Announce Type: replace-cross Abstract: Human oversight of AI is promoted as a safeguard against risks such as inaccurate outputs, system malfunctions, or violations of fundamental rights, and is mandated in regulation like the European AI Act. Yet debates on human oversight have largely focused on its effectiveness, while overlooking a critical dimension: the security of human oversight. We argue that human oversight creates a new attack surface within the safety, security, and accountability architecture of AI operations. Drawing on cybersecurity perspectives, we model human oversight as an IT application for the purpose of systematic threat modeling of the human oversight process. Threat modeling allows us to identify security risks within human oversight and points towards possible mitigation strategies. Our contributions are: (1) introducing a security perspective on human oversight, (2) offering researchers and practitioners guidance on how to approach their hum

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CY

Towards Structurally Explainable Machine-Generated Text Detection: A Graph-Perspective Framework

arXiv:2505.12507v2 Announce Type: replace-cross Abstract: Despite the success of machine-generated text detectors, the black-box nature remains a critical limitation. Traditional explainability methods rely on token-level saliency, insufficient to reveal the high-order structural dependencies that distinguish LLM outputs. In this paper, we propose \textsc{LM$^2$otifs}, a principled framework that shifts detection from linear sequences to graph-structured manifolds. We first provide a theoretical grounding based on probabilistic graphical models, demonstrating that detection performance is more distinguishable in the graph-topological space. Driven by this theory, \textsc{LM$^2$otifs} transforms text into lexical co-occurrence graphs to preserve latent structural fingerprints. The framework employs Graph Neural Networks for robust detection and utilizes graph-specific explainers to extract interpretable motifs. Crucially, our experiments reveal that these structural motifs achieve highe

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CY

AI From the Margins (AIM): Rethinking Participatory AI Design Through the Lived Experience of Minoritized Communities

arXiv:2606.01171v2 Announce Type: replace Abstract: Artificial intelligence (AI) can reproduce and amplify the structural inequities faced by minoritized communities. Participatory AI has been proposed as a response, but participation typically starts after problem definitions and success criteria have been set, leaving limited room for minoritized communities to reshape what an AI system is for. We propose AI From the Margins (AIM): a methodological stance that articulates the conditions under which lived experiences of minoritized communities can be elicited, centered, and carried forward to inform participatory AI design. AIM is not a fixed protocol; it articulates a set of preconditions that can be enacted through different techniques in different settings. We applied AIM in a Dutch healthcare context in eight sessions with 13 women and non-binary people of color and five municipal policy workers, namely through (1) narrative elicitation using the Biographic Narrative Interpretive

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Missing Variable: Socio-Technical Alignment in Risk Evaluation

arXiv:2512.06354v2 Announce Type: replace Abstract: This paper addresses a critical gap in the risk assessment of AI-enabled safety-critical systems. While these systems, where AI systems assist human operators, function as complex socio-technical systems, existing risk evaluation methods fail to account for the associated complex interaction between human, technical, and organizational components. Through a comparative analysis of system attributes from both socio-technical and AI-enabled systems and a review of current risk evaluation methods, we confirm the absence of explicit socio-technical considerations in standard risk expressions. To bridge this gap, we introduce a novel socio-technical alignment ($STA$) variable designed to be integrated into the traditional risk equation. This variable estimates the degree of harmonious interaction between the AI systems, human operators, and organizational processes. A case study on an AI-enabled liquid hydrogen ($LH_2$) bunkering system de

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CY

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

arXiv:2607.28617v1 Announce Type: cross Abstract: System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are rarely disclosed to the public or regulators, creating a serious trust and accountability gap in the wide deployment of AI systems. In this paper, we introduce Artificial Intelligence System Prompt Assurance (AISPA), a user-centric framework for systematically auditing system prompts in AI systems. AISPA examines specific parts of a system prompt and evaluates them along eight dimensions that matter to users. We then use this framework to review 3,249 instructions from system prompts in 88 commercial AI products, classifying each instruction as either protective (of users) or problematic. Our audit surfaces four core findings. First, system prompt design varies substantially across products and developers, with some organizations averaging over 60 protective in

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Role of Causality in Algorithmic Recourse

arXiv:2607.28497v1 Announce Type: cross Abstract: Algorithmic recourse aims to provide individuals with actionable changes to improve their predicted outcomes in high-stakes classification settings, such as loan and mortgage applications. However, most existing approaches focus only on flipping a model's prediction, without accounting for whether the recommended changes lead to genuine improvement in an individual's true qualifications or merely enable strategic gaming of the classifier. Consequently, deployed recourse policies can induce behavioral responses that degrade predictive accuracy and become ineffective after model retraining. In this work, we formalize this failure mode through a causal performative framework for recourse. We model how recourse actions propagate through a structural causal model, capturing interactions among features as well as their effect on the true label. These causal responses induce a non-convex optimization problem, even under standard convex losses.

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CY

Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations

arXiv:2607.28319v1 Announce Type: cross Abstract: This work presents Fairness Pruning, a lightweight structural intervention method designed for the management and future mitigation of demographic bias in large language models (LLMs). As a foundational empirical validation of this method, this work focuses on causal bias localization. Using minimally contrastive prompt pairs and inference-time activation capture, the method identifies neurons that react differentially when processing demographic attributes in GLU architectures, evaluating the signal at the down_proj input. Empirical evaluation was conducted on models of up to 3 billion parameters (Llama-3.2 family and Salamandra-2B), combining standardized benchmark evaluation with qualitative text generation experiments. Results demonstrate that zeroing the identified neurons alters how the model responds to associated demographic variables. However, rather than producing flat mitigation, the intervention causes bidirectional bias des

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CY

Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models

arXiv:2607.28128v1 Announce Type: cross Abstract: LLM tutoring poses a measurement problem: can a general-purpose helpfulness rubric distinguish direct answer-giving from pedagogical guidance? We audit this signal in a pre-registered study. Within each of three tutor bases, we compare conversational and pedagogical policies instantiated with the same underlying model and paired with one fixed weak simulated student. Deterministic detectors measure answer leakage and next-turn independent work. Claude Opus 4.8 is the frozen, condition-blind primary judge. After the Opus scores were fixed, GPT-5.6 Sol was prospectively specified for a post hoc robustness audit of the same 1,179 confirmatory answer-phase tutor turns under the frozen helpfulness and pedagogy rubrics. On the primary base under Opus, the policies do not differ significantly in helpfulness but are perfectly rank-separated under the pedagogy rubric (Cliff's $|\delta|{=}0.10$ vs. $1.0$). Across the two judges, pedagogy contrast

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CY

AI Literacy: An Exercise in Power-Knowledge

arXiv:2607.27547v1 Announce Type: cross Abstract: As generative artificial intelligence becomes one of the most significant systems of knowledge production in our society today, questions relating to who can access and shape that production grow increasingly important in our discourse. This paper argues that the existing frameworks for AI literacy, which are dominated by technical competency and responsible-use principles, are insufficient because they enforce a "consumer" orientation toward AI rather than fostering genuine epistemic agency. Based upon Foucault's concept of power-knowledge, Freire's pedagogy of critical consciousness, and scholarship of digital literacy, this paper proposes a reconceptualization of AI literacy as a critical practice that equips individuals not just to use AI systems, but to critically evaluate them, resist their structuring assumptions, and participate in their governance. The paper further argues that unequal access to AI tools in society recapitulate

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CY

Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups

arXiv:2607.27232v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview. This raises concerns beyond bias in AI: do LLMs grasp the emotional nuances conveyed via textual framing? In this work, we empirically evaluate how well an array of LLMs aligns with human emotional perception. Considering news headlines covering political and geopolitical conflicts, both human participants (n = 3011, a representative sample of the U.K. adult population, via a YouGov survey) and seven LLMs answered whether headlines evoked sympathy for a specified side in a conflict. We find that the correlation between AI and human evaluations varies across models, ranging from very high (0.789, GPT-5.2) to medium (0.4 ,Mistral Large 2512). Crucially, the leading models are broadly aligned with human judgments across all demographic subgroups, including age, gender, level of education, prior geopolitical knowledge, and participants'

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CY

Correcting Mode Collapse in Silicon Sampling with Semantic Similarity Rating

arXiv:2607.28550v1 Announce Type: new Abstract: Silicon sampling refers to the use of Large Language Models (LLMs) to generate responses to surveys. It has shown promise, but tends to generate response distributions with unrealistically low variance. We argue that this mode collapse is due to LLMs failure to generate numeric data, and that text responses may be better suited for this task. We analyze whether Semantic Similarity Rating can improve the fidelity of silicon sampling responses when asked about political attitudes. This method solicits text-only responses from LLMs, then maps this to a numeric scale using text embeddings. We find that this method both improves the fidelity of silicon sampling response distributions, and has few parameters to calibrate.

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CY

AIx4Soccer: A Unified Platform Architecture for Football Club Management and Structured Athlete Development

arXiv:2607.28531v1 Announce Type: new Abstract: Football clubs, academies, and federations operate a growing but fragmented portfolio of digital tools: separate systems for video analysis, GPS/performance tracking, medical records, scouting, and administration. This fragmentation is most acute outside the elite European clubs that can afford integration, producing a digital divide that disadvantages grassroots clubs in developing markets such as Brazil, paradoxically the world's largest exporter of professional players. This paper presents, at a conceptual level, the architecture of "AIx4Soccer One Platform," a multi-tenant cloud SaaS operating system that unifies club-management workflows and embeds a structured athlete-development methodology, the PDI Framework (Plano de Desenvolvimento Individual / Individual Development Plan). We describe two companion components: "Tak Tik," a certified two-sided marketplace connecting clubs with video analysts under a 75%/25% (analyst/platform) re

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CY

When AI Becomes Routine: A Decade of Public AI Mediation in Korean Go Commentary

arXiv:2607.28332v1 Announce Type: new Abstract: When AI systems surpass elite human performance and settle into everyday expert practice, the question that follows is how machine judgment is made publicly intelligible and attributable. We study Korean Go commentary on YouTube, where AI systems such as KataGo became standard analytic tools after AlphaGo. Our corpus spans a decade (2016--2025) and approximately $1{,}900$ hours of footage across institutional broadcasters and creator-led channels, in four phases of AI availability. We document a widening asymmetry between visual and verbal AI presence: AI winrate graphs are visible for about $98\%$ of late-period institutional broadcast time, yet AI-salient talk accounts for only $2.63\%$ of sentences. What recedes is the source label, not the metric: winrate and point-gap talk persists while ``AI'' itself goes unsaid. We read this recession as the communicative signature of domestication. Our strongest evidence is a compositional shift i

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CY

Technology-Enhanced Tabletop Exercises for Cybersecurity Education: Lessons Learned

arXiv:2607.28179v1 Announce Type: new Abstract: This innovative practice full paper examines the integration of technology-enhanced tabletop exercises (TTXs) into computing education, focusing on cybersecurity curricula. The motivation is to better prepare students for complex, collaborative problem solving typical of incident response and IT governance, where coordination, communication, and timely decision-making are essential. Although TTXs are well-established in professional practice, they remain underused in universities. We address this gap by augmenting TTX delivery and evaluation through the INJECT Exercise Platform (IXP), a web-based environment that automates scenario flow and enables data-driven assessment. Our practice implements IXP to automatically deliver scenario updates, facilitate team discussions, and collect interaction data to support automated assessment. This combination enhances realism, reduces instructor workload, and provides actionable insight into student

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CY

Asymmetric Communication: Large Language Models and Language Games

arXiv:2607.28137v1 Announce Type: new Abstract: Contemporary AI discourse attributes to language models properties they cannot bear: general intelligence as substrate-independent cognition, hallucination as cognitive failure, agency as autonomous goal-pursuit, sentience as emergent inner life, alignment as goal synchronization. This paper argues that these are instances of a single category mistake--properties constituted within human communicative practice are projected onto the machine side--and explains its structure. Human-LLM interaction constitutes a language game in which one side bears all normative activity. We call this configuration asymmetric communication since model outputs circulate communicatively, entering further exchanges, without the system undertaking commitments, bearing entitlements, or performing the assessment on which discursive standing depends. Three conditions define the asymmetry: (i) correctness is enforced exclusively by the receiver; (ii) accountability

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CY

When AI Does the Work, What Is Learning For? Post-Instrumental Learning and the Risk of Capacity Dissolution

arXiv:2607.28041v1 Announce Type: new Abstract: As AI systems become capable of producing the essays, code, reports, summaries, plans, and decisions through which institutions usually recognize competence, a familiar question becomes harder to answer: what is learning for? Existing AI ethics rightly emphasizes present failures--bias, opacity, hallucination, labor extraction, privacy risk, and weak accountability. But if the case for learning rests only on those failures, then each technical improvement appears to weaken it. This article develops a different answer. Using the idealization of AI that executes specified tasks flawlessly while lacking authority over purposes, legitimacy, and responsibility, we argue for post-instrumental learning: learning that preserves the capacities people and institutions need when many useful outputs can be delegated. We analyze five such capacities--end-setting, reason-giving, contestability, refusal/revision, and participation--and name their erosio

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CY

Scaling, Lock-In, and Proxy Compliance: A Political Economy of Responsible AI

arXiv:2607.28023v1 Announce Type: new Abstract: AI accountability at scale is an institutional problem: who can observe, verify, and change deployed systems. We develop a sequential political-economy model in which an AI vendor chooses auditability and substantive mitigation, a deployer monitors after adoption while facing switching costs, and enforcement depends on verifiable evidence. Anticipating the deployer's monitoring response, the vendor may stop at an observable procurement floor while mitigating below the social first best, producing a proxy-compliance equilibrium. We characterize the unique interior equilibrium and the corner in which harm is fully mitigated. Independent audit rights raise enforcement exposure directly; portability restores deployer leverage; incident reporting adds a regulator-visible evidence channel; and outcome-linked liability creates incentives that do not depend on vendor-controlled detection. The results explain why documentation and standardized eva

Source ↗
technology Fri, 31 Jul 2026 00:00:00 -0400
arXiv cs.CY

Is Solving Better Than Evaluating GenAI Solutions?

arXiv:2607.27586v1 Announce Type: new Abstract: As Generative AI (GenAI) tools become increasingly capable of generating solutions to computing assignments, the computing education community is exploring pedagogical approaches that emphasize solution evaluation, verification, and critique alongside traditional solution generation. However, evidence regarding the impact of such evaluation-centered tasks on student learning remains limited, particularly in upper-division, theory-heavy courses. We conducted a randomized A/B crossover study (N=220) in a junior-level algorithms course to compare evaluating GenAI-generated solutions with traditional problem solving. Across six assignments, student working groups either solved challenging algorithmic problems directly or evaluated often-flawed GenAI-generated solutions, with roles reversed midway through the semester. We found no statistically significant differences between groups in midterm scores, final exam scores, overall course grades,

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training

arXiv:2608.20574v2 Announce Type: replace-cross Abstract: Open-ended language-model evaluation often substitutes another model or a small preference panel for a missing answer key. We introduce FlavourBench, which instead compiles dense answer maps from a versioned culinary environment. Each task asks for a three-ingredient portfolio from eight candidates; before inference, Epicure scores all 56 portfolios. We evaluate 27 frontier endpoints on the same 534 substitution, pairing, and constraint tasks, yielding 14,418 complete model-task observations. Anchor-cluster bootstraps and multiplicity-controlled paired tests resolve 101 of 351 model contrasts. Grok 4.6 has the largest point estimate at 65.1, but the corrected evidence does not identify a unique best endpoint. The ranking replicates across independently compiled panels and remains similar under alternative metrics, task filters, family weights, and three public Epicure checkpoints. We then run a preregistered, three-seed post-tra

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

LLM Analysis of 150+ years of German Parliamentary Debates on Migration Reveals Shift from Post-War Solidarity to Anti-Solidarity in the Last Decade

arXiv:2509.07274v4 Announce Type: replace-cross Abstract: Migration has been a core topic in German political debate, from postwar expellee displacement to labor migration and recent refugee movements. Large-scale analysis of such political discourse has traditionally required extensive manual annotation, limiting coverage. Large language models (LLMs) offer a scalable alternative. Using a theory-driven annotation scheme, we examine how well LLMs annotate subtypes of solidarity and anti-solidarity in German parliamentary debates and whether the resulting labels support valid downstream inference. We first evaluate multiple LLMs across model size, prompting strategies, fine-tuning, historical versus contemporary data, and systematic errors. The strongest models, especially GPT-5 and gpt-oss-120B, achieve macro- F1 scores comparable to human agreement, although their systematic errors can bias downstream results. We therefore combine soft-label model outputs with Design-based Supervised

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

Small Changes, Big Impact: Demographic Bias in LLM-Based Hiring Through Subtle Sociocultural Markers in Anonymised Resumes

arXiv:2603.05189v3 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly deployed in resume screening pipelines. Although explicit PII (e.g., names) is commonly redacted, resumes typically retain subtle sociocultural markers (languages, co-curricular activities, volunteering, hobbies) that can act as demographic proxies. We introduce a generalisable stress-test framework for hiring fairness instantiated in the Singapore context: 100 neutral job-aligned resumes are augmented into 4100 variants spanning four ethnicities and two genders, differing only in job-irrelevant markers. We evaluate 18 LLMs in two settings: (i) Direct Comparison (1v1) and (ii) Score & Shortlist (Top-Score Rates), each with and without rationale prompting. We find that even without explicit identifiers, models recover demographic attributes with high F1 and exhibit systematic disparities, with models favouring markers associated with Chinese and Caucasian males. Ablations show language mark

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

Exploring the Role of Automated Feedback in Programming Education: A Systematic Literature Review

arXiv:2602.00089v2 Announce Type: replace Abstract: Automated feedback systems have become increasingly integral to programming education, where learners engage in iterative cycles of code construction, testing, and refinement. Despite its wider integration in practices and technical advancements into AI, research in this area remains fragmented, lacking synthesis across technological and instructional dimensions. This systematic literature review synthesizes 61 empirical studies published by September 2024, offering a conceptually grounded analysis of automated feedback systems across five dimensions: system architecture, pedagogical function, interaction mechanism, contextual deployment, and evaluation approach. Findings reveal that most systems are fully automated, embedded within online platforms, and primarily focused on error detection and code correctness. While recent developments incorporate adaptive features and large language models to enable more personalized and interactiv

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

Testing Fairness with Utility Tradeoffs: A Wasserstein Projection Approach

arXiv:2505.11678v4 Announce Type: replace Abstract: Ensuring fairness in data driven decision making has become a central concern across domains such as marketing, lending, and healthcare, but fairness constraints often come at the cost of utility. We propose a statistical hypothesis testing framework that jointly evaluates approximate fairness and utility, relaxing strict fairness requirements while ensuring that overall utility remains above a specified threshold. Our framework builds on the strong demographic parity (SDP) criterion and incorporates a utility measure motivated by the potential outcomes framework. The test statistic is constructed via Wasserstein projections, enabling auditors to assess whether observed fairness-utility tradeoffs are intrinsic to the algorithm or attributable to randomness in the data. We show that the test is computationally tractable, interpretable, broadly applicable across machine learning models, and extendable to more general settings. We apply

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit

arXiv:2608.27309v1 Announce Type: cross Abstract: Audits of LLM judges certify a bias by contrasting matched conditions, and the strongest designs difference twice: a within-item contrast between two candidate responses, differenced again across a manipulated attribute, read off a bounded rating scale. We show that this endpoint is not identified on the scale that reports it. Each term of the double difference is censored by its own share, so the observed statistic confounds differential preference with differential attenuation: a severity shift common to both responses manufactures an interaction whenever the two censor it unequally, as unequal distances from the bounds make them, exactly where good stimuli place them. We exhibit the failure inside a pre-registered audit of a frozen pedagogy judge, sealed before the first of its 990 calls. The registered primary endpoint, the effect of a stated learner profile on the judge's scaffolding preference, is null: $+0.085$ points (95\% BCa $

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

Research Design Tracking and Assessment for the Social Sciences

arXiv:2608.27049v1 Announce Type: cross Abstract: Reliable assessment of causal research designs in the social sciences is critical for evidence-based policy-making, yet has so far relied entirely on manual expert analysis. We introduce Automated Research Design Tracking and Assessment (ARDTrA), a task that involves detecting the research design used in a paper and assessing the quality of its application. We create an expert-annotated dataset of papers covering six families of counterfactual research designs and evaluate the task using a multi-turn RAG-based conversational pipeline. Across four retrieval strategies, four LLMs and six embedding models, we find that passage length is the main driver of performance, explaining 52-66% of the variance. A per-research-design analysis also shows that human and machine difficulty do not align: the designs that prove hardest for the system are not those on which expert annotators disagree most, pointing to two independent sources of task diffi

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

LiveSim: Simulating Environment-Shaped Users in Multi-Agent Live-Stream Ecosystems

arXiv:2608.26849v1 Announce Type: cross Abstract: User behavior simulation with large language models~(LLMs) is increasingly used to support multi-agent ecosystem simulation. Existing simulators typically rely on static user profiles inferred from historical observations, which become inadequate in socially intensive environments such as live streaming where interaction dynamics continuously reshape user behavior. We propose \textbf{LiveSim}, an LLM-based framework for live-stream ecosystem simulation. It represents users as editable behavioral hypotheses and progressively refines them through trajectory-grounded interactions, where discrepancies between simulated and observed trajectories reveal missing environmental shaping effects. These signals are further extracted as transferable environment-behavior patterns and accumulated in a collective behavioral memory to improve user-level behavioral fidelity and support ecosystem-level simulation. Experiments on real-world live-stream ris

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

DIRECT: Decomposing Audience Preference and Creative Effect in Visual Content Analytics

arXiv:2608.26584v1 Announce Type: cross Abstract: Which visual choices make a post perform better? A growing literature answers this question with pooled coefficients estimated across many creators, which platforms translate into creative recommendations. We show that these coefficients blend two distinct patterns that can point in opposite directions for the same attribute. The first, audience preference, arises because creators who favor a style attract differently composed audiences, so their posts perform differently because of who is watching, not what any single post does. The second, creative effect, captures how a creator's audience responds when she departs from her usual look. Pooled estimation averages the two, and audience preference can be large enough to reverse the signal that creative direction requires. We propose DIRECT (Decomposed Identification of Response Effects via Causal Tools), a panel-based causal-inference framework that separates them, combining the Mundlak

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

Decolonial Discourse in Postcolonial Contexts: How YouTubers Negotiate Audience Tensions, Platform Governance, and State Influence

arXiv:2608.26351v1 Announce Type: cross Abstract: Decolonial discourse on online platforms is often framed in terms of creator motivations and expressive possibilities. In this paper, we examine what it takes to sustain such discourse under layered sociotechnical constraints. Drawing on semi-structured interviews with YouTubers engaging in Bengali decolonial discourse, we analyze how audience publics, platform governance, and state influences shape what becomes sayable, visible, and viable. We show how fragmented postcolonial identities among audiences produce legitimacy policing, harassment, and coordinated backlash, requiring ongoing relational labor from creators. At the platform level, differential monetization, opaque moderation, and copyright regimes reorganize which publics are economically viable and reinforce existing hierarchies. Further, intermediaries such as multi-channel networks mediate regulatory pressure, introducing political risks and constraints on participation. In

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

Assessing Company Contributions to Societal Resilience: Extending the Societal Capacity Assessment Framework to Agentic AI

arXiv:2608.27238v1 Announce Type: new Abstract: Companies that deploy AI agents and make them available to others are creating the sociotechnical circumstances under which this technology integrates into existing social and economic structures. AI-deploying companies are institutional actors that actively shape society's capacity to withstand and govern the consequences of agentic AI. In view of these societal impacts, companies can build societal resilience by designing and promoting safer implementations of AI agents. To operationalize this goal, this paper adapts the indicator-based Societal Capacity Assessment Framework (SCAF) to measure how a company's deployment decisions contribute to societal resilience, inverting its original measurement of societal resilience as a backdrop for deployment decisions (Gandhi et al., 2025). Our procedure has two steps: a conceptual step in which we design a suite of indicators that define what SCAF's vulnerability, coping, and adaptive capacities

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

Animarium: an open, reproducible pipeline for synthetic populations of Italian cities, from ISTAT sources to open data (Tech Report v1)

arXiv:2608.27111v1 Announce Type: new Abstract: Synthetic populations of eleven Italian municipalities (1,814,317 individuals in 887,937 households) generated from published aggregates alone: ISTAT census and register tables, census-section counts, the national civic-address register, public-use survey microdata, and six municipal open-data portals, every source certified in a registry with licence, fingerprint and declared affordances. Four rings give every attribute a declared place: a maximum-entropy joint model of up to nine demographic attributes; whole-vector donation of twenty-three attitudinal and health variables from survey respondents; placement to census section, single year of age and address; and households constrained by the census size distribution per section. Every downstream layer (detailed titles, work, names, biographies) is a declared derivation adding no information. The pipeline is deterministic to the byte: regenerating all eleven municipalities from the tagged

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

Reclaiming Epistemic Agency: A Critical Framework for Human-Generative AI Co-Agency in Education

arXiv:2608.26937v1 Announce Type: new Abstract: Generative artificial intelligence (GenAI) has been primarily framed as an impartial educational tool. However, this framing overlooks an even larger shift: the reassignment of epistemological authority from teachers to students to machines. This paper presents a conceptual evaluation of the extent to which GenAI redistributes students' and teachers' ability to act in classrooms to produce knowledge, validate each other's claims, and create evidence of student learning while collaborating with and competing against humans. This evaluation draws on various theoretical paradigms, including Distributed Agency, Self-Determination Theory, Society 5.0, and Technology Integration Paradigms, including TPACK and SAMR. While all of the theoretical paradigms evaluated are relevant to the role of agency within education mediated by AI, none of them address the ongoing disparity regarding equitable distribution of power, ownership of the data used to

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

How Does Science Education Research Respond to Sociopolitical Change? A BERTopic Analysis of Korean Research

arXiv:2608.26675v1 Announce Type: new Abstract: Research fields do not evolve in isolation: their questions and priorities shift with policy, curriculum reform, and broader social change. Analyzing published literature can reveal not only how a field matures but also how it responds to these conditions. Prior work in science education has focused on identifying research topics and their trends, but paid less attention to the external conditions in which research is produced. We examine Korean science education research from 2008 to 2025, a case in which centralized curriculum revision, government education initiatives, and demographic decline are prominent. Using BERTopic, an embedding-based topic modeling technique, we identify major topics and temporal trends, and analyze their associations with selected sociopolitical factors. We interpret each topic and distinguish three groups: sociopolitical, subject-specific, and student-related topics. Within the first group, science teacher pr

Source ↗
technology Fri, 28 Aug 2026 00:00:00 -0400
arXiv cs.CY

ClassVision: AI-Powered Classroom Attendance System

arXiv:2608.26173v1 Announce Type: new Abstract: Students and working professionals have to go through the attendance process every day. Traditional methods of marking attendance using pen and paper or online platforms are human-intensive and time-consuming. To address the challenges in manual attendance processes, this research explores the use of face detection (FD) and face recognition (FR) technology to automate the attendance process, particularly in educational settings, and build a ClassVision course attendance system. We also propose an automated attendance system featuring a human-computer interaction (HCI) and user-friendly web interface that utilizes real-time image processing to identify and recognize students in classrooms and automatically record their attendance. We identified RetinaFace as the best face detection model, and when combined with Face Recognition for verification, it provided the most promising results with a cropped embedding of 50x50 pixels.

Source ↗
Showing 1301–1350 of 1593 signals
← Prev Page 27 of 32 Next →