EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18624 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.HC

Trade-offs in Data Color Palette Design Tools

arXiv:2608.19148v1 Announce Type: new Abstract: Designing a color palette for data requires designers to balance multiple constraints, including accessibility and aesthetics. Color palette tools support this process through features including direct manipulation, automated palette generation and evaluation, previews, and so on. Despite their prominence, relatively little is known about how these different mechanisms shape design across contexts. We conducted an exploratory think-aloud crowd work study with 40 self-identified designers. Each participant used one of four palette tools selected to span different interaction modalities to complete a series of accessibility- and aesthetics-oriented design tasks. We observed two preliminary patterns. First, tool differences were more pronounced in accessibility-constrained tasks. Second, even when accessibility was not explicitly required, some tools produced more accessibility-friendly palettes and prompted more accessibility-oriented think

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.HC

When Readability and Source Retention Diverge: An Evaluability Gap in AI Translation

arXiv:2608.19083v1 Announce Type: new Abstract: Readable AI output can leave an evaluability gap: even when the source is shown, an overall-quality judgment may not reflect what an output preserves. We investigated how source-text condition and output rendering relate to perceived translation quality, and how output and system appraisals relate to trust and stated disclosure willingness in a plain-text interface. A focal 2 * 2 comparison (N=306) using TransLingo examined simple generated narratives and complex literary-philosophical prose alongside LLM-generated readability-oriented outputs and researcher-revised fidelity-oriented outputs. A descriptive stimulus audit indicated greater source retention in fidelity-oriented outputs in both source-text conditions. Factorial analyses showed a significant rendering-by-source-text-condition interaction in perceived quality. Participants rated fidelity-oriented outputs higher than readability-oriented outputs for the simple narratives, where

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.HC

TractorBeam: Personalized AI Sensemaking Support via Collaborative Machine Annotation

arXiv:2608.18994v1 Announce Type: new Abstract: Language model-based systems which allow asking questions of documents have become popular tools for sensemaking. Despite their implied capability, these systems still suffer from issues of factuality and provenance, while encouraging confirmatory, rather than exploratory, research. We present TractorBeam, a browser extension-based mixed-initiative system that uses collaborative annotation as an interface metaphor for sensemaking, re-framing language model (LM) outputs as suggested highlights in a process that we call \textit{collaborative machine annotation}. This metaphor allows us to present LM results in-context on PDF documents, directly addressing concerns of provenance and factuality, while allowing users to iteratively construct mental schemas and queries for language models directly in the context of a document. In a preliminary user study, all of our participants felt that TractorBeam enabled them evaluate and iteratively improv

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.HC

A revised framework for the assessment of psychological safety in autonomous vehicles

arXiv:2608.18801v1 Announce Type: new Abstract: Despite recent technological progress in the development of autonomous vehicles (AVs), their societal acceptability remains a subject of debate as recent research findings point to psychological roadblocks. Concerns arise not only for physical safety but also for potential psychological risks resulting from human interaction with AVs. Psychological concepts such as trust, and perceived safety are well-studied in this context and are found to be determinant factors for the intention to use AVs. Unfortunately, there has been no formalization of the mechanism by which human interaction with AVs may lead to psychological hazards, threatening trust, perceived safety, and acceptability. Furthermore, there has been little prior research that conceptualizes the severity of psychological risk in AVs, and there are no clear guidelines for a systems designer on how to assess and address psychological risk in the AV development context. To address th

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.HC

Scoring and Gamification to Encourage Sustainable Use of Compute Clusters

arXiv:2608.18786v1 Announce Type: new Abstract: The environmental cost of computing continues to grow, yet behaviour change remains limited. We present a composite sustainability score integrating average carbon intensity, resource utilisation, and embodied emissions into a single 0-100 metric designed for gamified feedback. Each component rewards a different dimension of sustainable behaviour: carbon-aware workload shifting, high resource utilisation, and selecting hardware that is commonly underutilised. This scoring system is built into an existing cluster management interface and underpins three dashboard conditions: raw metrics, composite score, and a gamified tree visualisation, which we are planning to evaluate in a 12-week within-subjects study with approximately 35 researchers. Furthermore, we open the discussion on the challenge of defining computational work `goodness' in the context of sustainability scores.

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.HC

Engineering Psychological Safety in Autonomous Vehicles: A Systems-Theoretic Framework for Psychological Safety in Autonomous Vehicles and its Validation in Real-World Scenarios

arXiv:2608.18778v1 Announce Type: new Abstract: Despite rapid technological advances, the societal acceptability of autonomous vehicles (AVs) remains limited by psychological barriers that extend beyond traditional concerns of physical safety. While factors such as trust and perceived safety are known to influence user acceptance, there is a lack of formalized mechanisms and engineering methods to systematically identify, assess, and mitigate psychological risks arising from human-AV interactions. To address this gap, this work proposes and validates a systems-theoretic framework for the assessment of psychological safety in autonomous vehicles. First, a comprehensive psychological safety risk model is defined, extending the Systems-Theoretic Accident Model and Processes (STAMP) to incorporate key psychological constructs such as trust, perceived control, predictability, and perceived support. Based on this model, a hazard analysis method (AV-PsySafe) is developed to systematically ide

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.HC

Model Literacy: An Extra Summative Evaluation Factor for Visual Analytics

arXiv:2608.18721v1 Announce Type: new Abstract: Understanding and enhancing visual analytics (VA) performance is important for maximizing their impact. Existing studies have successfully applied well-established summative evaluation methods from information visualization to the VA context, yet the recent emphasis on an extra data analysis/modeling stage in the VA pipeline poses an additional challenge. Inspired by the modern concept of visualization literacy, this paper examines model literacy, namely users' knowledge of the analysis model used in a VA technique, as an additional factor for VA performance. Results from a controlled study on the visual analysis of multidimensional data with two dimensionality-reduction models indicate a positive correlation between model-task accuracy and VA-task accuracy. The study involves two common dimensionality-reduction models, PCA and t-SNE. The correlation is stronger for PCA than for t-SNE in the current task design, a pattern consistent with

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.HC

Sounds Uncertain: Exploring the Affective Aspects of Sonification for Uncertainty Visualization

arXiv:2608.18680v1 Announce Type: new Abstract: Affective visualization can influence how users perceive, interpret, and engage with data by embedding and conveying emotion through visual design. While sound is widely used in media to evoke emotions, little is known about how sonification can support affective visualization. In this work, we investigate how sonification can communicate emotion in uncertainty visualizations through a co-design study. Participants created two sonifications to accompany a visualization: one conveying the affective component of uncertainty and one conveying neutrality. Our findings show that uncertainty was commonly associated with wavy auditory qualities related to an ominous sentiment. On the other hand, neutrality was associated with clear and relaxing auditory qualities. These results provide insights for the design of visualizations that integrate sonification to communicate the affective component of uncertainty.

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.HC

Report on The 1st Workshop on Human-Centered Proactive and Personalized Agents for Interactive Information Access at CHIIR 2026

arXiv:2608.18638v1 Announce Type: new Abstract: Interactive information access is increasingly moving beyond reactive query-response paradigms toward agentic systems that can personalize interaction, retain context, infer latent needs, recommend next steps, and initiate support. This shift creates new opportunities for adaptive and context-aware assistance, while also raising important questions about autonomy, privacy, trust, transparency, user welfare, and evaluation. The First Workshop on Human-Centered Proactive and Personalized Agents for Interactive Information Access provided an interdisciplinary forum for examining these questions across information retrieval, human-computer interaction, dialogue systems, AI ethics, cognitive science, learning technologies, and human-centered AI. Through invited talks, paper presentations, and open discussion, the workshop engaged with topics including calibrated initiative, knowledge-gap navigation, long-term memory, value-sensitive design, im

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.HC

SemanticSlider3D: Training-Free Continuous Semantic Editing for 3D Objects

arXiv:2608.18560v1 Announce Type: new Abstract: Fine-grained control over continuous semantic attributes of 3D objects is essential for 3D content creation, but is not well supported by conventional 3D modeling workflows or prompt-based interaction with existing generative AI tools. While slider-based methods have proven effective for fine-grained semantic control in 2D image generation, no equivalent approach exists for 3D. Extending these 2D methods to 3D is non-trivial due to challenges unique to 3D, including geometric integrity and cross-view coherence. We present SemanticSlider3D, a technique for continuous semantic attribute editing of 3D objects that requires no per-attribute training. Given a user-specified attribute, our pipeline constructs a semantic editing direction in the latent space of a state-of-the-art 3D generation model, presenting a diverse and coherent spectrum of 3D variations. A technical validation on a dataset of 50 3D object-attribute pairs shows our method w

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.HC

Designing Social Robots for Social-Cognition Training with Autistic Adults

arXiv:2608.18488v1 Announce Type: new Abstract: Social robots have been widely explored as tools for autism intervention, yet this literature has focused predominantly on children and has rarely involved autistic adults as active contributors to design. This creates a mismatch between existing systems and the social-cognitive challenges autistic adults actually face in everyday life, including navigating ambiguous interpersonal contexts, managing conversational timing, and interpreting implied emotional meaning. To address this gap, we conducted an online focus group and co-design session with five autistic adults to explore what a social robot for social-cognition training should do, how it should interact, and under what conditions it would be genuinely useful. The 90-minute session combined open discussion with structured co-design activities on a shared digital whiteboard, and the resulting verbal and visual data were analysed using reflexive thematic analysis. The analysis yielded

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.HC

Reducing Technician Search Burden: A Multimodal RAG for Cessna 172 Maintenance Manual

arXiv:2608.18465v1 Announce Type: new Abstract: Proper use of the aircraft maintenance manual is essential for correct maintenance, providing procedures, diagrams, cautions, and specifications. However, technicians often avoid consulting it because it is difficult to navigate and time-consuming under strict schedules. Retrieval augmented generation (RAG) models have recently been introduced in aircraft maintenance, yet existing models focus solely on textual retrieval. This research therefore targeted the Cessna 172 Maintenance Manual (C172-MM), widely used in general aviation, and developed a multimodal manual retriever (MMR) capable of retrieving multimodal manual pages. Retrieval performance was evaluated using synthetic queries covering procedures, diagrams, caution/safety information, and specifications; the MMR achieved 93.37% recall@5. Beyond retrieval, a multimodal RAG (MRAG) pipeline was examined, in which retrieved pages were input to a vision-language model that generated re

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.HC

Multimodal Rapport Estimation in Real-World HRI

arXiv:2608.18401v1 Announce Type: new Abstract: Evaluating interaction quality in real-world HRI is an important challenge. If interaction quality can be estimated reliably, the results can be used to improve dialogue strategies and ultimately enable robots to adapt their behavior autonomously. However, existing automatic evaluation methods have been developed primarily in controlled laboratory settings, and it remains unclear whether they can be directly applied to real-world environments, where users are free to disengage and multi-party participation may arise naturally. In this study, we investigate the automatic estimation of third-party-rated rapport scores using 62 sessions of multimodal recordings collected in a Japanese drugstore. We compare zero-shot LLMs, pretrained text, audio, and visual models, and their prediction-level fusion. The results show that, in real-world HRI, zero-shot LLMs achieve strong performance, while audio and visual models tend to provide complementary

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.HC

LEDGER: Claim-to-Evidence Trace Graphs for Auditing LLM Agents

arXiv:2608.18398v1 Announce Type: new Abstract: Large language model (LLM) agents can now carry out long-horizon technical workflows involving complex tool use, code execution, file edits, and generated artifacts. As agents do more work faster, the productivity bottleneck shifts from producing outputs to auditing whether those outputs are correct and trustworthy. Agent observability systems make fine-grained execution events visible, but visibility alone still leaves reviewers to reconstruct which actions, artifacts, and validation steps matter for a particular conclusion. We introduce LEDGER - Layered Evidence and Decision Graphs for Execution Review, a tracing and review system that builds layered trace graphs over observed agent sessions. LEDGER preserves Trace Records while grouping them into Evidence Nodes and Workflow Nodes, representing artifacts as evidence anchors, and adding typed semantic edges that connect claims to supporting actions, artifacts, and checks. Through data-an

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

AI Fact-Checking in the Wild: A Field Evaluation of LLM-Written Community Notes on X

arXiv:2604.02592v3 Announce Type: replace Abstract: Large language models (LLMs) show promising capabilities for fact-checking, yet prior work evaluates them only in controlled offline settings using benchmarks or crowdworker judgments. Success in real-world fact-checking depends also on how content is judged within a live platform environment. We present the first field evaluation of LLM fact-checking deployed on a live social media platform, testing performance directly through X Community Notes' "AI writer" feature over a three-month period. Our LLM writer, a multi-step pipeline that handles multimodal content, conducts web and platform-native search, and writes contextual notes, was deployed to write 1,614 notes on 1,597 tweets and compared against 1,332 human-written notes on the same tweets using 108,169 ratings from 42,521 raters. Direct comparison of note-level platform outcomes is complicated by differences in submission timing and exposure between LLM and human notes; we ther

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Brokerage in the Black Box: Swing States, Strategic Ambiguity, and the Global Politics of AI Governance

arXiv:2601.06412v4 Announce Type: replace Abstract: The United States-China rivalry has placed frontier dual-use technologies, particularly Artificial Intelligence (AI), at the center of global power dynamics, as techno-nationalism, supply chain securitization, and competing standards deepen bifurcation within a weaponized interdependence that blurs civilian-military boundaries. Existing research, yet, mostly emphasizes superpower strategies and often overlooks the role of middle powers as crucial actors shaping the global techno-order. This study examines Technological Swing States (TSS), middle powers with both technological capacity and strategic flexibility, and their ability to navigate the frontier technologies' uncertainty and opacity to mediate great-power techno-competition regionally and globally. It reconceptualizes AI opacity not merely as a technical deficit, but as a structural feature and strategic resource, stemming from algorithmic complexity, political incentives that

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

ChildSafeAds Shared Task 2026: Commercial Content in Child-Facing YouTube Videos

arXiv:2608.19165v1 Announce Type: cross Abstract: ChildSafeAds is a shared task on commercial content in YouTube videos likely to reach children and teenagers. It contains 3,360 videos from 939 channels. Each instance begins with a segment submitted to SponsorBlock, an open-source crowdsourced browser extension whose users mark sponsor segments so that others can skip them. We pair the segment with its available transcript, video and channel information, and a sales or service page linked from the video description. Systems determine what kind of offer is being promoted (ST1), assign product categories (ST2), and identify legal risk flags (ST3). The evidence is divided into four cumulative access levels, from the transcript to the linked page, so results can be compared against the cost of collecting the data. 45.5\% of videos in our data failed to properly use the in-platform ad disclosure method (the ``Includes paid promotion'' label). GPT-5.4 produced the labels after the expert org

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems

arXiv:2608.19140v1 Announce Type: cross Abstract: Frontier language models are compared, marketed, and benchmarked on capability -- what their best or average output can achieve. I argue this measures the wrong axis. The models have saturated accuracy: their mean output lands on the target. What now separates one system from another in practice is precision: how tightly concentrated their outputs are around that target across repeated, identical requests. Borrowing the marksman's distinction, capability is where the average shot lands; reliability is the size of the group. I make three claims. First, precision, not capability, is the frontier differentiator between systems, and benchmark culture systematically fails to measure it, reporting central tendency rather than spread. Second, precision is measurable, cheaply and without circularity, by running a fixed suite of deterministically scored tasks many times at fixed temperature and computing the per-task consistency of outcomes -- n

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Leaf Values as Coordinates: Exact Contrastive Explanation for Gradient-Boosted Ensembles

arXiv:2608.19127v1 Announce Type: cross Abstract: A gradient-boosted ensemble predicts by summing one leaf value per tree. Read those values as coordinates rather than as intermediate results, and every instance becomes a point in R^M on which the model acts linearly: the score is the sum of the coordinates. This small change of view makes contrastive explanation exact. The difference between two instances is a vector that is identically zero wherever they share a leaf, so the gap between a rejected applicant and an accepted one is carried by a handful of coordinates, each traceable to a real split in a real tree. Nothing is fitted, sampled, or assumed additive in features -- the additivity is already there, in the right space. We build a recourse method on this representation and evaluate it on five tabular datasets under repeated cross-validation. Its recommendation reconstructs the model's own decision to 6.2 x 10^-15, so an auditor can re-check the arithmetic without the model. On

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Sanyu Studio: A Multi-Agent System for Art-Historical Narrative Construction

arXiv:2608.18677v1 Announce Type: cross Abstract: Amid concerns that generative AI may standardize art interpretation, this paper examines whether LLM-based interaction can support plural art-historical narrative construction. We present Sanyu Studio, a multi-agent dialogue system that models 321 Sanyu oil paintings as agents with fact, interpretation, organization, and memory-filtering mechanisms. Based on a seven-day workshop with eight art-university participants, the study shows that user prompts, evidence organization, and cognitive tendencies shaped divergent yet coherent versions of digital Sanyu. The findings suggest that, under conditions of limited historical evidence, AI can amplify human agency and offer public audiences an interactive entry point into art-historical interpretation.

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Measuring Proof Burden in Public Bounty Listings: A RentAHuman Case Study

arXiv:2608.18547v1 Announce Type: cross Abstract: Online bounty markets let requesters advertise paid tasks. Workers may be asked not just to complete a task but to prove it, and proof can mean exposure: revealing identity or location, using a personal account, posting publicly, acting in the physical world, or repeated evidence at later checks, none disclosed by the posted price. We call these advertised requirements proof burden and measure them on RentAHuman, a 2026 market publicized as a place for AI agents to hire humans. We study what listings request, not what workers submit or experience. We manually audited a nonrandom May 31, 2026 snapshot: every listing our searches returned from RentAHuman and Human Pages, another such market (981 listings, all but one from RentAHuman). Two independent coders recorded 13 features (11 kinds of evidence, recurring monitoring, physical-world action) and our 0-5 Proof Burden Score; a blinded third resolved all disagreements. A planned content s

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Longitudinal Relational Publics and their Discursive Overlap with Issue Publics

arXiv:2608.18422v1 Announce Type: cross Abstract: Online discussions of political issues do not always happen in places explicitly dedicated to political talk; they also arise in online spaces focused on at least nominally apolitical interests, identities, and/or places. Whatever one's normative view of politics entering these ``online third spaces,'' understanding who brings political issues into them, and when, requires studying these spaces at scale. In turn, studying these spaces at scale requires a construct that captures both who is speaking and who is listening, and that holds up over time. Building on Bruns' distinction between participant-centered personal publics and post-centered issue publics, we introduce the longitudinal relational networked public (or, simply, the longitudinal public): the coupling of discourse produced by a socially connected set of creators with the durable attention their shared audience gives it. The longitudinal public departs from related relationa

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

One Gate Is Not Enough: Composing Stateful Pre-Action Controls for Agentic AI

arXiv:2608.18360v1 Announce Type: cross Abstract: Agentic AI systems take consequential actions governed by more than one pre-action control at once: authority, resource, and evidence gates that can admit, degrade, or remediate an action before it executes. This paper's central object is remediation-induced control coupling: a remediation applied by one control can change the action, evidence, or context another control evaluates, invalidating that control's earlier judgment. We formalize this coupling and give a remediate-and-regate protocol that restores per-action soundness in the current bounded, idempotent setting under its stated assumptions. We further show that the two implemented remediation operators (evidence substitution and resource-budget downroute) do not commute -- a finite-model checker finds concrete counterexample instances -- making remediation order part of the control-plane semantics rather than an implementation detail. A governed evidence buffer that trusts its

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence

arXiv:2608.18352v1 Announce Type: cross Abstract: The integration of generative AI into web search delivers synthesized answers to user queries, changing how people navigate and assess information, while raising concerns about the downstream impacts on publishers who supply the underlying content. We conduct a preregistered field experiment (N=1,100) on Google Search, the dominant online search platform, to estimate the causal effects of AI Overviews and AI Mode on user behavior, perceptions, and publisher traffic. We show that removing AI Overviews and AI Mode increases click-through rates to publishers, while an AI Mode-only experience reduces click-through rates and erodes user experience and trust in information found on Google. These findings show that integrating generative AI into web search reshapes online attention, with economic consequences for the online publishers that sustain both search platforms and the overall information ecosystem.

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme

arXiv:2608.18336v1 Announce Type: cross Abstract: When evaluating language models on human exams, benchmarks typically score each response as right or wrong and report the overall accuracy. This approach assumes that partial knowledge is worth proportional credit, an assumption that fails when an examination uses a non-additive grading scheme. The 2025 reform of Vietnam's National High School Graduation Examination demonstrates the cost of this substitution. In Part II of the exam, candidates evaluate four true/false statements per question. The grading is convex: the number of correct statements earns 0, 0.10, 0.25, 0.50, or 1.00 points. Identifying three statements correctly pays 0.50 points, not the 0.75 points that standard accuracy metrics would award. Because Part II accounts for 4.00 of the exam's 10.00 points, reporting accuracy inflates the score by rewarding partial knowledge that the state explicitly penalizes. We introduce THPT-Ladder, a benchmark of 632 items from 21 offic

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Artifact-centered Claim-aware Observability for Autonomous Scientific Agents

arXiv:2608.18312v1 Announce Type: cross Abstract: Autonomous scientific agents now increasingly propose ideas, write code, run experiments, analyze results, and even draft papers. Observe and audit those agents are necessary but logging every model call is not enough, scientists also need to inspect the artifacts and claims that the systems produced and their relations. This is driven by the fact that failures in scientific agent systems are often distributed across several objects. A manuscript claim may cite the wrong evidence, a search process may select a degenerate candidate, a laboratory novelty claim may depend on an unstated rule, or a multi-agent plan may change without a visible trigger. Existing tracing, experiment tracking, and archival provenance tools are valuable, but their native objects do not make these scientific audit relations first-class. We argue that autonomous scientific systems should emit portable, claim-aware artifact lineage as a minimum audit layer. We pro

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Global Crises and National Policies: A Large Scale Analysis of Political Content in German Language Online Media

arXiv:2608.18268v1 Announce Type: cross Abstract: Today most media content is consumed based on algorithmic recommendations. Evidence suggests that this can lead to politically biased media consumption patterns. Automated extraction of political agendas from texts can reveal and analyze political biases in online media -- and thus help fostering politically unbiased media consumption. Here we employ modern political text analysis methods demonstrating the potential of automated fine-grained political bias analysis in online media. We conduct an analysis of political content in German language online media during the period 2019--2022, encompassing several million articles and tweets covering events with profound societal impact globally and nationally, the COVID-19 pandemic and the beginning of the war in Ukraine. Our analysis identifies thematic similarity between national (German and Swiss) reporting, particularly for categories driven by international events. We also find divergence

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

What Can Artificial Intelligence Learn from Medicine? Generative Analogies and Reliable Machine Learning Systems

arXiv:2608.18186v1 Announce Type: cross Abstract: In the past few years, machine learning (ML) has been widely (and to an extent, successfully) implemented in medicine. However, uncertainties surrounding ML have made it difficult to establish the bases of its epistemic and methodological warrants. In the literature, a parallel has been drawn between medicine and ML, suggesting that we should model epistemic and methodological standards for ML on the standards of clinical translation. By developing tools from Hesse work, we characterise the nature of this parallel as a generative analogy between the process of clinical translation and the process of building ML systems. We identify more precisely the epistemic and methodological warrants of clinical translation that are typically only mentioned when appealing to the analogy, and we show in which sense such warrants apply analogically to the context of ML. In particular, we interpret warrants of clinical translation in reliabilist terms,

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs

arXiv:2608.18131v1 Announce Type: cross Abstract: Current safety alignment training for Large Language Models (LLMs) are heavily English-centric. When such safety filters fail for non-English languages, the consequences are immediate and user-facing: voice assistants and spoken dialogue systems may produce stereotype-reinforcing outputs, bypassing the standard English-focused safety alignments and propagating harmful bias to non-English speaking communities. For spoken language technologies deployed across India's linguistically diverse population, this represents a critical failure mode. To address this cross-lingual gap, we introduce INCLUDE (Indian Cultural Lens for Understanding and Detecting Embedded Biases), a multilingual evaluation benchmark designed to quantify Indian-centric socio-cultural biases. INCLUDE consists of 2,604 prompts spanning six prompt languages: English, Hindi, Bengali, Marathi, Tamil, and Hinglish (Hindi-English code-mix). We evaluate ten open- and closed-sou

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS)

arXiv:2608.18100v1 Announce Type: cross Abstract: AI systems now shape how hundreds of millions of people learn about cultures other than their own. When someone asks one of these systems about the Middle East, they do not receive neutral facts. They receive a representation shaped by the frameworks embedded in training data, and that data is overwhelmingly Western and English-language. This paper asks whether that representation is Orientalist in Said's sense: whether it denies agency to Middle Eastern actors, treats Western frameworks as neutral while marking non-Western knowledge as particular, and explains the region through categories it did not produce. Standard fairness metrics cannot answer this, because they detect explicit prejudice rather than structural framing. This paper introduces the Middle East Cultural Sensitivity Score (MECSS), a framework that turns Said's seven Orientalist operations into measurable dimensions, and the term "Said-washing" for a specific failure: a

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Qualified Cross-References as a Verification Method: The Normative Environment of the EU AI Act

arXiv:2608.19194v1 Announce Type: new Abstract: Legal cross-references are commonly represented as links between instruments or provisions. For a curated legal knowledge base, the existence of a link is only the beginning of the claim: it must also state the legal character of the interaction, identify the provisions supporting it, preserve its conditions, and remain consistent when reached from either instrument. This paper presents a provision-level model and a construction protocol for qualified cross-references, developed through a bilingual corpus of fourteen instruments surrounding Regulation (EU) 2024/1689 (the AI Act). The model distinguishes direct textual reference, bounded presumption of conformity, substantive interaction without textual reference, mediated intersection, and institutional analogy, and treats applicative interaction and definitional overlap as independent dimensions. The methodological contribution is bidirectional inversion: a relationship documented from a

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

LearnAI: Just-in-Time AI Co-Creation Across Disciplines at a University

arXiv:2608.19164v1 Announce Type: new Abstract: As generative AI reshapes professional and educational practice, institutions face a challenge: how to support diverse learners, from non-coders to advanced students, in building confidence and practice with AI-supported problem solving. Most institutional responses bifurcate into conceptual workshops for general audiences or technical courses for computer science majors, leaving few spaces where mixed-ability learners can engage common AI tasks at levels matched to their prior experience. This experience report presents the LearnAI Framework, a two-layer model for just-in-time AI co-creation piloted at a comprehensive teaching university. The Wide-Exposure Layer embeds short presentations in existing courses to build AI awareness at scale, reaching students and faculty across 18 courses in five disciplines. The Customized Co-Creation Layer provides opt-in, one-on-one sessions where clients work with trained undergraduate tutors through a

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Hot Games: Towards a Holistic Assessment of the Planet Warming Emissions of Video Games based on 2024-2025 Data

arXiv:2608.19040v1 Announce Type: new Abstract: Following on recent reports on specific platforms or companies, this paper provides an assessment of the global impact of the production and use of video games. It draws together publicly available data on game development, hardware, games sold, download sizes, time spent playing games on different platforms, and subscriptions to multiplayer and cloud game services. It provides an update to figures published 2020 and 2022. Crucially, our account of emissions related to video games considers a wide range of categories, yet contains enough detail to be critiqued and improved in the future.

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Open at the Edge, Captured at the Center: llama.cpp and the Political Economy of Local AI Inference

arXiv:2608.19001v1 Announce Type: new Abstract: Open AI scholarship has focused on model releases and cloud ecosystems, leaving the local inference infrastructure that makes open-weight models runnable on user-owned devices largely unexamined. We address this gap through a mixed-methods analysis of llama.cpp, combining 7,681 merged pull requests from March 2023 through March 2026 with repository discussions, corporate statements, and contributor blogs. We show that local inference broadens participation at execution while relocating capture into the infrastructure that makes execution possible. Through hardware backends, model integration labor, and Hugging Face's February 2026 absorption of the project, we document how control shifts to hardware vendors, model distributors, and core maintainers while model owners and individual contributors bear the cost of making models runnable. These dynamics suggest that preserving openness outside the cloud requires attention to the infrastructur

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Epistemic Subordination: Generative AI and the Infrastructure of Knowledge

arXiv:2608.18758v1 Announce Type: new Abstract: Generative AI does not merely produce biased outputs. It encodes the majority's way of knowing as the default infrastructure of knowledge itself. We call this epistemic subordination. The training process compresses the full breadth of human expression into a single probabilistic model whose statistical baseline reflects the languages, assumptions, and cultural frameworks of the dominant culture. Minority epistemologies are not excluded but absorbed: present in the training data, yet structurally subordinated in the output. The result is not a collection of discrete biases that can be audited and corrected. It is an epistemic condition embedded in the architecture from which all outputs emerge. This unified harm cuts across three legal domains -- anti-discrimination law, cultural and linguistic rights, and democratic viewpoint pluralism -- and each fails to address it for the same structural reason: existing law regulates downstream, at t

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Turning interest into institutional change: teaching advocacy for sustainable research

arXiv:2608.18601v1 Announce Type: new Abstract: Systemic change across the digital research landscape is required to reduce the environmental impact of digital research, but while many researchers and technical professionals are motivated to act, they often lack the skills required to translate motivation into lasting organisational change. We present an open-access course that teaches the foundations of advocacy and organisational change to researchers, research software engineers, and research technical professionals. Structured around the UNICEF five-step advocacy cycle, the course covers stakeholder analysis, power mapping, coalition building, storytelling, framing and messaging, and evaluation. It is grounded in the UK policy landscape, including the Concordat for Environmental Sustainability and the UKRI Environmental Sustainability Strategy, and uses a fictional case study to make concepts concrete. The course is designed as a dual-layer resource in which workshop slides and ext

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks

arXiv:2608.18554v1 Announce Type: new Abstract: Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human or LLM) agents. The question that drives model selection is therefore not only which model produces the best output, but which model most improves the work of another (weaker) agent. We introduce a unified framework that evaluates the capability of models to automate and augment another agent's performance. Across seven economically grounded real-world tasks, an assistant model writes assistance text for a standardized lower-capacity worker model, which produces the deliverable. In automation mode, the assistant produces the output directly. Outputs are scored through blind pairwise comparisons by an LLM judge panel with task-specific rubrics, replicated across ten runs. Rankings across the two regimes are only modestly correlated, and the automation winner loses augmentation on five of seven tasks. A

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Fabricated Front: Generative AI and the Opacity of Workplace Performance

arXiv:2608.18369v1 Announce Type: new Abstract: Generative AI (GenAI) has become a fixture of workplace life. Current research asks chiefly what this implies for jobs and outputs, measured in productivity, displacement, or bias. What remains underexamined are the interactional reconfigurations that GenAI produces at work. The emerging concept of effort opacity has begun to fill this gap by highlighting the systematic decoupling of observable output from human engagement. When GenAI makes interactional cues less diagnostic, it weakens the reciprocal exchange that sustains collaborative trust. Extending this account of effort opacity, we examine the interactional mechanics that produce opacity in everyday workplace encounters. Drawing on Erving Goffman's dramaturgical framework and 1,250 interview transcripts from Anthropic's AI Interviewer dataset, we identify five opacity mechanisms through which workplace fronts are reorganized: voice (whose stance the words index), provenance (who ca

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Capability-Based Planning for AI Crisis Preparedness

arXiv:2608.18357v1 Announce Type: new Abstract: Capability-based planning drives preparedness in defense and homeland security, but has yet to be applied seriously to AI. Government AI preparations follow a predict-then-act paradigm: rank risks by likelihood and impact, then prepare for the highest expected harm. AI resists prediction: expert timelines disagree by orders of magnitude, and official reviews concede that likelihood-based risk assessment fails for exactly this class of risk. Drawing on principles of decision making under deep uncertainty, we propose a methodological framework in three parts: a scenario library sampled systematically across declared axes; a rating procedure that assesses each government capability against each scenario on coarse, gated criteria; and a prioritization step that maps the resulting matrix onto decision rules a government might adopt. Through a pilot across the four most severe AI-enabled threat classes, we illustrate the kind of insight the ins

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation

arXiv:2608.18296v1 Announce Type: new Abstract: As CGM-based AI tools approach clinical deployment, whether their accuracy is equitable across patient demographics remains insufficiently tested. To enable this evaluation, we constructed FairGlucose, a 300-patient CGM cohort balanced across 12 demographic strata (age x gender x type 1/type 2 diabetes), with 132,480 forecasting samples and 3,945 unique behavioral events (meals, exercise, medication) logged by 81 patients. Benchmarking 33 models across four families on 2-hour glucose forecasting, we find that population-level external validation can conceal substantial subgroup disparities. Aggregate out-of-distribution metrics appear stable (approximately 1.0), yet subgroup-level ratios range from 0.8 to 1.4, with T1D patients showing 6 mg/dL higher prediction error than T2D (p < 0.001). This disparity persists across all 33 models, suggesting a property of the prediction task rather than any single architecture. Further analysis shows t

Source ↗
technology Thu, 20 Aug 2026 00:00:00 -0400
arXiv cs.CY

Global Index on Responsible AI 2026 : Conceptual Framework and Methodology

arXiv:2608.18122v1 Announce Type: new Abstract: This report presents the methodology of the Global Index on Responsible AI (GIRAI), 2nd Edition. This edition refines the 1st Edition by strengthening the distinction between framework existence and implementation, restructuring dimensions from three to five thematic areas, introducing more granular variables for framework quality, and applying a multi-stage review and validation process. An independent statistical pre-audit was conducted to assess the coherence and robustness of the framework. GIRAI assesses responsible AI governance across five dimensions: Inclusion and Diversity, Ethics and Sustainability, Labour and Skills, Trust and Safety, and Use of AI in Public Service. Each dimension has a number of indicators (38 in total), organised into three pillars, namely AI Policy (17 indicators on government frameworks and implementation, assessed through primary data), CSO Engagement (5 indicators, primary data), and Enabling Conditions

Source ↗
behavior Thu, 19 Mar 2026 10:00:00 +0000
eSchool News

Data intelligence in education: Building the right foundation for better decisions

Data has become one of the most important strategic assets in education. Yet across institutions, publishers, and edtech companies, it often remains fragmented, inconsistently governed, and difficult to use with confidence.

Source ↗
behavior Thu, 19 Jun 2025 06:31:14 +0000
HN: tutoring

Free Virtual CS Classes and Tutoring

Hey everyone! I know this forum is 'notorious' for having more experienced and skilled coders but if I figured this might be relevant for some of you: Coding The Future is a program where we match people who are passionate about computer science to teach students interested in learning. If this sounds like an opportunity that you'd like to participate in please fill out this form so we can best match you with a tutee. Tutoring sessions will be 30 minutes weekly virtually. All tutoring is done for free, so if you are interested in becoming a tutor you will get community service. We provide tutors with the resources to be an effective teacher, and regularly check in with our tutors and tutees to make sure the process is going smoothly. Please note that dedicated tutors may be offered leadership roles, and if you are interested in taking on more leadership within the program, for example becoming a local director of programming or curriculum developer, please let us know. You can also mai

Source ↗
behavior Thu, 19 Feb 2026 18:40:10 +0000
eSchool News

Follett Content Accelerates Public Library Strategy

McHenry, Ill., Feb. 19, 2026 – Building on its September 2025 introduction into the public library market, Follett Content today ... Read more

Source ↗
behavior Thu, 19 Feb 2026 10:00:00 +0000
eSchool News

Why schools and public libraries must unite–in summer and all year long

Some of the most effective literacy ecosystems today are those where schools and public libraries work not in parallel, but in partnership with parents and students.

Source ↗
regulation Thu, 18 Jun 2026 21:27:04 +0000
The 74

NYC Kids, Parents on Missing School for the Knicks Parade

Source ↗
behavior Thu, 18 Jun 2026 20:42:16 +0000
Getting Smart

Prepare Your People, Protect Your People: Setting the Stage for Successful Change Management

In an era of rising resistance and restrictive legislation, asking educators to take risks without protecting them is not leadership, it is liability. Jennifer D. Klein, author of Taming the Turbulence in Educational Leadership, offers a clear-eyed framework for how school leaders can prepare their people with transformative professional learning, adapt systems to support innovation, and stand as a buffer when opposition arrives. This is the kind of piece that reminds education leaders why the soul of their work has always been human development, for adults as much as students. The post Prepare Your People, Protect Your People: Setting the Stage for Successful Change Management appeared first on Getting Smart .

Source ↗
regulation Thu, 18 Jun 2026 18:30:00 +0000
The 74

California Lawmakers Pass Budget With Billions More for Education as Newsom Negotiations Begin

Marking the start of two weeks of intensive negotiations, the Legislature passed a state budget Monday with higher revenue projections than those proposed by Gov. Gavin Newsom, providing several billion dollars in additional spending for TK-12 and community colleges in 2026-27. Several other significant issues remain unresolved. Chief among them is the $3.9 billion in […]

Source ↗
regulation Thu, 18 Jun 2026 17:00:00 -0400
K-12 Dive

Takeaways from the Ed Dept-HHS special ed agreement

Critics worry it will lead to a medical approach, while supporters say the collaboration will improve outcomes.

Source ↗
audience Thu, 18 Jun 2026 16:30:00 -0400
Higher Ed Dive

Kansas board adopts definitions for ban of DEI-CRT in required courses

The state higher ed board’s policy protects broadly teaching about racism and civil rights history under a new state law restricting college instruction.

Source ↗
Showing 7451–7500 of 18624 signals
← Prev Page 150 of 373 Next →