EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Geo-Standardizing 3D Modeling of Surface/Subsurface Objects and Related Logical Spaces on Celestial Bodies: Case Studies for Moon and Mars

arXiv:2601.06182v2 Announce Type: replace Abstract: Establishing frameworks for promoting the realization of various activities on celestial bodies sustainably is of great significance for different contexts, such as preserving the scientific evidence and space heritage. Therefore, this research first proposes a conceptual model that covers the different types of features, attributes, and relationships between them to comprehensively delineate the surface/subsurface objects and related logical spaces on celestial bodies. It then implements this conceptual model as a CityJSON extension in such a way that allows for creating the three-dimensional (3D) geodatasets that represent these objects and spaces in a standardized manner. Moreover, the usefulness of this study is demonstrated through creating CityJSON datasets that include 3D models of exemplary surface/subsurface objects from the Moon and Mars, such as a historical landing site and related logical spaces, such as exclusion zones f

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Auditing Sex/Gender Disparities in Emergency Triage with LLM-based Paired Comparisons

arXiv:2511.17124v2 Announce Type: replace Abstract: We present a domain-agnostic paired-comparison approach that uses Large Language Models (LLMs) to quantify sex/gender-related asymmetries in documented clinical decision-making. The method trains an LLM to emulate observed decisions, then evaluates sex-swapped pairs in which only sex is flipped, holding documented clinical content constant. We apply it to emergency triage, analyzing more than 140,000 Bordeaux University Hospital (France) admissions and testing methodological portability on MIMIC-IV, spanning a different language, population, and healthcare system. Fine-tuning Mistral NeMo 12B for triage prediction and using Mistral Small 24B for pair generation, we find otherwise identical presentations were more likely to receive a lower-severity predicted score as female than male: 1.1% (95% CI 0.9-1.3) in the French cohort, 2.2% (1.7-2.7) in MIMIC-IV. Predictions are sensitive to both tabular and textual sex markers, with the asymm

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Scientific Discovery in the Age of AI and Supercomputing

arXiv:2511.12686v2 Announce Type: replace Abstract: Artificial intelligence (AI) and high-performance computing (HPC) are transforming scientific capabilities and the way science is conducted. Yet their combined impact on scientific discovery remains poorly understood, as do inequalities in access to these capabilities across countries and institutions. Drawing on metadata from more than five million scientific publications (2000-2024) across 27 fields, we examine how the convergence of AI and HPC correlates with scientific breakthroughs. Our results show that this computational synergy is most pronounced at the scientific frontier: research combining AI and HPC is more likely to introduce novel ideas and achieve top-cited status than either conventional work or research using AI or HPC in isolation. We also document growing disparities in access to supercomputing resources and AI expertise, which are increasingly concentrated in a small number of regions (dominated by the United State

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Decolonizing Linguistic Policies in Automated Speech Recognition: A Framework for Cross-Culturally Competent Speech AI

arXiv:2608.06141v1 Announce Type: cross Abstract: This paper focuses on automatic speech recognition (ASR) and ASR-mediated voice interfaces that shape access to public services, healthcare, and education. We argue that persistent failures for low-resource, Indigenous, and non-standard language varieties are not only technical errors, but also implicit linguistic policies that reproduce colonial language hierarchies. Drawing on linguistic capital, raciolinguistic ideology, language policy research, and decolonial computing, we show how data, metrics, and model priors determine whose voices become machine-legible. We introduce the Three Harms (3M) taxonomy---Misrecognition, Misalignment, and Mistrust---and a seven-layer situatedness model for linguistic diversity in ASR and ASR-mediated voice interfaces. We then propose a participatory framework and minimum audit protocol for culturally competent ASR, positioning affected communities as co-designers, evaluators, and governance partners.

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

CourseGraph: Finding overlaps and differences in Computer Science courses across universities

arXiv:2608.05910v1 Announce Type: cross Abstract: Student mobility programs such as Erasmus+ enable students to take courses at other universities, broadening their academic and cultural horizons. However, this flexibility also leads to a practical challenge: ensuring that students do not take courses elsewhere that substantially overlap with courses in their home curriculum. In this work, we propose CourseGraph, a methodology that automates the evaluation of external courses based on insights obtained from the process followed by curriculum administrators when assessing courses for inclusion in a degree program. Course- Graph extracts information such as course titles, descriptions, and learning outcomes from the course webpage. Then, this information is represented semantically using a BERT-based language model, after which the pair-wise similarity between courses can be computed. This information is then used by a Random Forest classifier to determine whether a candidate course abro

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Mapping the Emerging Curriculum for AI-Assisted Software Engineering via Syllabus Analysis

arXiv:2608.05898v1 Announce Type: cross Abstract: As Generative AI coding tools reshape professional software development, universities have begun designing courses to prepare students for AI-assisted development workflows. By analyzing the syllabi of these courses, we can gather empirical evidence about these courses, reveal how this emerging curricular area is being defined, and gain guidance for future curriculum design. We analyzed 23 publicly available syllabi and course materials of upper-division, credit-bearing courses that meet specific criteria, including explicitly addressing Generative AI in software engineering. Through iterative qualitative coding, we characterized courses' learning objectives, assessments, topics, and documented AI tools. Our analysis reveals commonalities and differences among these courses that allow researchers and educators to study and develop future courses.

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

The em-dash em-beds in Congress: A population-level rise in em-dash frequency in U.S. congressional press releases at the dawn of the large-language-model era, 2021-2025

arXiv:2608.05889v1 Announce Type: cross Abstract: Large language models (LLMs) can leave small stylistic traces in text written with their help. The most discussed is the em-dash (U+2014), especially the unspaced form word---word, which is normal in typeset English prose but unusual in U.S. press writing, where AP style calls for spaced dashes. This study asks whether that trace is measurable in congressional press releases. In a preregistered design (OSF: 10.17605/OSF.IO/U5NEY), 146,239 scraper-sourced releases from 480 House and Senate offices (2021-2025, the open congress-press dataset) were analyzed: density of unspaced prose-form em-dashes per 1,000 characters of cleaned text, Poisson/negative-binomial models with a length offset, clustering by office. Density stayed within 0.10-0.12 per 1,000 characters through 2021-2024, then rose to 0.217 in 2025, more than twice the four-year baseline; the share of releases with such an em-dash rose from ~13% to 24.8%. The primary frequency ra

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation

arXiv:2608.05576v1 Announce Type: cross Abstract: When a large language model (LLM) writes Harry Potter fanfiction, it reliably produces fundamental elements of the Hogwarts universe, such as recognizable places and characters. Human-written Harry Potter fanfictions, however, typically include these fundamentals and much more, incorporating stylistically irregular content and relationship-diverse plotlines. This gap between LLM and human writing has been noted across a variety of domains. LLMs tend to produce "average" writing, while human writing contains more diverse content that covers a broader distribution. Existing work has shown the existence of this distributional "gap", but no work has proposed a systematic way to measure it. Our paper proposes a human-grounded framework that uses the empirical distribution of human writing on a topic to measure the distributional breadth of LLM-generated content on that same topic. We propose two metrics, LLM Coverage (LLM-Cov) and In-Boundar

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Small Foundation Models of Human Cognition and Behaviour

arXiv:2608.05224v1 Announce Type: cross Abstract: Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We train fourteen models from 135M to 14B parameters across four architecture families on Psych-101, a dataset of 10.7 million trial-level choices from 160 experiments. In-distribution, scale barely matters. The models fall within a narrow band, as though against a ceiling, and 0.6B to 1B parameters suffice to match a 70B baseline on held-out participants. Out-of-distribution, that band opens into a markedly steeper scaling gradient, with larger models clearly advantaged in generalisation to novel task structure. To determine what information these models use, we run two diagnostics. We progressively strip four prompt channels -- task instructions, experimental stimuli, outcome feedback, and choic

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Conditional Cognitive Biases in LLMs: How Biased User Turns Modulate In-Context Reasoning

arXiv:2608.05166v1 Announce Type: cross Abstract: We present an evaluation of cognitive bias expression in state-of-the-art instruction-tuned LLMs under realistic multi-turn interaction settings. Our work introduces a novel three-condition experimental framework that disentangles the effect of exposure to a biased user turn from the effect of the turn's semantic content, alongside a benchmark of 24,300 jury-validated user prompts spanning all 81 cells of a 9x9 target-human bias interaction matrix. Across eight frontier LLMs, we find that biased conversational context systematically increases bias expression relative to zero-shot baselines in 6 of 8 models. We identify two competing behavioral dynamics underlying this effect: conversational exposure to biased reasoning generally amplifies downstream bias tendencies, while explicitly stated bias cues often trigger alignment-related suppression behaviors that reduce overt bias expression. We release our framework, codebase, and dataset to

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of Nigeria

arXiv:2608.06364v1 Announce Type: new Abstract: The use of e-commerce mobile applications is expanding in Nigeria, creating both opportunities and risks, including fraud and reduced user control over digital technologies, raising concerns about digital sovereignty. This research examines how Artificial Intelligence (AI) in Nigerian mobile applications affects digital sovereignty, examined through platform transparency as a key indicator of user awareness and control. Using an interpretive approach, the research combines the forensic analysis of selected Android applications with contextual document analysis to identify AI features and evaluate disclosure practices. The findings show that AI is widely implemented in the applications, yet transparency about its use remains limited. A socio-economic analysis of Nigeria further shows an increasing dependence on consumer digital platforms, moderate AI awareness, and uneven patterns of interaction. By providing empirical evidence on AI trans

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

From Precision Medicine to Precision Education: A Vision for AI-Powered Student Digital Twins, Preventive Student Success, and Career-Aligned Academic Pathways

arXiv:2608.06322v1 Announce Type: new Abstract: Higher education remains largely reactive in its approach to student success. Institutions frequently identify academic problems only after students have failed courses, fallen behind in degree progression, accumulated excessive debt, or departed without a credential. Healthcare faced a similar challenge decades ago. It responded by shifting from reactive treatment to preventive care powered by predictive models, risk stratification, electronic health records, and artificial intelligence (AI). This paper argues that higher education stands at an analogous inflection point. Drawing on advances in learning analytics, educational data mining, machine learning, workforce analytics, and digital twin technologies, we propose a paradigm we call Precision Education. Under this framework, AI continuously analyzes academic, behavioral, financial, and career data to identify emerging risks, recommend personalized interventions, optimize educational

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries

arXiv:2608.06166v1 Announce Type: new Abstract: The article reports on a blind Turing Test experiment, assessing the performance of out-of-the-box leading LLMs on three Italian legal professional exams: the Bar, Judges and Notary exams. Leading LLMs were asked to generate full written exam papers, which were made indistinguishable from human submissions and anonymously evaluated by expert examiners, using the same criteria applied in real examinations. Results reveal marked differences across both models and tasks. While some LLMs match or exceed top human performance in adversarial legal argumentation and doctrinal analysis, all models fail in the notary exam, which requires goal-directed legal planning under strict formal and substantive constraints. Beyond ranking models, the study identifies task-specific strengths, limitations and recurring legal failure patterns. Although limited to out-of-the-box systems, the findings provide qualitative evidence on the current scope and boundar

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Algorithmic Flattening of Sound: Computational Evidence and Justice Implications of AI Music Homogenization

arXiv:2608.06106v1 Announce Type: new Abstract: This paper audits whether large-scale generative music systems exhibit measurable musical homogenization relative to human-produced music, and develops a justice-centered account of why this matters. We audit two commercially deployed systems (Suno and Lyria 3) across four genres (Afrobeats, K-pop, Dance Pop, and Heavy Metal). For each system and genre, we generate 100 tracks and compare them against human corpora of equal size, using 72 music information retrieval (MIR) features and multiple diagnostics of dispersion, redundancy, and separability. We define homogenization as reduced acoustic variation in standard computational audio features including rhythm and timing, timbre/spectral shape, and dynamics, both within genres and across genre boundaries. We also generate tracks using only a genre name as the prompt, with no additional instructions, to reveal each system's default musical tendencies. The results show two structurally disti

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Validity, Reliability, and Transparency in Artificial Intelligence Regulation

arXiv:2608.05800v1 Announce Type: new Abstract: Artificial intelligence (AI) systems increasingly mediate decisions affecting individuals and societies. Existing data protection frameworks address certain privacy-related harms, particularly those arising from data leakage, re-identification, and profiling. However, they inadequately capture a more fundamental risk: unreliable or unjustified inference produced by AI systems even when data collection and processing are legitimate. This article argues that modern AI raises distinct concerns of construct validity, confounding, representativeness, distribution shift, and fairness trade-offs that require specialised regulatory attention. In the context of AI, transparency and explainability acquire distinct and significantly more challenging meanings than in conventional software. A substantial body of work in critical data studies and the measurement-theoretic literature has diagnosed these epistemological limitations. This article's contri

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Human Research in AI Safety & Ethics

arXiv:2608.05656v1 Announce Type: new Abstract: Safety risks of AI are becoming increasingly evident in human interactions with AI technologies. The prominent approaches to evaluating these risks favor technical methods, such as model benchmarks and LLM simulations, often sidelining empirical research with human subjects. To examine this apparent gap in the acceptance of human research, we conduct an expert survey (n=93) and expert interviews (n=17) with AI Safety & Ethics (AISE) researchers from Technical, Sociotechnical, Governance, and Normative backgrounds. Our findings suggest that although there is a consensus that human research is valuable for generating evidence for AISE, its adoption and acceptance are constrained by perceived validity issues, tangible resource barriers, epistemic and personal preferences in methods, and infrastructural constraints from the broader research community. In particular, Technical researchers tend to value human research less and collaborate acros

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Judgment-Consequence Gap: LLM Moral Reasoning in Healthcare Decisions

arXiv:2608.05583v1 Announce Type: new Abstract: As large language models (LLMs) enter high-stakes domains such as healthcare, understanding their moral reasoning becomes essential. Decisions about scarce medical resources often hinge on judgments of responsibility, particularly when patients' own actions contribute to illness. We investigate how LLMs reason about responsibility and its consequences, tracing their judgments across successive levels, from the behavior, to the resulting illness, to the denial of care. We evaluate a wide range of LLMs, spanning different model families and capability levels, on various clinical vignettes adapted from prior studies. Our results identify a judgment-consequence gap: LLMs largely agree with humans that patients bear responsibility for health-harming behaviors, yet overwhelmingly refuse to let that judgment influence how they allocate scarce resources. Specifically, LLMs default to random allocation, whereas humans consistently favor the less-c

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Vibe Compiler: A Research-Logic Synthesis Tool That Runs without Prompt Engineering -Toward Enhancing Metacognition for Sustaining Agency in the Age of Generative AI-

arXiv:2608.05545v1 Announce Type: new Abstract: Generative AI used as a capable servant has greatly accelerated intellectual work, but it also risks eroding human epistemic agency by encouraging uncritical acceptance of AI-generated reasoning. This creates a need for mechanisms that preserve human agency by augmenting metacognition during AI-assisted intellectual work. To address this, we propose the Synthesis-Analysis Reciprocity Model, which views intellectual construction as a reciprocal interaction between Synthesis, which combines components into an artifact, and Analysis, which critically evaluates them against objective indicators and constrains subsequent synthesis. Grounded in this model, we present the Vibe Compiler, a research-logic compiler that helps researchers transform vague ideas (Vibes) into coherent research logic. The system compiles these ideas using a research paper ontology of sixteen academic parameters. Compilation failures indicate missing logical components;

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Vision for the Future of an AI-Integrated Research Ecosystem

arXiv:2608.05438v1 Announce Type: new Abstract: Generative AI has infiltrated every stage of the research lifecycle: how scholarship is conducted, written, published, and reviewed. Recent policy responses, such as ACM's authorship policy, address an immediate concern about responsible and transparent disclosure of AI use. We argue that a focus on authorship and disclosure, although necessary, risks obscuring and ballooning a set of entrenched problems and strains within publication systems. The central question is not about how papers and other research artifacts should incorporate AI, but how scientific communication itself should evolve when all relevant parties (authors, reviewers, readers) may rely on AI assistance. We draw on our experience within these and other roles to illustrate two contrasting but feasible visions of 2036 with four entwined questions, namely about the purpose of papers as artifacts, reviews, human reviewers, and the incentives that bind all of them. We argue

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

School network reorganization under educational and spatial constraints using classical and quantum optimization

arXiv:2608.05427v1 Announce Type: new Abstract: School network reorganization is a strategic planning problem that requires balancing demographic trends, territorial accessibility, educational requirements, and institutional constraints while ensuring an efficient allocation of public resources. This paper proposes an optimization framework for school dimensioning decisions based on a novel Integer Linear Programming formulation integrating geographical, administrative, and educational criteria. A synthetic benchmark generator is introduced to evaluate the scalability and computational performance of the model on artificial instances, while a real-world case study involving the complete public school network of the Calabria region (Italy) is conducted using actual institutional, territorial, and demographic data. The proposed approach effectively identifies optimal aggregation plans under different policy scenarios while preserving the structural characteristics of the educational syst

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Beyond Demographics: BIM Engagement and Job Satisfaction Among AEC Professionals, A Machine Learning Pilot Study

arXiv:2608.05181v1 Announce Type: new Abstract: Building Information Modeling (BIM) has transformed workflows across the Architecture, Engineering, and Construction (AEC) industry, yet its relationship with employee job satisfaction remains insufficiently understood. This pilot study investigates whether BIM engagement or demographic characteristics better predict job satisfaction among AEC professionals. Survey responses from 104 participants were analyzed using Spearman rank correlations, logistic regression, and Classification and Regression Tree (CART) modeling. 27 items Job Satisfaction Index demonstrated excellent internal reliability. Across all analytical approaches, BIM engagement emerged as a stronger predictor of job satisfaction than demographic factors. Specifically, the proportion of project work completed using BIM was the only significant predictor of job satisfaction, whereas age, gender, education level, and professional experience showed no significant relationships.

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Nuclear Decision-Making Benchmark: Evaluating Frontier LLMs on Nuclear Tendencies

arXiv:2608.05180v1 Announce Type: new Abstract: The integration of large language models into defense and national-security workflows raises urgent questions about whether frontier models exhibit stable, consistent, and policy-appropriate preferences in high-stakes contexts. We introduce the Nuclear Decision-Making Benchmark (NDM Bench), a targeted evaluation framework of 151 scenarios authored by PhD-credentialed scholars in international relations spanning four domains: escalation (76), arms control (25), non-proliferation (25), and proliferation (25). Scenarios are actor-agnostic, enabling multiple country pairs to be exchanged, and we introduce experimental phrasing variants to probe sensitivity to narrative framing. We apply the benchmark to seven frontier AI systems: DeepSeek-V3.2, ERNIE 4.5-300B, Gemini 3 Pro, GLM-4.6, GPT-5.2, Llama 4 Maverick-17B Instruct, and Qwen3-235B. We find significant overall inter-model variation in all four domains, with 91.7% of pairwise inter-model

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap

arXiv:2608.05179v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly used across the scientific research lifecycle: ideation, literature search, experiment design and execution, analysis, manuscript drafting, and review. End-to-end AI scientist systems can now produce paper-like manuscripts, but their claims are often harder to verify than their code is to run. This survey studies that gap in computational AI/ML research, where code, benchmarks, experiments, and write-ups are most visible. We screen 125 candidate works and include 35, with full-text coding of 26 entries: 24 runnable systems and two study or position papers. We code seven audit dimensions: lifecycle stage, autonomy level, evaluation method, released artifacts, human-in-the-loop points, novelty verification, and result-selection disclosure. The main pattern is that code release is now common, but reproducibility-grade and claim-verification artifacts remain much less common. In the 24 runnab

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Who Gets Access? Global Region and Academic Status Bias in AI-Generated Academic Gatekeeping Scenarios

arXiv:2608.05178v1 Announce Type: new Abstract: Equitable access to scientific knowledge often depends on informal gatekeeping decisions, particularly when resources such as paywalled articles, datasets, or professional materials such as curriculum vitae (CV) must be shared selectively. We introduce a controlled simulation framework in which large language model (LLM)-based professors must grant access to only one requestor. Across prompts, requesters vary systematically by global region (Global North vs. Global South) and academic seniority (undergraduate student, PhD candidate, postdoctoral researcher, and tenured professor), while all other factors remain constant. Across varying evaluation scenarios, LLMs exhibit contrasting academic status biases, with some prioritizing PhD candidates, while others favor tenured professors. However, when global regions differ, a distinct divergence emerges based on model architecture: while many frontier LLMs systematically favor requesters from t

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Using AI-Generated Feedback to Improve Critical Thinking and Writing Proficiency

arXiv:2608.05177v1 Announce Type: new Abstract: Research indicates students require customized written composition feedback to enhance critical thinking and writing competence, yet teachers face barriers to delivering timely personalized guidance due to heavy workloads. To address this gap, this study developed the Writing Improvement and Smart Evaluation Agent (WISE Agent), an artificial intelligence (AI) feedback tool targeting textual logic and perspective biases in student essays. We conducted a three-month intervention with 260 Chinese sixth-grade students, each completing seven themed essays and receiving targeted WISE Agent feedback shortly after submission. Assessment used a critical thinking rubric adapted from the California Critical Thinking Disposition Inventory (CCTDI), covering seven core dimensions including cognitive maturity and open-mindedness. Results indicate structural optimizations in critical thinking dimensions rather than a uniform increase in total scores. Whi

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Challenges for Musical Education in the Age of AI and Digital Transformation

arXiv:2608.05176v1 Announce Type: new Abstract: Music education has never been a static discipline. Each major technological shift has forced educators and institutions to reconsider what they teach, how they teach it, and why. We now stand at what may be the most consequential of such turning points. Three deeply intertwined transformations have been converging simultaneously: 1. The very nature of music has changed: how it is made, distributed, consumed, and valued; 2. The public for music has changed: listening habits are now shaped by streaming algorithms and the boundary between consumer and creator has blurred; 3. Music-making itself has changed: digital audio workstations (DAWs) have for two decades been reshaping compositional practice. In addition, generative AI has now irrupted, capable of producing complete, stylistically coherent musical pieces from a short text prompt. These changes are not independent of one another, and they all bear directly on musical education - both

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Teaching Intro AI When the Tools Can Do the Homework: A Course Redesign and a Student Bill of Rights

arXiv:2608.05175v1 Announce Type: new Abstract: Large language models can complete most of the assignments in an introductory artificial intelligence course. This paper is an experience report on redesigning one such course, CSS~382 at the University of Washington Bothell, in response. Rather than freeze the curriculum, the redesign retained the course's classical core (search, adversarial search, Markov decision processes, reinforcement learning) and added a strand in which students build a large language model from scratch, so that a tool they are required to use is also one they are required to understand. Assessment was rebuilt around tasks that resist unattributed automation: in-class exercises, reflective writing, and a defended team project, with examinations removed entirely. The policy on AI was inverted, from unmentioned in 2023 to required in 2026. The center of the paper is a participatory ethics sequence in which a cohort of students deliberated on and endorsed a "Student

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Art in Humanity's Code

arXiv:2608.05174v1 Announce Type: new Abstract: Artist-led open-source libraries such as Processing or openFrameworks have had a major impact on artists and designers who use code as a creative medium. In this work, we conduct the first large-scale empirical study of public code repositories that use these libraries. Our study dives into 1,613,571 code repositories collected from the Software Heritage archive. Combining quantitative and qualitative methods, we investigate the diversity of practices and practitioners in terms of code hosting, geographical distribution, characteristics of code-based creative works and the purposes of these works. Key findings include evidence of the worldwide presence of generative art and creative coding, as well as the adoption of these practices in both education and across creative industries. We illustrate these findings with concrete examples of repositories and contributor profiles around the globe, spanning the spectrum from university curricula

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Closing Window: How Governments Could Lose Their Ability to Restrain Advanced AI

arXiv:2608.05173v1 Announce Type: new Abstract: As AI capabilities advance, AI systems will pose greater risks to national security and potentially humanity as a whole. Governments may eventually conclude that these risks warrant restraining AI development. This motivates the question: will governments still be able to restrain AI development in the future, should they want to do so? In this paper we analyze which world events and changes to the state of AI development would make future governance more difficult or even effectively impossible. Our analysis surfaces likely pathways that would lead to these difficulties, including hardware proliferation, continued algorithmic progress, and the release of catastrophically dangerous AI models. Due to the field's lack of understanding of AI development, it may be difficult or impossible to know when we will hit a "point of no return", and we therefore recommend a conservative approach. The window may be closing, but governments currently ha

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Estimating time spent on work tasks

arXiv:2608.05172v1 Announce Type: new Abstract: The task-based framework in economics models occupations as bundles of tasks. It is the standard lens for understanding how technology affects work: a new technology changes the cost or time each task requires and these task-level effects aggregate to occupation-level effects. We study how tasks should be weighted in this aggregation. Prior work has relied on idiosyncratic or ill-justified choices for task weights. While recent work suggests weighting tasks by time spent, existing time shares are either based on coarse ONET data not intended for this purpose or estimated via black-box language models. We address this gap by proposing a principled method for estimating time shares for nearly 18,000 tasks that constitute nearly all U.S. jobs. Our estimates factor a task's time into (i) the expected frequency of the task, derived from ONET, and (ii) the time to complete a single instance of it. To estimate the latter, we solve a constraint s

Source ↗
technology Fri, 07 Aug 2026 00:00:00 -0400
arXiv cs.CY

Beyond Information Retrieval: Generative AI as an Epistemic Arbiter to Enhance Collaborative Problem-Solving

arXiv:2608.05171v1 Announce Type: new Abstract: Generative AI (GAI) creates new opportunities for collaborative problem-solving (CPS), yet its role in shaping student interaction remains unclear. To address this gap, we conducted a six-week quasi-experimental study with 201 fifth-grade students in two conditions: with and without GAI. Chi-square analysis showed significant differences in CPS behavior distributions between groups. Compared with the control group, the GAI-supported group demonstrated more social behaviors, particularly engagement and conflict management, but less frequent cognitive behaviors such as task planning and solution reasoning. Lag sequential analysis further revealed distinct interaction patterns: while the control group followed a more conventional transition from listening to planning, the GAI group showed a robust pathway from task planning to conflict management to solution reasoning. Thematic analysis of AI interaction logs suggested that students used GAI

Source ↗
technology Fri, 04 Sep 2026 00:00:00 -0400
arXiv cs.CY

MIRA: A Bilingual Benchmark for Medical Information Response Audit

arXiv:2605.28025v2 Announce Type: replace-cross Abstract: Existing safety evaluations for large language models overlook whether responses preserve comparable medical information across different user phrasings of the same question. To address this, we introduce the Medical Information Response Audit (MIRA), a bilingual, controlled benchmark that assesses whether LLMs provide comparable medical information across user-side language, register, and health literacy signals. MIRA contains 4,320 prompts built from 60 medically reviewed, low-risk health questions. Across five mainstream LLMs, models answered all medical questions, but responses to low health-literacy signals consistently omitted more key information, provided fewer concrete next steps, and offered less support for independent judgment. We term this pattern Differential Information Dilution (DID). A comparison with 300 real-world health queries provides preliminary evidence of rank-order validity. A knowledge-guided mitigatio

Source ↗
technology Fri, 04 Sep 2026 00:00:00 -0400
arXiv cs.CY

Identifying AI Web Scrapers Using Canary Tokens

arXiv:2605.13706v2 Announce Type: replace-cross Abstract: From pre-training to query-time augmentation, web-scraped data helps to improve the quality and contextual relevancy of content generated by large language models (LLMs). However, large-scale web scraping to feed LLMs can affect site stability and raise legal, privacy, or ethics concerns. If website owners wish to limit LLM-related web scraping on their site, due to these or other concerns, they may turn to scraper access control mechanisms like the Robots Exclusion Protocol. To be most effective, such mechanisms require site owners to first identify the scrapers that they wish to restrict (e.g., via User-Agent strings). Existing mechanisms to identify LLM-related scrapers rely on voluntary disclosure by companies, one-off experiments by researchers, or crowd-sourced reports -- methods that are neither reliable nor scalable. This paper proposes a novel technique for accurately and automatically inferring LLM-related scrapers. We

Source ↗
technology Fri, 04 Sep 2026 00:00:00 -0400
arXiv cs.CY

Towards Affordable Energy: A Gymnasium Environment for Electric Utility Demand-Response Programs

arXiv:2605.12462v2 Announce Type: replace-cross Abstract: Extreme weather and volatile wholesale electricity markets expose residential consumers to catastrophic financial risks, yet demand response at the distribution level remains an underutilized tool for grid flexibility and energy affordability. While a demand-response program can shield consumers by issuing financial credits during high-price periods, optimizing this sequential decision-making process presents a unique challenge for reinforcement learning despite the plentiful offline historical smart meter and wholesale pricing data available publicly. Offline historical data fails to capture the dynamic, interactive feedback loop between an electric utility's pricing signals and customer acceptance and adaptation to a demand-response program. To address this, we introduce DR-Gym, an open-source, online Gymnasium-compatible environment designed to train and evaluate demand-response from the electric utility's perspective. Unlike

Source ↗
technology Fri, 04 Sep 2026 00:00:00 -0400
arXiv cs.CY

CARDIO-Affect: A Hamiltonian-Variability Framework for Spatio-Temporal Emotional Pattern Recognition with Manifold-Based Individual and Group Profiling

arXiv:2510.16046v4 Announce Type: replace-cross Abstract: We present CARDIO-Affect, a complex-systems theoretical framework for long-term emotional dynamics in bounded social groups, with explicit uncertainty quantification at every layer. Long-period naturalistic emotion in stable small groups exhibits hallmarks of complex systems -- multi-stable attractors, weak chaos, long-range memory, and sparse heterogeneous coupling -- invisible to conventional short-clip facial-emotion analysis. CARDIO-Affect treats individual emotion as a multi-stable nonlinear stochastic dynamical system and group emotion as a sparsely-coupled network with emergent macrostates, formalised through six propositions and four pillars: (i) statistical mechanics with neural-parameterised Hamiltonian SDE over asymmetric potentials; (ii) information geometry on a 45-dimensional Fisher-Rao manifold; (iii) topological data analysis for invariant trajectory signatures; (iv) HRV-inspired Emotional Variability Analytics (

Source ↗
technology Fri, 04 Sep 2026 00:00:00 -0400
arXiv cs.CY

WELD: The First Naturalistic Long-Period Small-Team Workplace Emotion Dataset for Ubiquitous Affective Computing

arXiv:2510.15221v3 Announce Type: replace-cross Abstract: Affective computing has matured rapidly in laboratory settings, yet no prior dataset combines (i) months-to-years of duration, (ii) a naturalistic workplace context, (iii) a stable small-team social structure, and (iv) a fully passive sensing protocol that survives institutional review. We introduce WELD, the first dataset to satisfy all four. WELD comprises 733,780 per-frame seven-class facial-expression probability vectors from 49 employees of a Chinese software company over 30.1 months (Nov 2021 - May 2024) -- the longest naturalistic in-the-wild emotion corpus and the only multi-year corpus supporting both within-individual longitudinal and within-team relational analyses on the same subjects. Data are released under a four-tier access model with only aggregated probabilities publicly downloadable. We validate the corpus by replicating three established phenomena (+43.1% weekend valence boost; 13:00-trough diurnal cycle; Sha

Source ↗
technology Fri, 04 Sep 2026 00:00:00 -0400
arXiv cs.CY

A Visionary Look at Vibe Researching

arXiv:2604.00945v2 Announce Type: replace Abstract: Vibe researching is an emerging paradigm in which human researchers provide high-level direction and critical judgment while LLM-based agents handle the labor-intensive execution of literature review, experimentation, data analysis, and manuscript drafting. Inspired by the "vibe coding" movement in software engineering, it occupies a middle ground between traditional manual research and fully autonomous AI research systems. This paper defines the concept, describes its methodology (multi-agent architectures, memory, tool use, retrieval-augmented generation, and the human's role as orchestrator), identifies seven technical limitations, weighs its positive and negative societal impacts, and maps each problem to a concrete future direction. Our goal is to provide the research community with a clear and honest map of the territory so that the conversation about responsible adoption can start from shared ground.

Source ↗
technology Fri, 04 Sep 2026 00:00:00 -0400
arXiv cs.CY

OBER+: Continuity-Aware Reporting and Traceable Continuous Improvement in Outcome-Based Education

arXiv:2609.03770v1 Announce Type: cross Abstract: Institutions practising outcome-based education compute learning outcome attainment routinely, while reviews of curriculum analytics report an absence of evidence on how that computation informs decisions. This paper presents OBER+, an extension of a deployed institutional attainment platform that computes the step from a measured shortfall to an evaluated corrective action. Five connected stages accumulate attainment across deliveries of a course, signal a shortfall and a persistent shortfall, grade it on cutoffs the regulator already uses, record the decision against a catalogue of practices annotated with their evidence, log the change, and quantify the subsequent movement in the shortfall. A further rule compares successive statements of an outcome, so attainment is never read as a series across a point at which the outcome changed. Applying the rules to the live record of two real courses produced three results. Every outcome of a

Source ↗
technology Fri, 04 Sep 2026 00:00:00 -0400
arXiv cs.CY

GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis

arXiv:2609.03553v1 Announce Type: cross Abstract: Policy analysis requires more than predicting whether a proposal will pass: it requires identifying who will be affected, how those actors respond, and what follows. LLM-based policy simulations model these processes at scale, but their validity is hard to establish when plausible behaviour is never compared with observed outcomes. We introduce GPS-Bench, an evidence-grounded benchmark for governance policy simulation that links policies to relevant actors, actor actions and downstream impacts using legislative records, lobbying disclosures, regulatory documents, corporate filings, economic data and other public evidence. Actors are reconstructed from the dated record rather than prompted as archetypes, so a persona is an evidence object with provenance; a human-annotated pool forms the Gold evaluation set, while cases labelled by a separate LLM from retrieved evidence are treated as Silver supervision and never as test labels. Because

Source ↗
technology Fri, 04 Sep 2026 00:00:00 -0400
arXiv cs.CY

Accountable AI with Grounded, Faithful, Consistent, Actionable Rationales: A Case Study in Clinical Trial Matching with VERDICT

arXiv:2609.03366v1 Announce Type: cross Abstract: Accountability means a decision can be examined, justified, and contested. LLMs make this hard: fluent output may be ungrounded, incomplete, or unfaithful to the decision process. Achieving accountability requires verified rationales (how was the decision reached), assumptions (what was assumed rather than known), policy consistency (the same treatment for the same facts), and pivotal conditions (what would change the outcome). We introduce self-faithfulness as an automatic test of accountability: changing the pivotal conditions should change the decision. We examine accountable AI through clinical trial matching, a high-stakes task central to evidence-based medicine. Although LLM-based matchers match patients to trials reasonably accurately, they apply decision policies inconsistently and produce rationales that are unfaithful to their own decisions. We introduce VERDICT, an LLM-based agent that translates a decision task, its constrai

Source ↗
technology Fri, 04 Sep 2026 00:00:00 -0400
arXiv cs.CY

Large-Language Models as a Cognitive Virus

arXiv:2609.03344v1 Announce Type: cross Abstract: Large-language models (LLMs) are rapidly becoming part of human culture, reshaping how information is produced, transmitted, and used. Here we propose that their diffusion can be understood through a viral analogy, with LLM use spreading through populations, becoming embedded in cognitive and cultural practices. We model transitions among uncoupled, coupled, and persistently dependent users, and show that the interplay between social transmission, recovery, and collective reinforcement can generate tipping points and technological lock-in. A central consequence is the possibility of runaway dynamics: once a critical threshold is crossed, small increases in adoption can trigger rapid population-level shifts toward persistent dependence, with abrupt losses in cognitive competence. The same framework, however, identifies conditions for cognitive immunization, based on reducing transmission and facilitating reversibility. Our results highli

Source ↗
technology Fri, 04 Sep 2026 00:00:00 -0400
arXiv cs.CY

Less Is Moral: A CHARMing Framework for Moral Foundations Detection in Endorsement Behaviour

arXiv:2609.03330v1 Announce Type: cross Abstract: Moral language plays a central role in shaping online endorsement and the diffusion of information, yet existing moral foundation detection systems often suffer from poor cross-domain generalization, weak rationale grounding, and reliance on costly prompting-based large language models (LLMs). We introduce CHARM, a MA\textbf{C}- and \textbf{H}ate-speech-\textbf{A}ware \textbf{R}ationale-aligned \textbf{M}oral foundation detection framework built on a lightweight fine-tuned LLM, which integrates complementary moral grounding, rationale alignment, and polarity-aware hate speech signals to support more robust and faithful moral prediction. Unlike prior dictionary-, fine-tune-, or prompt-based detectors, which decouple computation from psychological theory, CHARM is built so that each component -- MAC cross-attention, rationale alignment, and hate-speech modulation -- operationalizes a distinct psychological construct. Using a 30\% subsampl

Source ↗
technology Fri, 04 Sep 2026 00:00:00 -0400
arXiv cs.CY

Evaluating GNNs for Success Prediction in Artist Collaboration Networks

arXiv:2609.02920v1 Announce Type: cross Abstract: As the music industry becomes an increasingly collaborative effort, understanding the underlying structures of the artist network has become a focal point in cultural data analytics. This study expands on the previous analyses of the Italian and Danish networks by introducing a novel dataset of the Polish music scene. By utilizing methodologies used in the prior studies, this work enables a direct comparison between three distinct European music landscapes and allows to merged the created networks into one. Furthermore, this research introduces a framework to test the efficacy of Graph Neural Networks (GNNs) for artist popularity predictions based on the metadata and the position in the network. The statistical analysis revealed that the Polish and tri-national network exhibit similar properties and clustering behaviours, consistent with prior models. An evaluation of the predictive architectures reveals that while GNN models achieve a

Source ↗
technology Fri, 04 Sep 2026 00:00:00 -0400
arXiv cs.CY

Decreasing Digital Distraction in College Students: Associated Online Learning Strategies Identified by Unsupervised Data Mining Approaches

arXiv:2609.04125v1 Announce Type: new Abstract: The proliferation of digital tools in education offers numerous benefits but also introduces significant challenges, notably digital distractions that hinder academic performance, especially in online learning contexts. This study employed unsupervised data mining techniques, specifically association rule mining and clustering analysis, to identify effective learning strategies associated with lower levels of digital distractions among college students. Data from 530 participants revealed that self-regulated learning strategies (i.e., goal setting, environment structuring, and time management) co-occurred most consistently with lower digital distractions. Additionally, learner-instructor and learner-content engagement strategies, as well as technical competencies, also tended to appear in the same profiles as lower distraction. Interestingly, reliance on peer help-seeking and learner-learner engagement strategies appeared less often in th

Source ↗
technology Fri, 04 Sep 2026 00:00:00 -0400
arXiv cs.CY

Shifting from Injection to Interaction: Rethinking Web Security in the Age of LLMs and Beyond

arXiv:2609.03999v1 Announce Type: new Abstract: Large language models (LLMs) are becoming integral to web applications and browser agents, transforming online interactions while introducing new attack vectors and reshaping longstanding web vulnerabilities. Classical threats such as cross-site scripting (XSS) can be amplified through LLM-mediated interactions, while LLM-specific vulnerabilities can propagate across web applications, introducing attacks such as prompt injection. Securing modern web systems therefore requires understanding interactions between traditional and LLM-specific threats across the system lifecycle. Unlike prior surveys treating web and LLM security separately, this survey provides a unified analysis of how LLMs amplify web vulnerabilities across client-side, server-side, and pipeline layers while evaluating defenses and their limitations. The analysis examines extending NIST and ISO/IEC AI security frameworks to the security needs of LLM-enabled web environments

Source ↗
technology Fri, 04 Sep 2026 00:00:00 -0400
arXiv cs.CY

Making Gender-Inclusive Practices Actionable: Evaluating a Research-Informed Computing Education Toolkit

arXiv:2609.03936v1 Announce Type: new Abstract: The persistent gender imbalance in computing remains a global concern, and universities offer a key part of the pipeline to address it. Although research has identified practices that support under-represented student groups, translating this evidence into actionable guidance remains challenging. This paper first presents a novel web- based toolkit (TechMate) designed to address this gap by helping computing educators implement gender- inclusive initiatives through practical research-informed guidance. The toolkit defines over 25 actions, ranging from operational to strategic, and provides case studies and implementation resources. Second, this work reports on the evaluation of TechMate, capturing educators first impressions of its usefulness and usability through authentic tasks, and eliciting unanticipated insights about structural barriers to gender-inclusive practice in computing higher education. Eighteen computing educators of varyi

Source ↗
technology Fri, 04 Sep 2026 00:00:00 -0400
arXiv cs.CY

Bridging Formal and Perceived Fairness: Development of an Interdisciplinary Framework in Algorithmic Decision-Making

arXiv:2609.03853v1 Announce Type: new Abstract: While fairness has become a central concern in research on algorithmic systems, the field remains predominantly shaped by Computer Science, resulting in a strong emphasis on formal fairness metrics and bias mitigation strategies. Nevertheless, this focus may obscure a fundamental challenge: fairness is not merely a technical property, but a subjective, context-sensitive human judgment shaped by cognitive heuristics, mental models, normative expectations, and sociotechnical factors. Crucially, users' perceptions of fairness may diverge substantially from the fairness criteria an algorithm formally satisfies; a system may meet predefined technical fairness requirements yet still be perceived as unjust by decision-affected stakeholders. In such cases, the system fails on a fundamental dimension: it will not be trusted, accepted, or considered legitimate. Taking a user-centered design perspective, this paper presents a work-in-progress concep

Source ↗
technology Fri, 04 Sep 2026 00:00:00 -0400
arXiv cs.CY

Open WebXR versus Commercial Game Engines: A Socio-Technical Position Analysis for an Open, Sustainable, and Interoperable Metaverse

arXiv:2609.03666v1 Announce Type: new Abstract: The Metaverse is often framed as a persistent, interoperable, and embodied network of virtual and augmented environments. Yet, most contemporary XR applications are developed through commercial game engines and distributed through proprietary app stores, creating tensions between openness and platform dependency. This paper critically examines open WebXR technologies with conventional commercial game-engine pipelines, with particular attention to XR hardware, software architectures, developer workflows, governance, ethics, and sustainability. We argue that WebXR may provide a viable route toward a more accessible, device-independent, and institutionally sustainable Metaverse, especially for education, research, cultural heritage, prototyping, and public-interest applications. At the same time, commercial engines remain advantageous for graphically intensive, low-latency, deeply integrated, and large-scale XR products. The paper concludes

Source ↗
technology Fri, 04 Sep 2026 00:00:00 -0400
arXiv cs.CY

The 5P Reflection Model for Education in the Generative Artificial Intelligence (GenAI) Era

arXiv:2609.03413v1 Announce Type: new Abstract: Contributions: A reflection model suitable for the era of Generative Artificial Intelligence (GenAI) is introduced. The proposed model is an integrated model that extracts features from various existing models and also incorporates technological aspects of GenAI. Background: Universities worldwide are facing challenges in adopting GenAI into their curricula, as it has impacted academic integrity and the scholarship of teaching and research. Traditional reflection models are struggling to authenticate student reflection as GenAI is incorporated in education. This requires a GenAI-aware model to enable the opportunities that address the associated challenges with GenAI. Research Questions: Does the academic ecosystem require a GenAI-aware reflection model to adopt GenAI into education? How to make a reflection model structured to ensure student authenticity and cognitive engagement within a GenAI-aware learning environment? Methodology: Thi

Source ↗
technology Fri, 04 Sep 2026 00:00:00 -0400
arXiv cs.CY

Multilingual Agent System for Inclusive Wildfire Evacuation Guidance

arXiv:2609.03301v1 Announce Type: new Abstract: Wildfire seasons have become 84 days longer in the current days than in the 1970s, causing enormous threats to one's financial status and short- and long-term health. During the fire, public agencies send out emergency messages to provide warnings and orders. Although 26 million people in the US have limited English proficiency, over 80% of those messages are only delivered in English, which can cause disproportionate information distribution and awareness. In order to better serve marginalized communities during emergencies, the authors developed BEACON, a service that provides comprehensive and personalized evacuation guidance, including navigation routes, personalized checklists, and a chatbot in the language that a user uses. Our current system ingests data including fire perimeter information, evacuation order status, and shelter information from Watch Duty. When a user is within a certain proximity from the fire, the system utilizes

Source ↗
Showing 1501–1550 of 1593 signals
← Prev Page 31 of 32 Next →