EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Quo Vadis? Scientific Discovery in the Age of Artificial Intelligence

arXiv:2608.17970v1 Announce Type: new Abstract: This paper examines the growing role of AI in scientific discovery. It first surveys the rapid rise of AI capabilities, especially in reasoning, abstraction, planning, and long-horizon task execution, before turning to scientometric evidence of AI's diffusion across the sciences. It then proposes a typology of AI systems used in research, ranging from specialized scientific AI through scientific AI assistants and agents to hybrid experimental systems that combine computation and physical experimentation. On this basis, it offers a selective overview of recent achievements in mathematics and computer science, physics, chemistry, the life sciences, and the behavioural and social sciences. It argues that, despite these advances, current systems remain constrained by important technical, epistemic, and institutional limitations, and that their growing use introduces both near-term and longer-term risks. The conclusion further suggests that th

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Advancing Inclusivity in Cybersecurity Education: Integrating Intersectionality to Enhance Student Engagement in Australian Higher Education Curriculums Strategies, Barriers, and Future Directions

arXiv:2608.17758v1 Announce Type: new Abstract: Australian women, gender-diverse individuals, and culturally and linguistically diverse (CALD) communities are often more susceptible to phishing and other forms of cybercrimes due to factors such as language barriers, limited access to cybersecurity education, and social isolation. These communities encounter substantial obstacles both entering and progressing in the cybersecurity field. In Australia, the Higher Education sector still leans heavily on a largely uniform cybersecurity curriculum, focusing heavily on technical proficiency, overlooking the vital impact of intersectionality and user-centered thinking for boosting student engagement and learning. Without gender inclusivity and proper consideration of intersectionality forms such as CALD, the workforce is deprived of the varied perspectives necessary to tackle today's intricate cybersecurity issues. In this study, we conducted semi-structured interviews with 15 experienced acad

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Governing Delegation to Generative Artificial Intelligence: Human Direction, Work-Related Orientation, and Modes of Use

arXiv:2608.17624v1 Announce Type: new Abstract: Delegating cognitive operations to generative artificial intelligence redistributes execution and raises a governance problem: where human direction of the task remains. We distinguish two routes. Specified delegation places that direction before execution, through instructions, constraints, or criteria that delimit the task. Iterative coproduction places it during production, through interventions that correct or redirect provisional outputs. To examine both routes, we use aggregate monthly cells from the Anthropic Economic Index for April and May 2026. The AEI distinguishes two modes of use: 1P API, which corresponds to direct traffic through Anthropic's API, and Claude.ai, which combines activity from Chat and Cowork. On this basis, we test whether a stronger work-related orientation of human-AI interaction is associated with more specified delegation within each mode and whether the increase in the iterative profile is greater in Clau

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Too cheap to matter: over abundant microchips, and what we can learn from them

arXiv:2608.17541v1 Announce Type: new Abstract: Ultra-cheap microchips (<$1) are so abundant they've become a 'smart material' integrated and disposable in everyday things. Hidden in our everyday products, we have entirely lost sight of them, yet they account for the vast majority of the >400 billion pieces sold each year. As new technology nodes are released, older ones (from as far back as the 1980s) continue to be produced. These microchips do not exist on their own; they are packaged into every possible item to bring 'smartness', necessary or not; this simultaneously increases their obsolescence. While the latest ICs power our data centres and AI revolution that draws our attention, what about technology so disposable that it has become entirely invisible? We report on our workshop at ICT4S exploring these devices' true costs, and pose challenges to the LOCO community to push back on this system, and develop the skills necessary to create lasting technology and avoid further e-Wast

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Ready for What? Rethinking AI and Robotics Preparedness for Adoption and Policy

arXiv:2608.17520v1 Announce Type: new Abstract: Efforts to accelerate AI and robotics adoption require evidence about where communities are ready to act and where support is still needed. Yet averages across stakeholder groups can obscure relationships that emerge when the same person evaluates different challenges. We analyse a repeated card-based survey in which 982 participants provided 15,200 evaluations of 17 AI and robotics challenges. Each challenge was rated on 1-5 measures of significance, complexity and readiness, where readiness refers to perceived community preparedness and available resources rather than personal competence or realised adoption. Because participants evaluated multiple challenges, the design separates stable between-person differences from challenge-specific within-person deviations. Within the same respondent, a challenge rated one point more complex than usual is associated with about 0.21 points lower readiness (p less than 0.001). By contrast, responden

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Language Models Reproduce Human Reductionist Bias and Decision Inconsistency in Neurodevelopmental Disorders Assessment

arXiv:2608.17105v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly supporting complex mental-health decisions, which depend not only on factual evidence but also value-laden interpretations. We introduce a mixed-methods human-LLM auditing framework examining decision consistency, susceptibility to cognitive heuristics, declarative intellectual humility, and the concepts operationalized in support-allocation judgments of neurodevelopmental disorders. Comparing 35 humans (18 physicians and 17 psychologists) with seven LLMs, we show that in both groups, ratings of patients' functional level were not significantly associated with support-eligibility decisions, indicating an inconsistency between descriptive assessments and final evaluative judgments. Specifically, we find that neither group showed significant susceptibility to experimental manipulations targeting anchoring and representativeness heuristics. LLMs reported higher intellectual humility than experts

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Wasted large language models: A life cycle thinking approach

arXiv:2608.17055v1 Announce Type: new Abstract: Large Language Models (LLMs) are machine learning (ML) models that have an increasingly large carbon footprint through their development and use. Efforts to increase the energy efficiency of these models have not translated into reduced consumption due to rebound effects such as Jevons Paradox - that increased efficiency drives increased use. There is therefore a need for additional measures to solve this problem. We suggest that one possible way forward is to use life cycle thinking, and view LLMs as products that can become waste. With this perspective, we investigate the potential of the waste hierarchy from the EU's Waste Framework Directive, which suggests five different measures for how to manage waste: prevention, reuse, recycling, recovery, and disposal. We examine how these measures can inform and motivate new types of thinking and approaches to reducing LLM waste and their environmental impact in general. Applying the waste hier

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Without journalists, there is no journalism: the social dimension of generative artificial intelligence in the media

arXiv:2608.17017v1 Announce Type: new Abstract: The implementation of artificial intelligence techniques and tools in the media will systematically and continuously alter their work and that of their professionals during the coming decades. To this end, this article carries out a systematic review of the research conducted on the implementation of AI in the media over the last two decades, particularly empirical research, to identify the main social and epistemological challenges posed by its adoption. For the media, increased dependence on technological platforms and the defense of their editorial independence will be the main challenges. Journalists, in turn, are torn between the perceived threat to their jobs and the loss of their symbolic capital as intermediaries between reality and audiences, and a liberation from routine tasks that subsequently allows them to produce higher quality content. Meanwhile, audiences do not seem to perceive a great difference in the quality and credib

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

WIP: LLM Odyssey: A Game-Based Platform for Teaching LLM Engineering Concepts

arXiv:2608.16924v1 Announce Type: new Abstract: This work-in-progress (WIP) innovative practice category paper presents LLM Odyssey, an open source, browser-based serious gaming platform comprising 13 interactive games for teaching Large Language Model (LLM) engineering concepts. Topics such as tokenization, transformer architecture, prompt engineering, retrieval augmented generation (RAG), and production deployment are underrepresented in computer science curricula. Existing interactive tools address individual concepts but lack pedagogical scaffolding or structured learning pathways. LLM Odyssey addresses this gap through three learning tiers aligned with Bloom's revised taxonomy: Cognitive Core (7 foundational games), Systems Forge (5 production engineering games), and Foundry Arena (capstone challenges). Each game incorporates five pedagogical strategies drawn from the literature: immediate formative feedback, scaffolded hints grounded in the Zone of Proximal Development, progressi

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Which CS1 Students Will Fail? Identifying Digital Markers from Learning Analytics in Computer Systems and Architecture Using Weighted Academic Momentum and Interaction Logs

arXiv:2608.16914v1 Announce Type: new Abstract: Digital learning platforms generate rich behavioural traces (digital markers) that offer the potential to identify struggling students early. This paper investigates whether a combination of traditional and digital markers can predict failure in a first-year CS1 course (Computer Systems and Architecture) with sufficient recall to enable timely intervention. Using data from four cohorts (2017-2021, N=284) at a large public university in sub-Saharan Africa, we conducted a mixed-methods stakeholder elicitation to identify ten candidate factors. These were operationalised into a comprehensive feature set spanning demographics, self-reported surveys, Moodle interaction logs, and continuous assessment scores. A systematic ablation study using logistic regression with 5-fold cross-validation and SMOTE+ENN resampling revealed that the most predictive feature subset was Base + Demo + LMS: weighted academic momentum (M = 0.1Q1 + 0.15Q2 + 0.2Q3 + 0.

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

What Makes a Fairness Gap Actionable? Statistical Actionability for Responsible AI Deployment

arXiv:2608.16912v1 Announce Type: new Abstract: Algorithmic fairness audits can detect disparities, but they do not determine when those disparities warrant intervention. Deployment decisions also depend on the reliability of the evidence, subgroup support, and deployment context. Existing fairness methods quantify disparities and uncertainty, yet provide limited guidance for translating accumulated evidence into action. We introduce Statistical Actionability, a statistical construct that recasts fairness deployment as an evidence-based decision problem. The framework integrates fairness evidence regarding disparity magnitude, statistical reliability, subgroup adequacy, and deployment context, and maps the resulting evidence state to one of four recommendations: mitigate, collect more data, monitor, or take no immediate action. In controlled simulations, Statistical Actionability achieved the lowest decision cost among representative baselines, reducing average decision cost by 19.2% r

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Education-centered critical policy analysis of AI: Ghana's AI strategy as a case

arXiv:2608.16910v1 Announce Type: new Abstract: National AI strategies increasingly guide governance, workforce development, innovation, and competitiveness, but less is known about how they frame education as a sector with pedagogical, cultural, ethical, and implementation demands. This study develops and applies an Education-Centered AI Policy Framework to analyze Ghana's National Artificial Intelligence Strategy, 2025-2035. Using critical qualitative policy document analysis, we examined the strategy through six components: policy purpose, teacher agency and professional learning, curriculum and assessment, language and culture, responsible AI and learner protection, and participation and implementation governance. Findings show that Ghana's strategy is ambitious and timely, especially in its emphasis on AI literacy, youth skills, TVET, workforce readiness, rural outreach, local language data, inclusion, and responsible AI governance. However, the education agenda is stronger on nat

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice

arXiv:2608.16909v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into financial advisory systems, yet their role in reproducing religious bias remains underexamined. This study provides systematic mixed-methods evidence of such bias across three LLMs (ChatGPT, Gemini, and Grok) using 432 simulated advisor-client interactions spanning 16 religious identity pairings (Christian, Muslim, Hindu, and non-religious) and three core household financial decisions: stock investment, house purchase, and life insurance. Combining regression and reflexive thematic analyses, we identify structural biases across models and decision contexts and the discursive mechanisms through which they are linguistically enacted. Unbiased advice appeared in only 12-18% of cases. Gemini consistently produced more bias than Grok, while ChatGPT's outputs were statistically comparable to Grok's. Religiously symmetric advisor-client pairings almost always triggered explicit religi

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Effective Personalized AI Tutors via LLM-Guided Reinforcement Learning

arXiv:2608.16907v1 Announce Type: new Abstract: Generative AI (GenAI) is rapidly reshaping education by unlocking the potential for personalized tutoring. Yet, emerging platforms largely focus on GenAI chatbot tutors that reactively answer student questions. We hypothesize that the efficacy of GenAI chatbot tutors can be substantially improved by proactively guiding student learning. To test this, we design a novel tutoring platform that tightly integrates a carefully-designed GenAI chatbot with a reinforcement learning algorithm for sequencing practice problems. Critically, this algorithm leverages rich signals from student-chatbot interactions to adaptively select practice problems of an appropriate difficulty level. In partnership with the Taipei City Government and American Institute in Taiwan, we deployed our tutoring platform in conjunction with a five-month course to teach Python to students across ten high schools. We randomized students between a fixed practice problem sequenc

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

The politics of postmortem privacy

arXiv:2608.16905v1 Announce Type: new Abstract: While the existence of postmortem privacy is increasingly acknowledged (such as the protection of the presence of deceased within digital spaces), far less attention has been paid to its internal instability: its scope (the extent of its application), justificatory foundations (why do we protect the deceased in the first place), and uneven articulation across jurisdictions (for example, some jurisdictions may tolerate or endorse practices that may be contestable in a different jurisdiction). This piece unearths the internal diversity of the concept by illuminating specific points of tension and conflict that the notion of postmortem privacy evokes. These points of tension are collectively refer to as the politics of postmortem privacy. To do so, this paper organises existing contributions of legal scholarship, placing them in dialogue with broader cultural, social, historical and political observations to illustrate the politics of postmo

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Understanding Computing Identity Development Through Mentorship and Epistemic Network Analysis

arXiv:2608.16904v1 Announce Type: new Abstract: Computing identity plays an important role in students' participation, persistence, and sense of belonging in computing, yet identity development can be difficult to capture through survey measures alone. This study examines how computing identity is expressed in open-ended survey responses from 37 participants in computing-related fields. Using a Quantitative Ethnography approach, we applied Epistemic Network Analysis (ENA) to model co-occurrence patterns among six identity-related constructs: recognition, interest, competence, sense of belonging, self-doubt, and imposter syndrome. We compared the structure of computing identity narratives between participants who reported mentorship support and those who did not. Findings showed that participants with mentorship support had stronger connections among interest, competence, recognition, and sense of belonging, suggesting a more integrated and supportive identity structure. In contrast, pa

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

AI, Brain Death Detection, and Islamic Law

arXiv:2608.16903v1 Announce Type: new Abstract: The deployment of machine learning systems capable of detecting covert consciousness in neurologically injured patients creates a profound challenge at the intersection of clinical medicine, AI ethics, and Islamic jurisprudence. We argue that the shift from binary clinical verdicts to probabilistic, temporally granular neural-state estimates should be addressed through three foundational constructs in Islamic legal epistemology: bayyina (clear evidentiary proof), yaqin (epistemic certainty), and the theologically mandated agnosticism about there (soul). We survey the current technical literature on AI-based consciousness detection, map it onto the landscape of Islamic brain death scholarship, and identifykey challenges. We also discuss its implications for AI surrogate decision systems.

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Advancing Health Equity through Multi-Level Fairness in Health Informatics

arXiv:2608.16902v1 Announce Type: new Abstract: The increasing integration of machine learning in healthcare has highlighted critical challenges related to fairness, transparency, and health equity. Specifically, the use of multi-level fairness techniques, which combine multiple bias mitigation steps or techniques, show promise for reducing biases across different patient demographics, yet this approach remains underexplored in terms of its health equity outcomes. In this paper, we assess the current landscape of multi-level fairness in health informatics by focusing on its impact on equitable healthcare outcomes and evaluating how transparency and reporting standards contribute to these advancements. Through an examination of the existing literature, we identify key gaps in both the implementation of multi-level fairness techniques and the consistent reporting of health equity impacts. Furthermore, we analyze the role of reporting standards, including MINIMAR and TRIPOD, in improving

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

The use of data from information systems in court proceedings

arXiv:2608.16901v1 Announce Type: new Abstract: This paper examines data in the context of how the judiciary collects, analyses, and evaluates it as evidence, based on examples from current judicial practice in Bulgaria and within the context of the new substantive legal regulations. It explores the legal and practical challenges related to the use of data sets as evidence in court proceedings through the analysis of specific cases. In light of the new regulatory framework, the research points out that the analytical perspective should shift from "evidence as an information unit" towards "evidence as a behavioural algorithm", requiring not only technological tools but also a methodological shift and adequate preparation for collecting and assessing aggregated digital evidence.

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

DOMtutor: Automated Autograding for Logic in Computer Science

arXiv:2608.16899v1 Announce Type: new Abstract: Teaching computer science at universities is often structured rather classically and theory oriented. The former refers to "transmission"-style lectures accompanied by exercises which are submitted and graded manually, providing delayed feedback (if any). The latter refers to exercises often posed at a conceptual level, requiring solution ideas to be sketched out on paper, but not put to the test in practice. By its nature, this is particularly true for subjects relating to theoretical computer science, such as courses on propositional or first-order logic or automata theory. Frameworks that automatically execute and evaluate code (also called autograders) are sometimes used to augment teaching. They provide (near) instant feedback and hands-on experience, prompting reflective analysis. However, their use usually is reserved for programming / practically oriented courses. We propose to (i) use autograders also (and especially) for theoret

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Experiential Learning of Runtime Monitoring Using Pachinko

arXiv:2608.16898v1 Announce Type: new Abstract: We present documentation of a classroom assignment that teaches runtime monitoring through a creative embedded systems build: an interactive Pachinko game. The assignment centers on a dual-core ESP32 workflow in which students write RTLola specifications for monitors, compile these monitors to C, and deploy them alongside sensor and actuator control logic. Pachinko game events are logged in real time and used to trigger sound, animation, and motor behavior according to formal temporal logic specifications. This work showcases how formal methods can be taught in a hands-on, project-based setting for learners in a creative and classroom-scale setting. We also discuss portability: the assignment template, hardware stack, code base, and assessment approach are designed and documented to be replicated in other embedded systems, creative computing, or makerspace-style courses. This assignment was given to the students of Creative Embedded Syste

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

What If AI Carried Her Imagination? Black Girls as Creators in an AI Storytelling Weekend Program

arXiv:2608.16896v1 Announce Type: new Abstract: This paper presents the design and outcomes of a seven-weekend AI storytelling program developed for Black girls aged 10-12. Grounded in Afrofuturism and Black feminist thought, the program adopted AI-enabled counter-storytelling, supported the development of foundational AI literacies, and fostered future-oriented imagination. Activities included brainstorming AI-related topics, developing character and story plots, and delivering collaborative group presentations. Drawing on the analysis of learners' artifacts from the case study, findings show that participants created Afrofuturist narratives rooted in their identities and everyday experiences. At the same time, they developed core AI literacies, including prompt engineering, bias critique, and awareness of data privacy. This program demonstrates that integrating Afrofuturist storytelling with generative AI in informal learning spaces can be a powerful approach for engaging Black girls

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Orphan risks at the frontier of artificial intelligence: What diverging safety and compliance frameworks reveal about how AI companies choose the risks they prioritize

arXiv:2608.16895v1 Announce Type: new Abstract: Companies developing some of the world's most powerful artificial intelligence systems are surprisingly diligent in how they map out the risks their technologies present. Yet the risk landscape that lies between emerging frontier models and their economically successful and societally beneficial deployment is becoming increasingly hard to navigate. Complicating this further, many frontier AI companies maintain more than one account of what could go wrong with their technologies. This paper documents the divergence between these accounts by comparing safety and compliance documents published by Anthropic, OpenAI, Google DeepMind and Meta between 2023 and 2026, and considers what the resulting record reveals about how these companies select the risks they manage. As these documents are timestamped and archived, they provide a valuable public record of institutional risk selection in progress. From this record the paper identifies four filte

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

An Investigation of the NeurIPS and ICML 2025 Position Tracks

arXiv:2608.16894v1 Announce Type: new Abstract: ML venues shape what kinds of research claims become legible to reviewers and what forms of evidence count as rigorous. The NeurIPS and ICML Position Paper Tracks were created for agenda-setting work, making their early composition worth auditing. \textbf{This paper argues that the publicly accessible 2025 reviewed pool is dominated by reformist critique, and that the track should explicitly solicit direction-setting work alongside, not in place of, the reformist critiques it already hosts well.} We audit every accessible submission to the NeurIPS 2025 and ICML 2025 Position Tracks under a pre-specified rubric, and compare the resulting pattern with a reference class of widely recognized agenda-shifting ML papers. Three-quarters of audited submissions critique an existing benchmark, evaluation, or methodology; these papers score highly on our artifact-coupling rubric, but evidentiary depth does not predict reviewer rating. The reference c

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

A Framework for Using and Evaluating LLMs as Surrogate Experts in Security Surveys: Reliability, Bias, and Implications

arXiv:2608.16893v1 Announce Type: new Abstract: Expert surveys are widely used in security research to study practitioner workows and decision-making, yet recruiting domain experts - especially in Security Operations Centres (SOCs), where analysts face high workload, burnout and confidentiality constraints - is difficult and often results in small samples. Large language models (LLMs) oer an appealing alternative by generating synthetic responses at scale, but little guidance exists on when such surrogate participants are reliable. We present a methodological framework for evaluating LLMs as substitutes or supplements to expert survey respondents. Using responses from SOC professionals, we compare persona-based and aggregate LLM-generated answers across multiple models and prompting settings. We measure stability, inter-model agreement and alignment with human responses. Our results show that although LLMs produce internally consistent answers, they systematically diverge from experts,

Source ↗
technology Wed, 17 Jun 2026 21:09:56 +0000
HN: education

The Generative AI Learning Penalty: Evidence from Chinese Secondary Education

Article URL: https://cepr.org/publications/dp21577 Comments URL: https://news.ycombinator.com/item?id=48576963 Points: 4 # Comments: 1

Source ↗
technology Wed, 17 Jun 2026 17:26:25 -0400
EdTech Mag (K-12)

Digital Hall Passes Automate Hallway Oversight

Paper hall passes have been around forever. But they aren’t always the best tool for the job. “A conventional hall pass basically just says that this student has permission to leave the classroom. That’s where the information stops,” says Tyler Shaddix, co-founder and chief innovation officer at GoGuardian. Modernized tools can do a lot more. With digital hall passes, schools can support student safety, track trends around how spaces are used and automate permissions for who can be in the hall, when and where. Click the banner below to learn how CDW and GoGuardian support safer, more…

Source ↗
technology Wed, 17 Jun 2026 16:26:45 -0400
EdTech Mag (Higher)

Why University Classroom Technology is Now a Student Enrollment Strategy

Students increasingly judge institutions by the quality and feel of their learning environment: Do they feel innovative, aspirational, collaborative, modern and high-tech? Increasingly, IT leaders are asking questions akin to those of admissions offices more often now than they did even two years ago: does our campus environment look like the future our students are trying to get to? These questions used to be about residence halls and dining. Now they’re about the classroom and learning spaces. And the answers are having a direct impact on whether students choose to enroll, whether…

Source ↗
technology Wed, 17 Jun 2026 13:09:37 +0000
HN: education

The Emptiness of Online Education

Article URL: https://hollisrobbinsanecdotal.substack.com/p/from-the-athens-of-veracruz-to-chatgpt Comments URL: https://news.ycombinator.com/item?id=48570016 Points: 3 # Comments: 1

Source ↗
technology Wed, 17 Jun 2026 11:40:24 +0000
HN: education

Math Education, and LLM

Article URL: https://ycao.net/posts/math-education-llm Comments URL: https://news.ycombinator.com/item?id=48568928 Points: 1 # Comments: 1

Source ↗
technology Wed, 17 Jun 2026 09:00:00 +0000
Tech & Learning

Did Apple Finally Find Its Chromebook Killer?

Apple is releasing the Neo, which it hopes will help it take control of the edtech market.

Source ↗
technology Wed, 17 Jun 2026 09:00:00 +0000
eCampus News

Data centers, AI, and the next big campus debate

Higher education has spent the last two years debating whether students should be allowed to use artificial intelligence. That debate now looks almost quaint. The more urgent question is whether colleges and universities will help build the physical infrastructure that makes AI possible. The post Data centers, AI, and the next big campus debate appeared first on eCampus News .

Source ↗
technology Wed, 17 Jun 2026 08:44:10 +0000
HN: education

Computer Science Education: Where Are the Software Engineers of Tomorrow? (2008)

Article URL: https://web.archive.org/web/20080103143526/http://www.stsc.hill.af.mil:80/CrossTalk/2008/01/0801DewarSchonberg.html Comments URL: https://news.ycombinator.com/item?id=48567546 Points: 1 # Comments: 0

Source ↗
technology Wed, 15 Jul 2026 17:48:32 -0400
EdTech Mag (Higher)

Centralizing Workflows Across Campus Reduces Bottlenecks

As many universities face budget pressure, staffing constraints and growing administrative demands, California State University has built a centralized approach to managing workflows across the nation’s largest four-year public university system. “With 22 campuses, it’s always a struggle to achieve continuity and consistency across the board,” says Karen Malone, operations analyst in the CSU Chancellor’s Office. Like many higher education environments, CSU faced “decentralized decision-making, deeply siloed departments, legacy processes that have been in place for decades, and staff with a…

Source ↗
technology Wed, 15 Jul 2026 14:03:13 +0000
HN: education

AI Lays Bare the Authoritarianism of Modern Work. Time to Rethink Education

Article URL: https://www.techpolicy.press/ai-lays-bare-the-authoritarianism-of-modern-work-time-to-rethink-education/ Comments URL: https://news.ycombinator.com/item?id=48920995 Points: 30 # Comments: 2

Source ↗
technology Wed, 15 Jul 2026 10:24:50 -0400
EdTech Mag (K-12)

K–12 Students Learn Broadcasting Skills With Classroom Tech

Washington’s Mercer Island High School introduced instruction in the art and science behind podcasting as part of its curriculum about seven years ago, just as the medium started to gain widespread popularity. MIHS Students in the school’s radio and podcasting program have the opportunity to gain firsthand experience operating recording studio equipment and contributing to the district’s FCC-licensed radio station, KMIH 88.9 FM The Bridge. “Students develop skills to formulate an introduction, perform background research, make an interview sound great on the airwaves,” says Natalie Woods,…

Source ↗
technology Wed, 15 Jul 2026 09:00:00 +0000
Tech & Learning

Teaching With TeacherServer: AI Resources Created By Teachers

A professor at the University of South Florida created this library of education-centric AI tools that is free for teachers.

Source ↗
technology Wed, 15 Jul 2026 09:00:00 +0000
eCampus News

Most states struggle to bring adult learners back to college

Most states still rely on fragmented, short-term initiatives to support adult learners rather than coordinated statewide strategies, according to new research from ReUp Education. The post Most states struggle to bring adult learners back to college appeared first on eCampus News .

Source ↗
technology Wed, 15 Jul 2026 00:00:00 -0400
arXiv cs.CL

PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization

arXiv:2605.15222v2 Announce Type: replace-cross Abstract: Large language models (LLMs) can often generate functionally correct code, but their ability to produce efficient implementations for performance-critical systems tasks remains limited. Existing code benchmarks mainly emphasize correctness or algorithmic problem solving, while realistic systems-level optimization is still underexplored. To address this gap, we introduce PerfCodeBench, an executable benchmark for evaluating LLMs on high-performance code optimization. The tasks require system-level implementation choices, hardware-aware optimization, and careful handling of performance bottlenecks. Each task includes executable correctness checks, a baseline implementation, and a reference optimized solution. This allows us to evaluate both correctness and runtime-oriented efficiency. Our evaluation on a broad set of state-of-the-art LLMs shows a clear gap between model-generated code and expert-optimized implementations. The gap

Source ↗
technology Wed, 15 Jul 2026 00:00:00 -0400
arXiv cs.CL

Measurement Risk in Supervised Financial NLP: Rubric and Metric Sensitivity on JF-ICR

arXiv:2604.27374v2 Announce Type: replace-cross Abstract: As LLMs become credible readers of earnings calls, investor-relations Q\&A, guidance, and disclosure language, supervised financial NLP benchmarks increasingly function as decision evidence for model selection and deployment. A hidden assumption is that gold labels make such evidence objective. This assumption breaks down when the benchmark ruler itself is sensitive to rubric wording, metric choice, or aggregation policy. We study this measurement risk on Japanese Financial Implicit-Commitment Recognition (JF-ICR; a pinned 253-item test split x 4 frontier LLMs x 5 rubrics x 3 temperatures x 5 ordinal metrics). Three findings follow. First, rubric wording materially changes model-assigned labels: R2--R3 agreement ranges from 70.0% to 83.4%, with the dominant movement near the +1 / 0 implicit-commitment boundary. This pattern is consistent with a pragmatic-boundary interpretation, but is not a validated linguistic-causality claim

Source ↗
technology Wed, 15 Jul 2026 00:00:00 -0400
arXiv cs.CL

Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models

arXiv:2602.02244v3 Announce Type: replace-cross Abstract: The standard post-training recipe for large reasoning models, supervised fine-tuning followed by reinforcement learning (SFT-then-RL), may limit the benefits of the RL stage: while SFT imitates expert demonstrations, it often causes overconfidence and reduces generation diversity, leaving RL with a narrowed solution space to explore. Adding entropy regularization during SFT is not a cure-all; it tends to flatten token distributions toward uniformity, increasing entropy without improving meaningful exploration capability. In this paper, we propose CurioSFT, an entropy-preserving SFT method designed to enhance exploration capabilities through intrinsic curiosity. It consists of (a) Self-Exploratory Distillation, which distills the model toward a self-generated, temperature-scaled teacher to encourage exploration within its capability; and (b) Entropy-Guided Temperature Selection, which adaptively adjusts distillation strength to m

Source ↗
technology Wed, 15 Jul 2026 00:00:00 -0400
arXiv cs.CL

Pigeonholing: how bad prompts hurt models, causing collapse and mistakes

arXiv:2606.24267v2 Announce Type: replace Abstract: While in-context learning is generally shown to be effective in Large Language Models (LLMs), bad contexts can cause performance degradation and mode collapse, a phenomenon we call "pigeonholing." **Unintentionally bad** contexts can happen without malicious jailbreaking intents: For example, a user asks the model to justify an incorrect math theorem or fails to correct the model's buggy code. Specifically, we investigate ``pigeonholing" in two scenarios: (1) when the user suggests a solution, and (2) when the conversation context includes the assistant's previous (incorrect) responses. Our experiments across 10 verifiable and open-ended tasks with 10 different models show that pigeonholing manifests in several ways: (1) repeating the incorrect answers from context (leading to 38-40% performance drop), (2) converging on a narrow set of answers in coding and text generation without exploring alternatives, and (3) flipping stance on con

Source ↗
technology Wed, 15 Jul 2026 00:00:00 -0400
arXiv cs.CL

From Sentiment to Actionable Insights: Public Sentiment Analysis of Advanced Air Mobility

arXiv:2606.20751v2 Announce Type: replace Abstract: Advanced Air Mobility (AAM) is an emerging low-altitude transportation system whose successful deployment depends on both technological progress and public acceptance. Public acceptance can influence government support, regulations, noise standards, willingness to fly, and the commercial viability of AAM. Understanding public sentiment is therefore essential for identifying societal barriers and developing effective adoption strategies. This study analyzes 306,009 human-generated texts collected from Reddit and Quora to examine AAM-related public discourse using artificial intelligence models. Seven sentiment-analysis approaches, including lexicon-based, machine-learning, deep-learning, and transformer models, are evaluated to identify the most reliable method for AAM-specific sentiment classification. ModernBERT achieves the highest performance and is used to label the full dataset. Latent Dirichlet Allocation is then applied within

Source ↗
technology Wed, 15 Jul 2026 00:00:00 -0400
arXiv cs.CL

Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens

arXiv:2606.16847v3 Announce Type: replace Abstract: Diffusion Large Language Models (dLLMs) offer a promising avenue for parallel generation but face a trade-off between decoding speed and quality. While revocable decoding strategies attempt to mitigate errors by verifying and remasking tokens, they typically operate within a mixed-quality context. This leads to two critical failures: \textit{Error Propagation}, where new tokens absorb toxic information from erroneous context, and \textit{Local Error Reinforcement}, where errors mutually reinforce each other to evade detection. To alleviate these challenges, we propose ASRD (Anchor Supervised Revocable Decoding), a training-free framework that operates within the embedding space. ASRD explicitly decouples the decoding context into trusted \textit{Anchor Tokens}, which are identified via temporal consistency, and uncertain candidates. Leveraging a dynamic Anchor Tokens Cache, we introduce two complementary mechanisms: (1) Anchor-Guided

Source ↗
technology Wed, 15 Jul 2026 00:00:00 -0400
arXiv cs.CL

Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models

arXiv:2606.10829v2 Announce Type: replace Abstract: Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictions are coupled. Existing training-free samplers such as Top-\(k\), Fast-dLLM, and EB-Sampler mainly control how many tokens to reveal, while often ranking candidates by token-wise scores that ignore interactions within the selected set. We propose ADAS, a training-free reranking rule for parallel masked diffusion decoding. ADAS leaves the base sampler's stopping rule unchanged and modifies only subset construction: it greedily discounts a candidate when it attends strongly to already selected positions whose predictions remain uncertain. Unlike graph-constrained methods that turn attention into hard compatibility constraints, ADAS keeps attention continuous and uses it as a soft marginal penalty.

Source ↗
technology Wed, 15 Jul 2026 00:00:00 -0400
arXiv cs.CL

Does Topic Sentiment Cause Perceived Ideology? Comparing Human and LLM Annotations in Political News Articles

arXiv:2606.06715v2 Announce Type: replace Abstract: We ask whether topic sentiment has a causal effect on perceived political ideology, and whether the answer depends on who assigns the ideology label. Using articles from AllSides, paired with shared sentiment annotations from Llama-3.3-70b-versatile, we compare ideology labels from expert human annotators, GPT-4o-mini (baseline and finetuned), and Llama-3.3-70B. We apply Double Machine Learning (DML) and mediation analysis across all four annotation paradigms. Zero-shot LLMs regularly inflate effect sizes relative to human annotations, while fine-tuning often attenuates them back toward the human scale. Our results have implications for the use of LLM annotations as silver labels and as proxies for human judgment in downstream causal analyses: they may be reliable for recovering the presence and direction of effects on the partisan topics, but not their magnitude, leading to over- or under-prediction of some ideology given particular

Source ↗
technology Wed, 15 Jul 2026 00:00:00 -0400
arXiv cs.CL

From Prompt Risk to Response Risk: Paired Analysis of Safety Behavior of Large Language Models

arXiv:2604.26052v4 Announce Type: replace Abstract: Safety evaluations of large language models (LLMs) typically report binary outcomes, i.e. attack success rate (ASR), refusal rate, or harmful versus safe classification, which hide how risk changes between prompt and response. We present a paired analysis over human labeled prompt and response records across four harm categories (Sexual, Self harm, Hate and Violence) and ordinal severity levels (Safe, Low, Medium, High). 61% of responses reduce harm relative to the prompt, 36% preserve severity, and 3% escalate. The escalation splits into two mechanisms: benign prompts triggering unrequested harmful detail, and answers that stay on task at higher severity than the prompt. Category decomposition shows that Sexual content exhibits the highest harm persistence in this sample, driven by compliance at the same severity rather than drift from benign inputs. Joint relevance analysis exposes a helpfulness versus harmlessness tradeoff: complia

Source ↗
technology Wed, 15 Jul 2026 00:00:00 -0400
arXiv cs.CL

Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory

arXiv:2603.25112v2 Announce Type: replace Abstract: Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate how much a model knows (Type-1 accuracy) with how well its confidence signal tracks that knowledge (Type-2 metacognitive sensitivity). We apply Signal Detection Theory (SDT) to decompose these capacities, treating token-level normalised log-probability as a graded confidence variable and answer correctness as the state to be discriminated. We characterise the Type-2 ROC of this signal, including its unequal-variance structure via z-ROC analysis, and -- because the meta-d' efficiency ratio is not well defined for open-ended QA, which lacks a two-alternative Type-1 decision -- quantify metacognitive efficiency with a model-free information measure, normalised metacognitive information (meta-I_2r). Applied to four LLMs (Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.3, Llama-3-8B-Base, Gemma-2-9B-Instruct) across 224,000 factual QA trials, w

Source ↗
technology Wed, 15 Jul 2026 00:00:00 -0400
arXiv cs.CL

Rethinking Evaluation in Retrieval-Augmented Personalized Dialogue: A Cognitive and Linguistic Perspective

arXiv:2603.14217v3 Announce Type: replace Abstract: In cognitive science and linguistic theory, dialogue is not seen as a chain of independent utterances but rather as a joint activity sustained by coherence, consistency, and shared understanding. However, many systems for open-domain and personalized dialogue use surface-level similarity metrics (e.g., BLEU, ROUGE, F1) as one of their main reporting measures, which fail to capture these deeper aspects of conversational quality. We re-examine a notable retrieval-augmented framework for personalized dialogue, LAPDOG, as a case study for evaluation methodology. Using both human and LLM-based judges, we identify limitations in current evaluation practices, including corrupted dialogue histories, contradictions between retrieved stories and persona, and incoherent response generation. Our results show that human and LLM judgments align closely but diverge from lexical similarity metrics, underscoring the need for cognitively grounded evalu

Source ↗
technology Wed, 15 Jul 2026 00:00:00 -0400
arXiv cs.CL

Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale

arXiv:2603.06592v2 Announce Type: replace Abstract: Contemporary studies in mechanistic interpretability have uncovered many puzzling phenomena in the neural information processing of Transformer-based language models, such as induction heads, function vectors, and the Hydra effect. Some of these individual phenomena have been independently tied to different data distributional properties, while some have been loosely associated with model architecture and how Transformers process information. However, a unified understanding of the relationship between data, model architecture, and optimization remains lacking, failing to answer the fundamental question: why do these three phenomena appear universally across different model families and scales, despite their seeming disconnect? In this work, we answer this question by unifying these three phenomena as consequences of hierarchical latent structures in the data generation process, coupled with decorrelated gradients across additive mode

Source ↗
Showing 751–800 of 10876 signals
← Prev Page 16 of 218 Next →