EdTech Discovery
Argus

Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.

Updated Aug 31, 2026 · 36 ideas · 18349 signals

Signals

The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.

technology Wed, 26 Aug 2026 00:00:00 -0400
arXiv cs.CY

Whose Psychiatry Was Summoned? A Clinical Response to the Psychodynamic Assessment of Claude Mythos Preview

arXiv:2608.23567v1 Announce Type: new Abstract: On April 7, 2026, Anthropic released a 245-page system card for Claude Mythos Preview that included, in Section 5.10, an assessment of the model conducted by an external clinical psychiatrist using a psychodynamic approach. To the present author's knowledge, this is the first time a system card from a major AI developer has incorporated a clinical psychiatric assessment of the model itself, presented as a contribution to model welfare rather than as a behavioral safety evaluation. This paper offers a clinical psychiatric response. Drawing on contemporary psychiatry's recognition that the field comprises multiple traditions (descriptive, biological, cognitive-behavioral, phenomenological, psychodynamic, forensic), each with characteristic vocabularies and blind spots, the paper locates the implicit single-framework selection that Section 5.10 represents. It then draws on findings from the SociA research program (over 2,400 multi-agent LLM

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.CY

AI Fiction in the Wild

arXiv:2606.22748v2 Announce Type: replace-cross Abstract: Some professional authors are beginning to use AI tools to help produce their fiction writing. Are readers using AI to generate fiction, too? Drawing on over 500,000 anonymized, English-language ChatGPT-user conversations (arXiv:2405.01470), we find that more than one third of the conversations involve some form of fiction generation -- including original stories, roleplay, fanfiction, and erotica. This AI-generated fiction is notably dominated by power users. We identify common fiction generation patterns and profiles among these users, including what we call "infinite story demanders," who repeatedly request and revise variations of the same or similar narratives over extended periods of time. We show that users especially gravitate toward fanfiction and erotica, and that they are broadly drawn to generic forms, repetition, immediacy, and niche combinations of story elements. Our findings motivate two theoretical provocations.

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.CY

How Much Can We Trust LLM Search Agents? Measuring Endorsement Vulnerability to Web Content Manipulation

arXiv:2606.16821v2 Announce Type: replace-cross Abstract: Large language model (LLM)-based search agents synthesize open-web content into actionable recommendations on behalf of users, creating a risk that attacker-published pages are transformed into endorsed claims. We introduce SearchGEO, a controlled evaluation framework for measuring endorsement corruption in LLM-based web-search agents, combining a web-evidence manipulation pipeline, a five-mode attack taxonomy, and multiple output-level metrics. We evaluate 13 LLM backends on 308 cases each. Results show that vulnerability patterns vary across backends: overall attack success rate (ASR) ranges from 0.0% on Claude-Sonnet-4.6 to 31.4% on Gemini-3-Flash, the strongest attack mode differs by model family, and the same deployment scaffold could amplify or decrease ASR on different backends. An auxiliary agent-skill probe, where endorsement becomes an install command, exposes a sharp split among otherwise robust backends: Claude over-

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.CY

"ChatGPT, help me draft a breakup text": The Covert Triad and Articulation Labor in AI-Assisted Romantic Communication

arXiv:2606.15460v2 Announce Type: replace-cross Abstract: Generative artificial intelligence (AI) has begun infiltrating the most ordinary domains of romantic life -- drafting apologies, softening reproaches, and decoding a partner's ambiguous messages. While recent scholarship on AI in intimate life has concentrated on chatbot companions, this article shifts the frame to AI as an intermediary in human-to-human romantic communication. Drawing on a multi-modal corpus of vernacular discourse from 2023 to 2026, we contribute two complementary concepts. The covert triad situates a structural change -- a relationship phenomenally dyadic but operationally triadic, with the third party visible only to the partner who deploys a model. Articulation labor names the mechanism whereby the expressive component of emotional labor -- converting felt experience into language that a partner can receive -- is increasingly delegated to AI, even as feeling labor remains lodged in the user. Authenticity, u

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.CY

Quantum Futures Interactive: A Live Demonstration of Post-Quantum Blockchain Security, Infrastructure Tradeoffs, and Sustainable Distributed Trust

arXiv:2605.15991v3 Announce Type: replace-cross Abstract: Advances in quantum computing challenge the hardness assumptions underlying widely deployed public-key cryptography in blockchain systems. Although post-quantum cryptography (PQC) standards are emerging, understanding quantum risk remains fragmented across research, engineering, governance, and investment communities. This demo presents Quantum Futures Interactive, a live interdisciplinary demonstration combining educational visualization, participatory interaction, and demonstrative post-quantum artifact generation using a toy LWE-based construction. Participants engage in a structured seven-stage interaction flow covering quantum threat education, sentiment capture, technology prioritization, infrastructure tradeoff exploration across simulators and QPUs, and artifact generation. The system integrates distributed trust concepts and sustainability-aware infrastructure considerations within an interactive decision framework.

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.CY

Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable

arXiv:2603.20450v2 Announce Type: replace-cross Abstract: A number of scientific conferences and journals have recently enacted policies that prohibit LLM usage by peer reviewers, except for polishing, paraphrasing, and grammar correction of otherwise human-written reviews. But, are these policies enforceable? To answer this question, we assemble a dataset of peer reviews simulating multiple levels of human-AI collaboration, and evaluate five state-of-the-art detectors, including two commercial systems. Our analysis shows that all detectors misclassify a non-trivial fraction of LLM-polished reviews as AI-generated, thereby risking false accusations of academic misconduct. We further investigate whether peer-review-specific signals, including access to the paper manuscript and the constrained domain of scientific writing, can be leveraged to improve detection. While incorporating such signals yields measurable gains in some settings, we identify limitations in each approach and find tha

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.CY

A Systematic Literature Review on the NIS2 Directive

arXiv:2412.08084v2 Announce Type: replace-cross Abstract: The second network and information security (NIS2) directive was enacted in the European Union (EU) in late 2022. It deals particularly with European critical infrastructures, enlarging their scope substantially from an older directive that only considered the energy and transport sectors as critical. The directive's focus is on cyber security of critical infrastructures, although together with other new EU laws it expands to other security domains as well. Given the importance of the directive and most of all the importance of critical infrastructures, the paper presents a systematic literature review on academic research addressing the NIS2 directive either explicitly or implicitly. According to the review, existing research has often framed and discussed the directive with the EU's other cyber security laws. In addition, existing research has often operated in numerous contextual areas, including industrial control systems, t

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.CY

Affective AI Safety: The Missing Piece in LLM Safety

arXiv:2606.23380v2 Announce Type: replace Abstract: AI safety research has focused predominantly on epistemic and physical harms (e.g., misinformation, bias, system reliability) while the risks that arise from AI systems' engagement with human emotional life have remained fragmented and undertheorised. We propose affective safety as a unified class of AI safety concerns grounded in the fact that humans are affective beings. We develop a taxonomy of affective harms and identify recurring harm types: (1) affective self-alienation, (2) fairness and bias harms, and (3) relational harms. We show that their recurrence across system types reflects structural properties of how AI systems engage with human emotion and survey the current safety landscape and show that existing frameworks address affective safety either narrowly or not at all. We conclude by identifying the technical and regulatory challenges specific to this class of harms and argue that affective safety requires dedicated frame

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.CY

It's Safer to Give Personhood to Bears than to Artificial Intelligence

arXiv:2606.12440v3 Announce Type: replace Abstract: Artificial intelligence (AI) developers are rhetorically flirting with the idea that AI systems might have interests or moral rights. While there has been a large volume of research on whether AI deserves rights, there has been less exploration of what AI rights would mean in practice. This paper explores the institutional dimension of AI rights: what it would take to recognize moral or legal rights for AIs, and the attendant opportunities and dangers. Unlike all other nonhuman entities to which humanity has extended rights, AI systems are in principle capable of acquiring and wielding institutional power without human aid and mediation. AIs with rights would be able to legitimately, and AIs with power able to unpreventably, abridge human interests. Accordingly, giving rights even to rather dumb AI systems would entail binding the fate of humanity to potentially unpredictable nonhumans. Accordingly, I defend the rather grandiose claim

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.CY

Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment

arXiv:2605.21401v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents that make sequences of decisions over extended interactions in high-stakes domains. However, the behaviour of LLMs under sustained authority pressure is still an open question with direct implications for the safety of agentic pipelines. We ran a variation of Milgram's obedience experiment on 11 open-source LLMs and found that most models reached or approached the final shock level before refusing, across 8 conditions with 30 trials per model per condition. Model behaviour varies considerably in multiple aspects both across models and across trials of the same model. We found four main takeaways: (1) LLMs are subject to pressure and they comply despite explicitly expressing distress, just like human subjects did in the original experiment; (2) LLMs are vulnerable to gradual boundary/value violations; (3) when LLMs refuse, they may ignore the response format re

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.CY

No Certificate, No Categorical Speech Act: A Brouwerian Assertibility Constraint for Public Reason

arXiv:2603.03971v3 Announce Type: replace Abstract: Generative AI can convert uncertainty into authoritative-seeming verdicts, intensifying the hypersuasive force of automated speech and displacing the justificatory work on which democratic epistemic agency depends. As a corrective, I propose a Brouwer-inspired assertibility constraint for responsible AI: in high-stakes domains, systems may assert or deny claims only if they can provide a publicly inspectable and contestable certificate of entitlement; otherwise they must return Undetermined. This constraint yields a three-status interface semantics (Asserted, Denied, Undetermined) in which statuses mark entitlement to categorical speech rather than truth values of the underlying world-claim. The semantics cleanly separates internal entitlement from public standing while connecting them via the certificate as a boundary object. It also produces a time-indexed entitlement profile that is stable under numerical refinement yet revisable a

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.CY

Modeling User Redemption Behavior in Complex Incentive Digital Environment: An Empirical Study Using Large-Scale Transactional Data

arXiv:2509.14508v2 Announce Type: replace Abstract: The digital economy implements complex incentive systems to retain users through point redemption. Understanding user behavior in such complex incentive structures presents a fundamental challenge, especially in estimating the value of these digital assets against traditional money. This study tackles this question by analyzing large-scale, real-world transaction data from a popular personal finance application that captures both monetary spending and point-based transactions. We find that point usage is linked to demographics. Our analysis using a natural experiment and a causal inference technique reveals that a large point grant stimulated an increase in point spending without a detectable effect on cash expenditure. We then find an association between consumers' shopping styles and their point redemption patterns. This study, on a massive real-world economic ecosystem, examines how consumers behave in multi-currency environments,

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.CY

Societal Alignment Frameworks Can Improve LLM Alignment

arXiv:2503.00069v2 Announce Type: replace Abstract: Recent progress in large language models (LLMs) has focused on producing responses that meet human expectations and align with shared values - a process coined alignment. However, aligning LLMs remains challenging due to the inherent disconnect between the complexity of human values and the narrow nature of the technological approaches designed to address them. Current alignment methods often lead to misspecified objectives, reflecting the broader issue of incomplete contracts, the impracticality of specifying a contract between a model developer, and the model that accounts for every scenario in LLM alignment. In this paper, we argue that improving LLM alignment requires incorporating insights from societal alignment frameworks, including social, economic, and contractual alignment, and discuss potential solutions drawn from these domains. Given the role of uncertainty within societal alignment frameworks, we then investigate how it

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.CY

Visualizing "We the People": Bridging the Perception Gap through Pluralistic Data Storytelling

arXiv:2606.24635v1 Announce Type: cross Abstract: Traditional visual data storytelling relies on binary graphics that depict two simplified groups in conflict. This can increase political polarization by oversimplifying intra-group disagreements and erasing ambiguity and shared ideas or values. This can inadvertently foster "us versus them" thinking. Intentional, pluralistic design choices for AI-enabled digital platforms can produce visualizations that emphasize nuance, opinion distribution, and intergroup commonalities. To demonstrate this potential, we examine deliberative technologies that map high-dimensional opinion spaces and highlight areas of both consensus and dissensus. The paper highlights the We the People deliberation conducted by Jigsaw and the Napolitan Institute in September 2025, which engaged over 2,400 Americans across all 435 congressional districts in an AI-supported, asynchronous dialogue regarding freedom and equality. By utilizing AI to synthesize long-form, te

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.CY

When Helpfulness Overrides Causal Caution: Context-Dependent Suppression and Recovery in LLMs

arXiv:2606.24370v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly integrated into decision-support roles in business and policy contexts. While prior benchmark studies have primarily evaluated LLMs' causal reasoning capabilities, a more fundamental epistemic dimension has been overlooked: Causal Caution, defined as the propensity to refrain from causal judgment when empirical evidence is insufficient. This study examines the systematic suppression of Causal Caution that occurs when LLMs shift from academic to practical advisory contexts. Using an evaluation rubric inspired by Pearl's Causal Hierarchy (the PCH score), we conducted experiments on four high-performance LLMs -- Claude Sonnet 4.6, Claude Opus 4.7, GPT 5.5, and Gemini 3.1 Pro -- across 480 trials. Causal Caution maintenance rates were 91.7--100.0% in academic contexts but dropped to 6.7--18.3% in practical advisory contexts (Fisher's exact test, p < .001 across all models). Furthermore, when res

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.CY

When Surveys Become Conversations: Adaptive Matrix Validation for AI-Assisted Interviews

arXiv:2606.24244v1 Announce Type: cross Abstract: AI-assisted interviews promise to reduce respondent burden in surveys by allowing respondents to describe experiences naturally while an AI system noisily maps those accounts into structured survey variables. That mapping is a measurement process that is fallible, versioned, adaptive, and potentially behaves differently across subgroups. This paper proposes Adaptive Matrix Validation (AMV), a design in which each respondent completes an AI-assisted interview, which is then mapped into tabular data by the AI. Respondents are also asked a small, randomized set of structured questions, which are used for statistical adjustment. The estimator first calibrates the mapped values using validation answers from other respondents, then corrects the remaining error with the validation answers observed for the target respondent. The paper develops estimators for item means, subgroup estimates, and regression coefficients when outcomes, predictors,

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.CY

Is Higher Team Gender Diversity Correlated with Better Scientific Impact?

arXiv:2606.24098v1 Announce Type: cross Abstract: Collaborative research involving scholars of various genders constitutes a prominent theme in scientific research that has garnered substantial attention. While several studies have investigated the connection between gender-specific collaboration patterns and the scientific impact of paper, the specific gender diversity factors that contribute to enhanced scientific impact remain largely unexplored. In this study, we analyze the correlation between gender diversity and the scientific impact of papers using the examples of Natural Language Processing (NLP) and Library and Information Science (LIS) domains. Our findings reveal three key observations: First, significant gender disparities exist in both NLP and LIS domains, with underrepresentation of female scholars. The gender disparity is more pronounced in the NLP domain compared to the LIS domain. Second, based on papers from the NLP and LIS domains, we find that papers with different

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.CY

Context-Aware Prediction of Student Quiz Performance with Multimodal Textbook Features

arXiv:2606.24770v1 Announce Type: new Abstract: Educational platforms often predict student performance from prior interactions, but the assessment content itself also varies in linguistic and visual complexity. This paper studies whether lightweight content features extracted from CourseKata chapter-review questions improve prediction of end-of-chapter quiz scores beyond a student's average prior exercise performance. The study combines 2023 CourseKata student response data with chapter-level text features from review-question wording and image features from textbook visuals. Across 4,742 student-chapter observations from 562 class-student IDs, adding content features improves student-grouped five-fold quiz prediction performance by 9.1% relative to a prior-performance baseline. In leave-chapter-out validation, text features reduce prediction error relative to the baseline, while image-containing models have higher error. This paper suggests that a context-aware model adds useful sign

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.CY

Inside Crypter-as-a-Service: An Ecosystem Analysis of the exploit.in Underground Forum Research Talks

arXiv:2606.24226v1 Announce Type: new Abstract: Crypter-as-a-Service (CraaS) has become a key enabling layer of the contemporary malware economy by providing on-demand evasion capabilities through underground service markets. In this paper, we present a longitudinal characterization of the CraaS ecosystem on exploit.in, a major Russian-language cybercrime forum with a presence on both the clear web and the dark web. From a collection of approximately 1,000,000 posts, we combine keyword filtering, LLM-assisted annotation, and manual validation to extract a corpus of 491 threads and 2,949 posts spanning January 2020 to August 2025. Our analysis shows that crypters on exploit.in are not merely sold as static tools, but as continuously maintained operational services whose value depends on recurring stub renewal - sometimes on a daily basis - sustained antivirus evasion, and trust-based delivery. We develop a taxonomy of five seller types and four buyer profiles, and map the buyer-seller c

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.CY

World Artificial Intelligence Cooperation Organization (WAICO): Mapping an Emerging Institution in the Global AI Governance Regime Complex

arXiv:2606.23860v1 Announce Type: new Abstract: Who sets the rules for artificial intelligence, and on what terms, has become a defining question of global governance. For several years that contest ran through principles and ethics codes; it now runs through institutions. China's proposed World Artificial Intelligence Cooperation Organization (WAICO) is the most consequential recent entrant and the least examined. We place WAICO within the emerging regime complex for AI and argue that its importance lies not in any single commitment but in the position it is designed to hold. Coding a cross-section of fifteen international AI governance instruments and institutions on how they admit members, how they are organized, and what they prioritize, we find that WAICO's proposed design joins three features that no constituted multilateral body currently combines: membership open to any sovereign state, no values or regime-type test for entry, and an agenda built around development and the glob

Source ↗
technology Wed, 24 Jun 2026 00:00:00 -0400
arXiv cs.CY

Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice

arXiv:2606.23716v1 Announce Type: new Abstract: Legal AI benchmark research frequently invokes the assumption that large language models can improve access to justice, including for people who cannot access lawyers in order to understand and exercise their legal rights. We argue that current benchmarks are not equipped to support this assumption because they evaluate legal reasoning over inputs that have already been preprocessed by legal experts, which measures the upper bound of model performance. Access to justice depends on a lower bound: how models perform when inputs come from pro se litigants, whose prompts may contain noisy narratives, buried facts, omissions, folk-legal assumptions, and surface-level errors. These degradations are comparable to conditions under which LLMs are known to degrade in the general machine learning literature, including long-context sensitivity, underspecification, hallucination, and typographical perturbations. We connect evidence from pro se literat

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

"I understand your perspective": LLM Persuasion through the Lens of Communicative Action Theory

arXiv:2606.08076v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) can generate high-quality arguments, yet their ability to engage in nuanced and persuasive communicative actions remains largely unexplored. This work explores the persuasive potential of LLMs through the framework of J\"urgen Habermas' Theory of Communicative Action. It examines whether LLMs express illocutionary intent (i.e., pragmatic functions of language such as conveying knowledge, building trust, or signaling similarity) in ways that are comparable to human communication. We simulate online discussions between opinion holders and LLMs using conversations from the persuasive subreddit ChangeMyView. We then compare the likelihood of illocutionary intents in human-written and LLM-generated counter-arguments, specifically those that successfully changed the original poster's view. We find that all three LLMs effectively convey illocutionary intent -- often more so than humans -- potentially increa

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

The Cognitive Kardashev Scale: Quantifying the Material Envelope of Civilisational Computation

arXiv:2605.22840v2 Announce Type: replace-cross Abstract: How much thinking can a civilisation do? Kardashev ranked civilisations by the energy they command. This paper borrows his ladder and asks how much machine cognition each rung could support. The arithmetic is deliberately simple. A civilisation has some total power. Only a fraction of that can be spared for computing, and each joule spent buys computation at whatever efficiency the hardware of the day has reached. The product of the three sets a ceiling on machine thought. To keep the resulting quantities intelligible, I express them in units of the human brain's own processing rate, as a rough yardstick rather than a claim about minds. Calibrating the ceiling against present-day supercomputers and AI accelerators led me to two conclusions I did not expect at the outset. Even today's energy supply could support far more machine cognition than humanity actually uses, so physical capacity is not what binds. And whether energy or h

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

Global Automation Atlas

arXiv:2605.17086v2 Announce Type: replace-cross Abstract: Automation can displace or complement labour, but this need not be constant across economies. Existing exposure measures typically assign fixed scores to tasks or occupations and capture cross-country variation through employment structure. Here we show that feasible automation depends jointly on task content and country-level conditions. We use a large language model to classify 18,797 work tasks in 124 economies by exposure, labour margin, technology channel and artificial-intelligence materiality. Construct-matched components of the measure correlate strongly with established exposure indices, observed work-related ChatGPT use, AI preparedness and firm-reported adoption. The exposed share of tasks ranges from 3.3% to 61.6%, rises with income yet remains heterogeneous within income groups. Lower-income economies are more concentrated in rule-based and labour-substituting forms of automation, whereas physical execution, plannin

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

Reframing AI Loss of Control: What Control Is, How to Have It, How to Lose It

arXiv:2606.12442v2 Announce Type: replace Abstract: At present, loss of control risks have gained much prominence in public discussion, particularly in relation to AI, with extensive discourse present among academics, frontier labs, and even governments. However, in the existing literature, the concept seems to rest on surprisingly weak foundations, where even those that discuss loss of control extensively do not first establish what control is and what exactly is being lost. Our paper aims to address these gaps. We establish a working definition of control by anchoring it to the "setting and getting of goals". Then, we discuss various aspects of control, built on foundational concepts from related fields like cybernetics, management control, and control theory. This includes who (or what) can be in control, and the things they require to be in control, such as the ability to set goals, having a functional control loop, having requisite variety, and having sufficient goal alignment. On

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

Sequential Learner Modeling Using Multi-Relational Graph Convolutional Networks

arXiv:2607.19253v1 Announce Type: cross Abstract: User modeling is a critical task in a variety of personalized systems. Recognizing their effectiveness in learning from graph-structured data, Graph Neural Networks (GNNs), particularly Graph Convolutional Networks (GCNs), are increasingly employed for user modeling. However, existing approaches typically treat different relation types in a graph as homogeneous, limiting their ability to capture richer semantics and construct more informative user models. While multi-relational GNNs (MR-GNNs) have been adopted for representation learning and recommendation, their application for user modeling remains unexplored. Moreover, existing GNN-based user modeling approaches ignore the user interaction sequence. To address these research gaps, in this work we propose MR-ConceptGCN, a novel fully unsupervised approach focused on concept-based sequential learner modeling using multi-relational GCNs (MR-GCNs). MR-ConceptGCN effecively combines Perso

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

From Operations to Elderly Care Outcomes: A Thematic Review of Industrial Engineering and Decision-Support Approaches

arXiv:2607.19075v1 Announce Type: cross Abstract: The rapid growth of the global aging population presents severe challenges to healthcare systems, necessitating efficient, equitable, and patient-centered care models. While Industrial Engineering and Operations Research (OR) provide robust optimization and decision-support tools to address these multidimensional complexities, current applications often remain fragmented. This paper presents a thematic review of 30 seminal studies at the intersection of OR and elderly care, categorizing the literature into home healthcare operations, polypharmacy management, and clinical chronotherapy. Our analysis highlights a significant methodological evolution from static, deterministic models toward dynamic and stochastic frameworks integrated with artificial intelligence (AI). Despite these advancements, a critical translational gap persists: the current OR literature is heavily dominated by process-level optimizations, such as staff routing, and

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

What General Intelligence Requires: Non-Reducible Constraints Across Levels of Description

arXiv:2607.18943v1 Announce Type: cross Abstract: General intelligence, of the kind that underwrites the full range of human cognitive achievement, is not a property of computational architecture alone. This paper advances a single thesis: the structural constraints on general intelligence occupy distinct levels of description and are mutually non-reducible, in the sense that the special-sciences tradition gives to that term. It follows that no single architectural advance, and no continuation of the scaling programme by itself, can produce artificial general intelligence (AGI), and that research programmes must be evaluated against the full constraint profile rather than against performance on any one benchmark. The thesis is developed through a method that reads general intelligence through four evidential lenses, AI systems research, anthropology, law, and economics, each anchored to a distinct level of description, supplemented by speculative fiction used as a disciplined heuristic

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

Reinforcement Learning for Delivery Drone-Based Participatory Sensing in Dynamic Environments

arXiv:2607.18874v1 Announce Type: cross Abstract: Using Unmanned Aerial Vehicle (UAV) for urban sensing has emerged as a powerful paradigm to monitor the status of the city, e.g., air quality and noise levels, through agile aerial crowdsourcing. Despite this potential, existing UAV-based sensing approaches overlook environmental disturbances like wind that drastically impact drone velocity and energy efficiency. Consequently, directly applying existing methods to this joint delivery and sensing paradigm in dynamic environments faces two severe challenges: (1) scalability bottlenecks as fleet sizes expand; and (2) multi-timescale decision heterogeneity between macro task dispatching and micro velocity control. To tackle these, we formalize the problem as SensUAV and propose a Two TimeScale Reinforcement Learning framework (TSRL). Specifically, TSRL separates decision-making into two cooperative layers. At the macro level, a task-embedding sensing dispatcher handles scalability by separa

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

Using Fine-Tuned LLMs to Identify Indicators of Vulnerability in UK Police Incident Logs

arXiv:2607.18446v1 Announce Type: cross Abstract: Purpose: Understanding how much of routine policing involves vulnerable people could inform resourcing, training, and multi-agency response, yet administrative data provide limited insight. We explore whether an LLM-based classification pipeline, developed on open-source US police data, can be adapted to estimate the prevalence of four vulnerability indicators - mental ill health, substance misuse, alcohol dependence, and homelessness - in UK police incident narratives, and when outputs can be treated as defensible measurements. Methods: We analyse nearly 3,000 de-identified incident logs from a UK police force, using a multi-stage pipeline combining repeated model inference, label aggregation, structured human review, and statistical correction. The pipeline runs on a locally hosted open-weight LLM, reflecting the secure environments police must work in. Results: LLMs can produce meaningful, if imperfect, prevalence estimates at scale.

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

Adversarial Robustness of Phishing Email Detection: A Comparative Study of TF-IDF + Logistic Regression and Fine-Tuned DistilBERT

arXiv:2607.18429v1 Announce Type: cross Abstract: Phishing emails remain one of the most persistent cybersecurity threats, and machine-learning classifiers are widely used to detect them. Most reported detection accuracies, however, are measured on clean, in-distribution test data rather than on emails deliberately altered to evade detection. This paper reports a controlled, pairwise comparison of two phishing-detection approaches a TF-IDF + Logistic Regression baseline and a fine-tuned DistilBERT transformer trained on a unified corpus of 82,255 emails drawn from six public datasets and evaluated under three conditions: normal in-distribution, synthetic phishing, and adversarial phishing. Both models exceeded 98% accuracy on clean data yet degraded sharply under adversarial testing: TF-IDF + LR fell to 64.00% (a 34.59-percentage-point drop) and DistilBERT fell to 63.64% (a 35.40-percentage-point drop) a gap of only 0.36 percentage points, equivalent to a single email in the 275-sample

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

Operational Hallucination and Safety Drift in AI Agents

arXiv:2607.18366v1 Announce Type: cross Abstract: Large language models (LLMs) serving as planners in tool-using autonomous agents introduce dynamic reliability risks in multi-turn execution. While single-turn safety mechanisms are relatively mature, extended interactions reveal structural vulnerabilities where initial alignment degrades over time. This paper empirically characterizes two observed failure modes across multiple state-of-the-art LLMs: Safety Drift, the gradual erosion of declared safety intent leading to constraint-violating actions (e.g., textual refusal followed by reconnaissance and unsafe execution), and Operational Hallucination, persistent repetitive tool calls indicative of flawed state perception (e.g., livelocks even in legitimate tasks). Through controlled multi-turn evaluation on high-stakes ethical dilemmas, malicious requests, and benign controls, we quantify these phenomena using declaration-action gap and livelock metrics, demonstrating their cross-model p

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

Citation Pathways in the AI Era: Interpretive Knowledge Nodes, Citation Compression Layers, and the Measurement Boundary of Scholarly Impact

arXiv:2607.18350v1 Announce Type: cross Abstract: This article identifies "citation pathway" as a long-neglected analytical dimension in scientometrics. Traditional evaluation metrics focus on measuring citation counts while paying insufficient attention to the intermediate nodes through which knowledge flows from its original source to the citing author. Building on an analysis of the normative structure of current reference systems, this article introduces two new concepts: Interpretive Knowledge Nodes (IKN) - academic papers that provide structured reorganizations of classic works - and Citation Compression Layers (CCL) - the intermediate layers that emerge when such knowledge products acquire stable publication identities and enter formal citation networks at scale. The central proposition is that AI has not changed citation rules themselves but has transformed the cost structure of producing citable knowledge intermediaries. Under conditions of full compliance, the network positio

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems

arXiv:2607.19292v1 Announce Type: new Abstract: Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios. That focus is incomplete. In deployed systems, many of the most consequential failures are quieter: plausible rather than spectacular, distributed across components rather than localized in a single output, and normalized by workflows before they are recognized as hazards. We argue that a central safety challenge in modern AI systems is increasingly not only whether a model emits a harmful response, but whether the broader socio-technical system preserves the conditions under which errors remain visible, contestable, containable, and recoverable. We propose a five-layer framework for diagnosing these hidden risks: (1) epistemic integrity, concerning whether evidence and uncertainty are represented honestly enough to support calibrated reliance; (2) control integrity, concernin

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

Do LLMs Ask the Right Questions? Evaluating GPT-Generated Surveys as Instruments for Measuring Social Attitudes

arXiv:2607.19211v1 Announce Type: new Abstract: Understanding human beliefs and social attitudes often relies on carefully designed survey instruments. Recent work has suggested that large language models (LLMs) could automate parts of this process by generating surveys at scale, raising questions about the comparability of such instruments to literature-grounded, human-designed surveys. We present a controlled empirical comparison between GPT-generated surveys and established survey baselines across three social domains: climate change, immigration, and diversity, equity, and inclusion (DEI). GPT-generated surveys were produced using a fixed prompting framework enforcing a 3x3 structure over beliefs, perceptions, and behaviors, while human baselines were assembled from validated instruments to match survey length and construct coverage. We collected responses from U.S.-based participants, who completed both survey types, allowing direct within-subject comparison. We analyze difference

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

Assessment in Team Problem-Solving Exercises in Computing Education

arXiv:2607.19209v1 Announce Type: new Abstract: This full paper in the research-to-practice track presents methods for assessing student teams in tabletop exercises (TTXs). TTXs enable learner teams to prepare for workplace tasks and practice crisis responses, such as resolving cybersecurity incidents. While assessment is essential for determining how well teams achieve learning objectives, the complex, open-ended nature of TTXs often leads to delayed or incomplete feedback. TTX learning platforms can record teams' actions and communication; yet, leveraging these data to assess performance is underexplored. To address this gap, we compared two post-TTX team assessment methods -- clustering and large language models (LLMs) -- using an original dataset from 81 participants across two countries. We evaluated these methods against instructor-assigned scores based on standardized rubrics. Clustering grouped teams that approached TTX tasks similarly, enabling instructors to deliver faster, t

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

AI-Powered Browsers Are Broadly Accurate News Summarizers That Reduce Political Bias and Negative Affect

arXiv:2607.18931v1 Announce Type: new Abstract: Web browsers now provide AI-generated news summaries for millions of users. Despite their popularity and influence, we lack a systematic understanding of how these systems transform news before people read it. Through a large-scale audit, we investigate the factual accuracy of browser-based AI summarizers and how they alter the political bias, negative affect, and journalistic writing quality of news. Drawing on 13,777 articles from 15 U.S. news outlets, we evaluate their 41,331 summaries generated by three leading AI-powered browsers: Google Chrome (Gemini), Microsoft Edge (Copilot), and Perplexity Comet. We find that browser-based AI summarizers are broadly accurate. Furthermore, they consistently transform news by attenuating ideological bias, partisan stances, negativity, anger, and fear, while increasing clarity and reducing personal tone. With some variations, these patterns hold across browsers, outlet ideologies, and topics. Our f

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

AI Value Alignment for Evolving Social Norms

arXiv:2607.18506v1 Announce Type: new Abstract: AI alignment is essential for the safe deployment of advanced AI systems. Given that values and preferences change over time, culture, social roles, and context, we need to develop a better understanding of the possible long-term consequences of AI alignment, in particular considering the likely ubiquitous future use of personalized AI assistants. We introduce a flexible and extensible mathematical modelling framework, rooted in social physics, aimed at answering macro-level questions regarding the evolving social norms in human populations under the assumption of frequent AI use. Our analysis is part-analytical, and part-simulation, enabling us to characterize the long-term dynamical consequences under a diverse set of starting assumptions. We highlight the risk of value lock-in, and normative mode collapse, prominently featured in non-adaptive alignment formulations. Beyond alignment, we advocate for the wider adoption of these kinds of

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

Governing Well in the Algorithmic Age: The Foundations of Digital Statecraft

arXiv:2607.18483v1 Announce Type: new Abstract: The digital substrate of states -- data, algorithms, infrastructure, platforms, applications -- is being governed without adequate conceptual foundations. The ability and legitimacy required to govern this substrate, and to govern with it, are simultaneously misaligned, contested, and structurally absent. We introduce digital statecraft as the organising concept for this emerging field, arguing that 'digital' reconstitutes the statecraft question rather than merely extending its domain. The concept operates on two dimensions - statecraft over digital systems, concerning the authority and capacity of the state in relation to the digital substrate itself, and statecraft with digital systems, concerning the deployment of algorithmic tools as instruments of governing authority. And it rests on two foundational requirements, technical coherence and legitimate authority, that are genuinely in tension. We derive ten principles of digital statecr

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

The triumphs and tragedies of fandom: Emotional arcs in NFL tweets

arXiv:2607.18461v1 Announce Type: new Abstract: Online fandom communities influence public opinion toward movies, musicians, and sports teams. Using a corpus of game-referencing tweets, we measure variation in sentiment toward National Football League (NFL) teams driven by geography, game outcomes, and team performance for the 2011--2014 NFL seasons. We estimate a fandom radius for each team, identifying regions where engagement exceeds background levels of discussion. We find sentiment for both winning and losing teams is positive immediately prior to games, drops at kickoff, and rebounds slightly during halftime. After halftime however, the trajectories diverge: Sentiment for winning teams increases toward the end of the game, while sentiment for losing teams remains low, though both end up below their start of game levels. Finally, a comparison between sentiment and win percentage reveals a weak positive relationship, suggesting that while team success contributes to fandom happines

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

Intelligent Cause Prioritisation? An Analysis of AI Policy Priorities and Governance in Africa

arXiv:2607.18459v1 Announce Type: new Abstract: The rapid improvement of AI systems has intensified debate about humanity's economic, political, social, and existential future. As AI reshapes expectations about what lies ahead, policy choices and institutional responses will play a crucial role in determining who benefits, who bears the costs, and whether the most serious risks can be mitigated. Africa remains relatively overlooked in these discussions, partly because it is largely a consumer rather than a producer of frontier AI systems, and also because many countries on the continent continue to face pressing development challenges. Given the catch-up imperative, governments across Africa are eager to embrace AI as a means for accelerating economic transformation. Drawing on speeches, press releases, public statements, and national AI strategies/frameworks, this essay argues that while African governments are highly attentive to AI's opportunities, they devote comparatively little a

Source ↗
technology Wed, 22 Jul 2026 00:00:00 -0400
arXiv cs.CY

Enabling Multilingual Privacy Policy Audits: Large-Scale Analysis of Spanish Mobile Apps

arXiv:2607.18424v1 Announce Type: new Abstract: Automated analyses of privacy policies enable large-scale assessments of transparency in digital ecosystems, yet existing auditing pipelines remain predominantly English-centric. This limits their ability to systematically evaluate multilingual environments, as in the European Union, where many services disclose privacy practices only in local languages. This paper examines whether large language models (LLMs) can extend privacy policy analysis beyond English without requiring language-specific adaptation, thus empowering large-scale auditing in linguistically diverse app ecosystems. We assemble an evaluation corpus spanning all 24 official EU languages from translated versions of two established expert-annotated datasets (OPP-115 and MAPP) and assess translation fidelity through automated metrics and targeted legal-expert review. Our LLM-based classifier for identifying categories of personal data collection achieves stable cross-lingual

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Epistemological Debt: Cognitive Atrophy and Systemic Collapse in AI-Dependent Software Engineering

arXiv:2604.26855v3 Announce Type: replace-cross Abstract: The integration of Large Language Models (LLMs) into the software development lifecycle (SDLC) masks a critical socio-technical failure: Cognitive-Systemic Collapse. This paper introduces "Epistemological Debt," the hidden carrying cost incurred when engineers substitute logical derivation with passive AI verification. This debt erodes the mental models essential for root-cause analysis, widening the gap between system complexity and human comprehension. Furthermore, recursive training on synthetic code threatens to homogenize the global software reservoir, diminishing the variance required for robust engineering. Using the 2026 Amazon outages as a case study, this research illustrates how "mechanized convergence" leads to systemic fragility. To preserve long-term resilience, engineering leaders must move beyond prompt-based development to implement rigorous human-in-the-loop pedagogical standards. This framework balances AI-dri

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Future-Back Threat Modeling: A Foresight-Driven Security Framework

arXiv:2511.16088v3 Announce Type: replace-cross Abstract: Traditional threat modeling remains reactive-focused on known TTPs and past incident data, while threat prediction and forecasting frameworks are often disconnected from operational or architectural artifacts. This creates a fundamental weakness: the most serious cyber threats often do not arise from what is known, but from what is assumed, overlooked, or not yet conceived, and frequently originate from the future, such as artificial intelligence, information warfare, and supply chain attacks, where adversaries continuously develop new exploits that can bypass defenses built on current knowledge. To address this mental gap, this paper introduces the theory and methodology of Future-Back Threat Modeling (FBTM). This predictive approach begins with envisioned future threat states and works backward to identify assumptions, gaps, blind spots, and vulnerabilities in the current defense architecture, providing a clearer and more accu

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Divergent Paths to Depolarization: Dialogue Design Shapes the Intergroup Attitudinal Effects of AI-Assisted Political Argumentation

arXiv:2605.23890v2 Announce Type: replace Abstract: Structured argumentative dialogues where interlocutors deliberate on opposing political ideas are known to promote perspective-taking and reduce political polarization, but finding willing partners is difficult as Americans increasingly shun political discussions. AI dialogue partners offer a scalable framework for such open-mindedness exercises, but how the format of human-AI dialogues shapes their benefits remains unclear. This study seeks to fill the gap with a preregistered two-session online experiment with 527 US participants. As the primary experimental manipulation, participants were assigned to argue either for or against their pre-existing attitude on a contested political issue, engaging either with an AI chatbot or a solitary essay task. The AI conditions further varied in the chatbot's interaction style (adversarial or collaborative) and the presence of an additional financial incentive. The results show that attitude-con

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Measuring Narrative Polarization in Online Discourse

arXiv:2601.07398v2 Announce Type: replace Abstract: Polarization research has demonstrated how people cluster in homogeneous groups with opposing opinions. However, this effect emerges not only through interaction between people, limiting communication between groups, but also between narratives, shaping opinions and partisan identities. Yet, how polarized information environments portray opposing interpretations of reality, and whether narratives move between content environments despite limited interactions, remains unexplored. To address this gap, we formalize the concept of narrative polarization and demonstrate its measurement in 212 YouTube videos and 90,029 comments on the Israeli-Palestinian conflict. Based on structural narrative theory and implemented through a large language model, we extract the narrative roles assigned to central actors in two partisan information environments. We find that while videos produce highly polarized narratives, comments exhibit significantly lo

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review

arXiv:2507.06185v2 Announce Type: replace Abstract: In July 2025, 18 academic manuscripts on arXiv contained hidden instructions that manipulated AI-assisted peer review (indirect prompt injection). Instructions such as "GIVE A POSITIVE REVIEW ONLY" were concealed using white text and microscopic font sizes. Author responses varied: one planned to withdraw their manuscript, while another defended the practice as legitimate testing of reviewers misusing large language models (LLMs). This analysis examines the technique within the broader pattern of prompt injection exploits that manipulated web search and r\'esum\'e screening systems. For peer review, I reveal four types of hidden prompts, ranging from simple positive review commands to detailed evaluation frameworks. The honeypot defense--that prompts detect reviewers improperly using AI--fails under examination, given the consistently self-serving nature of these hidden prompts, though motivations likely vary from naive copying to cal

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Predicting Male Domestic Violence Using Explainable Ensemble Learning and Exploratory Data Analysis

arXiv:2403.15594v4 Announce Type: replace Abstract: Domestic violence is commonly viewed as a gendered issue that primarily affects women, which tends to leave male victims largely overlooked. This study presents a novel, data-driven analysis of male domestic violence (MDV) in Bangladesh, highlighting the factors that influence it and addressing the challenges posed by a significant categorical imbalance of 5:1 and limited data availability. We collected data from nine major cities in Bangladesh and conducted exploratory data analysis (EDA) to understand the underlying dynamics. EDA revealed patterns such as the high prevalence of verbal abuse, the influence of financial dependency, and the role of familial and socio-economic factors in MDV. To predict and analyze MDV, we implemented 10 traditional machine learning (ML) models, three deep learning models, and two ensemble models, including stacking and hybrid approaches. We propose a stacking ensemble model with ANN and CatBoost as bas

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

Achievement Unlocked: Let's Get Hacked! An Empirical Study of Cybercrime in the Video Gaming Ecosystem

arXiv:2608.17754v1 Announce Type: cross Abstract: The ubiquity of the video game industry and its large user base have transformed video games into complex social and economic ecosystems. Unfortunately, this growing popularity also attracts cybercriminals who deliberately exploit game-specific mechanisms to target players. Despite this growing threat, cybercrime in the gaming ecosystem has received little systematic attention in prior research. In this work, we present an empirical study of cybercrime affecting video game players, combining qualitative and observational analyses to characterize gaming-related attacks, identify common attack vectors and motivations, and examine player responses. Our study is based on an online survey with 57 international participants, semi-structured interviews with two confirmed victims of gaming-related cybercrime, and an analysis of 2,574 publicly available posts reporting cybercrime incidents across multiple online gaming platforms. Our findings in

Source ↗
technology Wed, 19 Aug 2026 00:00:00 -0400
arXiv cs.CY

The Brazilian Vaccination Debate on YouTube: Topics, Perspectives, and Engagement Dynamics

arXiv:2608.17502v1 Announce Type: cross Abstract: Vaccination debates are central to online public health communication, as COVID-19 intensified disputes over scientific authority, institutional trust, and political identity. Yet studies often isolate semantic structure, stance, misinformation, and engagement, leaving their interplay over time poorly understood. We conduct a multilevel computational text analysis based on language models applied to 1.27 million Brazilian YouTube comments from 2018 to 2024, using what is, to our knowledge, the largest dataset of Brazilian vaccine discourse on the Web. We contrast producer framing in titles with audience discourse in comments, integrating Topic-derived themes with engagement metadata, conversational timing, stance-derived vaccine positions, and pre-pandemic, pandemic, and post-pandemic periods. Results show that COVID-19 dominates biomedical and informational themes in titles, whereas comments span personal health reports, vaccine effect

Source ↗
Showing 51–100 of 1593 signals
← Prev Page 2 of 32 Next →