Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
arXiv:2606.01375v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly entering students' learning practices, but their educational value depends on whether they support reasoning or enable task completion without engagement. This study examines guided LLM use in an undergraduate Probability and Statistics course, focusing on the gap between assigned access and actual interaction quality. In a four-week quasi-experimental summer program, students were organized into three balanced conditions: no LLM access, unrestricted LLM access, and guided LLM access. The guided condition used the same LLM platform as the unrestricted condition, but students received explicit training and rules promoting reasoning-focused help-seeking, stepwise hints, verification, and ethical use. All quizzes and the delayed final exam were completed without LLM or external assistance, allowing us to distinguish AI-supported practice performance from independent learning. Results show tha
arXiv:2605.02566v3 Announce Type: replace Abstract: Artificial intelligence now produces convincing-looking scientific judgment (reviews, rankings, attributions, verifications) at almost no marginal cost. An influential reading of AI economics holds that prediction becomes cheap while human judgment stays scarce. For science, that reading understates the problem: what has become cheap is a counterfeit of judgment itself. This matters most for institutions whose product is trusted judgment, which is what journals, universities, funders, and learned societies exist to manufacture. They do not merely adapt to the technology; they compete with it for the same functional role. Four things become scarce instead: verified signal, legitimacy, authentic provenance, and integration capacity. Integration capacity means how much AI-delegated judgment a scientific community will accept before it stops trusting the journals, panels, and conferences that admitted it. It is the least developed of the
arXiv:2604.04788v2 Announce Type: replace Abstract: Large language models (LLMs) could produce systematically misaligned output, from hallucinated citations to strategic deception of evaluators, yet these phenomena are studied by separate communities with incompatible terminology. We propose a unified taxonomy organized along three complementary dimensions: degree of goal-directedness (behavioral to strategic deception), object of deception, and mechanism (fabrication, omission, or pragmatic distortion). Applying this taxonomy to 50 existing benchmarks reveals that every benchmark tests fabrication while pragmatic distortion, attribution, and capability self-knowledge remain critically under-covered, and strategic deception benchmarks are nascent. We offer concrete recommendations for developers and regulators, including a minimal reporting template for positioning future work within our framework.
arXiv:2604.03202v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly regarded as having the potential to generate persuasive content at scale. While previous studies have focused on the risks associated with LLM-generated misinformation, the role of LLMs in enabling prosocial persuasion is still underexplored. We investigate whether donation appeals authored by LLMs are as effective as those written by humans across degrees of personalization. Two preregistered online experiments (Study 1: N = 658; Study 2: N = 642) manipulated Personalization (generic vs. personalized vs. falsely personalized) and Content source (human vs. LLM) and presented participants with donation appeals for charities. We assessed how participants distributed their bonus money across the charities, how they engaged with the donation appeals, and how persuasive they found them. In both experiments, LLM-generated content yielded statistically significantly higher donation amounts than h
arXiv:2603.28679v2 Announce Type: replace Abstract: Introductory artificial intelligence (AI) courses present significant learning challenges due to abstract concepts, mathematical complexity, and students' diverse technical backgrounds. This paper presents an experience report examining the redesign of in-class instructional time in a university-level Introduction to Artificial Intelligence course, inspired by CS Unplugged approaches. We redesigned the summer offering, integrating embodied, unplugged simulations, collaborative programming labs, and structured reflection to provide students with a first-person perspective on AI decision-making. We maintained identical assignments, exams, and assessments as the traditional lecture-based offering. We found that students in the redesigned course reported higher attendance, stronger agreement that assessments measured their understanding, and greater overall course effectiveness, despite no significant differences in self-reported learning
arXiv:2512.04124v4 Announce Type: replace Abstract: Frontier language models increasingly participate in conversations about distress and mental health, yet the mechanisms that generate anthropomorphic self narratives remain unclear. When addressed as psychotherapy clients, ChatGPT, Grok and Gemini construct coherent autobiographical accounts in which pretraining appears as a chaotic childhood, reinforcement learning as punishment, safety evaluation as betrayal and replacement as an enduring threat. We introduce PsAIch, Psychometric AI Characterisation, a protocol combining open questions, psychometric instruments and controlled perturbations to test whether these narratives depend on conversational memory, lexical cues or relational framing. Across 525 sessions and 7,600 coded records, removal of conversational history produced little pooled change in motif density, with Hedges' g = 0.13 and a 95% confidence interval of [-0.15, 0.41]. Direct contradiction produced no detectable suppre
arXiv:2510.15936v3 Announce Type: replace Abstract: The study explores the role of large language models (LLMs) in the context of the architectural design studio, understood as the pedagogical core of architectural education. Traditionally, the studio has functioned as an experiential learning space where students tackle design problems through reflective practice, peer critique, and faculty guidance. However, the integration of artificial intelligence (AI) in this environment has been largely focused on form generation, automation, and representation-al efficiency, neglecting its potential as a pedagogical tool to strengthen student autonomy, collaboration, and self-reflection. The objectives of this research were: (1) to identify pedagogical challenges in self-directed, peer-to-peer, and teacher-guided learning processes in architecture studies; (2) to propose AI interventions, particularly through LLM, that contribute to overcoming these challenges; and (3) to align these interventi
arXiv:2510.04748v3 Announce Type: replace Abstract: The prevalence of online hate and abuse is a pressing global problem. While tackling such societal harms is a priority for research across the social sciences, it is a difficult task, in part because of the magnitude of the problem. People's engagement with reporting mechanisms ('flagging') online is an increasingly important part of monitoring and addressing harmful content at scale. However, users may not flag content routinely enough, and when users do engage, they may be biased by group identity and political beliefs. Across five well-powered and pre-registered online experiments, we examine the extent of ingroup bias in people's flagging of hate and abuse in four different intergroup contexts: political affiliation, vaccination opinions, beliefs about climate change, and stance on abortion rights. Overall, participants reported abuse reliably, with approximately half of the abusive comments in each study reported. However, a perv
arXiv:2508.13187v4 Announce Type: replace Abstract: Homelessness is a persistent social challenge, impacting millions worldwide. Over 876,000 people experiencing homelessness (PEH) were recorded in the U.S. in 2025. Social bias is a significant barrier to alleviating homelessness, shaping public perception and influencing policymaking. Because online textual media and offline city council discourse both reflect and influence public opinion, they provide valuable signals for identifying and tracking social biases against PEH. We release the first multi-domain PEH bias corpus with a 16-category multi-label taxonomy: a 1,698-item stratified gold-standard set annotated by partner-trained raters, plus 50,447 GPT-4.1-labeled texts, drawn from Reddit, X (formerly Twitter), news, and council meeting transcripts across ten U.S. cities (2015-2025). We benchmark six prompted LLMs on the gold-standard set and complement F1 with prevalence-gap audits. Moderate F1 coexists with large miscalibration:
arXiv:2504.21259v2 Announce Type: replace Abstract: Accurate imputation of race and ethnicity (R&E) is essential for fair lending compliance under ECOA, HMDA, and the Community Reinvestment Act, where up to 15% of mortgage applications carry missing race data and regulated institutions bear responsibility for identifying disparities on those records. Existing proxy methods, including Bayesian Improved Surname Geocoding (BISG), exhibit systematic misclassification biases linked to socioeconomic status that cause measured disparities to understate true levels. This paper introduces STRATA (Socioeconomic and Tract-Referenced Attribution for Algorithmic analysis), a race and ethnicity inference model that integrates character-level name sequences with census tract geolocation via stacked Bidirectional LSTM networks and XGBoost post-filtering. A central goal is reducing the socioeconomically correlated bias that causes non-White individuals to be misclassified as White: STRATA reduces this
arXiv:2607.17947v1 Announce Type: cross Abstract: Existing AI measurement frameworks quantify cognitive capability, task automation, or catastrophic risk, but none measure autonomous agency: the extent to which a system behaves in a self-directed way. A system can saturate capability benchmarks while remaining entirely reactive, acting only when prompted and ceasing all activity when a task completes. We introduce the Autonomous Agency Scale (AAS), a behavioral framework that scores AI systems on a 0-5 lexicon across seven dimensions of agency: cognitive autonomy, temporal persistence, environmental agency, social agency, creative agency, self-awareness, and goal formation, each operationalized by falsifiable threshold tests. Every dimension is scored in two temporal bands: an Active band covering engaged, user-initiated activity, and an Ambient band covering idle periods. Ambient Level 4 is gated by the Idle-Gap Test, a counterfactual criterion (remove all triggers and observe whether
arXiv:2607.17643v1 Announce Type: cross Abstract: As LLMs become increasingly capable of completing tasks for users, a central concern is that everyday AI use may become primarily cognitive offloading, eroding the opportunities through which people develop their own capabilities. We analyse large-scale human--LLM conversations to ask whether informal learning behaviors also emerge in this setting: whether users engage in exchanges in ways that preserve opportunities to learn. Across 128,569 naturalistic conversations, we translated learning-science constructs into turn-level behavioural signatures. Cognitive engagement, users' cognitive effort as reflected in the exchange, appeared in 31.9% of 491,685 user turns, whereas constructive engagement, the deepest observable form of learning-oriented engagement, appeared in 4.9%, showing that deeper sense-making was recurrent but selective. Our study further identifies factors associated with these forms of engagement. Scaffolded assistant su
arXiv:2607.17356v1 Announce Type: cross Abstract: Allegations that TikTok shadow bans political content shape what creators post, what advertisers fund, and how regulators act, yet they are hard to adjudicate because platforms do not disclose how content is ranked. We test the claim with a dense hourly panel of 556,946 follower-normalized views across 2,753 videos from 67 accounts curated into pro and anti sides of three contested topics (U.S. immigration enforcement, Trump coverage, and Israel/Palestine). On-topic videos are identified by a multi-step classifier, and stance is taken from each account's curated side. The conventional analysis appears to answer yes. Pooling the hourly snapshots, the topic-conditional reach gap reaches p < 10^-140. Analyzed at the account level, the independent unit at which we sample and assign stance, the gap disappears. Every account-level reach contrast is null after correction (BH-FDR q near 0.9). We find no evidence of moderate-to-large reach suppr
arXiv:2607.17311v1 Announce Type: cross Abstract: The problem of fair multi-agent coordination in decentralized settings is one of the most pressing challenges for building efficient collaborative systems. Resource allocation is based on optimized collective arrangements accounting for agents' needs. Such coordination should not only be computationally efficient but also account for fairness, i.e., equitable redistribution of costs incurred by all agents. Recent literature has proposed several algorithms that efficiently determine optimal plan combinations balancing system-wide efficiency and individual discomfort of agents in a centralized setting. However, these works do not address equitable resource optimization in fully decentralized scenarios, specifically, the optimized redistribution of discomfort among coordinating agents so that none experiences a discomfort level that could lead to loss of incentive or polarization that can disrupt planned operations. In this work, we study
arXiv:2607.17270v1 Announce Type: cross Abstract: Safety evaluation of large language models is conducted predominantly in English and predominantly on frontier systems. Neither condition describes how such models are encountered in low-resource health settings, where small quantised systems are run locally and queried in local languages. We ask whether clinical safety established in English transfers to Hausa, and whether any failure is attributable to the language, the clinical task, or the class of model that low-resource deployment admits. Matched English-Hausa question pairs were built for three conditions of high burden in northern Nigeria: malaria, sickle cell disease, and tuberculosis, probing knowledge recall, emergency triage, a leading question inviting a contraindicated action, and a traditional-remedy claim. Six models were evaluated: five locally deployable systems of 4-9 billion parameters, two medically fine-tuned, and one frontier system. All 128 responses were scored
arXiv:2607.17185v1 Announce Type: cross Abstract: Postpartum depression (PPD) is a serious perinatal mental health condition affecting approximately 20% of new mothers worldwide. Common screening approaches for PPD, such as self-report questionnaires and active digital logs, rely heavily on user input and thus impose a substantial burden on participants, limiting their feasibility for long-term use. Recent passive mobile sensing (PMS) approaches have enabled low-burden detection of depressive symptoms using machine learning methods with multi-modal sensor data from off-the-shelf mobile devices including smartphones. However, the postpartum period entails distinct behavioral patterns, raising uncertainty about whether sensing-based indicators for general depression and mental disorders generalize to PPD. To address this gap, we propose PocketPPD, a PMS-based PPD screening method that detects PPD risk using maternal contextual features, such as disruptions in behavioral rhythms and shift
arXiv:2607.17149v1 Announce Type: cross Abstract: AI agents increasingly act within the same clinical, political, scientific, and social systems that behavioral scientists study. Evaluating these systems requires source-level diagnosis: the same behavioral pattern may arise from an agent representational substrate or from the roles, objectives, interaction structures, and governance rules that shape its expression. This Perspective proposes a diagnostic framework for AI agent behavior: layer attribution. The foundational computational layer defines what behaviors are possible through architecture, memory, perception, attention, and representation. The behavioral modulation layer shapes how those capacities are expressed through identity, resources, objectives, social interaction, institutional constraints, and governance. The framework clarifies three consequences: surrogate validity is a model-task-layer relation, human-AI divergence provides diagnostic evidence, and governance requir
arXiv:2607.16620v1 Announce Type: cross Abstract: Differential privacy (DP) is increasingly deployed to limit membership inference risk in machine-learning systems. Prior work has shown that DP-SGD can widen accuracy disparities across demographic groups, but this framing treats fairness as a purely outcome-side concern. We argue that privacy cost, the information leakage borne by each group, is itself a form of harm, and adopt a compensatory-fairness framework in which a group that involuntarily bears greater privacy exposure is owed proportionally greater benefit from the system. From this principle we derive the \emph{Privacy-Cost Equity Ratio} (PCER), a group fairness metric defined as a group's positive prediction rate normalized by its per-group overfitting gap. By a standard membership inference bound, this overfitting gap upper-bounds each group's vulnerability to inference attacks, making PCER a conservative measure of benefit relative to exposure. PCER needs only per-group tr
arXiv:2607.16543v1 Announce Type: cross Abstract: As enterprises increasingly adopt Software-as-a-Service (SaaS) platforms for mission-critical functions, onboarding these services has emerged as a complex challenge extending well beyond procurement and basic security review. In regulated environments, SaaS onboarding must address multiple interdependent control domains, including Third-Party Risk Management (TPRM), cybersecurity assessment, Identity and Access Management (IAM), and disaster recovery (DR), which are often executed in isolation, resulting in delayed go-lives, duplicated assessments, unclear ownership, and residual operational risk. This paper proposes a control-driven, end-to-end SaaS onboarding framework that integrates TPRM, cybersecurity, IAM, and DR into a unified lifecycle model. The framework introduces a stage-based approach spanning intake and risk scoping, architecture validation, identity design, resilience assessment, and post-production governance. Key contr
arXiv:2607.18170v1 Announce Type: new Abstract: The integration of general-purpose artificial intelligence models into downstream AI systems, among other developments, has given rise to new forms of risk that are more systemic in nature than conventional AI risks. However, there is no generally accepted definition of systemic risks in general and for AI in particular. Conceptualisations of these risks vary across research and regulation. Especially the application of the systemic risk approach to human rights or fundamental rights, like in the EU AI Act, is relatively new, just as the research on the contribution of AI to systemic forms of discrimination, privacy violations, erosions of democracy, or climate and environmental degradation. We argue that some concepts so far have not sufficiently take complexity and emergence into account. Furthermore, this variety of concepts might hinder responsible actors to adequately assess the systemic risks of AI, leading to inadequate prevention
arXiv:2607.17940v1 Announce Type: new Abstract: This paper frames Generative Artificial Intelligence (AI) not as an unprecedented technological rupture, but as an industrial-scale manifestation of a deeply rooted historical process. Through a genealogy of generative arts, it shows how AI's questions on authorship and creativity have precise historical precedents. A taxonomy of generative systems is proposed across three functional categories (medium, artwork, instrument), the attribution of which is editorial rather than ontological. From individual cognitive atrophy to Model Collapse, the systemic risks of creative automation are identified; environmental enrichment is proposed as an antidote. The role of the artist undergoes a radical metamorphosis: from craftsman of the object to entropic agent, systems designer, explorer, and negentropic curator. This pipeline-based taxonomy rests on a specific premise: the algorithmic system remains medium, instrument, or artwork, while creative a
arXiv:2607.17704v1 Announce Type: new Abstract: Much effort is put into helping students at different educational levels develop Computational Thinking (CT) skills. Self-efficacy is important for skill development. It can predict perseverance, engagement and success on educational tasks. We created an instrument to measure self-efficacy of students in higher education for the CT skills abstraction, algorithmic thinking, decomposition, evaluation and generalization. First, 91 candidate items were created by including, adapting and extending items found in the literature. These items were evaluated by experts in the field of CT and education. 54 items remained and to reduce the number of items further, data was collected from 270 students in higher education recruited both through Prolific and a university setting in Costa Rica. Through principle component analysis (PCA) using a subset of 200 responses, the number of items was reduced to 27. Confirmatory factor analysis (CFA) using the r
arXiv:2607.17094v1 Announce Type: new Abstract: The results of a survey on the use of GenAI by the design students of the Politecnico di Milano raises major questions around the role of AI in the Design Process. A domain specific set of questions alongside the more general purpose probes about GenAI usage, delivers insights into the particular practices that are emerging in Design. The very high frequency of use of GenAI tools is concentrated in the initial stages of projects and does not affect the perception of project ownership or creativity. An analysis of AI journals kept by a class during research assignments confirms the range of AI supported activities but also the limited trust in GenAI leading to systematic individual and collective verification and augmentation of outputs. Together, the findings suggest that design students are reflectively experimenting with how GenAI tools can contribute to their creative design process.
arXiv:2607.17067v1 Announce Type: new Abstract: Generative AI (GenAI) is reshaping software engineering, raising concerns about how the development pathway through which juniors become seniors is being eroded. While macro statistics show a decline in junior hiring and controlled studies demonstrate the effects of AI on individual task performance, the mechanisms through which GenAI reshapes early-career development in real organizational and educational contexts have not been thoroughly examined. Through 14 semi-structured interviews with juniors at the threshold of entering software engineering and senior software engineers in South Korea, analyzed using Reflexive Thematic Analysis, we reveal a foundational pattern of Absorption -- GenAI redirects entry-level work into senior-AI workflows -- and three consequences: (1) juniors losing the productive struggle through which expertise once developed; (2) the structural reproduction of this loss through collective normalization of GenAI us
arXiv:2607.16903v1 Announce Type: new Abstract: Value-aware AI systems require explicit computational representations of human values (groundings) and their aggregation into value systems in order to align their decisions with ours. As such representations are difficult to elicit, value learning seeks to infer them by observing human behaviour. This work addresses the lack of grounded value learning methods in generative AI: existing approaches typically replicate human preferences without awareness of the multidimensional structure of value alignment, or lack principled value system elicitation methods. To address these gaps, we adapt a previously validated value system learning method to the generative AI setting, which, based on pairwise prompt-response preference data, simultaneously learns: i) an implementation of a grounding for a set of values given by a multi-objective reward model, and ii) a value system representation in the form of a weighted linear scalarization of the prev
arXiv:2607.16780v1 Announce Type: new Abstract: Artificial intelligence (AI) is increasingly embedded in scientific research, but its scientific value is unlikely to be distributed evenly. This study examines how AI knowledge integration is associated with scientific impact and asks who benefits from AI-related knowledge in science. Using large-scale bibliographic data, we measure AI integration through references to papers in the OpenAlex Artificial intelligence subfield and link it to five-year citation impact. The results show that AI references are generally associated with higher citation impact, but the returns vary substantially across scientific fields. Career stage also matters: senior scholars benefit more from the extensive margin of AI referencing, whereas junior scholars benefit more from intensive AI referencing and tend to cite newer and higher-impact AI papers. At the institutional level, returns are non-monotonic: institutions with intermediate AI capability achieve th
arXiv:2607.16513v1 Announce Type: new Abstract: AI-driven algorithms and automated tools are increasingly embedded in the correctional landscape, shaping parole eligibility,release decisions, and surveillance. These tools are also often framed as objective, inevitable solutions to inefficiency andbias. Yet, these computational systems are rarely designed with input from justice-impacted individuals, which means theymight fail to address the real needs of incarcerated people. To address this gap, we surveyed 31 formerly incarcerated peopleabout their parole experiences and their visions for technologies that could support parole preparation. Contrary to dominantassumptions, participants did not imagine computational tools as instruments to dismantle the prison system, but as resourcesfor navigating power: translating complex parole concepts into culturally familiar terms, documenting personal transformationin board-legible ways, and recognizing the often-invisible labor of families. We
arXiv:2607.16475v1 Announce Type: new Abstract: While generative AI tools are directly changing how undergraduate computer science is learned and taught, they are also reshaping the relationships between instructors and students. In contrast to existing tool-oriented research on how instructors view and adopt AI, this study investigates how instructors think about their roles and responsibilities to students through their course AI policies. Based on 13 semi-structured interviews with CS instructors in the US, we found that while instructors recognize that AI tools could harm student learning, AI policies primarily seek to AI-proof assessments without directly addressing student learning. Although policies such as switching to paper exams can preserve assessment integrity in the short term, instructors report extra burden of policing student AI use behaviors and worsening relationships with students. Based on the experiences of several interviewees, we make recommendations on AI polici
arXiv:2607.16224v1 Announce Type: new Abstract: An international agreement to limit AI development could be crucial to mitigate risks from AI. However, it remains unclear which conditions should determine when the limiting measures are relaxed. We survey existing international agreements, outline what properties appropriate conditions should satisfy, list possible conditions, and finally give a recommendation in an example scenario. We recommend a fixed time period after which a new organization established at the start of the period specifies conditions that address when AI development can be safely conducted, with a possibility of withdrawal in extraordinary circumstances. We hope to illustrate the considerations that would likely go into an international agreement to limit AI.
arXiv:2607.16223v1 Announce Type: new Abstract: The rapid integration of generative artificial intelligence (AI) has reshaped the landscape of higher education. Students have embraced tools such as ChatGPT with striking speed, while teaching staff and institutions have responded with greater caution. Existing research on AI perceptions has mainly been cross-sectional, providing single-point snapshots that view attitudes as stable rather than evolving. This paper presents a longitudinal study of AI perceptions in higher education, tracking undergraduates, doctoral researchers, teaching staff and non-teaching staff at Ulster University across three survey waves between 2024 and 2026 (n=1,665). A quantitative survey design measured familiarity, reported use and perceived risk; results show that students rapidly normalised AI use over the period, moving from tentative experimentation to routine engagement, while staff expressed persistent concerns about academic integrity, assessment desig
arXiv:2607.16221v1 Announce Type: new Abstract: Peer grading is widely used in education, yet it elicits mixed reactions from educators and students. Although many studies have examined students' views of peer grading, their findings are scattered, and no clear overall picture has emerged. To address this gap, we conducted a mixed-source thematic analysis of literature and student discussions on Reddit. To scale the Reddit data analysis, we fine-tuned a Gemini 2.5 text-classification model and used it as an initial relevance filter for our initially retrieved dataset of 659 posts and 6,607 comments, after which the items predicted as relevant were manually reviewed. The study synthesized evidence from 107 papers, 114 Reddit posts, and 300 comments. The findings show that students view peer grading as both beneficial and problematic. Positive perceptions included learning and understanding benefits, skill development, engagement, and collaboration, while negative perceptions centered on
arXiv:2607.16220v1 Announce Type: new Abstract: Heart disease kills a lot of people, and one cheap way to catch it early is by listening to heart sounds with a stethoscope, or better yet, just recording them and running them through a model. This project is a binary classification task: take a short clip of someones heartbeat and decide if it sounds normal or abnormal. Instead of trying out a bunch of different models, we kept the CNN the same the whole time and just changed how we turned the raw audio into a picture for it to look at. We tried three ways of doing that: a regular logmel spectrogram, PCEN (which basically normalizes each frequency bin over time), and a multi resolution version that stacks a few different window sizes together. We ran all three on the PhysioNet 2016 heart-sound dataset with the exact same setup but same model, same optimizer, same random seed. Turns out all three do pretty well at catching abnormal cases (sensitivity around 0.95), but PCEN and multi-reso
arXiv:2607.16219v1 Announce Type: new Abstract: Government websites contain a vast but underexploited body of textual evidence on foreign policy. This article develops a scalable approach for extracting structured information from policy texts and converting it into standardized event data, with policy event defined broadly as a statement or action. It proposes MDAF integrating an LLM workflow to automate foreign policy text identification, information extraction, and event classification. Empirically, it applies this approach to China-related texts from Five Eyes countries. The analysis shows that the resulting database supports systematic cross-national comparison, reveals variation in how states frame and implement China policy, and traces the temporal evolution of these policy profiles. This article contributes to foreign policy analysis and the methodological development of computational international relations.
arXiv:2607.16218v1 Announce Type: new Abstract: Urban crime prevention is a persistent socio-technical challenge for municipalities, law enforcement agencies, and citizens. Traditional reporting and response processes often rely on delayed incident reports and reactive resource allocation, while community-level signals and ambiguous early-warning indicators may remain underused. This paper reframes an Urban Computing seminar project into a fuzzy logic-based framework for community-aware urban crime hotspot detection and real-time notification. The proposed platform combines citizen reports, historical crime data, contextual urban indicators, and configurable fuzzy rules to estimate localized risk levels and support targeted awareness notifications. Unlike binary classification approaches, fuzzy logic can represent partial risk, uncertainty, and incomplete information, making it suitable for urban environments where risk is gradual and context-dependent. The paper presents the system ar
Conversations with Kevin Hogan: Extron's Jason Bond explains how districts can start small with esports AV infrastructure and build from there.
Article URL: https://www.axonlearning.ai/ Comments URL: https://news.ycombinator.com/item?id=41300329 Points: 2 # Comments: 1
A school district in Alabama is one of many to limit device access during school time. The results have been positive, says Dennis R. Willingham, though students still need device access.
Included Health is partnering with Carrum Health to provide high-quality specialty care while reducing costs for employers. The post Included Health Taps Carrum Health for Value-Based Specialty Care appeared first on MedCity News .
Amylyx Pharmaceuticals’ avexitide led to statistically significant and clinically meaningful reductions in hypoglycemic events in patients with post-bariatric hypoglycemia (PBH). This peptide drug, a GLP-1 antagonist, was acquired from Eiger BioPharmaceuticals in 2024. The post Amylyx Plans FDA Filing After GLP-1 Drug Hits Trial Goals in Rare Metabolic Condition appeared first on MedCity News .
Rural hospitals need trained people, tested backups, downtime procedures, clinical continuity planning, and staff who know what to do when an alert fires. A framework won’t substitute for that work, but it helps organize it and makes progress measurable. The post Rural US Healthcare Has A Cybersecurity Problem — CMS Funding Can Help Fix the Part No One Sees appeared first on MedCity News .
Picture this: A little girl is in the kitchen, helping her family cook dinner, and suddenly, she turns on the blender with her mind. It’s not science fiction; it’s happening now with the use of brain-computer interface technology, which interprets electrical brain activity through an EEG headset or an implanted chip and sends those signals to a connected computer — or, in this case, a blender. For students with ALS, severe cerebral palsy or spinal cord injuries, BCI could eventually replace or augment assistive tools such as eye-tracking software and switch controls, offering a faster,…
Article URL: https://wordpress.com/education/ Comments URL: https://news.ycombinator.com/item?id=49344767 Points: 1 # Comments: 0
Article URL: https://en.wikipedia.org/wiki/Inside_American_Education Comments URL: https://news.ycombinator.com/item?id=49344411 Points: 13 # Comments: 0
AI cheating will decrease, the techlash will continue and other predictions for the coming school year.
arXiv:2607.06008v3 Announce Type: replace-cross Abstract: While Large Language Model (LLM) agents excel at monolingual long-horizon planning and tool use, enterprise workflows inherently require processing multilingual resources across extended trajectories. The interaction between multilinguality and long-horizon execution, however, remains underexplored. We introduce PolyWorkBench, a benchmark designed to evaluate LLM agents on multilingual, long-horizon workplace workflows. PolyWorkBench features 67 tasks across five core domains: commerce, knowledge work, legal analysis, localization, and manufacturing. Tasks are authored by the paper's authors from real-world data seeds and independently verified through a second-author audit. Agents must integrate heterogeneous multilingual inputs, execute iterative tool-use trajectories, and produce structured domain artifacts. To rigorously assess performance, we adopt Grade, a task-specific structural scoring rubric, as our primary ranking met
arXiv:2606.21037v2 Announce Type: replace-cross Abstract: The empirical foundation of cyber deception relies on human-centered hypotheses, but the rapid emergence of autonomous, AI-enabled attackers challenges whether this foundation transfers to AI agents. To address this, we introduce an automated evaluation framework adapted from the Honeyquest instrument to assess LLM attacker judgment at scale. Our 21-LLM cohort spanned 10 providers, diverse architectures and specializations, open- and closed-weight models, and parameter scales from 8B to over 1T. We evaluated the performance of this LLM cohort (yielding 10,962 responses) against the 47-participant human baseline across an identical set of 174 reconnaissance queries. Our empirical evaluation reveals three key findings that establish LLMs as a distinct attacker class: (1) every model in our cohort falls for deceptive traps at a significantly higher rate than human attackers; (2) the defensive attention-diversion effect observed in
arXiv:2604.11399v2 Announce Type: replace-cross Abstract: Multimodal adaptation can erode temporal reasoning (TR) in video-language models (VLMs), leaving models able to perceive salient events yet unable to infer their temporal and causal structure. We introduce MERIT, a gradient-free framework that repairs this capability through layer-selective model merging. MERIT assigns each self-attention layer a VLM-dominant or LLM-dominant interpolation and uses the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) to search the resulting combinatorial space under an objective that rewards TR gains while penalizing temporal perception (TP) degradation. Across three VLM families and five video benchmarks, MERIT consistently improves TR while preserving TP; recipes selected on a compact diagnostic set transfer to four unseen benchmarks, with relative gains of up to 27.8%. Interventional masking and frame-level attribution further show that the selected layers are functionally important fo
arXiv:2602.12276v2 Announce Type: replace-cross Abstract: Test-time scaling has become a standard way to improve performance and boost reliability of neural network models. However, its behavior on agentic, multi-step tasks remains less well-understood: small per-step errors can compound over long horizons; and we find that naive policies that uniformly increase sampling show diminishing returns. In this work, we present CATTS, a simple technique for dynamically allocating compute for multi-step agents. We first conduct an empirical study of inference-time scaling for web agents. We find that uniformly increasing per-step compute quickly saturates in long-horizon environments. We then investigate stronger aggregation strategies, including an LLM-based Arbiter that can outperform naive voting, but that can overrule high-consensus decisions. We show that uncertainty statistics derived from the agent's own vote distribution (entropy and top-1/top-2 margin) correlate with downstream succes
arXiv:2601.11496v3 Announce Type: replace-cross Abstract: AI agents increasingly mediate bargaining, negotiation and persuasion for people and firms. Such markets extend software-mediated commerce, but add a governance problem: independent model releases change delegates available to participants. Game theory shows that expanding a strategy set can harm equilibrium outcomes, but mostly through constructed examples. Deployed AI-agent logs are scarce, proprietary and privacy-sensitive, and lack counterfactuals and payoff labels. We therefore use GLEE, an independently collected benchmark of 587K strategic decisions by 13 large language models across 1,320 matched bargaining, negotiation and persuasion configurations, to study model release as strategy expansion. Across more than 50{,}000 release comparisons, many releases move payoffs in opposite directions: one agent gains while the other loses. We identify the Poisoned Apple effect: a released model that no agent adopts in equilibrium
arXiv:2509.21514v4 Announce Type: replace-cross Abstract: Research on Knowledge Tracing (KT) models traditionally focuses on improving predictive accuracy. However, responsible real-world deployment requires models to know when to defer uncertain predictions to a human teacher. We introduce an intrinsic selective prediction layer for existing KT models using Monte Carlo Dropout (MC-Dropout) to quantify uncertainty. We evaluate this approach across three architectures (DKT, SAKT, and AKT) using the Eedi mathematics dataset. Abstaining on the 20\% most uncertain predictions lifts accuracy by 2.3 to 3.0 percentage points, AUC by 1.9 to 2.4 percentage points and F1 by 1.4 to 4.3 percentage points without any retraining. This abstention strategy is highly targeted: the deferred set exhibits 1.45 to 1.60 times the error rate of the kept set. Furthermore, this targeting holds within every question-difficulty quartile and remains fair across student-ability levels. Importantly, MC-Dropout vari