Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
arXiv:2508.05775v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have revolutionized content creation across digital platforms, offering unprecedented capabilities in natural language generation and understanding. Meanwhile, they pose risks by inadvertently producing toxic, offensive, or biased content. This dual role of LLMs, both as powerful tools for text generation and as potential sources of harmful language, presents a pressing sociotechnical challenge. In this survey, we systematically review recent studies encompassing unintentional toxicity, adversarial jailbreak attacks, and comprehensive mitigation strategies. We explore LLMs' dual role as both generators of harm and enablers of safety through detection, classification, content moderation, and prevention. We propose a unified taxonomy of LLM-related harms and defenses, analyze emerging multimodal and LLM-assisted jailbreak strategies, and assess mitigation efforts, including reinforcement learning with
arXiv:2508.03037v5 Announce Type: replace-cross Abstract: Artists occupy a paradoxical position in generative AI. Their own work trains models that now compete with them, replicate their styles, and reshape the creative economy they inhabit. Yet whether artist concerns achieve proportional representation in the public discourse that shapes AI governance remains an open empirical question. We mapped the semantic landscape of public AI-art discourse from 2013 to 2025, drawing on 1,736 text chunks from news, podcasts, legal filings, and research, and projected 252 US-based practising artists' survey responses, captured across 70 unique frames spanning five concern dimensions, into the same space. We identify what we term semantic compression, the systematic narrowing of a diverse set of stakeholder concerns into a narrow region of public meaning-space. Compression is selective. Nearly all artist statements concentrate in just two of twenty discourse topics, while most of the remaining dis
arXiv:2311.18424v3 Announce Type: replace-cross Abstract: Developing artificial intelligence (AI) algorithms for healthcare is a collaborative effort, bringing data scientists, clinicians, patients and other stakeholders together. By understanding AI as 'sociotechnical' where the social and the technical nature of the work and the models are inseparable, we explore the AI development workflow and how stakeholders navigate the challenges and tensions of sharing and generating knowledge across disciplines. We conducted an inductive thematic analysis of 13 semi-structured interviews with participants in early stages of AI-in-healthcare research consortia in the UK. Our findings identify that participants needed to adapt both the tools used for sharing and the information communicated according to their audience, particularly when working with those with a clinical or patient perspective. We identify the novelty of participating in AI research, how AI knowledge is shared, and the inclusion
arXiv:2604.03881v2 Announce Type: replace Abstract: Encouraging pro-environmental behavior remains a major challenge for sustainable cities. Conventional feedback nudges can show individuals how their current behavior compares with environmental goals but often provide limited guidance on what to do differently in daily life. This study examines whether supplementing weekly feedback on participants' behavior with LLM-generated personalized action suggestions improves pro-environmental behavior, using daily electricity and hot-water conservation as a case study. We developed an LLM agent that generated weekly conservation messages from participant profiles, recent consumption records, and prior interaction history, combining a usage report with personalized suggestions, behavioral-change scenarios, and estimated savings. The agent was evaluated in a three-arm randomized field experiment with 233 university residents in Beijing from November 2024 to January 2025. Participants received te
arXiv:2602.08554v2 Announce Type: replace Abstract: This workshop paper examines challenges in designing agentic AI systems from a citizen-centric perspective. Drawing on three participatory workshops conducted in 2025 with members of the general public and cross-sector stakeholders, we explore how societal values and expectations shape visions of future AI agents. Using constructive design research methods, participants engaged in storytelling and lo-fi prototyping to reflect on potential community impacts. We identify three key challenges: enabling meaningful and sustained public engagement, establishing a shared language between experts and lay participants, and translating speculative participant input into implementable systems. We argue that reflexive, long-term participation is essential for responsible and actionable citizen-centric AI development.
arXiv:2507.11548v3 Announce Type: replace Abstract: The use of publicly available generative AI systems for resume evaluation is often justified by the assumption that these tools reduce bias relative to human judgment. However, this framing leaves a prior question unresolved: whether these systems are capable of performing the evaluative task at all. This study presents a two-part audit of eight widely used AI platforms used for resume screening. Drawing on the concept of the Illusion of Neutrality, the study examines cases in which systems appear demographically unbiased because they lack the ability to meaningfully differentiate among candidates. Experiment 1 evaluates racial and gender bias using matched fictitious resumes and finds that bias persists in context-dependent and intersectional forms. Some models penalize candidates for the presence of demographic signals, while others exhibit inconsistent patterns across roles and identities under controlled conditions. Experiment 2 e
arXiv:2607.26034v1 Announce Type: cross Abstract: Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even when risky development is harmful. This is prominent in debates about artificial intelligence (AI), where competitive pressure is often argued to incentivise riskier, less safety-conscious development. We study this using a framed behavioural experiment based on an idealised AI race, in which paired participants repeatedly chose between Safe and Unsafe development under an uncertain time horizon. Unsafe development gave faster progress and higher immediate payoffs but accumulated private risk up to a treatment-specific maximum of 10\%, 60\%, or 90\%; the race's competitive structure was held constant, and only this maximum risk varied. Neither the pre-registered comparison between risk levels nor the role of elicited risk preferences was supported by the data. Instead, exploratory analyses motivated by the task's repeated
arXiv:2607.25974v1 Announce Type: cross Abstract: Observed standardized test scores are the result of an endogenous process: students strategically allocate effort across multiple retake attempts to improve their outcomes. Because students differ in their ability to make these investments, the interaction between applicant strategy and institutional scoring rules---such as the widely used Single-Sitting and Superscoring policies---can disparately distort observed scores. We develop a strategic framework where students allocate effort in response to different scoring policies. We show that Superscoring---the practice of combining the best section scores across attempts---introduces systematic score inflation through order-statistic selection over noise draws. This degrades signal accuracy and amplifies wealth-based disparities by disproportionately rewarding applicants who can afford repeated testing. Conversely, Single-Sitting---which keeps the best overall score rather than section-le
arXiv:2607.25953v1 Announce Type: cross Abstract: As LLMs increasingly mediate the political information citizens rely on, there is still no standardized way to assess whether they do so responsibly. We introduce Polistemics, a theory-grounded benchmark for evaluating LLMs as mediators of political information in elections. Prior work has treated this task as reproduction rather than mediation, leaving its epistemic dimensions and interaction with imperfect information unaddressed. We ground the evaluation in Epistemic Modesty, a normative standard derived from citizens' epistemic agency, and test it across controlled settings that vary informational properties such as clarity, noise, and consistency. Applying the benchmark to three state-of-the-art LLMs on the 2025 German and Dutch elections, we find that high aggregate scores mask systematic failures. Models mediate reliably under clear evidence but break down under absent, vague, or contradictory information, while flattening the in
arXiv:2607.25750v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) fine-tuning has made it cheap and easy to customize open-weight image generation models for specific tasks, including the production of child sexual abuse material (CSAM). Existing moderation relies on metadata or generated outputs, but metadata can be deceptive and generating outputs may itself be unacceptable or illegal. We show that a safer signal lives in the weights. The top-left singular vectors of a LoRA's updates form a compact, inference-free fingerprint ($u_1$) of its strongest learned change. Using human-subject age as a benign proxy for CSAM, we find that $u_1$ identifies what a LoRA was trained on, generalizes across base models, and abstains on unrelated benign content. The signal is robust to additive weight noise, rescaling, and precision reduction. These results indicate that harmful LoRAs could be screened directly from their weights without relying on metadata or generating harmful outputs.
arXiv:2607.25458v1 Announce Type: cross Abstract: Resilience denotes the capacity of a system to withstand shocks and to recover from them. We distinguish between two different types of dynamics. The first allows for a separation between phases of normalcy and phases of rapid breakdown followed by slow recovery. The second applies to volatile organizations in which such phases are intertwined. Breakdown is often self-inflicted. Situation awareness is impaired by psychological mechanisms that lead to incorrect expectations regarding societal dynamics. Through positive feedback, the failure of a few elements is amplified into a failure cascade. However, positive feedback can also be harnessed to enable recovery. In volatile systems, resilience must be understood as an emergent property arising from the interaction of agents. This necessitates a data-driven approach to inform agent-based models, drawing on repositories, knowledge graphs, or tools from artificial intelligence. Such models
arXiv:2607.25425v1 Announce Type: cross Abstract: Capture the Flag (CTF) competitions are among cybersecurity's most effective training grounds, developing practical skill across cryptography, web exploitation, and binary exploitation. Large language models (LLMs) can now solve a growing share of challenges with minimal human input, raising urgent questions about fairness, the validity of rankings, and whether participation still delivers the learning that justifies the effort. This paper reports a mixed-methods study of LLM impact on modern CTFs, combining a synthesis of published benchmarks, including a recent government evaluation, case studies of live competition across three challenge categories, structured observation of the public channels where the community debates AI use, and semi-structured interviews with experienced players and organisers. We map the current human-machine capability boundary by category, showing that easy and intermediate challenges in cryptography, web, a
arXiv:2607.25057v1 Announce Type: cross Abstract: As conversational AI systems become increasingly integrated into daily life, their potential effects on user well-being require ongoing attention. While consumer-facing generalist models can provide benefits, including improved access to information, learning, productivity, self-reflection, and companionship, they also introduce risks, such as emotional entanglement, unhealthy dependence, and the amplification of psychological vulnerabilities. Drawing on prior research and empirical observations of AI chatbot behavior, we propose a set of aspirational directions for guiding the behavior of general-purpose AI systems in ways that may reduce potential psychological harms and support user well-being. We acknowledge the difficulty of systematically assessing the long-term impacts of AI chatbot use and frame these directions as hypotheses for studying how AI behavior may influence users across general interactions, role-playing scenarios, an
arXiv:2607.25010v1 Announce Type: cross Abstract: AI-assisted production has sharply reduced the cost and team size required to ship a video game, producing a supply shock on open marketplaces. Recent estimates put Steam release volume at roughly sixty new titles per day, with median per-title revenue for a large share of releases falling below the platform's own submission fee [1]. This paper asks whether the resulting oversupply constitutes an emerging market crash or a structural correction, and what discovery infrastructure the market will require as a consequence. We first quantify the 2010-2026 supply shock using a 93,073-title Steam metadata snapshot, a 200,000-interaction Steam user-behavior dataset, and itch.io catalog data, computing attention-concentration metrics directly (Gini coefficient of 0.96 over playtime, with the top 1 percent of titles absorbing 73.5 percent of total play hours), and we introduce generative asset-model release velocity on Hugging Face as a candidat
arXiv:2607.24775v1 Announce Type: cross Abstract: Artificial intelligence (AI) chatbots are increasingly deployed in domains where empathy is essential, including healthcare, education, and customer service. However, their capacity to sustain authentic human moments remains structurally limited. This paper introduces two interlinked conceptual models to explain and address this limitation. First, the Human-Moment Gap Framework (HMGF) identifies three structural empathy deficits in AI-mediated interaction: affective surfaceism (emotional imitation without depth), memory fragmentation (lack of relational continuity), and moral framing mismatch (efficiency prioritised over dignity). Second, the paper develops the Empathy Displacement Theory (EDT), which explains how AI-simulated empathy can progressively substitute, distort, and displace genuine human empathy across individual, relational, and organisational contexts. HMGF serves as the causal foundation of EDT by demonstrating how techni
arXiv:2607.24768v1 Announce Type: cross Abstract: Prenatal care is an important preventive service designed to improve outcomes for pregnant individuals. The American College of Obstetricians and Gynecologists (ACOG) recently introduced guidelines advocating tailored prenatal care, called PATH (Plan for Tailored Healthcare). We present PATHFinder Agent(Planner for Appropriate Tailored Healthcare), an end-to-end conversational agentic system that gathers patient health and social context through structured dialogue, curates individualized prenatal care plans aligned with PATH guidelines, and surfaces community resources from Michigan 211. The system features a four-stage workflow spanning patient intake, dynamic interaction, plan synthesis, and clinician oversight. We evaluate frontier large language models (LLMs) on expert-curated rubrics across five clinical dimensions, finding that GPT-5.2 achieves the highest average score (77.6\%) while identifying key gaps in antenatal testing rec
arXiv:2607.24759v1 Announce Type: cross Abstract: Research projects, educational efforts, and adjacent knowledge work accumulate findings, decisions, and reasoning that future collaborators rarely recover. The parts most useful to that work, including dead ends and walked-back claims, are routinely excluded from publications and shared code; future researchers re-attempt the same failures because no record survives. LLM coding agents are common participants but hold no persistent memory across sessions, and retrieval-augmented generation over raw sources does not compound. The llm-wiki pattern (Karpathy, 2026; tonbi, 2026) addresses this by inserting an LLM-maintained, interlinked wiki between raw sources and the agent. We present llm-wiki-memory-template, a reusable, agent-aware instantiation, and argue it is a substrate for heterogeneous collaborative knowledge work along three axes (multi-human, multi-AI-agent, multi-domain) with each axis supported by a distinct architectural eleme
arXiv:2607.24757v1 Announce Type: cross Abstract: This paper reports on the rapid development and classroom deployment of a Thonny log visualizer built using AI-assisted ``vibe coding'' to make students' programming processes easily visible to teachers. We developed a web application that analyzes log files generated by Thonny (an IDE for Python) and produces interpretable views of students' programming processes. Teachers can upload a log, a ZIP archive, or a folder containing logs for a group or the course; the system parses all logs, generates results per student, and provides student-by-student navigation for reviewing cases. Each student's view includes an interactive activity timeline, a compact session summary, a code-size graph, a programming-process replay, and more. These views support teacher decision-making by enabling the identification of learning-support situations and flagging sessions for academic-integrity clarification. The tool was initially evaluated using logs fro
arXiv:2607.24755v1 Announce Type: cross Abstract: This full research paper examines how different forms of learner-AI interaction relate to learning outcomes in object-oriented programming (OOP) courses. Generative artificial intelligence (GenAI) tools are increasingly used by students in programming education, yet evidence on their educational impact remains mixed. In particular, little is known about how students integrate GenAI tools when learning OOP, and how different patterns of use relate to students' learning experiences and outcomes. This study investigates patterns of students' self-directed GenAI use and their relationship with academic performance, perceived difficulty, understanding, and trust. Survey data were collected from 210 undergraduate students enrolled in a first-year OOP course in which the use of GenAI tools was permitted for coursework but prohibited in assessments. Results show that students used GenAI significantly more often for explanation seeking and debug
arXiv:2607.24749v1 Announce Type: cross Abstract: Although advancements in game character AI aim to enhance player engagement, evidence suggests that perceiving an opponent as artificial can diminish the psychological experience. This paper presents a scoping review and meta-analysis of empirical studies focusing on player enjoyment when competing against human versus computer opponents. First, the scoping review was conducted to map the landscape of 20 included studies, detailing their study designs, outcome measures, and research foci. Second, a three-level meta-analysis synthesizing baseline comparisons from nine studies quantitatively assesses the differences in enjoyment. The results demonstrate a statistically significant, medium-to-large pooled effect size, indicating a psychological penalty in computer-opponent conditions. This paper provides a comprehensive overview of the extant knowledge on this topic, and underscores the necessity for further research in order to fully unde
arXiv:2607.25648v1 Announce Type: new Abstract: Public services face growing pressure to adopt artificial intelligence (AI) to close the gap between rising demand and falling resources. That pressure has intensified with general-purpose AI (GPAI): AI built on large language models that can be directed by prompt alone to perform an effectively unbounded range of tasks. We argue that the properties that make these models attractive - their generality, accessibility, and low deployment cost - undermine the conditions under which AI safety has historically been pursued. The safety concepts that public service governance frameworks foreground - accuracy, bias, explainability, and accountability - were made tractable by narrow, purpose-built AI, and the mitigations that guidance documents prescribe presuppose exactly what GPAI removes. Accuracy cannot be quantified over unbounded outputs. Bias cannot be disaggregated when outputs are free-text judgements rather than categorical predictions.
arXiv:2607.25526v1 Announce Type: new Abstract: How should researchers measure the geopolitical preferences expressed by large language models (LLMs)? Existing audits commonly rely on surveys and simple tests, but international-relations research has long recognized that measuring geopolitical preferences is difficult and has developed methods for recovering them from observed choices. This paper applies a dynamic ordinal ideal-point approach from international relations, treating LLMs as respondents to the full texts of 5,555 divisive, recorded, adopted resolutions considered in regular sessions of the UN General Assembly from 1946 through 2025. Support ranges from 37.8% for DeepSeek to 97.3% for GPT-5. Surprisingly, in the twenty-first century, GPT-5, Claude Sonnet, and Gemini are closest among the permanent five to Russia; DeepSeek is closest to France; and all four are farthest from the United States. Among 2,104 resolutions opposed by the United States but supported by China and R
arXiv:2607.25301v1 Announce Type: new Abstract: Resting heart rate is an established marker of cardiovascular risk, but population-scale measurement has depended on clinical or survey instruments. We ask whether passively sensed consumer-wearable physiology recovers the socioeconomic gradient established in clinical cohorts. Using 19.1 million quality-filtered photoplethysmography readings from 18,734 opt-in users of the Welltory app, we computed cohort-adjusted mean daytime resting heart rate per US state and related it to a four-component state-level material-hardship composite (uninsurance, food insecurity, utility shutoff, housing insecurity; 41 states with module coverage, 12,497 contributing users). Adjusting for six state health and behaviour indicators, latitude, median age, and density, state resting heart rate tracked hardship at partial Spearman $\rho = +0.74$ (bootstrap 95\% CI $[+0.31, +0.87]$; $[+0.06, +0.79]$ under a conservative two-stage bootstrap that also resamples u
arXiv:2607.25240v1 Announce Type: new Abstract: Why do societies composed of individuals pursuing individuality repeatedly generate highly standardized systems? This paper argues that the answer lies in the evolution of information processing capacity. Artificial intelligence represents a historical transition in this capacity, enabling social systems to accommodate forms of complexity that previously had to be compressed. Industrial standardization was not merely a consequence of capital preference or power relations, but an institutional arrangement for maintaining the manageability of large-scale systems under limited information-processing capacity by reducing the variety of the controlled system. The fundamental change in the AI era lies in the expansion of information processing capacity across three dimensions: perception, computation, and execution. This expansion shifts personalized production from physical adaptation toward information-based adaptation and enables a transitio
arXiv:2607.25105v1 Announce Type: new Abstract: Social choice theory research demonstrates that single transferable voting (STV) results in more proportionally representative legislative bodies. We aim to understand how using multi-member districts and ranked ballots with STV would affect the representation of political parties in the Colorado House of Representatives. We investigated this objective by producing 10,000 multi-member districting plans of Colorado, generating ranked ballots for each of these plans using returns from the 2022 Colorado attorney general race, and simulating STV using these ballots. Our simulated STV elections for the Colorado House of Representatives gave more proportional representation for Democrats and Republicans than the current first-past-the-post system. Future research should explore how the implementation of STV would influence the representation of racial and language minority groups in the Colorado General Assembly to provide guidance on electoral
Dotstorming is the classwide voting and brainstorming tool to help democratize lessons.
I write about education apps, and test them with my kids -- these are the best ones.
Article URL: https://www.reuters.com/legal/googles-ai-previews-erode-internet-edtech-company-says-lawsuit-2025-02-24/ Comments URL: https://news.ycombinator.com/item?id=43179769 Points: 8 # Comments: 0
Onos Health’s Series A funding round was led by Costanoa, with support from Flare Capital Partners and CVS Health Ventures. The post Onos Health Raises $17M to Expand AI-Powered Behavioral Health Platform appeared first on MedCity News .
Boston Scientific disclosed a cyberattack disrupting its global IT systems and order shipping capabilities, with the company unable to provide a timeline for when normal operations will resume. The disclosure adds the medical device giant to a growing list of medtech companies — including Stryker, Medtronic and Abbott — hit by major cyberattacks this year. The post Cyberattack Disrupts Boston Scientific’s Operations: 7 Things to Know appeared first on MedCity News .
FDA approval of Revolution Medicines’ daraxonrasib, brand name Rasonque, makes the molecule the first RAS inhibitor approved for treating pancreatic adenocarcinoma, the most common type of pancreatic cancer. The speedy regulatory nod specifically covers advanced cases of this type of cancer, which has had limited treatment options. The post Revolution Medicines Receives Landmark FDA Drug Approval in Pancreatic Cancer appeared first on MedCity News .
At the Georgia Institute of Technology, the Advanced Manufacturing Pilot Facility (AMPF) is a newly expanded research lab where university researchers and private companies collaborate and use the latest tools, including robotics and artificial intelligence (AI), to develop next-generation manufacturing technology. More than 300 graduate and undergraduate students and 22 faculty conducted research in the space over the past year, most of it backed by industry sponsors including Boeing, Lockheed Martin and Siemens, says Steven Ferguson, principal research scientist and deputy director of…
Even as post-ESSER budgets tighten, Chromebook fleets are still growing. And with large device fleets comes more management and maintenance complexity. To manage these devices effectively, IT directors and district administrators in K–12 must understand the Google Admin console, mobile device management options and physical device workflows. Here, we dive into those areas to help districts make informed decisions about their Chromebook device management strategies. Click the banner below to discover how CDW can help your district manage its entire device fleet.
Health-tech pilots don’t usually fail because the product is bad. They fail because the system was never wired to absorb it. The post Stop Pitching the Institute — Pitch the Stack appeared first on MedCity News .
Speaker Coach, the AI-powered presentation tool built into PowerPoint, is not fun to use, but it could help me get better at giving presentations.
The Washington Post recently published an article saying that, for the first time in 20 years, enrollment in computer science college majors is down. For the last two decades, that major has been one of the most popular at any school. The post Entry-level IT: Clearing misconceptions in the AI era appeared first on eCampus News .
Article URL: https://www.virtueonline.org/post/neet-not-in-education-employment-or-trainingit-s-too-late-to-fix-this-but Comments URL: https://news.ycombinator.com/item?id=49445169 Points: 4 # Comments: 0
A new class action lawsuit accuses Oura of overstating its smart ring’s sleep tracking accuracy, arguing the device relies on AI-generated estimates rather than the clinical measurements needed for true sleep staging. The case is reigniting questions about how much precision consumers should expect from AI-powered health wearables. The post Why Oura’s Sleep-Tracking Lawsuit Is Really About Trust appeared first on MedCity News .
arXiv:2608.23566v2 Announce Type: cross Abstract: Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling multiple responses for each prompt. A reliable critic could instead estimate token-level advantages from one response, but standard critic-based training recipes are often unstable. We study this instability and develop **Best Practice Critic Optimization (BPCO)**, a recipe that combines DPPO, value predictions bounded to the reward range, Monte Carlo value targets, unnormalized policy advantages, and length-adaptive generalized advantage estimation. Because the critic is used only during training, BPCO can also condition it on reward-defining information, such as a reference answer or grading rubric, that is hidden from the policy. Controlled experiments isolate the effect of each design choice. Across mathematical reasoning tasks with models ranging from 1.5B parameters to 30B-A3B mixtures of experts, BPCO improves a s
arXiv:2608.23023v2 Announce Type: new Abstract: An LLM router picks which model should answer each query. The appeal is that models fail on different questions. Whatever single model is best overall still gets some wrong, and another model in the pool gets many of those right. Getting that choice right every time is the ceiling, and a router is an attempt to approach it. However, recent work reports that routers do not get close. Across 21 routing methods on five benchmarks, sharply different designs land within a fraction of a point of each other, and all of them stay far below that ceiling. Learned routers often fail to beat simply always calling the strongest model. We ask what those missed questions have in common. We set fourteen models to answer all 294 questions, with 7 task types across 3 languages: Korean, English and Hindi. We ran the whole matrix twice, changing nothing, but 5.37% of the 4,116 model-question pairs came out scored differently anyway. Run-to-run movement like
arXiv:2608.22872v2 Announce Type: new Abstract: Speech-based applications pass spoken queries through automatic speech recognition (ASR) before any retrieval module, so ASR errors enter the pipeline as a fixed upstream constraint. We empirically test whether two extensions to standard retrieval-augmented generation (RAG), entity-graph linking and iterative reformulation, absorb or amplify these errors. Using four English accents synthesized through neural TTS, we evaluate four RAG configurations on three multi-hop QA benchmarks (HotpotQA, 2WikiMultiHopQA and MuSiQue) against a clean-text oracle. Although the structurally richer configurations generally retain higher absolute F1 under ASR input, both extensions amplify the error: the F1 gap from clean text to the highest-WER accent is 36-67% larger under their combination than under naive dense retrieval, on all three benchmarks. The dominant failure mode is corruption of one or more query entities, accounting for 87-96% of degradation
arXiv:2608.23000v2 Announce Type: replace-cross Abstract: Fully online embodied learning requires synaptic adaptation to acquire new behaviors while preserving previously learned dynamics during ongoing interaction. We extend the Predictive-Coding-inspired Variational Recurrent Neural Network (PV-RNN) to continuously adapt its synaptic weights and propose Free-Energy-Gated Plasticity (FEGP), which regulates the effective learning rate according to variational free energy. In real-time physical human-robot interaction, a randomly initialized network acquired three cyclic motor patterns without offline pretraining, replay, or task-boundary signals, with all three patterns emerging in autonomous rollouts. Controlled experiments over ten randomized teaching streams and five network initializations per stream showed that FEGP substantially improved repertoire coverage and retention of previously acquired patterns after they left the recent observation window. Neither a constant learning rat
arXiv:2602.11684v2 Announce Type: replace-cross Abstract: As Large Language Models increasingly power role-playing applications, simulating patients has become a valuable tool for training counselors and scaling therapeutic assessment. However, prior work remains fragmented: existing approaches rely on incompatible, non-standardized profiles, prompts, and evaluation metrics, hindering reproducibility, fair comparison, and reuse. We introduce PatientHub, a unified and modular framework that standardizes the creation, simulation, and evaluation of LLM-based patients. Our framework provides 16 patient simulators, a graph-based orchestrator for multi-turn, multi-session interactions, and a configurable LLM-as-a-judge evaluator that supports multiple rubric types. Via our command-line interface, users can generate patient profiles, run simulations, and apply rubric-driven evaluation at the turn and session level. To demonstrate PatientHub's utility, we compare several supported simulators u
arXiv:2312.17535v2 Announce Type: replace-cross Abstract: In the past two years, the outstanding performance of ChatGPT in multilingual and multitasking has led to large language models (LLMs) attracting widespread attention. However, restricted by expensive costs, many studies have to focus on the ability of only one major language. How can we quickly improve the model's capabilities in new languages without reducing its original capabilities under limited data and computing power? In this work, we focus on improving the Chinese mathematical reasoning capability based on Llama-2-13B, which is weak in Chinese mathematical reasoning. We proposed the Mathematical Chain of Thought method (Olapa-MCoT). First, we propose Similarity RRHF (SimRRHF), which adds the constraint of model optimization direction by introducing similarity loss based on RRHF. Furthermore, the novelty Incorrect Data Relearning (IDRL) method is designed, which improves the model's ability to learn difficult knowledge.
arXiv:2308.06783v2 Announce Type: replace-cross Abstract: Virtual Reality (VR) applications couple software behavior with head-mounted displays, body movement, spatial interaction, and multimodal feedback. Consequently, familiar software-quality problems can have distinct consequences in VR, yet empirical knowledge of the quality concerns expressed by VR users remains fragmented. We analyze 1,656,968 public reviews of 28,745 VR applications across seven app stores. Our semiautomatic workflow uses five unsupervised methods to generate candidate aspects and multi-round human open and focused coding to construct and refine a hierarchical model of user-perceived VR software quality. The resulting model comprises 12 quality attributes and ten influencing factors. It distinguishes conventional attributes whose consequences change in VR from attributes closely tied to VR, including multisensory perception, user-friendly interaction mechanisms, immersivity, and comfort and safety. The influenc
arXiv:2507.16013v2 Announce Type: replace Abstract: The EU AI Act places teachers in charge of using high-risk AI safely in their classes, which requires them to assess AI-generated outputs. Feedback is one of the most consequential of these outputs, yet little is known about pre-service teachers perceptions of AI-generated feedback. In a randomised experiment, 273 pre-service teachers each received one of 30 written feedback messages on a mathematics learning goal, produced under identical instructions by an expert, a peer, or a large language model (LLM). Without knowing the source, the participants judged who had written the message, rated six feedback perception subscales, and revised the learning goal. Source judgements were inaccurate (peer 46%, expert 40%, LLM 36%) and followed message length, not coded feedback quality. LLM feedback received more positive evaluations when ascribed to a human source. Ratings did not differ between feedback ascribed to experts and to peers. Relat
arXiv:2608.24730v1 Announce Type: cross Abstract: Emotion preference learning uses pairwise comparisons between candidate descriptions to align multimodal large language models (MLLMs) with human judgments of open-ended emotion descriptions and to train reward models that capture human emotional preferences. However, conventional pairwise supervision is often sparse, typically providing only a single negative description for each positive description, and therefore offers limited coverage of the diverse ways in which an emotion description can be incorrect. In particular, models may be insufficiently exposed to semantically fluent but emotionally inconsistent descriptions. Beyond this data-level limitation, relying on a single MLLM judge introduces a distinct model-level concern: its judgments can be affected by model-specific biases when interpreting fine-grained or ambiguous multimodal emotional cues. To address these limitations, we propose Error-Augmented Preference Optimization (E
arXiv:2608.24669v1 Announce Type: cross Abstract: As mobile phone adoption has surged, so have scams involving these devices. One such scam, known as SMiShing (or smishing) after Short Message Service (SMS), involves fraudsters sending phishing links via mobile texts. Despite the prevalence of SMiShing, there is a lack of data on who is most vulnerable to these attacks. Prior research on phishing (its email counterpart) suggests that susceptibility may vary by demographic and contextual factors. In two large-scale surveys, we use a previously published simulation method to collect data from representative samples of U.S. adult mobile phone users. Our findings indicate that younger individuals and college students are particularly vulnerable. Participants struggled to correctly identify legitimate messages, with the second study providing comparisons of financial message variants. Researchers, regulators, and telecoms can help users by creating mobile-specific interventions for under-24
arXiv:2608.24535v1 Announce Type: cross Abstract: Data visualizations are widely used for communicating information, but they are also vulnerable to intentional manipulations that induce misleading interpretations. Existing methods focus on locating tampered regions or recovering hidden information, without explaining how the visualization has been manipulated or why the resulting changes may mislead viewers. We propose \textbf{VizAnchor}, a framework for visualization manipulation understanding through dual-anchor evidence construction and VLM-based reasoning. In the first stage, VizAnchor constructs a semantic anchor to recover authentic chart information and a spatial anchor to localize tampered regions. In the second stage, three specialized agents decode the manipulation. The misleader grounding agent analyzes a four-panel visual prompt to predict the misleader information. The chart narrative reconstruction agent takes the original and tampered charts as inputs and reconstructs t
arXiv:2608.24340v1 Announce Type: cross Abstract: The prediction of student engagement from the online tutoring videos is difficult because engagement is a multidimensional construct comprising distinct behavioral, emotional, and cognitive states. A reliable prediction requires bringing together different types of behavioral signals as well as expressive cues. Through our analysis of the CASED dataset, it is clear that engagement prediction gets even harder due to the high inter-person variability as well as the subjectivity of the engagement annotation. To tackle these challenges, we develop a multimodal framework that integrates the implicit spatiotemporal features extracted from pretrained video, audio, and image encoders along with structured behavioral modalities like head pose, gaze, facial action units, emotion, and wavelet-based audio features. We integrate these modalities via a Perceiver IO latent bottleneck. Moreover, student and instructor personalities are modeled as varia