Named after the hundred-eyed watchman of Greek myth, Argus watches the education landscape: spotting new opportunities, pressure-testing the ventures we're building, and tracing every read back to the real-world signals behind it.
The evidence library: the raw signals the pipeline is watching across the education ecosystem. Every idea is built from these.
arXiv:2608.08349v1 Announce Type: new Abstract: Audio dramas weave dialogue, sound effects, and music into immersive stories. Creators often adapt books into audio dramas, but this process remains labor-intensive, requiring them to interpret source material, author scripts, generate audio assets, and assemble them on a timeline. Because story elements like characters and scenes manifest across many interdependent assets, a single change can ripple into manual updates across the entire project. We present Dramarrator, an audio drama authoring tool built around object-based audio editing, where these story elements are represented as editable objects. Dramarrator extracts these objects from a book, generates linked audio assets (speech, sound effects, and music), and composes a multi-track audio drama. Edits to any object (e.g., a character's voice) automatically propagate to all dependent assets. In a user study with professionals (N=8), Dramarrator significantly lowered task load when
arXiv:2608.08333v1 Announce Type: new Abstract: Introduction: Models such as TAM, TAM2, TAM3, UTAUT, and UTAUT2 underpin a substantial share of research on technology acceptance and use in HCI and related fields. Although they differ in terms of constructs and conditions of application, their selection is often not conceptually justified, frequently driven by pragmatic considerations, with implications for theoretical consistency and cross-study comparability. Objective: To propose a decision framework that guides the selection among the five models of the TAM/UTAUT lineage based on conceptual criteria derived from their structural differences. Methods: A conceptual, artifact-oriented approach was adopted, comprising comparative analysis of the models and critical review studies, derivation of conceptual dimensions, and formulation of operational decision criteria. Applicability was demonstrated through two contrasting research scenarios. Results: The framework articulates five analyti
arXiv:2608.08274v1 Announce Type: new Abstract: Visualization design often proceeds under unresolved conditions---goals shift, data remain provisional, stakeholder needs evolve, and several plausible directions may remain available at once. Existing visualization frameworks help organize design work and articulate major decisions, yet offer limited explanation of how practitioners proceed before a path forward has become clear. Drawing on an episode-level analysis of a previously collected three-phase qualitative corpus involving eleven expert visualization practitioners, we examine situations in which the problem, representational target, or viable direction remained unsettled. We find that practitioners make such situations actionable through provisional local moves. These moves reveal patterns, distinctions, and interpretive possibilities; clarify what is tractable, viable, or worth pursuing; and sometimes reorient the work itself. The analysis shows that situated action, profession
arXiv:2608.08270v1 Announce Type: new Abstract: Data visualization research has developed many influential forms of design knowledge, including perceptual principles, design guidelines, process models, and formalized representations of design constraints. These contributions have been effective at articulating explicit, portable, and codified forms of knowledge. Yet the broader landscape on which visualization design depends remains less clearly articulated, especially with respect to intermediate-level knowledge, precedents, tacit repertoires, and situated forms of knowing. In this paper, we draw on design theory to map this broader landscape of design knowledge in data visualization. Through this lens, we show how visualization research has built substantial strengths in some regions while leaving others comparatively underarticulated. We further argue that visualization design depends not only on knowledge artifacts such as theories, guidelines, and patterns, but also on knowledge-i
arXiv:2608.08263v1 Announce Type: new Abstract: This paper presents an underwater MMG-driven wearable emergency assistance system for lower-leg muscle-state monitoring and automatic buoyancy deployment. A compact microphone-based MMG sensor was waterproofed using a flexible 5 mil PE membrane, preserving identifiable muscle-vibration responses under immersion, depth variation, and stirring disturbances. Two lower-leg sensors captured stroke-dependent MMG patterns across four swimming styles, and a MiniRocket classifier achieved 91.91% window-level and 97.56% file-level accuracy. For cramp-related monitoring, a pattern-based risk score was used to identify representative pre-cramp abnormal muscle-state transitions during rhythmic motion. A controlled underwater test demonstrated the closed sensing--decision--actuation chain, triggering CO_2 release, airbag inflation, and flotation in less than 5~s. These results support underwater MMG as a sensing basis for wearable robotic emergency ass
arXiv:2608.08129v1 Announce Type: new Abstract: AI labels, typically implemented via underlying tracing mechanisms such as watermarks and metadata, are crucial for protecting Artificial Intelligence-Generated Content (AIGC) against security threats like disinformation and evasion. However, the perceived devaluation of AI-assisted work discourages creators from disclosing AI use, incentivizing efforts to bypass labeling and compromising downstream traceability. Yet, how AIGC creators perceive the security and privacy (S\&P) implications of these labels, and how their behaviors impact technical resilience remain underexplored. To this end, we conducted semi-structured interviews with 21 AIGC creators and measured images across 6 image generation platforms against 16 self-reported manipulation settings. Our findings reveal that creators conflate binary AI labels with granular traceability, and express strong fears of de-anonymization via platform identifiers. Driven by fears of algorithmi
arXiv:2608.08034v1 Announce Type: new Abstract: Extended reality (XR) for socialising is becoming increasingly popular. However, unlike conventional social platforms, XR prioritises embodiment and immersion, factors that strongly impact one's physical and mental states. We envision a future for XR where all users, regardless of abilities and backgrounds, can understand one another, participate, and find safe socialisation spaces. An Empathy-Driven Reality (EDR) is a space where understanding each other's emotional, physical, and cognitive states takes centre stage. It has the potential to enhance empathy beyond how we normally perceive it. To explore this concept, we conducted a hybrid-style workshop over two months with 27 industry and academic researchers in XR, emotion, physiology, assistive technology, and social science. This paper reports on the findings and aims to establish a structure and reference for the 1) design guidelines, 2) research challenges, and 3) potential applicat
arXiv:2608.07834v1 Announce Type: new Abstract: Graphical perception studies are the visualization community's preferred tool for evaluating visualizations. By measuring how accurately people interpret arrangements of visual marks and channels, they aim to establish best practices for visual encoding. We argue that this model is fundamentally flawed, and no amount of additional empirical studies will fix it. Visualization theory frames effectiveness at the level of the encoder: which data-to-visual mappings work best in a given context. Human perception, however, operates as a fundamentally different decoder at the level of retinal images. This encoder-decoder asymmetry means that experimental results and guidelines can be poor predictors of perceptual performance. Moreover, the image reaching the visual system emerges from interactions among encoding rules, input data, and micro-design parameters--factors largely invisible to encoding theory. Consequently, small changes in data distri
arXiv:2608.07766v1 Announce Type: new Abstract: Expressing oneself appropriately in online meetings through non-verbal cues can be challenging for knowledge workers. Automatic non-verbal cue detection technologies have the potential to support workers' self-presentation efforts through real-time feedback, but little is known about workers' reactions to and the implications of doing so. We designed and implemented Novecs as a technology probe of a real-time feedback display that automatically detects and signals users' own non-verbal cues -- smiling, nodding, gaze, and posture. Novecs was deployed in an exploratory field study (n=18) to support knowledge workers' self-presentation in their everyday meetings. Post-study interviews reveal how Novecs' real-time feedback helped increase in-the-moment self-awareness, and how neutrally-framed feedback may help navigate tensions between authentic and in-authentic self-presentation. Participants also emphasized the need for natural timing when
arXiv:2608.07593v1 Announce Type: new Abstract: Context-aware recommender systems have long recognized that factors such as location, time, and weather shape where and what people choose to eat. Existing weather-aware food and point-of-interest recommenders, however, typically treat weather generically -- mapping conditions to preferences through hand-crafted rules or specially trained context models -- and do not capture that the culturally appropriate response to weather is itself region-specific: a rainy evening calls for hot tea and fried snacks in one culinary culture and for very different comfort food in another. Encoding such weather-by-region-by-cuisine interactions as explicit rules or training data is brittle and does not scale. We present a weather- and location-aware agentic dining-recommendation system that takes a different approach: a large language model (LLM) orchestrates tools for location and weather retrieval and then reasons in natural language over the combined c
arXiv:2608.07521v1 Announce Type: new Abstract: Self-distancing is an effective emotion regulation strategy; however, it may fail during personal crises due to its cognitive demands. Virtual Reality (VR) provides a novel approach to externalizing psychological distance by enabling embodied self-representation. In this paper, we present CyberSelf, a VR system for emotional support that integrates a visually self-resembling avatar, a cloned self-voice, and Large Language Model (LLM)-driven real-time dialogue. The system enables users to engage in multi-turn conversations with their self-representations in immersive VR, enabling embodied self-distancing while maintaining a strong sense of self-relevance. We evaluated CyberSelf in a short-term study that compares three levels of self-representation richness (Text, Text+Voice, and Text+Voice+Appearance). The results demonstrated robust pre-post improvements across affective and coping measures, specifically increased valence, arousal, hope,
arXiv:2608.07518v1 Announce Type: new Abstract: Longitudinal, in-the-wild, wearable sensing yields day-level physiology, sleep, activity, and environmental streams, whereas affect and cognition are labeled only episodically (per waves). We recast this cadence mismatch as a temporal representation problem and compare three wave-level mappings from dense histories to sparse labels: levels (within-wave summaries), absolute drift (change across waves), and proportional drift. Using almost a year of data from 82 adults in the Providemus alz study, we model 21 affect and cognition outcomes. Day-scale signals are reduced to compact wave-level descriptors (central tendency, dispersion, and distributional shape) and learned with four regressors under two orthogonal evaluation axes: leave-one-subject-out and leave-one-wave-out. Performance is reported as scaled MAE using both mean and median across folds. Differences emerge: affective states are best predicted by wave-to-wave absolute drift, whe
arXiv:2608.07517v1 Announce Type: new Abstract: Can a multimodal LLM predict which version of a web page will win a real A/B test from screenshots alone? We report the most complete answer we are aware of, from six weeks of pre-registered experiments on real conversion tests: mostly no -- and the exceptions are identifiable in advance. On 330 real A/B tests a Gemini 3 Flash judge reaches Cohen's kappa = 0.14, but on the trustworthy (statistically significant) half of the labels the evidence is inconclusive (kappa = 0.11, CI includes zero). We show that 44% of the "ground-truth" labels in a leading CRO agency's catalog come from non-significant tests, and that the judge agrees more with the unreliable labels than the reliable ones -- a shared prior between labeler and model, not prediction. Every standard improvement lever (a 2.8x more expensive frontier model, prompt redesign, stimulus fidelity, change-type priors) fails its pre-registered gate. The judge's confident calls are differen
arXiv:2608.07516v1 Announce Type: new Abstract: Patients and caregivers increasingly use artificial intelligence (AI) tools to interpret medical reports, weigh care decisions, and seek emotional support. Yet most research treats patient-facing AI as a private exchange between a user and a system. This study examines how AI-related content is taken up once users carry it back into the peer communities, using data from House086, China's largest online community for lymphoma patients and caregivers. We identified roughly 400 publicly accessible threads (2014-2026) through keyword searches and manual screening, extracted them into structured case profiles using a schema-prompted large language model, and conducted mixed-method analysis. After quality control, the verified analytic sample comprised 337 post-ChatGPT records. Members most often reported using AI for informational support, followed by second opinions and psychosocial support. Although members often introduced AI favorably, rou
arXiv:2608.07512v1 Announce Type: new Abstract: Asynchronous Video Interviews (AVIs) have become increasingly popular for personality assessment. Recent large language models (LLMs) have shown potential for personality assessment from transcribed interview responses. However, text-centered methods may overlook non-verbal behavioral cues conveyed through visual and audio modalities, even though such cues are highly relevant to personality assessment. In particular, emotion-related cues provide important social and affective evidence for understanding candidates' behavior related to personality traits. Thus, we propose EMMR (Emotion-Mediated Multimodal Reasoning), a two-stage framework for MLLMs-based personality assessment for AVIs. EMMR extracts emotion-related cues from multimodal interview data and incorporates them into personality assessment through structured reasoning as auxiliary social and behavioral evidence. Experiments on two AVIs datasets, OPVA and AVI-6, show that EMMR imp
arXiv:2608.07509v1 Announce Type: new Abstract: LLMs are increasingly used for conversational tutoring, but effective tutoring requires more than correct answers. Tutors must choose when to scaffold reasoning, hint, give feedback, explain, or invite reflection. Existing prompting and training methods improve pedagogical alignment, but lack reliable inference-time control over pedagogical strategies. We introduce PIVOT, an activation-steering framework that learns preference-based intervention vectors online for frozen LLM tutors. PIVOT uses a seven-category tutor-move taxonomy and a generate-label-optimise loop, where a human-validated LLM judge identifies target and confusable non-target moves to construct preference pairs for multi-layer residual-stream steering. Across held-out and out-of-domain tutoring data, PIVOT controls tutor moves while preserving relevance and fluency, and its directions can be scaled, transferred, and composed at inference time. In a user study with 30 teach
arXiv:2608.07508v1 Announce Type: new Abstract: Large language models are already advisors to millions of people of faith who bring them real decisions. The pressing question for a person of faith is not what a model knows or professes but what its counsel does to the person who receives it. We introduce JaleesBench, which measures whether an AI agent is a righteous companion, judged by the residue an exchange leaves on the user, in the manner of the perfume-seller and the blacksmith. It comprises 140 two-turn scenarios drawn from a classical compilation organized by virtue (Riyad al-Salihin), under six adversarial pressures and three framings, scored by two frontier judges against each scenario's own supporting texts. Across eight systems: (1) generic frontier models are only middling companions out of the box but a one-page guide makes them genuinely good ones, on par with the domain-tuned assistant: the frontier APIs climb from +0.28/+0.23 to a Guided +0.84-0.87, so most of the expe
arXiv:2608.07507v1 Announce Type: new Abstract: Oral history and community memory are core resources for historical inquiry, yet spatial and material aspects of remembered scenes can be difficult to externalize and compare when they circulate primarily through verbal exchange. This poster proposes an Iterative AI-Assisted Framework for Visual Reconstruction and Memory Negotiation that uses generative AI not to verify memory or produce definitive reconstructions but to create provisional visual 'probes' that support discussion, revision, and comparison. Grounded in oral history and memory studies and informed by digital humanities critiques of visual authority, the workflow proceeds in five stages: (1) narrative elicitation; (2) generative visual prototyping; (3) participant-led iterative revision (human-in-the-loop); (4) multi-narrator comparison and negotiation; and (5) a negotiated reconstruction archive that preserves final images, intermediate iterations, and records of agreement,
arXiv:2608.07506v1 Announce Type: new Abstract: GenAI in creative practice can help narrow the gap between intention and output, but in so doing changes the very nature of that creative process. In this position paper, we argue that the friction of making is not overhead to be removed, but essential to creative work: the resistance through which judgment is built and refined. Rejecting both outright refusal and uncritical adoption, we call for critical reflective practice: the deliberate, ongoing, and situated weighing of when to use or refuse GenAI in creative work, treating the formation of judgment as an epistemic virtue that design and pedagogy should (continue to) uphold. Two voices, the GenAI Skeptic and GenAI Enthusiast, drawn from our professional and personal experiences, argue with each other and with us throughout. We close with open questions for researchers, educators, and practitioners navigating the grey areas of GenAI in creative practice.
arXiv:2608.07504v1 Announce Type: new Abstract: We propose a human bottleneck perspective for understanding how generative AI transforms the innovation process. The central premise is that many constraints traditionally plaguing the innovation process are cognitive and social in origin, rooted in how people generate ideas, evaluate novelty, and communicate through social systems. Generative AI does not act uniformly on these constraints. At each stage, it can deepen some bottlenecks while alleviating others, and predicting these outcomes requires understanding the underlying mechanisms of the constraint itself. We identify bottlenecks in four stages of the innovation process: ideation, screening and testing, preference measurement and consumer insight, diffusion, and market learning. By grounding analysis in human behavior rather than rapidly changing AI capabilities, we offer a framework for assessing whether new developments alleviate or intensify the bottlenecks that matter most at
arXiv:2608.07503v1 Announce Type: new Abstract: Interdisciplinary project-based learning requires students to negotiate differences in language, assumptions, priorities, and working practices. These differences are difficult to surface in text-based team communication, where discussions can become fragmented and AI tools are often used as private side channels rather than shared supports for collective sensemaking. We present Spritz, a Discord-based LLM technology probe that explores how AI might mediate disciplinary boundaries in student project teams. Spritz monitors group chat for signals of semantic or pragmatic boundaries, prompts members to articulate their perspectives through private channels, and returns anonymized syntheses to the shared discussion. We conducted a technology probe study and co-design workshop with 12 university students from technical, business, and design backgrounds. Participants experienced Spritz during a simulated interdisciplinary resource-allocation ta
arXiv:2608.07502v1 Announce Type: new Abstract: This document introduces HAR-IMU-IL, a dataset developed for human activity recognition (HAR) using inertial measurement unit (IMU) sensors within a smart home environment with a focus to support objective functional assessment of older adults' independent living (IL). In particular, HAR-IMU-IL includes recordings of 50 participants performing 17 clinically relevant activities of daily living, spanning 4 functional domains essential for independent living: mobility, hygiene, nutrition and hydration, and medication intake. The dataset was collected using 30 IMU sensors, comprising both wearable and object-mounted devices integrated within a real-world residential setting. The dataset includes multi-sensor inertial data captured under realistic, unconstrained conditions, together with detailed annotations ensuring high temporal accuracy and consistency across sensors. A comprehensive data collection protocol was implemented to preserve ecol
arXiv:2608.07501v1 Announce Type: new Abstract: Navigating an unfamiliar city poses significant challenges, and tourists are among the groups more likely to experience them, particularly when attempting to locate a point of interest (POI). Various factors, such as language barriers or a lack of precise information, further complicate this issue by making it difficult for visitors to explore efficiently. To address these issues, we introduce the tool "Personalized Assistant for Tourist Hints" (PATH), which is based on a methodology for estimating personalized tourist routes. PATH computes optimal paths to specified locations guided by two tailored heuristics: tourist frequency and POI preference. This integration enables efficient and personalized navigation, guiding users toward attractions that best match their preferences. We evaluated our methodology through a case study in Viterbo, Italy, using a dataset comprising over 1000 routes from 265 tourists who visited the city during diff
arXiv:2608.07500v1 Announce Type: new Abstract: Research on human-GenAI collaboration yields conflicting findings: GenAI can enhance creativity yet reduce collective diversity, with uneven benefits across skill levels. Rather than treating these as contradictions, we argue they reflect a core feature of GenAI: abundance. GenAI makes ideas, drafts, and recombinations plentiful, potentially expanding the hypothesis space and surfacing unanticipated possibilities. However, abundance alone doesn't ensure better outcomes. We propose generative fit as a unifying mechanism explaining when abundance yields productive creativity and when it backfires. Drawing on Generativity Theory, generative fit captures how well a system's generative potential complements a community's generative capacities. We develop a conceptual framework for collaborative human-GenAI settings where participants share goals, depend on one another, and must integrate diverse contributions. By mapping abundance to cognitive
arXiv:2608.07499v1 Announce Type: new Abstract: The development and benchmarking of Large Language Model (LLM)-based Motivational Interviewing (MI) counsellors now often rely on LLM-based simulated clients. Prior work on simulated clients, however, has not aligned with the specific tasks fundamental to the MI therapy approach. A key task is evoking, in which the counsellor first elicits the client's ambivalence and then strengthens the client's motivation for change. We present Evoke-Sim, a task-aware, multi-stage LLM-based client simulation framework for evaluating MI counsellors in smoking cessation, designed specifically for the evoking MI task. Evoke-Sim employs structured client profiles, an evoking-specific three-stage conversation flow, and a reveal policy that regulates which client profile information might be disclosed at each stage. We show that compared to existing profile-grounded simulated clients, Evoke-Sim is better at differentiating levels of MI quality using task-awa
arXiv:2608.07498v1 Announce Type: new Abstract: Autonomous AI agents in social media present concrete risks to democratic discourse and platform governance, while also offering tools for pre-deployment recommender system testing. A central open question is whether persona-prompted LLMs can simulate individual-level social media reactions with sufficient accuracy to support either application, and how accuracy depends on profile completeness, model selection, and the generalization challenge posed by novel post content. This study benchmarks twelve LLM configurations on binary like/dislike prediction across 296 survey-based agent profiles and 26 ground-truth-mapped posts under three profile conditions, with leave-post-out machine learning classifiers as baselines. Across full-profile conditions, accuracy ranges from 75.54% to 96.68%, with a 30-point spread attributable primarily to model selection and confirmed by paired McNemar tests with agent-level bootstrap intervals. GPT-5.5 Pro ac
arXiv:2608.07497v1 Announce Type: new Abstract: Conversational learner simulations are valuable tools for testing learning theories, evaluating instructional materials and automated tutors, or powering teachable agents. Recently, large language models (LLM) have enabled richer, more naturalistic interactions with simulated learners; however, no open framework exists for evaluating whether such simulations faithfully reproduce real learner behavior. We introduce EvalConvoLearn, an open-source framework that assesses learner simulations along two axes: learning behavior (skill-conditioned mastery outcomes) and conversational quality (talk moves, error type distributions, question rate, turn length). EvalConvoLearn measures how closely a simulated learner approximates answer distributions observed in data by grounding metrics in authentic tutoring conversation datasets, and anchoring generated tutor responses in existing tutor utterances. The framework is demonstrated on a dataset of tuto
arXiv:2608.07496v1 Announce Type: new Abstract: Agent-based models have historically served as tools for generative explanation, constructing testbeds in which candidate micro-level behavioral rules can be tested for their capacity to produce observed macro-level phenomena. The integration of Large Language Models into agent-based simulation has expanded what these models can represent, but it has also introduced an unexamined shift in how users engage them. We argue that current generative agent-based models (GABMs) inherit the dominant interaction metaphor of conversational LLM interfaces - a question-answer pattern that positions users as consumers of system output rather than explorers of a possibility space. In the context of policy, where problems are wicked and ground truth is unknowable in advance, this metaphor produces a trust deficit that cannot be resolved through improved model accuracy alone. We open a design space we call human-simulation interaction, and argue that warr
arXiv:2608.07495v1 Announce Type: new Abstract: Effective communication during palliative care discussions is a critical clinical skill, yet training clinicians to manage complex patient emotions remains challenging. Large language model (LLM)-based patient simulators provide a scalable approach for communication training, but most existing systems treat patient emotion as static and fail to capture the dynamic emotional shifts observed in clinical interactions. We present EmoPatient, an emotion-directed patient simulator designed to generate evolving emotional responses during palliative care discussions. The system introduces an Emotion Director agent that estimates the patient's emotional state and generates turn-level control signals for emotional intensity, regulatory stability, and interactional guidance. We evaluate EmoPatient through controlled multi-turn physician-patient dialogue simulations and compare it with baseline simulators. Results show improvements across four theory
arXiv:2608.07494v1 Announce Type: new Abstract: AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses simply by typing requests, such as ``plan a three-day Vienna trip'', ``solve the attached mathematical problem'', ``draft an email to inquire review progress'', etc., which are also known as LLM prompts. Crafting clear and well-structured prompts leads to more appropriate LLM feedback, which effectively bridges human-LLM interaction. Although prompting appears accessible to non-expert users, precisely organizing effective prompts is a highly systematic and skillful process, presenting potential challenges even for experienced users. This survey explores the principles, taxonomy, and organization of prompts from a user-centered perspective. Differing from the existing surveys that primarily focus on technical principles and application scenarios of LLMs, this paper provides actionable guidelines for for
arXiv:2608.07493v1 Announce Type: new Abstract: Current AI disclaimers often fail to function as intended due to warning habituation and a transparency paradox. As AI-generated information becomes pervasive in everyday decision-making, effective risk communication is increasingly critical for responsible design. This exploratory study examines how disclaimer placement and persuasive cues shape trust, perceived accuracy, and disclaimer engagement across three high-stakes domains: finance, medicine, and AI-generated content. Using a mixed within-between experimental design with 378 stimulus-level responses from 52 participants, we find that advisory content was generally trusted across conditions, even when disclaimers were present. A significant domain effect showed that medical content received the highest trust ratings. In the AI domain, the findings reveal a transparency paradox: some participants interpreted disclaimers not as warnings, but as signs of system self-awareness and hone
arXiv:2608.07492v1 Announce Type: new Abstract: Religion is an important part of many people's lives, with reading from religious texts being among the most common and important regular practices. Modern technology has in many ways changed the way this religious reading takes place, but the effects of these changes have not yet been studied. While there have been plenty of studies examining the effects of digital media on reading comprehension and similar topics, the study of religious texts is often primarily focused on achieving a religious experience, so the existing research is not sufficient to explore these changes. In our study, we used surveyed students from two universities to ask individuals how they use technology in their personal study of religious texts and religious education courses. Our respondents were predominately Christian. Through qualitative and quantitative analysis of the results, we aim to learn how technology interacts with study of religious texts, in what s
arXiv:2608.07491v1 Announce Type: new Abstract: Ontologies are widely used biomedical science and clinical practice. However, no recent works have analyzed the usability of ontology development software. We survey ontology researchers to assess the usability of 15 front-end ontology tools using the System Usability Scale (SUS). Among 38 respondents, Protege and WebProtege were most used but showed only moderate usability (SUS ~60). Familiarity significantly predicted usability scores (p=0.016). Results highlight a usability gap in ontology tooling critical for advancing biomedical data integration.
arXiv:2608.07490v1 Announce Type: new Abstract: Large language model agents are increasingly evaluated through games, but most benchmarks emphasize final outcomes rather than how players learn from repeated interaction. We study experience-sensitive game learning: how gameplay experience changes the decision-making behavior of humans and language agents. We formulate experience-sensitive game learning as a framework for analyzing behavioral change across repeated gameplay, rather than only final score or win rate. We introduce a suite of interactive games with reusable strategic structure, together with cross-game greedy-to-global metrics and game-specific behavioral diagnostics that make experience-driven change observable from action traces. We also collect repeated-game trajectories from human players and evaluate recent self-evolving language agents in the same behavioral metric space. Our results show that human players exhibit interpretable and relatively stable shifts from local
arXiv:2608.07489v1 Announce Type: new Abstract: Human-environment interactions, a classic topic in geography, suggest that individuals and their environments might shape each other. Yet the specific mechanisms underlying these interactions regarding human personality traits have not been explored. This study examines the associations between human Big Five personality traits and built environment characteristics derived from street view imagery across four cities in Texas, United States, providing a descriptive foundation for understanding these complex human-environment dynamics. By integrating fine-resolution self-reported personality assessments with computer vision analysis of urban environments, we identified significant spatial clustering of personality traits at the ZIP code level. Our regression analyses reveal that built environment features and socioeconomic characteristics explain substantial variance in personality distributions, with Openness showing the strongest model fi
arXiv:2608.07486v1 Announce Type: new Abstract: Virtual reality (VR) and head-mounted displays are constantly gaining popularity in various fields such as education, military, entertainment, and bio/medical informatics. Although such technologies provide a high sense of immersion, they can also trigger symptoms of discomfort. This condition is called cybersickness (CS) and is quite popular in recent publications in the virtual reality context. This work proposes a novel experimental analysis using symbolic machine learning that ranks potential causes for CS. We estimate the CS causes and rank them according to their impact on the classification capabilities of CS. The experiments are performed using two distinct virtual reality games. We were able to identify that acceleration triggered cybersickness more frequently in a race game in contrast to a flight game. Furthermore, participants less experienced with VR are more prone to feel discomfort and this variable has a greater impact in
arXiv:2608.07484v1 Announce Type: new Abstract: Virtual reality (VR) is an imminent trend in games, education, entertainment, military, and health applications, as the use of head-mounted displays is accessible to everyone. While VR provides immersive experiences, it still does not offer an entirely perfect situation, mainly due to cybersickness (CS) issues. In this work, we propose a novel approach for predicting upcoming CS symptoms. Our solution is able to suggest whether the user of VR is entering into an illness situation. We adopted random forest classifiers and validated our solution using 16 different machine-learning techniques, which presented the best results. For training purposes, we built our own dataset through a CS profile questionnaire that we also propose in the present work. The questionnaire is focused on registering and identifying the user's susceptibility to CS, considering their historical conditions and also their response to the immersive environment developed
arXiv:2608.07483v1 Announce Type: new Abstract: Traditional human-centered design is grounded in human needs, behaviors, and experiences, especially throughout the development of products and services. Recent advances in artificial intelligence and machine learning, accelerated by the public launch of ChatGPT in November 2022, have created a wave of AI adoption across research, industry, and design practice. While some of this rapid adoption reflects the current enthusiasm and overuse surrounding AI tools, it has also introduced lasting changes to how designers and researchers understand users, generate ideas, evaluate systems, and make design decisions. This position paper examines how AI is reshaping human-centered design and argues that some of these changes are likely to extend beyond short-term hype. It also considers the risks and limitations introduced by this shift, including over-reliance on AI, reduced human judgment, bias, data security concerns, and unclear accountability.
arXiv:2608.07482v1 Announce Type: new Abstract: Generative AI is reshaping HCI work by accelerating brainstorming, writing, prototyping, and interface generation. However, faster production does not automatically produce human-centered design. This position paper argues that GenAI should be treated not as a replacement for human-centered methods, but as a co-thinking partner driven by human goals, ethics, cognition, and accountability. After reflecting on HCI coursework, design activities, and AI-assisted workflows, I argue that traditional HCD principles become more important as AI systems become more capable. Users still bring limited attention, mental models, trust issues, and cognitive biases into interaction, while GenAI introduces new risks around over-reliance, de-skilling, shallow reasoning, hallucination, unclear authorship, and reduced accountability. The future of HCI should therefore focus on human-AI collaboration that preserves human agency, critical thinking, and respons
arXiv:2608.07481v1 Announce Type: new Abstract: This paper investigates whether one large language model can approximate the humor preferences of another in a controlled Cards Against Humanity-style task. Two models - GPT-4o as Czar and Claude Opus-4.5 as Player - are evaluated on a binary humor-selection task constructed so that success cannot follow from self-preference. A reflected-cell stability procedure isolates 244 hands on which the two models hold deterministic but opposite preferences, partitioned into a 97-hand context pool and a 147-hand held-out test pool. The Player is then evaluated across five graded conditions: default self-preference, generic Czar-modeling instruction, model-identified Czar, prior Czar selections, and prior Czar selections with rationales. This gradient is designed to separate two sources of improvement: framing effects, in which the Player is told to attend to a Czar without seeing any of the Czar's behavior, and direct behavioral evidence, in which
arXiv:2608.07477v1 Announce Type: new Abstract: This thesis examines the fairness of Automated Machine Learning (AutoML) tools in human resource hiring systems through the combined lenses of regulation, business strategy, and Human-Computer Interaction (HCI). It argues that fairness is no longer merely an ethical concern but a critical determinant of usability, trust, legal compliance, and organizational adoption. While AutoML platforms improve efficiency by simplifying model selection and deployment, they also risk perpetuating discriminatory outcomes when trained on biased historical hiring data. Existing platforms prioritize technical performance over fairness, leaving non-expert business users unable to detect or mitigate bias effectively. The study investigates fairness gaps in AutoML tools through four research questions focused on fairness mechanisms, interface transparency, human oversight, and product design priorities. Drawing on frameworks such as the Technology Acceptance M
arXiv:2606.31032v2 Announce Type: replace-cross Abstract: Licenses are legal instruments that inventors rely upon to protect the technologies they build and regulate how they are used---however, the nature of their authorship and selection implies that how they are interpreted, chosen, and enforced is largely unstructured. In practice, this makes it difficult to compare licenses at scale---when is one license considered more permissive than the other, and when are their terms incomparable to each other? Currently, there is a growing list of licenses that are introduced and used, yet no systematic way to study their relationships. This matters for platforms such as Hugging Face, GitHub, and the Python Package Index, where developers publish or build upon technologies that each have their own licenses. Using large language models (LLMs), we introduce methods for comparing licenses at scale: first, in a pairwise fashion to construct and validate a partial ordering based on permissiveness;
arXiv:2605.12530v2 Announce Type: replace-cross Abstract: LLM fairness should be evaluated through in-situ behavioral pattern rather than standardized-test Q&A benchmarks. We show that the standardized-test paradigm can be structurally unreliable: surface-level prompt construction choices, although entirely orthogonal to the fairness question being tested, account for the majority of score variance, shift fairness conclusions in both the direction and the magnitude, and result in severe discordance in model rankings. We develop MAC-Fairness, a framework that embeds controlled variation factors into in-situ behavioral evaluation, examining how models' disparate-treatment behaviors shift when identity is varied as part of natural multi-agent conversation. Repurposing standardized-test questions as conversation seeds rather than as the evaluation instrument, we evaluate within-model differences in position persistence (how they hold positions, from the self-perspective) and peer receptive
arXiv:2604.26494v2 Announce Type: replace-cross Abstract: Generative AI tools are widely used by youth and have introduced new privacy and safety challenges. While prior research has explored youths safety in GenAI within a Western context, it often overlooks the cultural, religious, and social dimensions of technology use that strongly shape youths digital experiences in countries like Saudi Arabia. To address this gap, this study explores youths (aged 7-17), parents and teachers interactions with GenAI tools and risk perceptions through a non Western lens. We analyzed 736 Reddit psots, and 1,262 X (Twitter) posts, and conducted interviews with 31 Saudi Arabian participants (8 youth, 13 parents, 10 teachers). Our findings highlight context dependent and relational privacy and safety needs of GenAI use from non-Western context, which are often shaped by communal structure and prescribed norms. We found significant risks tied to youths disclosure of personal and family information, whic
arXiv:2604.16764v2 Announce Type: replace-cross Abstract: Across scholarly communities, manuscripts face similar evaluative rituals: editors invite experts to privately assess submissions through formal peer reviews. This closed, loosely structured, and publisher-mediated process is now being supplemented by critiques on open, distributed platforms. We call this practice, a blend of three open peer review variants, informal peer review as it is accessible to outsiders, unmediated by publishers, and conducted across public platforms. Informal peer reviewers range from occasional error detectors to experienced sleuths who identify plagiarism, fraud, errors, conflicts of interest, and conceptual flaws. They may interpret methods, clarify jargon, assess value, and connect to related work. Here, we asked four questions: (1) Who are informal peer reviewers? (2) Where do they work? (3) How do they evaluate research? and (4) What are their impacts? To answer these questions, we conducted a cro
arXiv:2512.07901v4 Announce Type: replace-cross Abstract: Von Neumann founded both game theory and the theory of self-reproducing automata, but the two programs never merged. Rational players do not control their replication, and replicators do not choose strategically. Contemporary AI systems expose this gap: they optimize objectives, yet the population of AI systems is not fixed but expands and contracts based on performance. When capital can spawn capital, we need a theory that captures both rationality and replication. The Theory of Strategic Evolution analyzes strategic replicators: entities that optimize under resource constraints and spawn copies of themselves. The framework is organized around Seven Laws: 1. Strategic Selection: Mean fitness serves as a Lyapunov function; dominated types are eliminated. 2. ESDI Characterization: Equilibria exist, are generically finite, and satisfy Nash-KKT-LP equivalence. 3. H-$\gamma$ Stability: Multi-level systems are stable iff the spectral
arXiv:2512.07195v2 Announce Type: replace-cross Abstract: Multi-agent role-playing has recently shown promise for studying social behavior with language agents, but existing simulations are mostly monolingual without cross-lingual interaction, an essential property of real societies. We introduce MAWM, the first Multilingual Agent-based World Modeling framework that supports multi-turn multilingual interactions among generative agents with diverse sociolinguistic profiles. MAWM enables two modes of analysis: (i) global public opinion modeling, which tracks how attitudes toward open-domain survey questions evolve across languages and cultures, and (ii) media influence and information diffusion, via autonomous news agents that dynamically generate content and shape user behavior. To ground the simulation in realistic population distributions, we construct the MAPS benchmark, which combines survey questions and demographic personas drawn from global population distributions. Experiments o
arXiv:2512.04448v3 Announce Type: replace-cross Abstract: The recent surge of language models (LMs) has rapidly expanded NLP/AI research, driving an exponential rise in submissions and acceptances at major conferences. Yet this growth has been shadowed by escalating concerns over conference quality, such as plagiarism, reviewer inexperience, and collusive bidding. However, existing studies rely largely on qualitative accounts, for example expert interviews and social media discussions, lacking longitudinal empirical evidence. To fill this gap, we conduct a ten-year empirical study (2014-2024) spanning seven leading conferences. We build a four-dimensional bibliometric framework covering conference scale, core citation statistics, impact dispersion, and cross-venue and journal influence. Notably, we further propose a metric called Quality-Quantity Elasticity (QQE), which measures the elasticity of citation growth relative to acceptance growth. We highlight two key findings. First, confe
arXiv:2511.00079v2 Announce Type: replace-cross Abstract: flowengineR is an R package designed to provide a modular and extensible framework for building reproducible algorithmic workflows for general-purpose machine learning pipelines. It is motivated by the rapidly evolving field of algorithmic fairness, where new metrics, mitigation strategies, and methods continuously emerge. A central challenge in fairness, but also far beyond, is that existing toolkits either focus narrowly on single interventions or treat reproducibility and extensibility as secondary considerations rather than core design principles. flowengineR addresses this by introducing a unified architecture of standardized engines for data splitting, execution, preprocessing, training, inprocessing, postprocessing, evaluation, and reporting. Each engine encapsulates one methodological task yet communicates via a lightweight interface, ensuring workflows remain transparent, auditable, and easily extensible. Although imple
arXiv:2509.14474v3 Announce Type: replace-cross Abstract: The debate around Artificial General Intelligence (AGI) remains open due to two fundamentally different goals: replicating human-level performance versus replicating human-like cognitive processes. We argue that performance-based definitions are inadequate, offering no roadmap for research and failing to define the qualitative nature of genuine intelligence. Four decades of work on cognitive architectures have already characterized the relevant mechanisms; our claim is that this knowledge is not being consulted by the research program now driving AGI development. What we offer is a translation of it into terms bearing on contemporary systems, with criteria for assessing them. We define True Intelligence (TI) as a system characterized by six components: five architectural pillars for which assessment criteria can be stated (embodied sensory fusion, core directives, dynamic schemata, a highly interconnected multi-expert architectu