Skip to main content

Focus on AI: Robot Conflict | Teammate Attacks | Algorithmic Harms

UMSI Research Roundup. Focus on AI. Robot Conflict, Teammate Attacks, Algorithmic Harms. Check out UMSI faculty and PhD student publications.

Friday, 09/25/2026

By Noor Hindi

University of Michigan School of Information faculty and PhD students are advancing the field of artificial intelligence through innovative research and impactful contributions. Here are some of their recent publications.

Publications 

Mitigating Human–Security Robot Conflict Through Fairness

International Journal of Human–Computer Interaction, September 2026 

Xin Ye, Lionel P. Robert

Security robots are increasingly deployed in public spaces to enhance safety and enforce rules. However, their exercise of authority often generates conflict with the public. Such conflicts may stem from incompatible goals, referred to as cognitive conflict, or from interpersonal tensions, referred to as affective conflict, both of which threaten public acceptance. Drawing on insights from policing research, this study investigates whether distributive fairness (outcome equity) and informational fairness (treatment equity) can mitigate these conflicts in human–robot encounters. In a 2 (informational fairness: low vs. high) × 2 (distributive fairness: low vs. high) between-subjects experiment (N = 464), we found that both forms of fairness significantly reduced cognitive and affective conflict, fostering greater acceptance of security robots. These findings highlight fairness as a critical mechanism for conflict mitigation and underscore its central role in promoting the public acceptance of robotic authority.


Personalized Trust Prediction in Human–Autonomy Interaction: A Discounting Model

IEEE Transactions on Human-Machine Systems, September 2026 

Hyesun Chung, Shreyas Bhat, X. Jessie Yang

The increasing adoption of autonomous systems has highlighted the need to understand and predict users’ trust in these technologies. Existing trust prediction models perform well for most users but fail to capture the volatile, highly fluctuating trust patterns of “oscillators.” This raises two key challenges: First, developing a model that can accurately approximate oscillators’ dynamics, and second, identifying which individuals are likely to be oscillators so the model can be selectively applied. To address the first challenge, we propose a discounting trust model that prioritizes recent interactions while attenuating older ones, improving prediction accuracy for oscillators (optimal discount factor λ=0.8). To address the second, we develop a classification model based on seven personal characteristics to screen for potential oscillators. Finally, we integrate the two approaches, applying the discounting trust model only to individuals predicted as oscillators. We used two datasets (N=130 development and N=41 validation). The development dataset was used to identify the optimal discount factor and train the screening classifier. The validation dataset was used to identify potential oscillators and to validate the performance of the discounting trust model on them. Four out of the 41 participants were identified as potential oscillators. We compared the discounting and nondiscounting baseline models using a linear mixed-effects model that accounted for autocorrelation. The discounting model significantly reduced prediction errors compared to the baseline model (p<.001). The findings showcase the value of incorporating both personal traits and a discount factor to enhance trust prediction for oscillators.


Effects of AI Assistance Timing on Pharmacists’ Trust in Automated Pill Recognition Technology: Within-Participants Experimental Study

JMIR Human Factors, September 2026

Jin Yong Kim;  Brigid Rowell, Megan Whitaker, Qiyuan Chen, Raed Al Kontar, Corey Lester, Xi Jessie Yang

Background: Image-based pill verification systems demonstrate high model accuracy. However, their effectiveness in pharmacy practice depends on how pharmacists interact with the AI output. The timing of AI advice is one design factor that influences these interactions, yet its impact on pharmacists’ moment-to-moment trust dynamics during medication verification requires further investigation. Objective: This study aims to investigate how the timing and conditionality of AI assistance shape pharmacists’ trust dynamics during medication verification. 

Methods: Between April and December 2024, 50 licensed pharmacists completed a browser-based simulated medication dispensing task with 2 AI types: ex-ante advice (AI advice given concurrently with clinical information) and ex-post advice (AI advice given after an initial diagnosis). Ex-post advice was further divided into the involved ex-post and the not-involved ex-post conditions. The experiment used a within-participants design with varying AI types and AI recognition patterns (right fill-correct recognition, right fill-incorrect recognition, wrong fill-correct recognition, and wrong fill-incorrect recognition). The primary outcomes were trust adjustment magnitude and trust adjustment, which were analyzed using mixed-effects linear regression models. 

Results: Trust adjustment magnitude differed significantly across AI assistance conditions. The involved ex-post condition led to the highest magnitude of trust adjustment, followed by ex-ante advice (mean difference 9.65, 95% CI 8.55-10.75; P<.001), with the not-involved ex-post condition showing the lowest magnitude (mean difference 3.33, 95% CI 2.63-4.03; P<.001). Analysis by each recognition pattern revealed significant differences in trust adjustment when the right drugs were incorrectly rejected. In this pattern, the involved ex-post condition led to larger trust decrements than both ex-ante advice (mean difference −2.13, 95% CI −3.83 to −0.44; P=.008) and not-involved ex-post conditions (mean difference −7.32, 95% CI −13.69 to −.95; P=.009). A marginal difference was observed between ex-ante advice and not-involved ex-post conditions (mean difference −5.19, 95% CI −11.55 to 1.17; P=.051). No significant differences were observed for other recognition patterns. 

Conclusions: Both the timing and conditionality of AI assistance influenced pharmacists’ trust dynamics. Disagreement-based AI interventions (involved ex-post) that incorrectly challenged pharmacists led to a substantial trust decrement, whereas the not-involved AI intervention resulted in more stable trust fluctuations. These findings highlight the importance of designing AI systems that align intervention strategies with user expertise and task demands to foster appropriate trust in safety-critical workflows.


Designing Healthcare Robots for Community-Dwelling U.S. Older Adults: A Kano Model Perspective

International Journal of Social Robotics, August 2026

Qiaoning Zhang, Feng Zhou, Lionel P. Robert Jr, X. Jessie Yang

Healthcare robots at home are increasingly essential for promoting the independence of older adults, yet their widespread acceptance is hindered by a lack of clarity regarding optimal design features, particularly among users with varying levels of knowledge and attitudes towards this emerging technology. To address this, this study applies the Kano model to classify and prioritize healthcare robot features based on their impact on user satisfaction and design decisions, factoring in older adults diverse robot-related knowledge and attitudes towards robots. Following a thorough literature review that highlighted 27 distinct robot features, we conducted a survey with 253 community-dwelling older adult participants and identified essential features such as ‘Medication Management’ and ‘Managing Illness and Monitoring Health’ as one-dimensional features, whereas ‘Animal-like Appearance’ was negatively received. The Kano model classifications including must-be, one-dimensional, attractive, indifferent, and reverse, offer direct guidance for design priorities by identifying which features are most likely to enhance satisfaction when included and cause dissatisfaction when absent or poorly implemented. The analysis also showed that user preferences vary significantly with their knowledge and perception of robots. These insights emphasize the need to tailor healthcare robots to the initial expectations of community-dwelling seniors, prioritizing functional features that support daily independence over therapeutic care.


CATEMS: A Lifecycle Model of Compromised AI Teammate Attacks, Effects, and Mitigation Strategies for Enhanced Resilience

Proceedings of the Human Factors and Ergonomics Society Annual Meeting, August 2026 

Sarah V. Mendoza, Yayun Tian, Beau G. Schelble, Allyson Hauptman, Lionel P. Robert

Human-AI teams are gaining traction as they combine the unique strengths of humans and AI. Very little research has examined the unique vulnerability of compromised AI teammates that work directly against the team’s shared goals. Compromised AI attacks combine the potency of technological attacks with teammate betrayal, potentially devastating the team’s ability to execute its goals. Such attacks by a malicious actor may misalign mental models, individual situational awareness, and shared situation awareness. This misalignment could potentially result in the complete loss of team coordination, endangering performance, resiliency, and security. Given the severity of a compromise, this paper presents a model to explain the types of attacks, when they may occur, their effects on the team, and potential approaches for attack identification and recovery.


Distributional Alignment for Social Simulation with LLMs: A Mixture Modeling Approach

KDD '26: Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining, August 2026

Yutong Xie, Ruoyi Gao, Qiaozhu Mei

Simulating human distributions (e.g., distributions of characteristics, behaviors, or preferences) in social contexts is essential for understanding complex dynamics across disciplines and drives many real-world applications. While large language models (LLMs) have advanced this area, a key challenge remains: capturing the inherent diversity within populations. We present a distributional alignment framework that models human heterogeneity as mixtures of system prompts. To efficiently estimate these mixtures, we adapt expectation–maximization (EM) and gradient boosting algorithms for LLMs. We evaluate our method across three typical social dimensions: personality traits, economic behaviors, and ideological values. Compared to existing approaches, our method achieves better alignment with observed human distributions. This framework offers a robust foundation for simulating diverse populations, supporting more accurate, scalable, and socially grounded applications of LLMs across the social sciences and beyond.


Development and Validation of the Perceived Personal Algorithmic Harms Scale: Algorithmic Stigma, Algorithmic Privacy Harms, Tangible Algorithmic Harms, and Race, Gender, and Disability Differences

ACM Journal on Responsible Computing, August 2026 

Nazanin Andalibi, Oliver Haimson

Algorithmic systems, including those using AI techniques, are increasingly deployed across domains from healthcare to law enforcement. Despite promises of societal good, these systems can inflict real harm to individuals and societies. Yet no validated scales exist to systematically measure individuals’ experiences and perceptions of algorithmic harm, which we refer to as perceived personal algorithmic harms. We develop and validate the Perceived Personal Algorithmic Harms Scale with a representative U.S. sample (n = 742) and oversampling of people of color, disabled individuals, transgender, and nonbinary participants. The scale comprises three subscales—algorithmic stigma, algorithmic privacy harms, and tangible algorithmic harms—and can be used to measure personal algorithmic harms across systems, contexts, and populations. Our findings empirically confirm that algorithmic stigma and privacy harms, previously theorized, are experienced in measurable ways, with minoritized groups (transgender, nonbinary, mixed race, disabled, and/or neurodivergent people) reporting higher levels of harm than majority counterparts. The scale provides a versatile tool—including as a means for critical refusal and repair—for researchers, lawyers, policymakers, technologists, civil society, and the public to assess past, present, or anticipated harms, evaluate specific systems, and identify vulnerable groups prior to deployment. In sum, this work advances both theoretical understanding and practical intervention in addressing the lived impacts of algorithmic systems.


Trusting Security Robotic Authority: The Impact of Interactional and Distributive Fairness

Proceedings of the Human Factors and Ergonomics Society Annual Meeting, August 2026

Xin Ye, Lionel P. Robert, Jr.

The growing deployment of autonomous security robots has raised critical questions about how people trust and accept robotic systems exercising public authority. Yet fairness, a central mechanism in public responses to human authority, has received limited attention in the context of robotic authority. Drawing on fairness theory from policing research, this study examines how interactional fairness and distributive fairness influence trust and acceptance of security robots. Using a 2 × 2 between-subjects experiment with 100 U.S. participants, we found that both interactional and distributive fairness significantly increased trust, whereas only interactional fairness significantly increased acceptance. These findings suggest that fairness serves as a critical mechanism shaping public responses to robotic authority. This research extends fairness theory to human-security robot interaction and highlights the importance of clear explanations and consistent rule enforcement for building public trust and acceptance.


Recognizing, Supporting, and Futuring Diverse Pathways in Computing Work

CHIWORK Adjunct '26: Adjunct Proceedings of the 5th Annual Symposium on Human-Computer Interaction for Work, July 2026 

Lara Karki, Aarti Israni, Tawanna Dillahunt, Julie Hui, Dana Priest, Asher Brown, Betsy DiSalvo

Professional computing work is no longer limited to traditional roles requiring computer science degrees. HCI has long characterized computing tasks in roles like administrative assistant or operations specialist as end-user programming. With generative AI and enterprise low-code/no-code tools, however, many of these roles now center on building and maintaining computational systems. These are middle-skill jobs crucial to financial security for adults without college degrees, and a clear concept of them is needed to shape and support the future of these occupations and the people in them as AI transforms the workforce. In this workshop, we will bring together scholars and practitioners to ask: How do we articulate a broader conception of computing work? How do we support the diverse needs and populations engaging in this work? What is our vision for the future of broader forms of computing work?


Small Language Models for Education: Opportunities, Challenges, and a Shared Research Agenda

Communications in Computer and Information Science, June 2026

Yumou Wei, Steven Moore, Paulo F. Carvalho, John Stamper, Christopher Brooks, Michael Liut

Small language models (SLMs) are emerging as a promising alternative to large language models for educational applications. This workshop aims to explore the potential of SLMs in education, focusing on their unique advantages and the challenges they present. By bringing together researchers from the AIED, EDM, and L@S communities, we hope to foster interdisciplinary collaboration and identify new research directions that leverage the strengths of each community to improve learning outcomes. The expected results of the workshop include a comprehensive understanding of the current state of small language models, the identification of key research challenges and opportunities, and the development of a shared research agenda that guides future work.


Self-Regulated Personal Contracts as a Harm Reduction Approach to Generative AI in Undergraduate Programming Education

ITiCSE 2026: Proceedings of the 31st ACM Conference on Innovation and Technology in Computer Science Education, July 2026 

Aadarsh Padiyath, Jessica Shen, Barbara Ericson

Students learning programming exercise agency in deciding when and how to use GenAI tools like ChatGPT. However, this agency is often implicit and shaped by deadline pressure and peer behavior rather than explicit and conscious learning goals. We designed a GenAI Contract grounded in harm reduction and self-regulated learning theory to scaffold intentional decision-making: students articulated personal learning goals, created usage guidelines, and reflected on alignment at strategic points across an eleven-week semester. The contract was non-binding and graded only for completion, emphasizing self-awareness over enforcement. We implemented this with N=217 students in an intermediate Python course. For students still forming their relationship with GenAI, it worked, as 58% of students reported the intervention changing their thinking and created helpful accountability structures. However, awareness did not always translate to sustained behavior change. Some students who valued their guidelines still abandoned them under various pressures. Maintaining guidelines required constant self-control across hundreds of decisions, while using GenAI freely requires none. Many students could not sustain this burden despite this self-awareness. We discuss supporting student agency when GenAI tools and learning goals create tension.


What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks

Conference on Language Modeling (COLM), October 2026 [selected for oral spotlight]

Meera Desai, Sang T. Truong, Hanna Wallach, Alex Chouldechova, A. Feder Cooper, Jean Garcia-Gathright, Daniel E. Ho, Abigail Z. Jacobs, Sanmi Koyejo, Nicholas Pangakis, Angelina Wang

Benchmarks play a central role in the development and governance of models, yet it is often unclear whether they actually measure the concepts they purport to measure (e.g., reasoning, refusal). We adapt convergent and discriminant validity from the social sciences into an approach for interrogating AI benchmarks, applying it to 56 capability and safety benchmarks across 53 models. We label benchmarks with substantively similar purported concepts to a shared assigned concept, and ask whether model rankings on benchmarks with the same assigned concept correlate more strongly than rankings on benchmarks with different assigned concepts. We ask analogous questions at the item level using item response theory (IRT) models. We find that correlations between model rankings on benchmarks with the same assigned safety concepts are often weak, suggesting these concepts may be conceptualized inconsistently across benchmarks. For assigned capability concepts (e.g., reasoning, knowledge), model rankings are often as strongly correlated among benchmarks with the same assigned concept as between benchmarks with different assigned concepts, suggesting these capability concepts may not discriminate well from one another. In some cases, benchmarks that share design elements (e.g., score format) correlate more strongly than benchmarks with the same assigned concept. Finally, some individual benchmarks correlate more strongly with benchmarks assigned a different concept than with benchmarks sharing their own assigned concept, suggesting they may measure a different concept than they purport to. For example, BBQ-accuracy correlates more strongly with benchmarks labeled reasoning than with benchmarks that share its assigned concept, bias. To support future empirical work on benchmark validity, we release our extensive dataset of model outputs and scores at the item- and benchmark-level.

Pre-prints, Working Papers, Articles, Workshops and Talks

Modeling the Structure of Human Behavior with AI Prompt Vectors

arXiv, August 2026 

Matthew O. Jackson, Benjamin S. Manning, Yutong Xie, Walter Yuan, Qiaozhu Mei

We introduce a general, easy-to-implement AI-based method for modeling and analyzing the structure and complexity of human behavior. We assign a large language model a "type vector" and then prompt it to choose actions across settings in which we observe human choices. For instance, the type vector (2, 4) becomes "You are a player characterized by the following profile: Altruism: 2 out of 5, Risk Aversion: 4 out of 5," after which it is prompted to make choices. We vary the dimensions (e.g., Altruism, Fairness, Trust,...) and values (e.g., 1-5) to minimize distance to human choices. Applying the method to 119,147 decisions made by 78,657 subjects from more than 35 countries across 10 classic economic game roles, we find that human behavior can be closely matched using three dimensions: Risk Aversion, Strategic Sophistication, and Trust. Moreover, the types needed to fit individuals across games cluster into fewer than a dozen groups, and can predict behavior in held-out games with different rules and available actions. The results suggest that behavior across diverse settings can be approximated by a low-dimensional, portable representation, supporting the possibility of general yet parsimonious theories across the behavioral sciences. More broadly, this new modeling method can provide insights into the structure of many human behaviors.


What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development

arXiv, August 2026 

Christopher Brooks 

Between AI-assisted item generation and expert review sits a computational evaluator whose decisions are usually treated as technical preliminaries. Yet representation, structural reduction, and selection policy determine which items and evidence psychometricians ever receive. Across two linked in-silico studies of 32,000 selected Big Five items, we followed fixed source populations from semantic representation through structural evaluation and candidate-form construction. Broad agreement in semantic geometry concealed consequential local differences: identical wording acquired different construct evidence, different items survived, and intended attributes could disappear even as community correspondence improved. These sensitivities also differed across generated source populations. At the final review boundary, both eligibility policies filled every content cell in every evaluable form, yet they presented different wording. Across embedding configurations, inclusive primary forms shared a median of only 6 of 40 items, reflecting the total downstream consequence of changing representation across structural evidence and ranking. The apparent stability of global summaries and complete forms therefore concealed instability in the content reaching psychometricians. The computational evaluator is not neutral infrastructure between generation and expertise; it is an inspectable and revisable part of measurement design.


Agentic AI: User Empowerment or Foreclosure?

arXiv, August 2026 

David Gamba, Daniel M. Romero, Grant Schoenebeck

Agentic AI promises systems that can act on users' behalf, from filtering content to negotiating prices to selecting services. Whether it will empower users is an open question, and one that depends on more than the technology. We conduct a comparative case analysis of four earlier, more mature domains in which similar forms of agency emerged: browser-based ad blockers, platform recommender systems, financial robo-advisors, and email spam filtering. Across the cases, questions about whose interests agents would serve were resolved through technical arrangements: API choices, protocol governance, industry standards, and default configurations. Beyond their technical form, these were political decisions. We identify this settling of contestable questions in a technical form as depoliticization, a concept from political theory, here at work in technological systems. Its most consequential effect is that individual outcomes and collective contestation capacity can move in opposite directions: spam inbox quality improved substantially while the organized capacity to contest spam governance collapsed. Where intermediary institutions sustained formal channels for challenge, user-aligned agency proved more durable; where proprietary infrastructure and closed standard-setting absorbed contestation, the material basis for user-aligned alternatives was dismantled, and the loss proved hard to reverse. Applying this lens to agentic AI, we find a similar pattern forming: governance is consolidating around the Model Context Protocol and the Agentic AI Foundation, an industry-governed venue already deciding what agents will be able to do. Unlike in the completed trajectories, these decisions have not yet hardened, and remain open to challenge by users and the public.


The Future of NLP may not be at NLP Conferences: Scholarly Migration Patterns in Natural Language Processing

arXiv, July 2026

David Jurgens 

Natural Language Processing (NLP) has traditionally been published in its core disciplinary venues like ACL. However, advances in Large Language Models (LLMs) has led to a blurring of the disciplinary lines between NLP and general Machine Learning (ML), with authors regularly publishing in venues from both fields. Here, we ask whether the disciplinary center of gravity is shifting. Using NLP research published from 2010 to 2026 and studies of both established and new authors, we find that a migration is taking place. First, comparing the pre- and post-LLM eras, established authors lost 19.2pp of share at flagship *ACL main-conference tracks while gaining 14.8pp in the newer Findings tracks, and general ML venues rose 8.6pp, even when adjusting for parallel growth in the fields. Second, among newer authors who debut with at least three first-author NLP-topic papers, the share whose work appears mostly at *ACL venues fell from 84% (2019) to 74% (2024), while the share appearing mostly at general ML venues rose from 5% to 21%. Using causal inference techniques, we estimate that these general ML venues confer a significant citation premium, which influences venue selection. Together, these results point to a significant shift in where NLP research is published.


Information Discernment in Large Language Models

arXiv, May 2026 

Joshua Ashkinaze, Laura Kurek, Alina Faisal, Tongyuan Miao, Mariam Joseph, Ceren Budak, Eric Gilbert

LLMs are increasingly used with external knowledge sources like the internet. Do they weigh information appropriately -- updating more for reliable sources (source discernment) and more when claims bring priors closer to the truth (truth discernment)? We formalize this as information discernment and introduce Learn2Discern (L2D), an experimental framework and benchmark grounded in three normative axioms with interpretable metrics. To establish external validity, a pre-registered, quota-matched user study (n=299) confirms that real LLM users endorse all three axioms and report that violations reduce their trust and usage intent. Across 13 models and nearly 670K trials, we find consistent failures across both dimensions: models perform near chance on source and truth discernment, rely on source popularity twice as much as source reliability, and update roughly equally whether a claim improves or worsens their position relative to the ground truth. Models integrate external knowledge most effectively on datasets where their priors are already the most accurate. Newer and larger models improve truth discernment but not source discernment, a blind spot that model complexity does not address. We identify simple inference-time interventions that improve both forms of discernment. We release our dataset and survey as a testbed for a core alignment property that scales in importance as LLMs replace traditional search.

RELATED

Keep up with research from UMSI experts by subscribing to our free research roundup newsletter!