Nice to Know You're Not Alone | Navigating Complexity: UMSI Research Roundup
Friday, 09/25/2026
By Noor HindiUniversity of Michigan School of Information faculty and PhD students are creating and sharing knowledge that helps build a better world. Here are some of their recent publications.
Publications
Navigating Complexity: How Context Shapes Debugging Strategy Choices Among Expert Developers
IEEE Transactions on Software Engineering, September 2026
Maryam Arab, Jenny T. Liang, Valentina Hong, Steve Oney, Thomas D. LaToza
Debugging is a central yet cognitively demanding part of software development, requiring problem-solving skills, expertise, and appropriate tools. Prior work has documented a variety of debugging strategies—such as hypothesis testing, simplification, and forward or backward reasoning—as well as individual factors that influence their use. However, we still lack an integrated understanding of how expert developers select and adapt these strategies in response to changing contextual factors during real-world challenging debugging. In this paper, we conducted a three-phase study: we first synthesized prior work on challenging debugging contexts, then used a short survey of 35 web developers to surface up-to-date examples of difficult defects, and finally conducted semi-structured interviews with 16 expert web developers to understand how they select and adapt strategies in these contexts. We identified a taxonomy of static and dynamic contextual factors that shape strategy selection. Dynamic factors can evolve during debugging and prompt strategy transitions, whereas static factors provide relatively stable constraints on developers’ choices. We also present a descriptive state-transition model showing how experts adapt strategies as contextual conditions like clarity, reproducibility, and constraints evolve during a debugging scenario. Our findings highlight the need for more contextual design of debugging tools as well as educational decision-making frameworks for choosing effective strategies.
Using emoji reduces dropout of remote workers: a causal analysis on GitHub
Humanities and Social Sciences Communications, August 2026
Xuan Lu, Wei Ai, Qiaozhu Mei
Remote work relies heavily on online communication methods including asynchronous text-based messaging, which is often associated with delayed feedback and reduced emotional support. Inefficient communication may ultimately lead to negative impacts on work-related outcomes such as a higher turnover rate. We propose that emojis can help mitigate such disadvantages and present a causal analysis on GitHub—a setting where remote workers work voluntarily and rely primarily on textual communication. Our analysis demonstrates that, after adjusting for confounding bias using propensity score-based methods, including emojis in at least one post in 2018 reduces the dropout risk in the following year by half (an absolute risk reduction of 0.052, p ≪ 0.001), and higher emoji usage intensity further leads to lower dropout risk. While the absolute effect of emoji use is greater for new contributors, the relative reduction in dropout risk is larger for developers with a longer contribution history. The effect of using emojis is also more pronounced for developers from collectivist or restrained cultural backgrounds, who may be more sensitive to group identity and social norms. Our analysis provides valuable insights for designing interventions to enhance the experience and work-related outcomes of remote workers.
Confronting complexity: toward equitable healthcare technology design and evaluation for socioeconomically marginalized patients
JAMIA Open, August 2026
Tiffany C Veinot, Alicia K Stone, Sage Davis, Rupa S Valdez, Lorraine R Buis, Vaishnav Kameswaran, Marcy G Antonio, Graciela Lagraba, Tawanna R Dillahunt
Objectives: Investigate how patients at a Federally-Qualified Health Center (FQHC) perform with, and perceive, complex telehealth tasks. Determine which complexity dimensions most affect patients’ performance and perceptions, and compare this to researcher-evaluators’ complexity walkthrough results. Materials and Methods: A novel complexity walkthrough inspection method was implemented by researcher-evaluators (n = 8) to identify theory-informed complexity dimensions in 6 FQHC-required telehealth tasks. Remote user testing (n = 24) where FQHC patients were observed performing tasks, then completed newly-developed complexity-focused interviews and surveys. Descriptive statistics regarding complexity dimension presence, task performance, cognitive load, and perceived difficulty were integrated with qualitative data analyzed using inductive and deductive coding.
Results: Patients completed 33.9% of the required subtasks without issues. Cognitive load and perceived difficulty were high for 2 tasks. Complexity dimensions of ambiguity (unclear inputs/processes; new concepts/words) and relationship (context switching; deep navigational hierarchies) most affected patient-perceived difficulty. Patients spent twice as long as walkthrough evaluators on tasks, and encountered broader complexity dimensions: new concepts/words, and errors. Many patients ended tasks early, asserting that they would abandon them outside of a study or had previously done so.
Discussion: Technology-mediated task complexity may explain some telehealth uptake inequities. The complexity dimensions that challenge patients extend known usability heuristics by enhancing their equity sensitivity. Complexity walkthroughs surface design patterns that challenge patients, but complexity-focused user testing with patients reveals additional difficulties. Findings support complexity reduction of tasks via structuring, familiar concepts/words, feature integration, and shallow/broad navigation.
Conclusion: This paper’s novel, theoretically-grounded complexity-focused methods and findings may inform future equitable design and evaluation of technology-mediated tasks for socioeconomically marginalized patients.
Acting in the Best Interest of the Other An Ethics of Care in Digital Curation
International Journal of Digital Curation, August 2026
Rebecca D. Frank, Katherine Polasek, Daniel Delmonaco, Kara Suzuka, Elizabeth Yakel
Digital repositories occupy a unique position in the data ecosystem, serving as intermediaries that preserve and provide access to valuable research data. This role carries ethical responsibilities toward multiple stakeholders, including participants represented in repository holdings. This responsibility is particularly salient for repositories of sensitive qualitative data, such as video records of practice (VROP), depicting teachers and students in classrooms, where participants remain identifiable and vulnerable to potential harm long after data collection. Drawing on Tarlow's dimensions of care, we conducted 44 semi-structured interviews with VROP data producers and reusers in education to examine how they perceive care in repository practices. Our findings indicate that data producers and reusers view repositories as sites where ethical practices can be modeled and learned, understand data curation as a form of care enacted on behalf of research participants and future reusers, and expect repositories to inherit and extend the ethical obligations originally held by data producers. We introduce the concept of anticipatory care as actions taken by repositories to shape how future reusers engage respectfully with data across temporal and social distance. These findings suggest that repositories must attend to the ethical epistemologies of their designated communities, not merely their technical and informational needs.
Nice to Know You're Not Alone: Co-designing Community-centered Online Safety and Privacy Education with Librarians
Proceedings of the Twenty-Second Symposium on Usable Privacy and Security, August 2026
Tanisha Afnan, Sheza Naveed, Griffin Christie, Jackie Hu, Byron M. Lowens, Allison McDonald, Florian Schaub
Online safety and privacy literacy education is often offered in community settings and spaces, such as libraries, which allow for educational offerings tailored to community needs. However, such efforts are often resource constrained and place substantial burdens and responsibility on individual educators. We present findings from a co-design study with librarians offering online safety and privacy education in public and academic libraries to address their challenges in bringing these trainings to their communities. Fourteen participants across four co-design workshops engaged in guided prompts, discussions, and activities to develop innovative and sustainable ideas for enhancing community centered online safety and privacy education efforts. Educators also ideated promising directions for independent learning resources that could deepen patrons’ understanding of relevant topics on their own time after a training workshop. We discuss the implications of our findings for strengthening community-centered online safety and privacy education and resources.
Exploring Privacy Negotiation Strategies for Camera-Equipped Smart Home Devices in Airbnb
Proceedings of the Twenty-Second Symposium on Usable Privacy and Security, August 2026
Zixin Wang, Sunyup Park, Haojian Jin, Yaxing Yao
Privacy conflicts among multiple users are prevalent in smart environments. To address such conflicts and mitigate the privacy concerns of bystanders, privacy negotiation has been proposed as a promising approach. Yet it remains unclear what makes such negotiations accessible to tenants. In this paper, using camera-equipped smart devices in Airbnb as a case, we explored different privacy negotiation strategies to shed light on future designs to support such actions. We developed 24 storyboards varying four factors, including device types (e.g., driveway camera, video conferencing device for TVs, smart display device), negotiation initiators (host vs. tenant), strategies (physical adjustment vs. monetary compensation), and outcomes (accepted vs. declined). Then, using a vignette survey study (N = 400), we find that device type and strategy significantly shape evaluations: physical adjustment was consistently preferred over monetary compensation, and driveway camera negotiations were evaluated more favorably than those involving indoor devices (video conferencing device for TVs, smart display device). We summarize design implications to support privacy negotiations.
Characteristics of Tailored Text Messages Associated With Increased Physical Activity Among Cardiac Rehabilitation Enrollees: Secondary Analysis of a Microrandomized Trial
JMIR Mhealth Uhealth, August 2026
Namratha Atluri, Kashvi Gupta, Tanima Basu, Evan Luff, Jieru Shi, Thomas Boyden, Bhramar Mukherjee, Sachin Kheterpal, Predrag Klasnja, Walter Dempsey, Brahmajee K Nallamothu, Jessica R Golbus
Background: Emerging data suggest that text message–based mobile health interventions may enhance physical activity levels in patients with cardiovascular disease enrolled in cardiac rehabilitation. The optimal characteristics of texts that lead to maximal patient engagement and drive meaningful behavioral change are not well understood.
Objective: This study aimed to understand how text- and participant-level characteristics impact physical activity levels after text delivery.
Methods: The VALENTINE (Virtual Application-Supported Environment to Increase Exercise) study was a randomized controlled trial designed to evaluate a mobile health intervention delivered to low- and moderate-risk adults enrolled in cardiac rehabilitation. Embedded within this study was a microrandomized trial focused on the effect of texts on physical activity levels among intervention participants. Participants in the intervention group received texts through a smartwatch (Apple Watch or Fitbit Versa) that were tailored to the time of day, day of the week (weekday vs weekend), weather, and time since enrollment in cardiac rehabilitation. Texts also differed in content type (walking vs antisedentary) and in the level of personalization (inclusion of the participant’s name or not). Delivery was randomized at 4 user-selected time points daily, with participants having a 25% probability of receiving a text at any time point. The primary outcome was step count 60 minutes after a decision point. This analysis focuses on the text- and participant-level factors that moderated the intervention’s effect on the primary outcome. Given potential measurement differences determined a priori, analyses were stratified by device type and phase of cardiac rehabilitation and adjusted for age, sex, and baseline activity status using a generalization of regression analysis.
Results: More than 70,552 randomizations occurred in 108 participants (mean age 59.5, SD 10.7 years; n=36, 33.3% female; n=19, 17.6% non-White; n=68, 63% Apple Watch users) over 6 months. Overall, no text characteristics (including personalization with the participant’s name) or participant characteristics (including baseline physical activity) consistently impacted text responsiveness for either device type. Although the findings were not consistently significant between device types and across phases of the trial, there was a trend toward increased responsiveness to texts that promoted walking (compared to antisedentary texts) and that were delivered to younger (aged <65 years) and male participants.
Conclusions: In this randomized clinical trial, we found that tailored texts improved physical activity levels among cardiac rehabilitation enrollees in the initiation phase, but this effect was not explained by text- or participant-level moderators. Additional work is needed to explore the impact of tailoring based on an extended set of personal and environmental factors to optimize the delivery and efficacy of text message–based interventions.
From Creative Engagement to Illustrative Insights: Exploring the Role of Design Probes for Implementation Research
Implementation Research and Practice, August 2026
Andrea J. Hoopes, Abigail Matson, E. Ruby Cramer, Rosemary D. Meza, Predrag Klasnja, Shannon Dorsey, Bryan J. Weiner, Aaron R. Lyon, Lorella Palazzo, Ruben G. Martinez
Background: To advance implementation of evidence-based interventions in healthcare, research teams need methods that center engagement, inclusivity, and creativity. We address this challenge by exploring design probes for implementation research. Design probes are packaged materials given to users that prompt them to asynchronously capture data about their context and experience, subsequently reflecting upon aspects of that data salient to the topic of study. This study had three objectives: (1) explore the potential utility of design probes in implementation research, (2) describe researcher and participant experiences with design probes, and (3) generate considerations for leveraging design probes to enhance engagement in implementation research.
Method: We used a multi-informant, multimethod approach. For objective 1, we undertook a literature scan and elicited expert (n = 8) input to explore how, when, and why design probes could be used in implementation research. For objective 2, we pilot tested the method with practitioners (n = 22) in an implementation research project exploring barriers and facilitators to implementing measurement-based care in community mental health settings. For objective 3, we sought feedback about the method in focus groups with youth (n = 8) and practitioners (n = 9).
Results: The literature scan and expert input identified five scenarios in which design probes may enhance implementation research, including enhancing engagement among implementation partners, accessing hard-to-reach populations, and identifying partner-centered implementation strategies. Our pilot experience demonstrated that design probes are feasible for implementation research from the perspective of research teams and participants. Focus group findings indicated that design probes hold promise for engaging youth but may have variable appeal and utility with practitioners.
Conclusions: We explore design probes for implementation research at a time of critical need for partner engagement. Future research will examine feasibility and value of design probes in a range of implementation initiatives.
Intersectionality and the Internet: Race, Class, Gender, Sexuality, and the Digital
Annual Review of Sociology, August 2026
Safiya U. Noble, André Brock, Kishonna Gray, Sarah T. Roberts
Over the past 20 years, the broad field of digital media and technology studies has been influenced by theories of power that include intersectionality as a key organizing theory and method. In this article, we review the major contributions across multiple fields of study, particularly critical race and digital studies, to the examinations of the Internet and technology that use intersectionality as an organizing logic. Applying intersectionality as a conceptual lens into digital technology design, dissemination, and use has aided in findings that the computing industries are ill-equipped to map the world, its conditions, or its inhabitants. In reframing this sociocultural complexity as solvable by mathematics, statistics, code, and data, these industries make dehumanizing political and ethical decisions through reduction—from people and culture to content and abstracting the world to data. In the process, they intentionally reify long-standing inequities in the name of efficiency, progress, and profit and fail to capture what is necessary to improve life in socially, politically, and environmentally just ways. While intersectionality theory is complex, its intricacies afford critical culture digital scholars the epistemological tools and methodological insights needed to highlight that people, communities, and the environment must be considered prior to, in the midst of, and downstream of technoculture's mechanistic, rationalist, extractive desires or capitalism's desires for profit through exploitation.
American Presidential Campaigns in the Attention Economy
Cambridge University Press, August 2026
Ceren Budak, Jonathan M. Ladd, Josh Pasek, Lisa Singh, Michael W. Traugott
In the social media era, people and organizations increasingly fight over the public's attention. This competition is especially intense in American presidential campaigns. As someone who used attention-grabbing strategies throughout his career, Donald Trump's communication skills were well suited to this environment. His campaigns illustrate the power and limits of attention-grabbing campaign tactics. While he successfully dominated the public's attention in his presidential campaigns, people consistently consumed more negative than positive information about him. Additionally, in this era, the campaign media system and the public continue to focus on incumbent performance, resisting efforts to change the subject. Finally, the editorial decisions that prestigious news outlets make about what to cover still seem to shape which stories people hear about in the crucial final weeks of the campaign.
A test of time: Modeling the long-term success of crowd collaborations
PNAS Nexus, August 2026
Abraham Israeli, David Jurgens, Daniel M Romero
The Internet has significantly expanded the potential for global collaboration, allowing millions of users to contribute to projects like Wikipedia. While some of these collaborative efforts see early success, their long-term success is key for lasting impact. Despite its importance, this dimension of success remains largely unexplored. Prior work has assessed the success of online collaborations, however, most approaches are time-agnostic, evaluating success without considering its longevity. Research on the factors that ensure the long-term preservation of high-quality standards in online collaboration is scarce. We address this gap. We propose a novel metric, “Sustainable Success,” which measures the ability of collaborative efforts to maintain their quality over time. We introduce the SustainPedia dataset, which compiles data from 48.7K Wikipedia articles. All articles in SustainPedia have reached the highest indicators of quality provided by the English Wikipedia, but a portion of them (7.5%) were later demoted from this high-quality status. SustainPedia includes each article’s label and more than 300 explanatory features such as edit history, user experience, and team composition. Using this dataset, we develop machine learning models to predict the sustainable success of Wikipedia articles. Our analysis reveals important insights. For example, we find that articles that take to be recognized as high-quality are more likely to maintain their status over time (i.e. be sustainable). Additionally, user experience emerged as the most critical predictor of sustainability. Our model provides the opportunity of automatically flagging articles at risk of being unsustainable. It also supports actionable factors that are significantly associated with (un)sustainable Wikipedia articles.
“Once you’re teaching, you’re consumed with so much!”: A longitudinal study of factors influencing teachers’ beliefs about teaching critical data literacy in social studies
Theory & Research in Social Education, July 2026
Tamara L. Shreiner, Mark Guzdial
Critical data literacy is essential in our data-laden society and should be taught in social studies, where data visualizations are prevalent, and students are supposed to develop skills that will prepare them for engaged citizenship. However, research indicates that social studies teachers typically have little formal preparation and a low sense of efficacy related to teaching data literacy. This longitudinal study seeks to identify factors that influence beginning elementary and secondary teachers’ sense of efficacy related to teaching critical data literacy in social studies. Through surveys and semi-structured interviews, we examined the beliefs and reported practices of teachers from a pre-service course designed to help them learn about teaching critical data literacy through fieldwork and their first year of teaching. We were interested in how their beliefs about critical data literacy and their self-efficacy changed, whether they ultimately reported teaching critical data literacy as first-year teachers, and what sources of efficacy information seemed to influence their beliefs and practices. Our findings suggest that pre-service coursework can have a positive influence on pre-service teachers’ beliefs, but they are likely to encounter challenges coupled with inconsistent sources of efficacy information in the field that will undermine their commitments to teaching critical data literacy.
Exploring the Design Space of Glanceable Smartwatch Feedback Displays: Experimental Study
JMIR Mhealth Uhealth, July 2026
Yuxuan Li, Mark Newman, Predrag Klasnja
Background: Self-monitoring technologies are commonly used to promote health behavior change, with glanceable displays offering continuous feedback throughout the day. Yet, it is still unclear how various aspects of these glanceable representations affect their interpretability and usability.
Objective: This study aimed to investigate the effects of 3 design factors—stylization, granularity, and salience—on users’ ability to understand glanceable smartwatch-based feedback on daily step goals. Methods: We conducted an online simulation study to examine how 3 design dimensions—stylization, salience, and granularity—influence the effectiveness of glanceable feedback displays. Stylization and salience were crossed in a 2×2 factorial design, while granularity varied from 1% to 20% progress increments. A total of 202 Amazon Mechanical Turk participants were randomly assigned to 1 of 16 smartwatch display conditions. In each condition, participants viewed feedback on daily step progress and estimated the level of progress shown. We measured estimation error and questionnaire-assessed perceived usability and acceptability. The collected data were analyzed using generalized estimating equations and linear regression.
Results: High stylization reduced accuracy (+4.52 error points; P<.001) and negatively affected perceptions across 6 dimensions, including comprehension (P=.003), complexity (P<.001), and usability (P=.001). Granularity had a nonlinear effect: error was lowest around 5%-10%, with sharp increases at 20%. The 10% level also received the most favorable ratings, for example, comprehension (+0.656; P=.003). Salience had no effect. Previous smartwatch users were less accurate than never-users (+7.46 points) but rated displays as more useful (P=.002) and easier to focus on (P<.001). Current users gave similarly positive ratings on attention and usefulness.
Conclusions: These findings could help researchers design effective glanceable smartwatch feedback displays and expand the design space for glanceable feedback.
Payment integrity in government programs: Takeaways from incorporating the behavioral sciences in US federal evaluations
Proceedings of the National Academy of Sciences of the United States of America, July 2026
Maya Duru, Hanna Hoover, Heather Barry Kappes, David Schwegman, Brigitte Seim, Mattie Toma Mary Clair Turner
A primary way the US federal government delivers public goods and services is via monetary payments. Ensuring that these payments are calculated accurately, delivered on time, and made to the correct recipients is important for government fiscal health. Inaccurate or delayed payments can weaken public trust in the government and undermine government accountability. In this article, we examine findings from a set of impact evaluations assessing interventions designed to improve payment integrity in US federal programs. The low-cost, evidence-based interventions draw on insights from the social and behavioral sciences and include modification of forms, changes to how and when agencies request information, and altering existing communications. The evaluations were conducted by the US General Services Administration’s Office of Evaluation Sciences in collaboration with agency partners. We extract three takeaways across four representative evaluations. First, the real-world evaluations validate a key implication of the behavioral science literature: interventions that reduce burdens for individuals have small effects that meaningfully improve payment integrity at scale. Second, effects attenuate across interventions and over time, suggesting a need for iterative evaluation. Finally, bureaucratic hurdles and administrative complexity are the main barriers to translating academic insights into real-world government programs. Addressing these challenges will require close collaboration between behavioral scientists and practitioners throughout the intervention design and evaluation process.
Creative Exploration Meets Time Constraints: Academic and Professional Approaches to Adversarial Thinking
Aadarsh Padiyath, Brooke Compton, Barbara Ericson, Mark Guzdial
Cybersecurity professionals must anticipate how attackers exploit system weaknesses -- a perspective known as ''adversarial thinking'' (AT). While cybersecurity education emphasizes the development of AT skills, there is little consensus on what these skills entail or how to assess them. We conducted a scoping review of computing education literature and semi-structured interviews (N=8) with cybersecurity professionals to better understand how adversarial thinking (AT) manifests in both academia and industry. Our analysis of 19 academic papers found that adversarial thinking is frequently described through system-centric approaches and/or hacker-centric definitions, with both presenting it as an open-ended creative exploration. However in practice, time constraints force cybersecurity professionals to develop a risk-based mindset and rely on institutionalized adversarial knowledge rather than a constant creative analysis. Our professionals strategically deploy adversarial thinking when standardized checklists feel inadequate or when potential risks warrant a deeper investigation. This paper contributes an empirically grounded account of how AT operates under professional constraints, and identify implications for how cybersecurity education can better prepare students.
Seeing like an API: Platform-mediated research and the politics of access
Big Data & Society, July 2026
Zoë Natalia Cullen, Nicole B Ellison, Irene V Pasquetto
This study examines the challenges that researchers faced while accessing, interpreting, and assessing data accessed via Meta's now-sunset CrowdTangle (CT) application programming interface (API). Drawing on interviews with 20 academic users, we introduce the framework of Seeing like an API to illustrate how CT's technical design, governance rules, and platform-mediated knowledge networks shaped what researchers could see, ask, and conclude about platform activity. Participants adapted to limitations such as follower-count thresholds, the exclusion of comment data, and opaque content removal through three recurring strategies: optimizing the API's capabilities, extending data visibility with supplementary methods, and verifying outputs against external sources. These adaptations were developed in response to ongoing uncertainty about data completeness and provenance, which constrained both the scope of feasible research questions and the reliability of resulting analyses. Our findings show how platform-controlled data infrastructures shape research design, collaboration, and epistemic norms, and they inform emerging models for independent, third-party-regulated data access, such as those envisioned in the European Union's Digital Services Act.
TubeStats and TokStats: Research Tools for Random Samples of YouTube and TikTok
Media and Communication, July 2026
Kevin Zheng, Reagan Keeney, Ryan McGrady, Vikramaditya Jaisingh, Ethan Zuckerman
YouTube and TikTok are two of the most popular digital communications platforms in the world, playing a disproportionately large role in global communications infrastructure in general and the consumption and dissemination of information in particular. As neither platform provides adequate mechanisms to produce representative samples of the content they host, researchers largely depend on opportunistic samples of popular, recommended, or otherwise known content. In this article, we present two dashboard-based tools, TubeStats and TokStats, built upon our recent research into random sampling techniques for each platform. These tools provide platform-wide statistics such as the number of hosted videos, view count distributions, linguistic distributions, and growth over time, which researchers can use to quantify and contextualize their research. We explain the architecture and sampling pipeline of each tool as well as the unique technical and methodological affordances and constraints involved with each. We document how these related techniques and tools have been applied by our lab, other scholars, and journalists to contextualize non-representative samples, compare platform use across languages and regions, and examine quotidian uses of the platforms that attention-optimized samples may obscure, as well as the broader range of methodological possibilities that representative sampling opens for platform research. Not to be taken for granted, we also explain the many challenges we face in developing and maintaining such tools, with implications for the practical development of open research infrastructures.
Language Disparities in Moderation Workforce Allocation by Social Media Platforms
FAccT '26: The 2026 ACM Conference on Fairness, Accountability, and Transparency, June 2026
Manuel Tonneau, Diyi Liu, Ryan McGrady, Kevin Zheng, Ralph Schroeder, Ethan Zuckerman, Scott Hale
Content moderation is one of the earliest and most consequential large-scale applications of artificial intelligence, shaping both online safety and the boundaries of permissible speech. In practice, moderation operates as a sociotechnical human-AI system, combining automated flagging with human review, as automated systems remain too error-prone to operate independently. Leveraging newly mandated transparency disclosures under the European Union's Digital Services Act (DSA), we conduct the first cross-platform audit of human content moderation workforce allocation across languages. Across six major platforms, we uncover substantial cross-lingual disparities in both language coverage and moderator staffing relative to the volume of user-generated content. While larger platforms such as YouTube and Meta employ moderators for many languages, millions of EU-based users on smaller platforms, including Twitter/X, post in languages without any dedicated human oversight. Among covered languages, staffing is often highly disproportionate to content volume, with English consistently prioritized over widely spoken Global Majority languages suchas Spanish, Portuguese, and Arabic. Where enforcement data permit analysis, we further document large cross-lingual disparities in individual moderator workload, with per-moderator daily decision volumes differing by more than an order of magnitude across languages. Together, these findings raise fundamental fairness concerns for both users and moderators, implying both an unequal user protection from online harms across linguistic communities, and an uneven distribution of the cognitive and emotional burdens of moderation labor. They further demonstrate that meaningful accountability for AI-assisted content moderation requires transparency not only about automated systems and enforcement outcomes, but about how human moderation capacity is allocated, justified, and sustained across languages.
Exposure to news and political content on TikTok: A Case Study Around the 2024 US Election
FAccT '26: The 2026 ACM Conference on Fairness, Accountability, and Transparency, June 2026
Chang Ge, Nilay Gautam, Shreyas Vajjhala, Christina NG, Ariel Hasell, Sabina Tomkins
TikTok is a major social media platform with a distinctive centralized recommendation system used by 56% of adults under 34 years of age in the U.S. In this study, we deploy approximately 200 sock puppet accounts during the 2024 U.S. election to examine how TikTok's recommendation algorithm might respond to users' interests in news and political content. Drawing on over 217,064 unique videos viewed by these accounts, we provide a data-driven analysis of the extent to which accounts with different interests are exposed to news media, and election-related political content. We find that TikTok rarely shows news or political content to accounts. Even the accounts that only engage with news and political content are recommended less than 0.30 total news videos per browsing session on average, though they are recommended slightly more non-news political content (1.60 total political videos on average). We find that users are recommended far more election-related political content posted by everyday users than by news organizations. The lack of news content on TikTok, especially for users who have demonstrated an interest in news, has implications for political knowledge amongst social media users.
From Moves to Pathways: Characterizing Pedagogical Discourse Dynamics in Online Tutoring with Bayesian Generative Modeling
Communications in Computer and Information Science, June 2026
Michelle Light, Michael Ion, Kevyn Collins-Thompson
What distinguishes tutoring conversations where students resolve their problem from those where they don’t? Prior work has established that learning depends not on any single tutoring strategy but on specific patterns of student impasse and self-resolution, and that these patterns form latent dialogue modes that predict learning better than individual moves. However, these findings typically come from small, controlled studies with manual annotation based on a relatively constrained move taxonomy. We introduce a two-stage generative model-based approach that analyzes 2,437 naturally-occurring math tutoring conversations (51,376 messages) from MathMentorDB, each labeled with a rich two-level discourse taxonomy of tutor and student moves obtained with high-quality Large Language Model (LLM) classification. With these move labels, we then fit a Bayesian Hidden Markov Model (HMM) that discovers conversation pathways comprising latent sequences of pedagogical states from discourse move sequences, and estimates outcome-level variation that characterizes how resolved and unresolved conversations differ. For example, while direct support from tutors (Lecturing state) has the same prevalence in resolved and unresolved tutoring sessions, resolved sessions are characterized by much more active student inference.
Position Paper: Replication Crisis in Human–Centered Security Research-Are We There Yet?
2026 IEEE Symposium on Security and Privacy Workshops, May 2026
Juliane Schmüser, Jan-Ulrich Holtgrave, Florian Schaub, Sascha Fahl
Replicability is critical for meta-science, verifying scientific knowledge, and advancing science as a whole. However, the broader security community and the human-centered security (HCS) research community, in particular, struggle with making studies replicable, and only a few replication studies have been published so far. In this position paper, we analyze Call for Papers (CfPs) and existing replication studies to identify pain points that hinder replication in HCS research, use those insights to discuss whether the community is already in a replication crisis, and assess the extent to which this contributes to or hinders meta-science in the field. We conclude with calls to action to improve reporting transparency and data sharing, better recognize replication studies as valuable scientific contributions, and develop community-driven understandings of how replications should be conducted for HCS research.
Pre-prints, Working Papers, Articles, Workshops and Talks
Mentored research experiences bridge the wealth gap in PhD admissions
Research Square, August 2026
Wenhao Sun, Nicholas David, Misha Teplitskiy, Sidney Xiang, Daniel Romero, Dallas Card
PhD admissions decide which aspiring researchers gain access to scientific careers. Socioeconomic background has long shaped academic opportunity, influencing entry into elite undergraduate institutions1–3 and the professorate4, yet its role in doctoral admissions has remained difficult to measure directly. Here we analyze 52,888 PhD applications submitted between 2013-2023 to one of the largest PhD-granting institutions in the United States, combining applicant credentials and recommendation letters with faculty evaluations and socioeconomic measures—including family home equity, parental education and reported financial hardship. We uncover a distinct wealth gap in doctoral admissions: applicants from higher socioeconomic backgrounds are both more likely to apply and to be admitted. Although faculty do not explicitly evaluate wealth, socioeconomic advantage systematically propagates through key admissions criteria including grades, test scores and undergraduate institutional prestige. In contrast, we find that mentored research experiences, and their resulting recommendations, provide an influential and wealth-agnostic signal of research potential. Expanding access to scientific training therefore calls for a stronger ecosystem of mentored research experiences, broader faculty participation in mentoring, and earlier encouragement for aspiring scholars to seek out these opportunities.
Mixed-Agent Museum Tour Guide Design Improves Gendered Learning Outcomes and Visitor Preferences
arXiv, July 2026
Annette M. Masterson, Wonse Jo, Helena C. Sieh, Lionel P. Robert Jr., Dawn Tilbury
Robots are increasingly integrated into everyday contexts, including museums, where they can both entertain and educate visitors. To enhance visitor experience and engagement, we present a novel mixed-agent tour guide system that combines a physical robot with a projected virtual agent that actively participates in the tour through conversation and interaction, achieving the interaction richness of two mobile agents from a single platform. We validate the system through a within-subjects study with 30 participants to assess engagement, quality of experience, and learning performance. Participants experienced different conversational styles and agent configurations, and data were collected via surveys, behavioral sensors, and interviews. Results showed that engagement and quality of experience remained consistent across conditions. Learning performance revealed a significant gender-moderated difference: the mixed-agent conditions improved learning performance for female participants. This suggests that the proposed dyadic conversational style in this paper influenced learning performance differently by gender. Nonetheless, in interviews, participants reported a greater preference for mixed-agent teams regardless of gender, citing interaction as a key factor in their experience.
Evaluating Affective Objectives: Statistical Numbing in Data Visualization
arXiv, July 2026
Elsie Lee-Robbins, Eytan Adar
Visualizations can help audiences understand the scale of tragedies, such as the consequences of natural disasters, war, genocide, and pandemics. In these cases, a visualization designer's default behavior may be to focus on communicating quantitative information: numbers, statistics, and trends. However, this may not reflect higher-level affective objectives to inspire their audience to care about an issue, empathize with others, or take action to help those in need. Worse, standard visualizations may conflict with these goals, as statistics can numb emotions and reduce prosocial feelings toward people in need. Designers have developed strategies to increase affective responses through data visualizations, such as blending data narratives and personal narratives about individuals. In this paper, we explore three design strategies for communicating a humanitarian crisis: data-driven, human-driven, or mixed narratives. We conducted an empirical study to explore the effect of statistical numbing in the context of these types of narratives in the format of data videos. In particular, we measure prosocial feelings and behaviors by giving participants the option of donating money as part of the study. We find that human-driven narratives (photographs and stories of individuals) elicited the highest donations and that the mixed narrative combination led to the lowest donations. We discuss the limitations of this study and the implications of pursuing affective objectives and the numbing of empathy in data visualization design.
Reducing the rate of personal insults in social media with bystander bots
arXiv, June 2026
Libby Hemphill, Lingyao Li, Ryan Burton, David Jurgens
Prompted by previous research on strategies for reducing interpersonal conflict and addressing problematic behaviors in online communities, a randomized controlled trial on Reddit compared various responses for reducing the rate of personal insults users post to the site. We generated replies from five deescalation strategies and used an automated procedure for posting them as replies to insulting comments. The findings reveal that automated replies to insults can effectively reduce their rate. Appreciation performed best. Not all strategies performed well, though. We conclude that automated responses are a viable tool for addressing some problematic behaviors. We discuss their potential utility and limitations.
BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks
arXiv, June 2026
Libby Hemphill, Lingyao Li, Ryan Burton, David Jurgens
Prompted by previous research on strategies for reducing interpersonal conflict and addressing problematic behaviors in online communities, a randomized controlled trial on Reddit compared various responses for reducing the rate of personal insults users post to the site. We generated replies from five deescalation strategies and used an automated procedure for posting them as replies to insulting comments. The findings reveal that automated replies to insults can effectively reduce their rate. Appreciation performed best. Not all strategies performed well, though. We conclude that automated responses are a viable tool for addressing some problematic behaviors. We discuss their potential utility and limitations.
Validating LLMs in social science: Epistemic threats and emerging norms
arXiv, July 2026
Meera Desai, Dallas Card, Abigail Z. Jacobs
Large language models (LLMs) are reshaping social science methodology. Researchers increasingly prompt language models to generate quantitative measurements of social concepts, for example labeling data or simulating survey responses. Yet LLMs pose methodological challenges including bias, hallucination, and brittleness across contexts, with unclear threats to validity. Standard practices and norms for addressing these challenges are still emerging. We collect and systematically analyze validation practices in a comprehensive corpus of papers from eight flagship social science journals that use LLMs as measurement instruments. We find that LLM-generated measurements frequently play a central role in empirical analyses, yet validation practices are inconsistent and limited. We outline complementary strategies for more robust validation, pointing toward better norms and standards around the use of LLMs in social science.
RELATED
Keep up with research from UMSI experts by subscribing to our free research roundup newsletter!