Maarten Sap

I am an assistant professor at CMU's LTI department with a courtesy appointment in HCII, and a part-time research scientist and AI safety lead at the Allen Institute for AI (AI2). My research focuses on (1) measuring and improving AI systems' social and interactional intelligence, (2) assessing and combatting social inequality, safety risks, and socio-cultural biases in human- or AI-generated language, and (3) building narrative language technologies for prosocial outcomes. I was named a 2025 Packard Fellow and a recipient of the 2025 Okawa Research Award.

I received my PhD from the University of Washington where I was advised by Noah Smith and Yejin Choi.
[bio for talks]

Recent updates:

December 2025 πŸ…πŸ“ƒ: Very excited to have our paper Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond) selected for a Best Paper Award at NeurIPS 2025 (Datasets and Benchmarks Track)!! Huge congrats to the first author Liwei Jiang!!!

November 2025 πŸ’ŽπŸš€: Honored to be a Spring 2025 recipient of the Amazon Research Award for our project on measuring AI agentic safety!

October 2025 πŸ…β­: I’m super excited and grateful to announce that I'm part of the 2025 class of Packard Fellows. The Packard Foundation and this fellowship will allow me to explore exciting research directions towards culturally responsible and safe AI 🌍🌈

October 2025 πŸ”πŸ§‘β€πŸŽ“: Due to my lab being quite full already, I'm not taking looking for any new students in this upcoming PhD application cycle 😟.

October 2025 πŸ‡¨πŸ‡¦πŸŽ‰: Excited to be attending COLM 2025 in Montreal this October! I'll be giving a talk at the Social Sim Workshop on Unlocking Social Intelligence in AI agents. I'm also thrilled that five papers I co-authored will be presented by my amazing collaborators at COLM: HAICOSYSTEM: An Ecosystem for Sandboxing Safety Risks in Human-AI Interactions (led by Xuhui Zhou et al.), ALFA: Aligning LLMs to Ask Good Questions: A Case Study in Clinical Reasoning (co-led by Jimin Mun et al.), PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages, Fluid Language Model Benchmarking, and The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains.

August 2025 🌟: Incredibly honored to be one of 7 US recipients of the 2025 Okawa Research Grant from the Okawa Foundation!

August 2025 πŸ§‘β€πŸŽ“: Welcoming my first postdoc, Vasudha Varadarajan, to the lab!

[older news]


My research group:

Dan Chechelnitsky

CMU Portugal LTI PhD student
co-advised with Chrysoula Zerva

Joel Mire

LTI PhD student

Karina Halevy

LTI PhD student
co-advised with Mona Diab

Malia Morgan

Pre-doctoral Young Investigator at Ai2

Jimin Mun

LTI PhD student

Jocelyn Shen

MIT PhD student
co-advised with Cynthia Breazeal

Kynnedy Smith

HCII PhD student
co-advised with Motahhare Eslami

Vasudha Varadarajan

LTI Postdoc

Akhila Yerukola

LTI PhD student

Mingqian Zheng

LTI PhD student
co-advised with Carolyn RosΓ©

Xuhui Zhou

LTI PhD student


Overarching Research Themes

Themes extracted and images generated with the OpenAI API; there may be inconsistencies.

Social Pragmatics in AI

My research group explores how to evaluate and improve AI systems’ ability to handle subtle social meaning, including theory of mind, non-literal intent, and context-sensitive interaction. Recent work such as [Cognitive Chain-of-Thought: Structured Multimodal Reasoning about Social Situations](https://arxiv.org/abs/2507.20409) and [SOTOPIA-ToM: Evaluating Information Management in Multi-Agent Interaction with Theory of Mind](https://arxiv.org/abs/2605.02307) pushes toward benchmarks that probe whether models can track beliefs, hidden goals, and shifting social context. Other studies like [When Should AI Read the Room? Public Perceptions of Social Intelligence in AI Agents](https://arxiv.org/abs/2605.29938) show that social intelligence is not just a capability question, but also a user-expectation question. Together, these papers suggest a move from isolated language tasks toward richer evaluations of socially aware behavior in interactive settings.

Measuring Safe Agency and Reliance

My research group explores new ways to measure the safety of increasingly agentic AI systems, especially when they influence, mislead, or over-dependably assist people. [OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety](https://arxiv.org/abs/2507.06134) broadens the safety lens to realistic agent deployments, while [The Hidden Puppet Master: Predicting Human Belief Change in Manipulative LLM Dialogues](https://arxiv.org/abs/2603.20907) focuses on the risk of subtle persuasion and manipulation. Work such as [Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance](https://aclanthology.org/2025.naacl-long.556/) highlights that safety must include human-centered failure modes like overreliance, not just model-side errors. Overall, the field is moving toward metrics that capture both agent autonomy risks and the downstream effects on user judgment and behavior.

Culturally Grounded AI Adaptation

My research group explores how to build AI that is culturally competent, adaptable, and fair across communities, languages, and social norms. [CCBENCH: Assessing LLM Cultural Competence via Implicitly Signaled Norms using Health Queries](https://arxiv.org/abs/2607.05405) and [NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models](https://aclanthology.org/2025.naacl-long.120/) exemplify the push toward benchmarks that test whether models understand local norms rather than just global defaults. At the same time, [Black LLMirror: User (Self) Perceptions in Black American English Interactions with LLMs](https://dl.acm.org/doi/abs/10.1145/3772318.3791111) shows how personalization and dialect handling can shape identity, trust, and user experience in ethically important ways. This line of research emphasizes that cultural adaptation is not a cosmetic feature: it affects inclusion, bias, and whether AI systems serve users equitably.

Story Understanding for Connection

My research group explores AI for human-human connection through story understanding, focusing on how models interpret narratives, empathy, and interpersonal meaning in stories. [Social Story Frames: Contextual Reasoning about Narrative Intent and Reception](https://arxiv.org/abs/2512.15925) highlights how the same story can be read differently depending on audience, intent, and context. [HEART-felt Narratives: Tracing Empathy and Narrative Style in Personal Stories with LLMs](https://arxiv.org/abs/2405.17633) and [Modeling Empathic Similarity in Personal Narratives](https://arxiv.org/abs/2305.14246) show growing interest in using LLMs to recognize emotional resonance and narrative style in personal accounts. Taken together, these papers point toward systems that support better listening, interpretation, and connection across human stories.