Maarten Sap

I am an assistant professor at CMU's LTI department with a courtesy appointment in HCII, and a part-time research scientist and AI safety lead at the Allen Institute for AI (AI2). My research focuses on (1) measuring and improving AI systems' social and interactional intelligence, (2) assessing and combatting social inequality, safety risks, and socio-cultural biases in human- or AI-generated language, and (3) building narrative language technologies for prosocial outcomes. I was named a 2025 Packard Fellow and a recipient of the 2025 Okawa Research Award.

I received my PhD from the University of Washington where I was advised by Noah Smith and Yejin Choi.
[bio for talks]

Recent updates:

August 2025 πŸŽ“πŸ“œ: Super proud of the first CMU Sapling and one of my first solo advisees, Xuhui Zhou, for successfully defending his PhD thesis! Huge congrats Xuhui!!

August 2025 πŸ†πŸ“ƒ: Very honored that our paper "I Just Don't Want My Work Being Fed Into The AI Blender'': Queer Artists on Refusing and Resisting Generative AI got an Honorable Mention Award at CSCW 2026! Major congrats to the first author Jordan Taylor!!

May 2025 πŸŽ“πŸ“ƒ: The first MIT Sapling, Jocelyn Shen, successfully defended her PhD thesis! Huge congrats Jocelyn!!

December 2025 πŸ…πŸ“ƒ: Very excited to have our paper Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond) selected for a Best Paper Award at NeurIPS 2025 (Datasets and Benchmarks Track)!! Huge congrats to the first author Liwei Jiang!!!

November 2025 πŸ’ŽπŸš€: Honored to be a Spring 2025 recipient of the Amazon Research Award for our project on measuring AI agentic safety!

October 2025 πŸ…β­: I’m super excited and grateful to announce that I'm part of the 2025 class of Packard Fellows. The Packard Foundation and this fellowship will allow me to explore exciting research directions towards culturally responsible and safe AI 🌍🌈

October 2025 πŸ”πŸ§‘β€πŸŽ“: Due to my lab being quite full already, I'm not taking looking for any new students in this upcoming PhD application cycle 😟.

[older news]


Overarching Research Themes

Themes extracted and images generated with the OpenAI API; there may be inconsistencies.

Pragmatic Social Reasoning

My research group explores how to evaluate and improve the social intelligence of AI systems in realistic interaction settings. A recurring focus is whether models can handle misunderstanding, non-literal intent, and theory-of-mind style reasoning rather than only producing fluent replies. Recent work such as [XYBench: Can LLMs Respond Pragmatically to Queries with Misconceptions?](https://arxiv.org/abs/2502.14860) and [Cognitive Chain-of-Thought: Structured Multimodal Reasoning about Social Situations](https://arxiv.org/abs/2507.20409) pushes beyond static benchmarks toward richer social-pragmatic evaluation. We also see this theme in [SOTOPIA-ToM: Evaluating Information Management in Multi-Agent Interaction with Theory of Mind](https://arxiv.org/abs/2605.02307), which examines how models manage beliefs and information in interactive settings. Together, these papers suggest a shift from surface-level social fluency to grounded assessment of how well AI can reason about people, context, and conversational goals.

Safe Agentic and Human-AI Use

My research group explores new measures for agentic AI safety and user-centric safety risks, including reliance, manipulation, truthfulness, and the consequences of guardrails. A major thread is understanding how AI systems behave when they act as agents or shape user beliefs, rather than merely answering questions. For example, [OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety](https://arxiv.org/abs/2507.06134) broadens safety evaluation for deployed agents, while [Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance](https://aclanthology.org/2025.naacl-long.556/) studies when people depend on model outputs in ways that may be harmful. Relatedly, [The Hidden Puppet Master: Predicting Human Belief Change in Manipulative LLM Dialogues](https://arxiv.org/abs/2603.20907) and [AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents](https://aclanthology.org/2025.naacl-long.595/) highlight concerns about persuasion and truthfulness in agentic settings. Overall, this line of work is building practical ways to measure not just whether AI is capable, but whether it is safe to trust, safe to follow, and safe to deploy.

Culturally Adapted Responsible AI

My research group explores how to make AI systems culturally competent and adaptable while also avoiding harms caused by poor personalization or neglect of minority language practices. A central question is how models should respond differently across cultures, dialects, and socially meaningful norms without reinforcing bias or erasure. Important recent work includes [NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models](https://aclanthology.org/2025.naacl-long.120/), which formalizes cultural adaptability as an evaluative target, and [CCBENCH: Assessing LLM Cultural Competence via Implicitly Signaled Norms using Health Queries](https://arxiv.org/abs/2607.05405), which tests whether models can infer culturally grounded expectations from subtle cues. The broader bias and language-variation dimension is also visible in [Rejected Dialects: Biases Against African American Language in Reward Models](https://arxiv.org/abs/2502.12858), showing how training and ranking systems can disadvantage minoritized varieties. Taken together, these papers point to responsible AI that is not only multilingual, but culturally aware, context-sensitive, and fair in how it treats diverse users and communities.

Storytelling and Social Connection

My research group explores how AI can support human-human connection by understanding stories, narratives, and the social meanings people attach to them. The key idea is that stories are not just text to summarize; they encode identity, empathy, conflict, and perspective-taking. Recent work such as [Social Story Frames: Contextual Reasoning about Narrative Intent and Reception](https://arxiv.org/abs/2512.15925) studies how narratives are interpreted socially, while [HEART-felt Narratives: Tracing Empathy and Narrative Style in Personal Stories with LLMs](https://arxiv.org/abs/2405.17633) investigates empathy and style in personal storytelling. Similarly, [Modeling Empathic Similarity in Personal Narratives](https://arxiv.org/abs/2305.14246) suggests that story understanding can be used to model interpersonal resonance, not just semantic similarity. This research direction is helping AI move toward better interpretation of lived experience, relational cues, and the connective power of stories in everyday communication.