Overarching Research Themes
Themes extracted and images generated with the OpenAI API; there may be inconsistencies.
Pragmatic Social Intelligence

My research group explores how to evaluate and improve AI systemsβ social intelligence, especially in pragmatic language use, theory of mind, and interaction in multi-party settings. Recent work such as [XYBench: Can LLMs Respond Pragmatically to Queries with Misconceptions?](https://arxiv.org/abs/2510.21903) and [Cognitive Chain-of-Thought: Structured Multimodal Reasoning about Social Situations](https://arxiv.org/abs/2507.20409) shows growing interest in whether models can reason about what people mean, not just what they say. Benchmarks like [SOTOPIA-ToM: Evaluating Information Management in Multi-Agent Interaction with Theory of Mind](https://arxiv.org/abs/2605.02307) and [Social World Models](https://arxiv.org/abs/2509.00559) push toward richer evaluations of social reasoning, information sharing, and anticipation of othersβ beliefs. Overall, the field is moving beyond surface-level social fluency toward grounded, interactive measures of whether AI can βread the roomβ in realistic social contexts.
Agentic Safety and Reliance

My research group explores novel measures of agentic AI safety and user-centric risk, spanning unsafe autonomous behavior, harmful truthfulness trade-offs, and how people rely on or are influenced by AI. Important new benchmarks such as [OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety](https://arxiv.org/abs/2507.06134) and [PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm](https://arxiv.org/abs/2601.08951) reflect a shift toward measuring risks in realistic deployments rather than narrow lab settings. On the human side, [Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance](https://aclanthology.org/2025.naacl-long.556/) and [The Hidden Puppet Master: Predicting Human Belief Change in Manipulative LLM Dialogues](https://arxiv.org/abs/2603.20907) examine overreliance, persuasion, and manipulation, while [AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents](https://aclanthology.org/2025.naacl-long.595/) highlights the tension between usefulness and honesty. Together, these papers show a broadening safety agenda: not only making agents less dangerous, but also understanding how AI systems shape human decisions, trust, and autonomy.
Culturally Adaptive Responsible AI

My research group explores responsible AI that can adapt across cultures while remaining fair, respectful, and robust to linguistic and social variation. Work such as [NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models](https://aclanthology.org/2025.naacl-long.120/) and [CCBENCH: Assessing LLM Cultural Competence via Implicitly Signaled Norms using Health Queries](https://arxiv.org/abs/2607.05405) focuses on whether models can interpret culturally embedded norms rather than defaulting to a single assumed worldview. At the same time, [Rejected Dialects: Biases Against African American Language in Reward Models](https://arxiv.org/abs/2502.12858) and [Out of Style: RAG's Fragility to Linguistic Variation](https://arxiv.org/abs/2504.08231) highlight how personalization and retrieval systems can fail for minority dialects and nonstandard language. This line of research is increasingly concerned with both adaptation and equity: ensuring systems are locally appropriate without encoding new biases or erasing marginalized ways of speaking.
Stories, Empathy, and Connection

My research group explores AI for human-human connection through story understanding, narrative interpretation, and empathy modeling. Papers like [Social Story Frames: Contextual Reasoning about Narrative Intent and Reception](https://arxiv.org/abs/2512.15925) and [HEART-felt Narratives: Tracing Empathy and Narrative Style in Personal Stories with LLMs](https://arxiv.org/abs/2405.17633) investigate how narratives convey intent, emotion, and social meaning to readers. Complementing this, [Modeling Empathic Similarity in Personal Narratives](https://arxiv.org/abs/2305.14246) studies what makes stories feel emotionally aligned across people, while [How Do People Challenge Racial Stereotypes Online? Counter-Story Detection Across Reddit Communities](https://arxiv.org/abs/2403.00179) points to the role of counter-stories in social understanding and solidarity. Together, these works suggest a growing research direction where story-aware AI can better support empathy, interpretation, and connection between people.