Skip to main content

Spring Deadline: Sunday, March 1 @ 11:59pm PT. Click here to apply.

UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs

UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs

December 1, 2025

Theory of Mind (ToM), the ability to understand the mental states of oneself and others, remains a challenging area for large language models (LLMs), which often fail to predict human mental states ac...

Accepted to NAACL SRW 2025

Authors: Prameshwar Thiyagarajan, Vaishnavi Parimi, Soumil Garg, Zhangir, Shamant, Nitin Yarlagadda

Theory of Mind (ToM), the ability to understand the mental states of oneself and others, remains a challenging area for large language models (LLMs), which often fail to predict human mental states accurately. We present UniToMBench, a unified benchmark that integrates the strengths of SimToM and TOMBENCH to systematically improve and assess ToM capabilities in LLMs by integrating multi-interaction task designs and evolving story scenarios. Supported by a custom dataset of over 1,000 hand-written scenarios, UniToMBench combines perspective-taking techniques with diverse evaluation metrics to better stimulate social cognition in LLMs. Through evaluation, we observe that while models like GPT-4o show consistently high accuracy in tasks involving emotional and belief-related scenarios, with results usually above 80%, there is significant variability in their performance across knowledge-based tasks.

Begin Your Journey

The application takes 10 minutes and is reviewed on a rolling basis. We look for strong technical signal—projects, coursework, or competition results—and a genuine curiosity to do real research.

If admitted, you will join a structured pipeline with direct mentorship to take your work from ideation to top conference submission at venues like NeurIPS, ACL, and EMNLP.