Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations

arXiv cs.AI·Nanxu Gong, Zixin Chen, Haotian Li, Zishu Zhao, Jianxun Lian, Huamin Qu, Yanjie Fu, Xing Xie

4d ago

·~2 min·5/18/2026·en·1

Quick Take

Improving Theory of Mind in LLMs enhances human-AI interactions but requires dynamic evaluation methods.

Key Points

Existing benchmarks overlook first-person, dynamic HAI interactions.
New interactive ToM evaluation paradigm proposed for better assessment.
Static benchmark improvements don't guarantee better dynamic performance.

📖 Reader Mode

~2 min read

[Submitted on 28 Apr 2026]

View PDF HTML (experimental)

Abstract:Improving the Theory of Mind (ToM) capability of Large Language Models (LLMs) is crucial for effective social interactions between these AI models and humans. However, the existing benchmarks often measure ToM capability improvement through story-reading, multiple-choice questions from a third-person perspective, while ignoring the first-person, dynamic, and open-ended nature of human-AI (HAI) interactions. To directly examine how ToM improvement techniques benefit HAI interactions, we first proposed the new paradigm of interactive ToM evaluation with both perspective and metric shifts. Next, following the paradigm, we conducted a systematic study of four representative ToM enhancement techniques using both four real-world datasets and a user study, covering both goal-oriented tasks (e.g., coding, math) and experience-oriented tasks (e.g., counseling). Our findings reveal that improvements on static benchmarks do not always translate to better performance in dynamic HAI interactions. This paper offers critical insights into ToM evaluation, showing the necessity of interaction-based assessments in developing next-generation, socially aware LLMs for HAI symbiosis.

Subjects:	Artificial Intelligence (cs.AI)
Cite as:	arXiv:2605.15205 [cs.AI]
	(or arXiv:2605.15205v1 [cs.AI] for this version)
	https://doi.org/10.48550/arXiv.2605.15205 arXiv-issued DOI via DataCite

Submission history

From: Nanxu Gong [view email]
[v1] Tue, 28 Apr 2026 15:38:31 UTC (7,139 KB)

— Originally published at arxiv.org

Continue reading on arxiv.org

Want this in your inbox every morning?

Daily brief at your local 8am — bilingual EN/中文, free.

Subscribe — it's free

Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations

Quick Take

Key Points

📖 Reader Mode

Submission history

Want this in your inbox every morning?

More from arXiv cs.AI

From Prompts to Protocols: An AI Agent for Laboratory Automation

Agentic Trading: When LLM Agents Meet Financial Markets

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents

Related in this space

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?