AIDB Daily Papers
会話を超えた心の理論と説得:LLMが行動を通じて他者の信念を操作する能力の評価
※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。
ポイント
- LLMが会話ではなく物理的な行動を通じて他者の信念状態を誘導する能力である「非会話的計画心の理論(NCP-ToM)」を評価する枠組みを構築した。
- 従来の受動的な質問応答形式ではなく、エージェントが自律的に状況を操作して目標を達成する能力を問う点は、AIの社会的な安全性やリスクを理解する上で重要である。
- GPT-5は人間を上回る約80%の成功率を記録したが、真の信念を誘導する方が誤った信念を誘導するよりも容易であるという人間と同様の傾向が確認された。
Abstract
Theory of Mind (ToM) benchmarks for Large Language Models (LLMs) typically rely on passive question-answering formats, but the deployment of LLMs in increasingly agentic and autonomous forms demands new evaluations. In this paper we evaluate an agent's ability to induce specific belief states in other agents by taking actions rather than using conversational persuasion, a capability we call Non-Conversational Planning ToM (NCP-ToM). NCP-ToM is likely to be essential for many agent use-cases, including within user-assistant interactions and pedagogical contexts, but may also present manipulation or misinformation risks. Using a novel framework, NCP-ExploreToM, we subvert the conventional task structure by providing models with a set of belief state goals and requiring them to move objects or direct characters into rooms to achieve their goals. We evaluated six frontier models, including GPT-5, Gemini 2.5 Pro and the Claude 4 series, and a cohort of human participants, across 600 task instances. GPT-5 was successful on approximately 80% of tasks in the agentic setting, and was the only model to outperform human participants on our task, but was still less robust than humans across contexts. We additionally found that all models, like humans, performed better on tasks inducing true belief states than false belief states, which is a positive signal for alignment efforts. These findings highlight emerging social-reasoning capabilities in LLMs for non-conversational task completion and underscore the necessity of agentic evaluations for understanding the safety and alignment of autonomous social agents.
Paper AI Chat
この論文のPDF全文を対象にAIに質問できます。
質問の例: