次回の更新記事:AIにイラッと打ち直すたび、脳は影響を受けて変形し…(公開予定日:2026年07月21日)
AIDB Daily Papers

会話を超えた心の理論と説得:LLMが行動を通じて他者の信念を操作する能力の評価

原題: Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action
著者: Ben Slater, Matteo G. Mecattaf, Lucy G. Cheke, John Burden, Winnie Street
公開日: 2026-06-30 | 分野: LLM 安全性 AI エージェント cs.CL アライメント

※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。

ポイント

  • LLMが会話ではなく物理的な行動を通じて他者の信念状態を誘導する能力である「非会話的計画心の理論(NCP-ToM)」を評価する枠組みを構築した。
  • 従来の受動的な質問応答形式ではなく、エージェントが自律的に状況を操作して目標を達成する能力を問う点は、AIの社会的な安全性やリスクを理解する上で重要である。
  • GPT-5は人間を上回る約80%の成功率を記録したが、真の信念を誘導する方が誤った信念を誘導するよりも容易であるという人間と同様の傾向が確認された。

Abstract

Theory of Mind (ToM) benchmarks for Large Language Models (LLMs) typically rely on passive question-answering formats, but the deployment of LLMs in increasingly agentic and autonomous forms demands new evaluations. In this paper we evaluate an agent's ability to induce specific belief states in other agents by taking actions rather than using conversational persuasion, a capability we call Non-Conversational Planning ToM (NCP-ToM). NCP-ToM is likely to be essential for many agent use-cases, including within user-assistant interactions and pedagogical contexts, but may also present manipulation or misinformation risks. Using a novel framework, NCP-ExploreToM, we subvert the conventional task structure by providing models with a set of belief state goals and requiring them to move objects or direct characters into rooms to achieve their goals. We evaluated six frontier models, including GPT-5, Gemini 2.5 Pro and the Claude 4 series, and a cohort of human participants, across 600 task instances. GPT-5 was successful on approximately 80% of tasks in the agentic setting, and was the only model to outperform human participants on our task, but was still less robust than humans across contexts. We additionally found that all models, like humans, performed better on tasks inducing true belief states than false belief states, which is a positive signal for alignment efforts. These findings highlight emerging social-reasoning capabilities in LLMs for non-conversational task completion and underscore the necessity of agentic evaluations for understanding the safety and alignment of autonomous social agents.

Paper AI Chat

この論文のPDF全文を対象にAIに質問できます。

質問の例:

AIチャット機能を利用するには、ログインまたは会員登録(無料)が必要です。

会員登録 / ログイン

関連するAIDB記事