次回の更新記事:【論文著者監修・コメント】AIエージェントへの人間…(公開予定日:2026年07月27日)
AIDB Daily Papers

LLMの心の理論を試す:3者間の駆け引きを導入した「人狼ゲーム」による評価

原題: Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs
著者: Avni Mittal
公開日: 2026-06-26 | 分野: LLM マルチエージェント cs.CL cs.AI cs.MA cs.GT

※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。

ポイント

  • 従来の2者間対立を超え、第3の陣営「道化師」を加えた複雑な人狼ゲームを用いてLLMの推論能力を検証した。
  • 複数の利害関係が絡む構造にすることで、単なる言語的バイアスに頼らない高度なマルチエージェント推論を評価可能にした。
  • GPT-4やDeepSeekなどのモデルは、複雑な状況下で自己敗北的な行動をとるなど、モデルごとの戦略的思考の差が浮き彫りとなった。

Abstract

Theory-of-mind evaluations of large language models typically use dyadic social-deduction games, where every observable cue points to a single hidden side, so a model with strong language priors can score well without ever simulating opponents' incentives. We extend the Werewolf game with a Jester, a third faction whose utility on peer suspicion is inverted because it wins by being voted out, so optimal play requires reasoning across three opposing utility functions. Across 60 games on GPT-4.1, DeepSeek-V3.1, and Llama-3.3-70B with Jester self-learning on and off, the Jester wins 60-70% of games while Werewolves never exceed 20%, and GPT-4.1 wolves vote the Jester out on day 1 in 60-70% of games, a strictly self-defeating action. Self-learning helps DeepSeek and Llama but hurts GPT-4.1, with the cost landing on Villagers rather than Werewolves. Only DeepSeek learns the subtle strategy of looking suspicious without looking intentionally suspicious, and it gains the most from the loop. Triadic incentive structure exposes a layer of multi-agent reasoning that dyadic deduction games leave invisible.

Paper AI Chat

この論文のPDF全文を対象にAIに質問できます。

質問の例:

AIチャット機能を利用するには、ログインまたは会員登録(無料)が必要です。

会員登録 / ログイン

関連するAIDB記事