次回の更新記事:AIエージェントのトークンコストを4割削る設計(公開予定日:2026年07月27日)
AIDB Daily Papers

AIの解釈性を根本から問い直す:ラディカル解釈のフレームワーク

原題: Radical AI Interpretability
著者: Daniel A. Herrmann, Benjamin A. Levinstein
公開日: 2026-06-25 | 分野: 解釈性 AI XAI cs.AI cs.LG AI安全性

※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。

ポイント

  • AIシステムを「エージェント」とみなし、その信念、欲求、意味を計算事実から解き明かすフレームワークを提案する。
  • AIの信頼性確保や欺瞞検出のために、解釈手法の成功基準を明確化し、全体論的なアプローチを提唱する点が重要である。
  • 提案手法は、AIの内部状態と外部行動の間の制約関係を利用し、解釈の精度を高めることを示唆する。

Abstract

We develop a framework for interpreting AI systems as agents, drawing on the philosophical tradition of radical interpretation and the tools of mechanistic interpretability. The core question is: given the computational facts about a system, how do we solve for its beliefs, desires, and meanings? This matters increasingly for safety. We want to be able to trust the systems we deploy, whether by understanding their goals or, more modestly, by reliably detecting deception. Interpretability researchers are building tools to read beliefs and desires off a model's internals, but there is no settled account of when such a tool has succeeded. This book supplies one. We propose criteria on both representationalist and interpretationist approaches, and tie each to tests current interpretability methods can carry out. A central lesson is that these attributions cannot be made piecemeal. Beliefs, desires, and the propositional structure they presuppose are jointly constrained, and a method that fixes one while measuring the others inherits whatever distortions that introduces. This holism becomes pressing for AI systems, which may not share the interpreter's concepts. However, it also provides leverage: a system's attitudes constrain its propositional structure, that structure constrains which attitudes can be attributed, and mechanistic interpretability can help us measure both.

Paper AI Chat

この論文のPDF全文を対象にAIに質問できます。

質問の例:

AIチャット機能を利用するには、ログインまたは会員登録(無料)が必要です。

会員登録 / ログイン

関連するAIDB記事