AIDB Daily Papers
Who&When Pro:LLMはAIエージェントの失敗原因を本当に特定できるか?
※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。
ポイント
- AIエージェントの失敗箇所と原因を自動特定するための大規模ベンチマークWho&When Proを構築した。
- 成功した手順を再現した後に意図的に失敗を注入する手法で、12,326件の高品質なラベル付きデータセットを作成した。
- モデルによる失敗原因の特定には体系的なパターンが存在することを明らかにし、今後の自動診断システムへの指針を示した。
Abstract
Automated failure attribution uses LLMs to identify where and why agentic systems fail. As agents become more capable, their failures become subtler, making automated attribution increasingly important. We introduce Who&When Pro, a large-scale benchmark for automated failure attribution in agentic systems. Using a strictly controlled pipeline that injects a failure only after exactly replaying a successful prefix, we construct 12,326 failed trajectories with golden labels across 3 modalities and 26 benchmarks covering various scenarios. Beyond benchmarking, we conduct extensive experiments and analyses, revealing systematic patterns in how models attribute failures across modalities, protocols, and model families, and providing empirical guidance for future automated failure attribution systems.
Paper AI Chat
この論文のPDF全文を対象にAIに質問できます。
質問の例: