AIDB Daily Papers
ターミナル環境における自律型コードレビュー:行動・コスト・人間とのアラインメントに関する軌跡レベルの分析
※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。
ポイント
- ターミナル環境における自律型コードレビューエージェントの動的な行動やコストを軌跡レベルで詳細に分析した。
- 従来の性能中心の評価では捉えきれなかった、リポジトリに根ざしたエージェントの成功・失敗要因や運用コストを明らかにした点が新しい。
- 高いレビュー精度を誇る一方で多大な検証オーバーヘッドが発生し、成功するレビューには強力な計画性が伴うことが判明した。
Abstract
Agentic code review in terminal-based environments enables early feedback during local development before pull request creation. However, existing evaluations remain performance-centric and fail to capture the dynamic behaviors of repository-grounded agentic reviewers. Understanding these behaviors is critical for identifying how agentic reviewers succeed, fail, and incur hidden operational costs in practice. Then, we analyze the reviewers' behavior based on their trajectories. Our results show that agentic reviewers achieve higher review precision but incur substantial exploration and validation overhead, while successful reviews are associated with stronger planning and less downstream validation. These findings highlight the potential benefits of trajectory-aware and cost-sensitive evaluation of future agentic code review systems.
Paper AI Chat
この論文のPDF全文を対象にAIに質問できます。
質問の例: