AIDB Daily Papers
再帰的ハーネス自己改善によるAIエージェントの性能向上
※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。
ポイント
- プロンプトベースの仕様としてエージェントのループを表現し、自己の履歴に基づくフィードバックで繰り返し改善する再帰的ハーネス自己改善(RHI)を提案した。
- モデルとハーネスの共進化パラダイムにおいて、軽量かつ少数の更新イテレーションでタスク特化型のハーネス最適化を実現できる点が新しい。
- わずか数回の反復で低推論コストモデルの性能上限を大幅に引き上げ、最大推論努力の設定を超える性能を発揮しつつ推論コストを最大60%削減した。
Abstract
Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This motivates harness-in-the-loop learning: optimizing harnesses for both immediate agent performance and the quality of traces used for future model training. However, continually updating provider-built scaffolds is costly and labor-intensive. We therefore investigate whether optimizing user-constructed harnesses in a task-specific manner can improve execution-trace quality while remaining computationally lightweight and requiring only a few update iterations. To this end, we introduce Recursive Harness Self-Improvement (RHI), which represents the harness as a prompt-level specification of the agent loop and iteratively refines it using pairwise feedback over its own revision history. Across 30 synthetic machine-learning research tasks spanning quantitative finance, robotics, and pharmacy, a few RHI iterations suffice to substantially raise the performance ceiling of low-reasoning-effort agents, exceeding the corresponding maximum-reasoning-effort setting while reducing inference cost by up to 60%. We show that these gains arise primarily from improved task-specific context management through more effective inter-agent information flow rather than longer reasoning traces. Finally, we formalize this behavior as an information-theoretic hypothesis for RHI's implicit optimization objective, suggesting RHI as a practical algorithm for continual learning within the paradigm of model--harness co-evolution.
Paper AI Chat
この論文のPDF全文を対象にAIに質問できます。
質問の例: