AIDB Daily Papers
MacAgentBench:macOSデスクトップでのAIエージェントのベンチマーク
※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。
ポイント
- macOSデスクトップ上でのAIエージェントの自動化能力を評価するための包括的なベンチマーク「MacAgentBench」を開発した。
- 既存のベンチマークでは捉えきれない、GUIとCLIの連携や複数アプリケーションに跨る長期間タスクの評価を可能にした点が重要である。
- Claude Opus 4.6とOpenClawの組み合わせが最高性能を示したが、その成功はフレームワーク設計よりもスキルライブラリに依存することが明らかになった。
Abstract
Computer use agents (CUAs) have advanced rapidly in desktop automation, and a growing number of users deploy CUAs such as OpenClaw on Mac Mini for always-on automation. However, existing benchmarks, including those for macOS, evaluate agents without framework augmentation and rely on binary evaluation. As a result, they fail to capture both the framework capabilities leveraged by modern CUAs and the partial progress on long-horizon, multi-application tasks. We present MacAgentBench, a comprehensive macOS agent benchmark comprising 676 tasks across 25 applications, with nearly 60% involving both GUI and CLI interaction. The benchmark adopts deterministic rule-based evaluation and introduces fine-grained multi-checkpoint scoring with capability annotations for multi-application tasks. Experiments across three frameworks and 16 models show that the best configuration, Claude Opus 4.6 on OpenClaw, attains 73.7% Pass@1, while this advantage is primarily driven by the skill library rather than by framework design. Fine-grained metrics further reveal that models with similar Pass@1 can differ substantially in sub-goal completion. Our code and data are publicly available at https://github.com/JetAstra/MacAgentBench.
Paper AI Chat
この論文のPDF全文を対象にAIに質問できます。
質問の例: