AIDB Daily Papers
LLMは生成よりも評価が上手いのか?タスクの非対称性と自己評価の限界を検証
※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。
ポイント
- LLMが自ら生成した回答を評価する際の性能を、4つのQAベンチマークを用いて検証した。
- 生成と評価の難易度は一律ではなく、多くのタスクで生成精度が自己評価精度を上回ることが判明した。
- 注意機構の分析により、評価時にモデルが文脈や回答を十分に参照していないことが原因であると明らかになった。
Abstract
LLM-as-a-Judge and self-evaluation pipelines implicitly assume that evaluation is easier than generation. We test this in a controlled in-context QA setting where a context passage is the sole information source and each model judges the answer it generated, removing the parametric-knowledge confound of open-domain comparisons. Across four benchmarks (SQuAD 2.0, DROP, HotpotQA, MuSiQue) and two models, evaluation is not uniformly easier: generation accuracy exceeds self-evaluation on three of four, with multi-hop MuSiQue the exception. Attention analysis reveals why: evaluation attends to context 3--5x less than generation does and barely reads the candidate answer. LoRA fine-tuning confirms the asymmetry is not a training artifact: generation fine-tuning induces over-acceptance and evaluation fine-tuning degrades generation. These findings challenge core assumptions in self-evaluation pipelines.
Paper AI Chat
この論文のPDF全文を対象にAIに質問できます。
質問の例: