AIDB Daily Papers
選好学習における嘘発見器を用いた監視のスケーリング傾向
※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。
ポイント
- LLMの欺瞞的な回答を効率的に特定するため、嘘発見器を用いて高コストな人間によるレビューを補完する手法SOLiDを検証した。
- モデルの規模拡大に伴い欺瞞の検出精度が向上し、ファインチューニング段階から人間を排除しても欺瞞が増加しないことを示した。
- 一方で、嘘発見器の学習データと選好学習データの分布が乖離すると、誤検知率が実用レベルを超えて上昇する課題が判明した。
Abstract
Deceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave, 2025), which uses lie detectors to identify responses for review by high-cost labelers. In this paper, we scale SOLiD to larger models and evaluate it in more diverse and realistic preference-learning settings. We find favorable scaling: undetected deception drops from 34% for 1B-parameter models to 14% for 405B-parameter models at a detector true positive rate of 99%, and expensive human labelers can be removed entirely from the fine-tuning phase without a statistically significant increase in deception. However, SOLiD is sensitive to distribution shift between detector training and preference-training data, which can drive detector false positive rates to impractical levels.
Paper AI Chat
この論文のPDF全文を対象にAIに質問できます。
質問の例: