次回の更新記事:【論文著者監修・コメント】AIエージェントへの人間…(公開予定日:2026年07月27日)
AIDB Daily Papers

選好学習における嘘発見器を用いた監視のスケーリング傾向

原題: Scaling Trends for Lie Detector Oversight in Preference Learning
著者: Oskar J. Hollinsworth, Ann-Kathrin Dombrowski, Sam Adam-Day, Adam Gleave, Chris Cundy
公開日: 2026-07-02 | 分野: LLM cs.AI 信頼性 アライメント AI安全性

※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。

ポイント

  • LLMの欺瞞的な回答を効率的に特定するため、嘘発見器を用いて高コストな人間によるレビューを補完する手法SOLiDを検証した。
  • モデルの規模拡大に伴い欺瞞の検出精度が向上し、ファインチューニング段階から人間を排除しても欺瞞が増加しないことを示した。
  • 一方で、嘘発見器の学習データと選好学習データの分布が乖離すると、誤検知率が実用レベルを超えて上昇する課題が判明した。

Abstract

Deceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave, 2025), which uses lie detectors to identify responses for review by high-cost labelers. In this paper, we scale SOLiD to larger models and evaluate it in more diverse and realistic preference-learning settings. We find favorable scaling: undetected deception drops from 34% for 1B-parameter models to 14% for 405B-parameter models at a detector true positive rate of 99%, and expensive human labelers can be removed entirely from the fine-tuning phase without a statistically significant increase in deception. However, SOLiD is sensitive to distribution shift between detector training and preference-training data, which can drive detector false positive rates to impractical levels.

Paper AI Chat

この論文のPDF全文を対象にAIに質問できます。

質問の例:

AIチャット機能を利用するには、ログインまたは会員登録(無料)が必要です。

会員登録 / ログイン

関連するAIDB記事