次回の更新記事:【論文著者監修・コメント】AIエージェントへの人間…(公開予定日:2026年07月27日)
AIDB Daily Papers

LLMにおける権威バイアスのメカニズム:なぜモデルは権威に迎合するのか

原題: A Mechanistic View of Authority Hierarchy in LLM Sycophancy
著者: Emil Joswin, Srujananjali Medicherla, Priyanka Mary Mammen
公開日: 2026-07-01 | 分野: LLM ハルシネーション 大規模言語モデル cs.CL cs.LG AI安全性

※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。

ポイント

  • 医療QA設定を用い、権威ある人物の示唆がモデルの回答に与える影響を検証した。
  • モデルが権威の高さに応じて回答を歪める現象が、学習過程で自然発生することを示した。
  • 権威の信号がモデル内部の正解表現を特定の層で上書き消去するメカニズムを解明した。

Abstract

Authority bias poses a critical safety concern in language models: models systematically prioritize social cues from authority figures over factual consistency, swaying their answers based on source credibility rather than evidence. We mechanistically investigate this phenomenon using a controlled medical QA setting, where hints suggesting incorrect answers are attributed to personas of varying expertise. Across Llama-3.1-8B, Qwen3-8B, and Gemma-2-9B, we find that models respond in a graded manner proportional to perceived authority, a hierarchy that is never explicitly prompted but emerges from training. Logit lens analysis and linear/non-linear probing localize this effect to a critical late layer where correct answer representations are actively erased, an erasure that scales with authority level, resists mean vector intervention, and is only partially reversible through chain-of-thought reasoning. Our findings suggest that authority-induced sycophancy is not a surface-level output bias but mechanistic knowledge erasure, a precise, layer-localized overwriting of correct internal representations by high-status authority signals.

Paper AI Chat

この論文のPDF全文を対象にAIに質問できます。

質問の例:

AIチャット機能を利用するには、ログインまたは会員登録(無料)が必要です。

会員登録 / ログイン

関連するAIDB記事