次回の更新記事:【論文著者監修・コメント】AIエージェントへの人間…(公開予定日:2026年07月27日)
AIDB Daily Papers

LLMの性格を操る:潜在特徴量への介入によるメカニスティックな性格分析

原題: Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions
著者: David Courtis, Ting Hu
公開日: 2026-06-27 | 分野: LLM 制御 性格 cs.AI

※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。

ポイント

  • LLMの潜在空間における性格特性に対応する特徴量を特定し、活性化値への直接的な介入で性格を制御する手法を提案した。
  • プロンプトや微調整に頼らず、モデル内部のメカニズムを直接操作することで、言語能力を維持しつつ性格を精密に調整できる点が新しい。
  • 特定の性格特性を強化する加算ベクトルを適用することで、タスク性能を損なわずに意図した性格を制御可能であることを実証した。

Abstract

Large Language Models (LLMs) have demonstrated the ability to simulate human-like OCEAN personality traits in generated text. Previous efforts have focused on prompt engineering or fine-tuning to shape LLM personality. In this work, we propose a mechanistic interpretability approach that directly intervenes on the model's latent features. Our method identifies latent directions in the residual stream corresponding to a target OCEAN trait using sparse autoencoders (SAEs) and contrastive activation analysis. We formalize an additive steering vector in activation space and demonstrate how applying a small additive shift to the hidden states enhances the target trait while preserving overall language modeling performance. To determine the optimal combination of feature shifts, we explore a linear weighting heuristic with grid search optimization that balances personality expression with task performance. Our approach shows promise in controllably steering personality traits at the mechanistic level while maintaining high performance on standard benchmarks.

Paper AI Chat

この論文のPDF全文を対象にAIに質問できます。

質問の例:

AIチャット機能を利用するには、ログインまたは会員登録(無料)が必要です。

会員登録 / ログイン

関連するAIDB記事