次回の更新記事:社内で量産されるAIエージェントが使い物にならなくなってしまう9つのパターン(公開予定日:2026年08月04日)

【8月31日まで】個人情報マスキングツール「PII Mask」を全会員に開放中 / ローカルで動作、文書は外部に送信されません

ダウンロード →
AIDB Daily Papers

表現から行動へ:大規模言語モデルにおける個人・状況・行動の三つ組みの探索

原題: From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs
著者: Ruikang Zhang, Shuo Wang, Qi Su
公開日: 2026-07-29 | 分野: LLM NLP AI 大規模言語モデル cs.CL cs.AI

※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。

ポイント

  • 大規模言語モデル内部におけるパーソナリティ関連の表現、状況適応、行動の結びつきを分析するフレームワークを提案した。
  • 人間心理学の三つ組みフレームワークを適用し、モデル内のスパースな特徴量から特性を同定して検証する手法が新しい。
  • 特徴量の介入によって異なる状況下でも一貫した特性表現や人間と同様のトレードオフを持つ行動変化が引き起こされた。

Abstract

Human personality theories characterize traits not as isolated attributes captured by a single score, but as stable individual tendencies expressed through the interplay among persons, situations, and behaviors. Existing studies of personality-related behavior in LLMs have primarily focused on outputs elicited under personality conditioning, characterizing observable trait-related expressions while lacking mechanistic evidence for the existence of internal personality-related representations, their cross-situational expression, and how these representations shape specific behaviors. Building on Funder's personality triad framework, we adapt its three components for LLM analysis: Person as personality-related internal representations, Situation as contexts that afford trait-relevant responses, and Behavior as response patterns on broader social tasks. We introduce a framework for discovering, controlling, and validating trait-like representations in LLMs. First, using contrastive behavior pairs grounded in shared situations, we identify sparse internal features associated with opposing poles of personality traits through SAE decomposition. We validate their trait relevance through effects on behavior to situation, token-level activation patterns, and robustness to paraphrasing. Second, feature-level interventions induce bidirectional trait-related shifts across a separate, diverse set of situations while preserving response validity, demonstrating consistent expression across contexts. Third, applying the same interventions to social intelligence tasks reveals behavioral changes with benefit-tradeoff patterns consistent with findings from human personality research, providing behavioral-level validation beyond personality scores. Our findings provide evidence that LLMs contain controllable trait-like representations linking internal states, situational expression, and behavioral outcomes.

Paper AI Chat

この論文のPDF全文を対象にAIに質問できます。

質問の例:

AIチャット機能を利用するには、ログインまたは会員登録(無料)が必要です。

会員登録 / ログイン

関連するAIDB記事