AIDB Daily Papers
大規模言語モデルはどの価値観を混同するのか:シュワルツの理論に基づく認識研究
※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。
ポイント
- 大規模言語モデルが具体的な状況において表現された価値観を正確に識別できるかを検証した。
- シュワルツの10の基本価値観に基づき、ロシア語の状況記述テキスト1,000件を用いて21のモデルを評価した。
- モデルは正しい価値観の領域を特定できる一方で、類似した価値観の間で不安定な順位付けを行い混同しやすいことが判明した。
Abstract
Large language models are increasingly evaluated through the values they endorse, but such evaluations presuppose that models can identify the value expressed in a concrete situation. We study this prerequisite as controlled top-1 recognition over Schwartz's ten basic values. Our evaluation set contains 1,000 Russian situational texts, balanced across the ten values and independently labeled by two human annotators per item. We evaluate 21 instruction-tuned LLM runs under a fixed ranked-response protocol; 20 runs with reliable outputs form the semantic panel. Pooled Acc@1 is 0.683 and Acc@3 is 0.892, showing that models often locate the correct motivational region while ranking close alternatives unstably. Adjacent values account for 50.9% of semantic errors, compared with 24.4% under a checkpoint-specific null. Eight directed confusions recur across checkpoints and human-confirmed subsets. Several are strongly asymmetric, including Universalism to Benevolence, Tradition to Conformity, and Security to Power, whereas Stimulation-Hedonism forms a bidirectional boundary. Their severity is checkpoint-specific and can bias higher-order value profiles. The results motivate value-recognition evaluation that combines exact accuracy, ranked recovery, and directed error analysis.
Paper AI Chat
この論文のPDF全文を対象にAIに質問できます。
質問の例: