AIDB Daily Papers
大規模言語モデルにおける生物性の所在:生物性概念の回路の追跡
※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。
ポイント
- 大規模言語モデルが生物と非生物の概念を区別するメカニズムを特定するため、制御されたデータセットを用いて回路探索を行った。
- 生物性の処理を担う因果メカニズムが存在することを実証し、モデル内に生物性回路を発見した点が新しい。
- 発見された回路は十分に局所化されておらず、モデルやタスク間で部分的にしか汎化しない分散的な性質を持つことが判明した。
Abstract
Distinguishing animate from inanimate concepts in written language requires more than shallow text processing, as it involves recognizing complex selectional constraints and contextual cues, such as verb-argument interactions. Yet, current large language models (LLMs) appear to be capable of doing it. We investigate whether this animacy-sensitive behavior of LLMs can be traced to a localized set of causally relevant components and connections. To do so, we construct a controlled dataset of minimal pairs and perform circuit discovery on four open-weight models. Through in-depth experiments and ablations, we show that a causal mechanism responsible for handling animacy in these models does exist, thus discovering an animacy circuit. At the same time, this circuit appears to be less localized compared to other known ones and generalizes only partially across models and animacy tasks, confirming the distributed, context-dependent, and somewhat graded nature of the animacy concept.
Paper AI Chat
この論文のPDF全文を対象にAIに質問できます。
質問の例: