AIDB Daily Papers
確率モデルによるインコンテキスト学習の理論的解釈
※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。
ポイント
- 大規模言語モデルのインコンテキスト学習を確率モデルとして定式化し、その性能を理論的に導出した。
- 汎用的なパラメータ分布や指数型分布族を対象に、インコンテキスト学習の仕組みを厳密に解析した点が新しい。
- デモンストレーションの数やパラメータの感度、クエリとの類似性が学習性能に与える影響を明らかにした。
Abstract
In-context learning (ICL) is an emerging paradigm that employs the semantic information inherent in large language models (LLMs) for generating answers to user queries. While the remarkable performance of ICL has been widely known, a general modeling and a rigorous theoretical analysis of this paradigm are still lacking. This work presents a probabilistic model for ICL and derives the performance of ICL for both general parametric distributions and exponential families. Based on the derived results, the work explains the impact of multiple factors such as the number of demonstrations, the sensitivity of the probabilistic model to the variation of its parameters, as well as the similarity between the demonstrations and the query on the performance of ICL.
Paper AI Chat
この論文のPDF全文を対象にAIに質問できます。
質問の例: