AIDB Daily Papers
LLMは理論的構成概念の測定器として妥当か?「グレイン・キャリブレーション」による検証手法の提案
※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。
ポイント
- LLMが人間と同等のコーディングを行っても、理論的根拠が不明確な「誤った理由による正解」のリスクを指摘した。
- 構成概念を節レベルの要素に分解し、理論に基づいたルールで統合する「グレイン・キャリブレーション」という手法を提案した。
- この手法により、LLMの推論プロセスを可視化し、出力結果が理論的定義に基づいているかを検証可能にした。
Abstract
When a large language model (LLM) codes a construct in text as a human annotator would, that agreement makes the LLM a reliable coder. Yet reliability leaves construct validity untouched. The instrument may be theory-naive, reaching the code through a correlate that meets none of the demands the construct's theory makes, and no current method tells that apart from genuine measurement. We propose grain calibration as a method that closes the gap. It decomposes a construct into clause-level components, tests each against the text with extractive evidence, and combines the results through an explicit, theory-derived rule. Because the rule is stated rather than lodged in one opaque pass, its structure is evidence about the process rather than the output. It shows which components settled a code, and, when the code is wrong, whether a component was missed or an adjacent construct mistaken for it. Validation shifts from scoring an instrument's outputs against an annotator to showing that the instrument runs on the construct its theory specifies.
Paper AI Chat
この論文のPDF全文を対象にAIに質問できます。
質問の例: