次回の更新記事:AIによる企業リスクを抑えるには、AIに対して性悪説…(公開予定日:2026年07月20日)
AIDB Daily Papers

LLMは理論的構成概念の測定器として妥当か?「グレイン・キャリブレーション」による検証手法の提案

原題: Correct codes for the wrong reasons? validating LLMs as measurement instruments for theoretical constructs
著者: Manuel Pita
公開日: 2026-06-26 | 分野: LLM NLP cs.CL cs.AI cs.CY キャリブレーション

※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。

ポイント

  • LLMが人間と同等のコーディングを行っても、理論的根拠が不明確な「誤った理由による正解」のリスクを指摘した。
  • 構成概念を節レベルの要素に分解し、理論に基づいたルールで統合する「グレイン・キャリブレーション」という手法を提案した。
  • この手法により、LLMの推論プロセスを可視化し、出力結果が理論的定義に基づいているかを検証可能にした。

Abstract

When a large language model (LLM) codes a construct in text as a human annotator would, that agreement makes the LLM a reliable coder. Yet reliability leaves construct validity untouched. The instrument may be theory-naive, reaching the code through a correlate that meets none of the demands the construct's theory makes, and no current method tells that apart from genuine measurement. We propose grain calibration as a method that closes the gap. It decomposes a construct into clause-level components, tests each against the text with extractive evidence, and combines the results through an explicit, theory-derived rule. Because the rule is stated rather than lodged in one opaque pass, its structure is evidence about the process rather than the output. It shows which components settled a code, and, when the code is wrong, whether a component was missed or an adjacent construct mistaken for it. Validation shifts from scoring an instrument's outputs against an annotator to showing that the instrument runs on the construct its theory specifies.

Paper AI Chat

この論文のPDF全文を対象にAIに質問できます。

質問の例:

AIチャット機能を利用するには、ログインまたは会員登録(無料)が必要です。

会員登録 / ログイン

関連するAIDB記事