AIDB Daily Papers
介入としてのLLM検出:戦略的ユーザー行動下における下流への影響
※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。
ポイント
- 不完全なLLM検出器がユーザーのワークフローにおけるインセンティブを歪める影響を理論モデルを用いて分析した。
- LLM検出の導入がかえって人間のLLM利用を増加させたり、出力品質を低下させたりするという直観に反する現象を示した点に新しさがある。
- 検出器の導入により検出対象の属性は一時的な増減パターンを示すが、下流の品質や利用行動には重大な歪みが生じることを発見した。
Abstract
As LLM adoption becomes more widespread, there is a growing interest in detecting LLM-generated content, for example through LLM detection tools and through heuristics based on language patterns. Detectors operate as an intervention that steers not only the detected attribute itself, but also downstream metrics such as LLM usage and output quality. In this work, we demonstrate how imperfect LLM detectors lead to counterintuitive impacts on these downstream metrics, by distorting how users are incentivized to use LLMs in their workflow. We develop a stylized model which captures how users strategically choose how much to use the LLM and how to post-process content to reduce the detected attribute. Using this model, we show that LLM detection can counterintuitively lead humans to increase their LLM usage. Moreover, even when reducing the detected attribute improves output quality, we find that introducing an LLM detector can lead users to produce lower quality outputs. In contrast, we show that detectors result in a clean "rise-then-fall" pattern for the detected attribute, which we empirically reproduce for word frequencies on arXiv abstracts. Altogether, our work illustrates how LLM detection can distort LLM usage and output quality, uncovering failure modes when LLM detectors operate as an intervention on these downstream metrics.
Paper AI Chat
この論文のPDF全文を対象にAIに質問できます。
質問の例: