AIDB Daily Papers
AIアシスタントは過剰に支援してしまう:問題解決におけるLLMの介入行動の評価
※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。
ポイント
- 学習中のLLMの介入タイミングや頻度を評価するためのシミュレーションベースのベンチマークであるInt-Benchを提案した。
- 人間と比較してLLMはより高頻度かつ早期に介入し、ヒントではなく完全な解答を提供する傾向があることを明らかにした。
- 現在のLLMアシスタントは、長期的な学習に必要な思考プロセスよりも短期的な成功を最適化していることが示された。
Abstract
Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems. While guidance from AI assistants can scaffold thinking and foster learning, such benefits depend on how they help--for instance, intervening too early or too frequently may hinder true learning and cognitive engagement. Yet how AI systems navigate intervention decisions during problem-solving remains poorly understood. Here, we introduce Int-Bench, a simulation-based benchmark for evaluating LLM interventions during learning. Int-Bench simulates a "student" solving a problem while a "teacher" monitors the student's reasoning and decides whether, when, and how to intervene. Across three domains--code debugging, mathematics, and brain teasers--we evaluate LLM teachers on the frequency and timing of interventions, as well as their impact on both immediate task success and generalization to new problems. We also compare LLMs to humans, finding that LLMs intervene more frequently and earlier than humans. Moreover, in contrast to humans, they tend to provide complete solutions rather than targeted hints. These findings suggest that current LLM assistants often optimize for short-term success rather than supporting the reasoning processes needed for deeper learning and long-term success.
Paper AI Chat
この論文のPDF全文を対象にAIに質問できます。
質問の例: