AIDB Daily Papers
アインシュタイン・ワールドモデル:言語を超えた視覚的思考実験でLLMの推論能力を拡張する
※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。
ポイント
- 本研究では、LLMが視覚的・時間的なロールアウトを推論プロセスに組み込む「アインシュタイン・ワールドモデル」を提案する。
- これは、言語だけでは捉えきれない複雑な思考を、視覚的な思考実験によって補完し、LLMの推論能力を向上させる点で重要である。
- 提案手法により、LLMはテキストだけでは困難な視覚的思考実験を可能にし、ツール呼び出し能力を拡張する結果となった。
Abstract
Does intelligence require the ability to reason about phenomena beyond direct experience? It is natural to suspect that some complex thought cannot be captured through language alone. However, of particular concern to this work, is whether visualising counterfactual events can complement language as a mechanism for complex thought. We ask whether LLMs can be trained to utilise such visualisation mechanisms, in a way that benefits their reasoning abilities. Motivated by this question, we propose Einstein World Models. EWMs are a blueprint for LLM-based reasoning systems that place visual-temporal rollouts inside the reasoning trace, allowing them to reason in ways that text alone may not support well. In an EWM, the LLM calls a world-module (not to be confused with a world model), to produce short rollouts of scenes under consideration. The returned rollout is treated not as the answer, but as an inspectable hypothesis that can support later reasoning. Einstein World Models extend the capability of LLMs for tool calling (such as web search or code execution), into the domain of visual thought experiments.
Paper AI Chat
この論文のPDF全文を対象にAIに質問できます。
質問の例: