次回の更新記事:【論文著者監修・コメント】AIエージェントへの人間…(公開予定日:2026年07月27日)
AIDB Daily Papers

言語とシンボル表現の切り替えによる空間推論の強化

原題: Spatial Reasoning via Modality Switching Between Language and Symbolic Representation
著者: Shreya Rajpal, Tanawan Premsri, Parisa Kordjamshidi
公開日: 2026-06-30 | 分野: LLM マルチモーダル 推論 空間 大規模言語モデル cs.AI

※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。

ポイント

  • 複雑な空間的物語を自然言語のみで処理せず、グリッドなどの幾何学的な構造へ変換して推論を行う手法を提案した。
  • 信頼性と複雑性に基づく指標を用いて、言語推論と構造化表現のどちらが適しているかをモデルが自律的に判断する点が新しい。
  • 空間的な情報を構造化表現に切り替えることで、大規模言語モデルの推論性能が最大42%向上することを実証した。

Abstract

Human reasoning is inherently multimodal: when problems become difficult, we rarely think in words alone. We often externalize our reasoning by sketching diagrams or drawing grids to understand the underlying conceptual structure and avoid mistakes. Building on this premise, our research investigates: (a) whether grounding multi-hop textual-spatial stories into geometry-aware modalities, such as layouts or grids, improves reasoning compared to natural language-based inference; and (b) whether a model can decide when to rely on natural language reasoning and when to switch to a structured modality. We address these questions by introducing a switching metric based on trustworthiness and complexity signals, which estimates when grounding a spatial story into structure is likely to improve performance. This takes a first step toward principled modality selection in Large Language Model (LLM) reasoning. Across our settings, switching from natural language-based reasoning to a grid-based representation improves LLM performance by up to 42%, highlighting the importance of modality choice in shaping reasoning outcomes.

Paper AI Chat

この論文のPDF全文を対象にAIに質問できます。

質問の例:

AIチャット機能を利用するには、ログインまたは会員登録(無料)が必要です。

会員登録 / ログイン

関連するAIDB記事