AIDB Daily Papers
非英語言語における推論のコスト:日本語を対象としたケーススタディ
※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。
ポイント
- 日本語で推論を行うLLMを構築し、コーディングや数学などのベンチマークで性能を評価した。
- 推論言語の制御は可能だが、日本語特化の推論モデルは英語ベースの強力なモデルと同等の性能に留まった。
- 日本語での推論能力の獲得が、必ずしも日本文化に関連するタスクの精度向上には直結しないことが判明した。
Abstract
Reasoning Language Models (RLMs) achieve their strongest performance when they reason in English, the language for which reasoning-oriented training data is most abundant. However, reasoning trace is a clue for model interpretability and safety, and useful in practice for both the model users and for model developers. Thus, it is desirable to be able to develop a model that reasons in a language of the user's choice, while still maintaining strong reasoning performance. To this end, we study the feasibility of training a model that reasons in Japanese. We develop a Japanese-reasoning variant of Qwen-3-Swallow-8B, which is a Japanese LLM continually pretrained from Qwen-3-8B, with GRPO and evaluate it across coding, math, and science benchmarks. The study shows that reasoning-language control is feasible by training a Japanese continually pretrained model with GRPO. However, its performance is at best on par with strong English-reasoning baselines on several benchmarks. We also evaluate the trained model on Japanese cultural benchmarks and observe that the model's performance is worse than the baseline models, suggesting that the reasoning in Japanese does not immediately improve performance on culturally relevant tasks for free.
Paper AI Chat
この論文のPDF全文を対象にAIに質問できます。
質問の例: