次回の更新記事:【論文著者監修・コメント】AIエージェントへの人間…(公開予定日:2026年07月27日)
AIDB Daily Papers

YOMI-Bench:大規模言語モデルの日本語漢字読みと音韻理解を評価するベンチマーク

原題: YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese
著者: Ryota Mibayashi, Hiroya Takamura, Hitomi Yanaka
公開日: 2026-07-01 | 分野: LLM NLP ベンチマーク 日本語 cs.CL 言語モデル

※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。

ポイント

  • 日本語の複雑な漢字の読みと音韻理解を評価するための新しいベンチマーク「YOMI-Bench」を提案した。
  • 漢字の読みが文脈に依存し多岐にわたるという日本語特有の難しさを考慮し、モデルの性能を詳細に測定できる点が重要である。
  • 評価の結果、日本語特化型モデルや商用モデルであっても、漢字の読みを考慮した生成タスクでは依然として性能が低いことが判明した。

Abstract

We propose YOMI-Bench, a benchmark for evaluating kanji reading and phonological understanding of large language models (LLMs) for Japanese. In Japanese, a single kanji character often has multiple possible readings, making it difficult to infer the correct reading from surface-level text alone. Due to these linguistic characteristics, it is empirically known that LLMs exhibit low performance in kanji reading for Japanese. The proposed YOMI-Bench consists of four tasks specifically designed to evaluate kanji reading performance in Japanese. In our evaluation using YOMI-Bench, we assessed one multilingual open LLM, four Japanese-specific open LLMs, and five commercial LLMs. As a result, we found that even Japanese-specific models show low performance, and that commercial models also perform poorly on generation tasks that require consideration of kanji readings.

Paper AI Chat

この論文のPDF全文を対象にAIに質問できます。

質問の例:

AIチャット機能を利用するには、ログインまたは会員登録(無料)が必要です。

会員登録 / ログイン

関連するAIDB記事