AIDB Daily Papers
YOMI-Bench:大規模言語モデルの日本語漢字読みと音韻理解を評価するベンチマーク
※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。
ポイント
- 日本語の複雑な漢字の読みと音韻理解を評価するための新しいベンチマーク「YOMI-Bench」を提案した。
- 漢字の読みが文脈に依存し多岐にわたるという日本語特有の難しさを考慮し、モデルの性能を詳細に測定できる点が重要である。
- 評価の結果、日本語特化型モデルや商用モデルであっても、漢字の読みを考慮した生成タスクでは依然として性能が低いことが判明した。
Abstract
We propose YOMI-Bench, a benchmark for evaluating kanji reading and phonological understanding of large language models (LLMs) for Japanese. In Japanese, a single kanji character often has multiple possible readings, making it difficult to infer the correct reading from surface-level text alone. Due to these linguistic characteristics, it is empirically known that LLMs exhibit low performance in kanji reading for Japanese. The proposed YOMI-Bench consists of four tasks specifically designed to evaluate kanji reading performance in Japanese. In our evaluation using YOMI-Bench, we assessed one multilingual open LLM, four Japanese-specific open LLMs, and five commercial LLMs. As a result, we found that even Japanese-specific models show low performance, and that commercial models also perform poorly on generation tasks that require consideration of kanji readings.
Paper AI Chat
この論文のPDF全文を対象にAIに質問できます。
質問の例: