AIDB Daily Papers
大規模言語モデルにおける日本語方言への頑健性評価:音声・テキスト両面から
※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。
ポイント
- 大規模言語モデル(LLM)と音声言語モデル(SLM)の日本語方言に対する頑健性を、方言と標準語の性能比で評価した。
- LLMのテキストベースでの方言理解度が、統合されたSLMの性能に影響することが明らかになった点が新しい。
- 実験の結果、SLMの方言頑健性は基盤となるLLMの頑健性と相関し、方言データでの学習や音声エンコーダのファインチューニングが頑健性向上に寄与した。
Abstract
Dialogue systems based on large language models (LLMs) have advanced significantly in recent years. However, dialectal variation remains a major challenge, particularly for systems that process spoken input. LLM-based speech language models (SLMs), which integrate LLMs with speech processing components, show promise for spoken language tasks, yet their ability to comprehend dialects has not been sufficiently studied. Moreover, it remains unclear how the dialectal understanding of the base LLM affects SLM performance. This study investigates the dialectal robustness of both LLMs and SLMs using Japanese dialects as a test case. We define robustness as the ratio of performance on dialectal versus standard inputs, enabling fair comparisons. Our experiments show that SLM robustness correlates with that of their text-based counterparts. Furthermore, training with dialectal data and fine-tuning the speech encoder each improves robustness in SLMs.
Paper AI Chat
この論文のPDF全文を対象にAIに質問できます。
質問の例: