AIDB Daily Papers
Re-Sonance: ASR-LLM-TTSの3段階カスケードアーキテクチャに基づく構音障害者向け非同期リアルタイム音声変換システム
※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。
ポイント
- 音声認識、大規模言語モデル、音声合成を統合した構音障害者向けのリアルタイム音声変換システムを開発した。
- 高い遅延や不自然な音声パターンという既存の補完・代替コミュニケーションシステムの課題を解決するアプローチとして新規性がある。
- 軽度から中等度の構音障害を持つ話者において、意味的一貫性を保ちながら音声の明瞭度が大幅に向上したことが示された。
Abstract
Individuals with dysarthria face significant challenges in professional speaking scenarios such as conferences, presentations, and meetings, where real-time communication is crucial. While existing Augmentative and Alternative Communication (AAC) systems provide basic support, they often fail to meet the demands of professional speaking environments due to high latency and unnatural speech patterns. This paper presents Re-Sonance, a novel LLM-enhanced speech-driven AAC system designed for real-time professional speaking scenarios. By integrating Whisper ASR, Qwen LLM, and CosyVoice TTS, Re-Sonance achieves improved speech intelligibility and naturalness while maintaining real-time performance. Both subjective and objective evaluations using a Mandarin dysarthric speech dataset demonstrate that our speech reconstruction approach significantly improved intelligibility while preserving semantic coherence for speakers with mild to moderate dysarthria. Although performance remains limited for severe dysarthria cases, our findings validate the potential of LLM-based methods for enhancing speech-driven AAC systems, paving the way for more effective and accessible communication technologies.
Paper AI Chat
この論文のPDF全文を対象にAIに質問できます。
質問の例: