AIDB Daily Papers
MeloDISinger:メロディと長さを維持した歌声編集のためのオーディオインフィリング手法
※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。
ポイント
- 元のメロディや楽曲の長さを保持したまま、歌唱歌詞を部分的に修正できる歌声編集モデルを提案した。
- 音素とメロディの対応を明示的に制御する機構を導入し、自然な歌声の編集を実現した点が新しい。
- 実験の結果、客観的および主観的評価の両面で、既存手法を上回る最高水準の性能を達成した。
Abstract
Text-based singing voice editing (SVE) aims to revise sung lyrics while preserving the original melody, total duration, and non-edited regions. In this paper, we propose MeloDISinger, a flow-matching-based SVE model for melody-aware and duration-preserving editing. Its core module, MeloDRP, predicts fixed-budget duration ratios, enabling explicit span-wise duration control. For melody-aware duration allocation, MeloDRP fuses phonetic cues with pseudo-MIDI melodic context through cross-attention, while temporal-overlap supervision encourages soft phoneme--note correspondences. We further use a flow-matching mel decoder for audio infilling to synthesize edited regions while preserving surrounding context. In addition, we introduce a duration-aware edited-lyric generation pipeline using WhisperX and an LLM to construct feasible evaluation scenarios. Experiments demonstrate state-of-the-art performance in both objective and subjective evaluations.
Paper AI Chat
この論文のPDF全文を対象にAIに質問できます。
質問の例: