次回の更新記事:【論文著者監修・コメント】AIエージェントへの人間…(公開予定日:2026年07月27日)
AIDB Daily Papers

MeloDISinger:メロディと長さを維持した歌声編集のためのオーディオインフィリング手法

原題: MeloDISinger: Melody-Aware & Duration-Preserving Singing Voice Editing with Audio Infilling
著者: Yoonjeong Park, Jaekwon Im, Juhan Nam
公開日: 2026-06-29 | 分野: 音声 モデル 生成AI cs.SD eess.AS 音声合成

※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。

ポイント

  • 元のメロディや楽曲の長さを保持したまま、歌唱歌詞を部分的に修正できる歌声編集モデルを提案した。
  • 音素とメロディの対応を明示的に制御する機構を導入し、自然な歌声の編集を実現した点が新しい。
  • 実験の結果、客観的および主観的評価の両面で、既存手法を上回る最高水準の性能を達成した。

Abstract

Text-based singing voice editing (SVE) aims to revise sung lyrics while preserving the original melody, total duration, and non-edited regions. In this paper, we propose MeloDISinger, a flow-matching-based SVE model for melody-aware and duration-preserving editing. Its core module, MeloDRP, predicts fixed-budget duration ratios, enabling explicit span-wise duration control. For melody-aware duration allocation, MeloDRP fuses phonetic cues with pseudo-MIDI melodic context through cross-attention, while temporal-overlap supervision encourages soft phoneme--note correspondences. We further use a flow-matching mel decoder for audio infilling to synthesize edited regions while preserving surrounding context. In addition, we introduce a duration-aware edited-lyric generation pipeline using WhisperX and an LLM to construct feasible evaluation scenarios. Experiments demonstrate state-of-the-art performance in both objective and subjective evaluations.

Paper AI Chat

この論文のPDF全文を対象にAIに質問できます。

質問の例:

AIチャット機能を利用するには、ログインまたは会員登録(無料)が必要です。

会員登録 / ログイン

関連するAIDB記事