Research project · Singing voice synthesis

VocalRender

Score-native singing voice synthesis for real-world composition.

VocalRender directly transforms composer-oriented symbolic scores—lyrics, MIDI pitches, note values, and tempo—into expressive singing audio without requiring phoneme-level durations or time-aligned acoustic guidance.

01

Score-native representation

Interleaved lyrics and pitch–duration pairs preserve note-to-word alignment and melisma.

02

Continuous acoustic latents

AudioVAE retains fine pitch, timbre, articulation, and local acoustic detail.

03

Autoregressive diffusion

Global prosody modeling and local reconstruction produce expressive, high-fidelity singing.

Audio comparisons

Listen side by side.

Each human-reviewed sample pairs one symbolic score with ground truth and six system outputs. Headphones recommended.

10human-reviewed
score excerpts
Ours Reference Baseline
Expanded music score