Score-native representation
Interleaved lyrics and pitch–duration pairs preserve note-to-word alignment and melisma.
Score-native singing voice synthesis for real-world composition.
VocalRender directly transforms composer-oriented symbolic scores—lyrics, MIDI pitches, note values, and tempo—into expressive singing audio without requiring phoneme-level durations or time-aligned acoustic guidance.
Interleaved lyrics and pitch–duration pairs preserve note-to-word alignment and melisma.
AudioVAE retains fine pitch, timbre, articulation, and local acoustic detail.
Global prosody modeling and local reconstruction produce expressive, high-fidelity singing.
Each human-reviewed sample pairs one symbolic score with ground truth and six system outputs. Headphones recommended.