YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Summary (Overview)
-
Unified Architecture: YuE2 introduces a single AR–NAR Mixture-of-Transformers (MoT) model that first generates a readable symbolic score (melody, harmony, rhythm, form), then expands it into semantic music tokens, and finally realizes it as full-song audio—unifying symbolic and audio music generation in one checkpoint.
-
Symbolic Planning Improves Quality: Expert listening tests show that symbolic planning improves perceived song quality, with 49.3% of overall preferences favoring planning versus 34.6% without planning, demonstrating that explicit composition before audio generation enhances musicality.
-
Frontier Performance: On WildSongBench, YuE2 scores 6.73 on SongBench Global Avg, exceeding all evaluated public baselines; best-of-8 selection reaches 6.96, the highest observed mean among all evaluated systems including proprietary ones.
-
New Supervision Methods: The paper introduces MERT2 (music representation learning) and SheetSage2 (lead-sheet transcription) to construct symbolic and semantic supervision from recordings, with MERT2 setting state-of-the-art results on 14 of 15 MARBLE metrics.
-
Versatile Score Interface: The same checkpoint supports creation, controlled editing, zero-shot cover generation (surpassing dedicated cover systems), and agentic music editing through external language models.
Introduction and Theoretical Foundation
Progressive Musical Commitment
The paper frames music generation through a hierarchy of representations organized by the decisions they settle:
- Text and lyrics state intent
- Score commits to melody, harmony, rhythm, and form
- Performance and audio representations resolve timing, articulation, timbre, expression, and production
This framework, called progressive musical commitment, suggests each stage settles additional decisions while leaving remaining choices to lower stages. The space between a score and its sound is room for interpretation.
The Gap in Existing Systems
The authors identify that existing generators expose only one side of this hierarchy:
- Symbolic models make composition readable but stop before finished recordings
- Audio models produce complete songs but leave composition implicit
Core Research Question
Can a single model traverse this hierarchy, making composition readable and editable while improving the quality of the finished song?
Key Contributions
- YuE2: Unifies symbolic and audio generation through symbolic planning, reaching frontier full-song quality with readable scores enabling editing, covers, and agentic editing
- Symbolic planning validation: Shows that composing a score first improves perceived song quality
- MERT2 and SheetSage2: Supply semantic and symbolic supervision from recordings
Methodology
Generation Framework
The model factorizes generation as:
where:
- = symbolic composition (ABC notation)
- = semantic tokens (25 Hz)
- = continuous acoustic latents (25 Hz, 64-dimensional)
- = text conditions and lyrics
AR–NAR Mixture-of-Transformers
The model uses a 28-layer backbone with separate AR and NAR normalization, projections, and MLP experts per layer, sharing one attention computation. The hybrid attention mask is:
| query\key | A | N |
|---|---|---|
| A | causal | blocked |
| N | full | bidirectional |
- AR stream: Predicts symbolic and semantic tokens causally
- NAR stream: Predicts flow velocity over acoustic latents with bidirectional latent attention
Training Objectives
Autoregressive loss (next-token cross-entropy):
Flow matching loss (conditional flow matching):
Joint objective:
Training Data and Tasks
- ~346,000 hours of music training data
- Songs packed into 24,576-position contexts without splitting
- Four training tasks: with/without score, with/without semantic tokens
- 3.58B parameters total
- Text and lyric conditioning can be dropped for classifier-free guidance
MERT2: Multi-View Target Synthesis
MERT2 uses a four-stage curriculum:
- Foundation Pretraining: Predict four code streams from masked audio using a ConvNeXt–Conformer encoder
- Causal Adaptation: Switch to causal attention for tokenizer branch
- Supervised Fine-Tuning: Add lyric CTC, mel, and chroma reconstruction objectives
- Semantic Quantization: Insert a 32,768-entry clustered vector quantizer between layers 13 and 14
The multi-view target synthesis combines frozen MuQ and Qwen2-Audio-Instruct encoders through a shared bottleneck with four-level residual vector quantization.
SheetSage2: Full-Song Transcription
- SheetSage2-Prober: Non-autoregressive model trained on human-annotated datasets and MIDI-rendered audio with task-specific CRF-based structured decoding
- SheetSage2-AR: Six-layer autoregressive RoFormer decoder with full-context MERT2-FS encoder, generating chronological event sequences converted to ABC notation
Empirical Validation / Results
Frontier Full-Song Quality (WildSongBench)
Table 1: Comparison with proprietary systems on WSB (192 prompts)
| Model | SongBench Mus.↑ | SongBench Avg↑ | SongEval Mus.↑ | SongEval Avg↑ | AudioBox PQ↑ | MuLan↑ | AMCaps↑ | Q3O↑ | PER↓ |
|---|---|---|---|---|---|---|---|---|---|
| Suno v5 | 5.9918 | 6.8721 | 4.3051 | 4.3579 | 8.1698 | 0.5428 | 0.4353 | 4.5907 | 0.0810 |
| Suno v4.5 | 5.8317 | 6.6995 | 4.3198 | 4.3666 | 8.2541 | 0.5022 | 0.3873 | 4.4149 | 0.0580 |
| Mureka 9 | 6.0488 | 6.9377 | 4.3619 | 4.4111 | 8.0226 | 0.4394 | 0.4102 | 4.6368 | 0.1169 |
| YuE2 | 5.9075 | 6.7316 | 4.2151 | 4.2625 | 8.2598 | 0.5068 | 0.4054 | 4.6819 | 0.0844 |
| YuE2 (Bo8) | 6.2666 | 6.9632 | 4.2506 | 4.2960 | 8.2714 | 0.5051 | 0.3980 | 4.7009 | 0.0979 |
Table 2: Comparison with publicly released systems on WSB
| Model | SongBench Mus.↑ | SongBench Avg↑ | SongEval Mus.↑ | SongEval Avg↑ | AudioBox PQ↑ | MuLan↑ | AMCaps↑ | Q3O↑ | PER↓ |
|---|---|---|---|---|---|---|---|---|---|
| YuE1 | 4.0847 | 4.9165 | 3.1524 | 3.2150 | 7.8683 | 0.2623 | 0.2882 | 3.7301 | 0.3638 |
| LeVo 2 | 5.4590 | 6.3247 | 3.9819 | 4.0234 | 8.3966 | 0.3542 | 0.2680 | 3.9458 | 0.2612 |
| HeartMuLa† | 5.4963 | 6.2483 | 4.5329 | 4.5519 | 8.2933 | 0.3823 | 0.2786 | 3.4907 | 0.1071 |
| YuE2 | 5.9075 | 6.7316 | 4.2151 | 4.2625 | 8.2598 | 0.5068 | 0.4054 | 4.6819 | 0.0844 |
| YuE2 (Bo8) | 6.2666 | 6.9632 | 4.2506 | 4.2960 | 8.2714 | 0.5051 | 0.3980 | 4.7009 | 0.0979 |
Expert Listening Results
- YuE2 (best-of-8) vs Suno v4.5: 57.3% favor YuE2 vs 30.5% for Suno v4.5
- YuE2 (best-of-8) vs Suno v5: Nearly balanced at 40.4% vs 39.9%
- Audio quality: Average tie-adjusted preference of 58.9% across six proprietary baselines (CR1 95% CI 55.6–62.2%)
Symbolic Planning Ablation
- Overall quality: 49.3% favor planning vs 34.6% without (p = 0.0070)
- Musicality: 45.0% vs 29.4% (p = 0.0141)
- Melody: 44.0% vs 29.5%
- Chord progression: 39.0% vs 21.0%
Unified MoT vs Separate LM+DiT
- Overall quality: 53.4% favor unified MoT vs 35.6% (p = 0.0084)
- Audio quality: 48.5% vs 29.6% (p = 3.06 × 10⁻⁴)
Score–Audio Agreement
Table 3: Score–audio consistency on WSB
| Audio | Melody Seq.↑ | Chords Seq.↑ | Rhythm F1↑ | Key Exact↑ | Key Wtd.↑ | Tempo Log err.↓ | Tempo Acc.↑ | Form Content↑ | Form Bound.↑ |
|---|---|---|---|---|---|---|---|---|---|
| Corresponding | 0.9464 | 0.9246 | 0.9396 | 0.9307 | 0.9515 | 0.0142 | 0.9869 | 0.7799 | 0.8724 |
| Mismatched | 0.2160 | 0.2473 | 0.6119 | 0.3004 | 0.3629 | 0.1131 | 0.6806 | 0.1547 | 0.1454 |
| Without score | 0.2055 | 0.1897 | 0.5875 | 0.2234 | 0.2831 | 0.1468 | 0.5366 | 0.1236 | 0.1245 |
Score Editing Results
Table 4: Score editing on WildSongBench (0–100 scale)
(a) Edit adherence
| Edit | Metric | Score↑ |
|---|---|---|
| Melody | Pitch accuracy | 84.17 |
| Harmony | Chord agreement | 79.54 |
| Rhythm | Relative onset accuracy | 73.43 |
| Key | Weighted key score | 90.58 |
| Tempo | Acc2 | 95.68 |
(b) Content preservation
| Edit | Melody↑ | Harmony↑ |
|---|---|---|
| Melody | 93.10 | 94.05 |
| Harmony | 93.43 | 90.16 |
| Rhythm | 94.34 | 93.08 |
Zero-Shot Cover Generation
Table 5: YuE2 with full scores leads both evaluated cover systems on all eight retrieval measures
| Method | CLEWS mAP↑ | CLEWS MRR↑ | CLEWS Hit@1↑ | CLEWS Hit@5↑ | Discogs mAP↑ | Discogs MRR↑ | Discogs Hit@1↑ | Discogs Hit@5↑ | MuLan↑ | Q3O↑ | PQ↑ | Mus.↑ |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SongEcho | 0.419 | 0.536 | 48.4 | 58.8 | 0.122 | 0.227 | 16.6 | 28.4 | 0.366 | 4.474 | 6.862 | 3.286 |
| ACE-Step 1.5 | 0.024 | 0.036 | 2.4 | 4.1 | 0.006 | 0.014 | 0.6 | 1.4 | 0.166 | 4.190 | 6.918 | 3.689 |
| YuE2 (full score) | 0.647 | 0.748 | 71.3 | 78.6 | 0.288 | 0.438 | 37.5 | 50.3 | 0.382 | 4.273 | 8.044 | 5.104 |
| Without chords | 0.598 | 0.715 | 67.3 | 76.1 | 0.179 | 0.304 | 23.6 | 36.6 | 0.417 | 4.482 | 8.117 | 5.490 |
| Without score | 0.006 | 0.008 | 0.3 | 0.9 | 0.004 | 0.008 | 0.2 | 0.8 | 0.474 | 4.837 | 8.186 | 5.691 |
MERT2 Music Understanding Results
Table 6: MERT2-30s leads on 14 of 15 MARBLE metrics
Key results (MERT2-30s vs best prior baseline):
- MTT Tagging ROC: 91.91 vs 91.7 (PupuJEPA-Large)
- GTZAN Genre: 91.72 vs 86.9
- EmoMusic : 63.23 vs 62.5
- EmoMusic : 80.01 vs 78.5
- MTG-Jamendo Instrument ROC: 80.27 vs 78.4
Semantic Tokenizer Evaluation (375 bits/s)
Table 7: MERT2 tokenizer leads five of six metrics among quantized representations
| Representation | Rate (Hz) | Dim. | kb/s | Genre Acc.↑ | Emotion R²↑ | Key Score↑ | MTT AUROC↑ | MTT AP↑ | Beat F1↑ |
|---|---|---|---|---|---|---|---|---|---|
| MERT2 pre-VQ | 25 | 1024 | - | 84.14 | 66.19 | 56.94 | 91.60 | 40.02 | 89.81 |
| MERT2 post-VQ | 25 | 1024 | 0.375 | 65.86 | 54.49 | 52.20 | 90.06 | 35.65 | 86.38 |
| LeVo 2 post-VQ | 25 | 1024 | 0.350 | 44.83 | 30.17 | 61.23 | 86.72 | 30.79 | 79.72 |
| EnCodec post-VQ | 75 | 128 | 1.500 | 31.72 | 26.51 | 15.76 | 80.84 | 22.32 | 76.50 |
SheetSage2-AR Transcription Results
Table 8: SheetSage2-AR achieves highest scores on 12 of 15 benchmark–metric pairs
Notable improvements:
- RWC-Pop vocal melody F1: 82.51 vs 62.71 (SheetSage1), +19.80 points
- Rock Corpus vocal melody F1: 67.08 vs 49.19, +17.89 points
- JAAH chord maj/min: 64.50% vs 59.45 (ChordFormer), +5.06 points
Theoretical and Practical Implications
Theoretical Significance
-
Progressive Musical Commitment validated: The paper provides empirical evidence that generating through hierarchical representations (score → semantic tokens → audio) improves perceived quality, supporting the theoretical framework that explicit composition aids generation.
-
Unified vs. separate architectures: The MoT architecture demonstrates that jointly learning semantic-token prediction and acoustic generation outperforms the widely used LM+DiT design, suggesting benefits from shared representations and attention.
-
Multi-View Target Synthesis: The Platonic Representation Hypothesis motivates the approach of constraining discrete targets by multiple encoder views, showing that complementary representations can converge toward shared musical structure.
Practical Implications
-
Editable music generation: The readable score interface enables:
- Local edits (melody, harmony, rhythm, key, tempo) with high adherence and preservation
- Zero-shot covers without cover-specific training
- Agentic editing through external language models
-
Scalable supervision: MERT2 and SheetSage2 provide a pipeline for constructing aligned symbolic and semantic supervision from unlabeled recordings, addressing the lack of aligned score-audio data.
-
Frontier quality at smaller scale: YuE2 achieves competitive results with proprietary systems at 3.58B parameters, fewer than six of eight public baselines.
-
Trade-off insights: The cover generation results reveal a tension between work-identity preservation and style adaptation, with harmony playing a crucial role in musical style.
Conclusion
YuE2 unifies readable composition and full-song audio generation at frontier quality through symbolic planning. The single AR–NAR Mixture-of-Transformers model first writes an editable score, then expands it into semantic tokens and acoustic latents. Expert listening confirms that symbolic planning improves perceived quality, and the unified architecture outperforms separate LM+DiT designs.
The model achieves state-of-the-art results on WildSongBench, competitive performance with proprietary systems, and enables novel capabilities including controlled editing, zero-shot covers, and agentic music editing—all with one checkpoint. MERT2 and SheetSage2 provide the semantic and symbolic supervision that makes this hierarchy learnable from recordings at scale.
Future directions suggested by the work include:
- Further exploration of the trade-off between score fidelity and style adaptation
- Scaling the approach to more diverse musical genres and styles
- Extending agentic editing capabilities for more complex compositional tasks
Related papers
- Hill Sampling for Test-Time Scaling: A Simple and Better Alternative to Repeated Sampling, Evolution, and Training
Hill Sampling, repeatedly sampling edits to the best program found so far, outperforms all complex evolutionary and weight-space methods, setting new state-of-the-art results on circle packing.
- FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders
FuseReg samples normalized subsets of frozen encoder layers during training, regularizing representation autoencoders against cross-layer disagreement and improving reconstruction and generation without architectural changes.
- Transferring the Intelligence of VLMs to Robotic Control
RoboDawn lets frozen vision-language models control robots zero-shot via a semantic action interface and in-context demos, beating robot-trained VLAs on RoboTwin and RoboDojo benchmarks.