Skip to main content
Back to timeline
arXivSource publication:

YuE2 writes a readable score before rendering, and experts prefer symbolic planning 49.3% to 34.6% on overall quality

Synopsis

YuE2 uses a single roughly 3.58B-parameter AR–NAR Mixture-of-Transformers to first write an ABC score specifying melody, chords, key, meter, tempo, and form, then expand it into semantic music tokens and render full-song audio, with companion models MERT2 and SheetSage2 constructing semantic and symbolic supervision from recordings; with the same checkpoint, experts prefer symbolic planning 49.3% to 34.6% for overall quality, YuE2 scores 6.7316 on SongBench Global Avg on WildSongBench (6.9632 with best-of-8), MERT2 surpasses previous best results on 14 of 15 MARBLE metrics, and SheetSage2-AR leads 12 of 15 benchmark–metric pairs.

AI-generated editorial illustration: YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Interpretation

YuE2 makes composition an explicit intermediate in generation: the model first emits a readable ABC score (melody, chords, key, meter, tempo, form), then generates 25-Hz semantic tokens and 25-Hz acoustic latents, which a separately trained 48-kHz stereo decoder turns back into a complete song. Earlier symbolic models such as Music Transformer and SymPAC stop at symbolic sequences, while audio models such as MusicLM, MusicGen, YuE, SongBloom, and LeVo 2 leave composition implicit in audio representations; YuE2 jointly models readable scores, semantic tokens, and acoustic latents inside one publicly available generator. The paper reports full architecture and training details: a 28-layer backbone with approximately 3.58B parameters, a 24,576-position context, roughly 346,000 hours of music training data, and four training tasks that support matched ablations from the same checkpoint.

Symbolic planning itself improves perceived song quality: with checkpoint, prompts, candidate budget, and decoder held fixed, experts give 49.3% of overall-quality preferences to planning versus 34.6% without it, with musicality at 45.0% versus 29.4%, melody at 44.0% versus 29.5%, and chord progression at 39.0% versus 21.0%. Prior work mostly treats symbolic conditions as external inputs (Music ControlNet and JASCO follow supplied controls) or plans only tempo, key, duration, and structure (ACE-Step 1.5); here writing a melody-and-chord score is a step the generator itself takes, tested against a matched no-planning condition. The direct planning comparison contains 212 retained expert evaluations, with 211 judged responses each for overall quality and musicality; melody and chord progression were assessed by four listeners from the same expert pool over 100 song pairs, 200 judgments per criterion, with CR1 standard errors clustered by evaluator and song and two-sided tests.

The unified MoT beats a separate LM+DiT: when both systems use melody-and-chord symbolic planning and share the same semantic tokenizer, acoustic VAE, and training-data volume, experts give 53.4% of overall-quality preferences to YuE2 versus 35.6% to separate LM+DiT, and 48.5% versus 29.6% on audio quality. Many recent song generators pair an autoregressive language model with a separate diffusion Transformer; this experiment fixes the representations and the planning method so that only the unified-versus-separate design choice is compared. The baseline LM and DiT each have 1.7B parameters (3.4B total), comparable to the 3.58B MoT, and the LM and MoT's AR component use the same token budget, batch size, and number of training updates; the comparison contains 209 retained expert evaluations with 206–208 judged responses per criterion.

The same checkpoint turns the score into an editing interface: corresponding recordings reach melody and chord sequence similarities of 0.9464 and 0.9246 versus 0.2160 and 0.2473 for mismatched recordings; local edits attain 84.17% target-pitch accuracy and 79.54% target-chord agreement while preserving 90.16%–94.34% of unedited melody and harmony; and zero-shot covers of 948 unseen works surpass both evaluated cover systems on all eight work-identity retrieval measures. Earlier cover systems need dedicated melody-conditioning modules or original–cover training pairs; YuE2 reuses the score-to-audio mapping learned from individual recordings without cover-specific training, and lets an external language-model agent translate user feedback into score revisions. The score–audio agreement study covers 384 generated scores and 768 recordings with mismatched and no-score controls; the editing study yields 3,844 recordings with all outputs retained; the cover evaluation uses 948 works from the official SHS100K test split, which the authors verified as absent from the training corpus using CLEWS and Discogs-VINet.

Perspective

This work targets full-song generation, score-level editing, and zero-shot cover generation, in settings where text and lyrics are the conditions and the composition needs to be inspectable. It enables follow-up work to treat the score as an interface that people and language-model agents revise together, to use MERT2's 375 bits/s semantic token stream for downstream autoregressive generation, and to build aligned symbolic supervision from recordings at scale with SheetSage2-AR. Beneficiaries include song-generation researchers, music information retrieval researchers, and developers of controllable composition tools. The cover findings apply to SHS100K works that are vocal, have complete automatic transcriptions, and have two annotated style alternatives; the expert-preference findings apply to the paper's listening protocol and candidate budgets.

Automatic metrics and expert preferences do not always agree: YuE2 best-of-8 has a higher SongBench Avg than Suno v6, yet experts still favor Suno v6 on overall quality 59.3% to 31.6%; on SongEval, HeartMuLa scores higher on Musicality and Avg, and the paper notes HeartMuLa used SongEval and AudioBox in post-training. In the cover experiments, relaxing score conditioning raises style alignment and quality scores while lowering work-identity retrieval, and the authors link this trade-off to harmony's role in musical style as a qualitative explanation. Agentic music editing is currently a single case study (nine curated steps and fourteen rendered versions of The Last Train) rather than a controlled evaluation. The evidence bundle is the full paper text, but some appendix tables and figures are not fully reproduced here, so details tied to specific numbers should be checked against the original.

Sources