AutoSynth generates editable Helm synthesizer programs directly with an audio-conditioned autoregressive model, achieving competitive results in both synthesizer inversion and text-driven generation
Synopsis
AutoSynth serializes MIDI performance events, fixed synthesizer parameters, and variable-length modulation routes into a unified sequence and learns their dependencies with an audio-conditioned autoregressive model, trained first by supervised learning on large-scale automatically constructed audio-program pairs and then by group-relative policy optimization with a mixed reward, so that it predicts a synthesizer program directly from reference audio and converts text into a program via a pretrained audio generation model, achieving competitive results in both synthesizer inversion and text-driven generation.
Figure 1: Overview of AutoSynth.
arXivInterpretation
A unified autoregressive program representation organizes MIDI performance events, fixed synthesizer parameters, and variable-length modulation routes into one sequence generated by a single decoder. Prior symbolic representations such as MIDI and ABC express only pitch and timing and do not determine timbre or synthesis; this work brings performance, timbral configuration, and modulation structure into one executable description and models their dependencies in the order MIDI, fixed parameters, modulation routes. The paper reports the sequence format and an ordering ablation: the current order generally achieves lower in-domain validation loss than a fixed random permutation of parameters and than a fixed-parameters-first order, indicating that autoregressive ordering affects training performance.
A two-stage training pipeline requires no paired text-target-program annotations: large-scale supervised data are constructed automatically from a small set of native presets, followed by mixed-reward GRPO post-training from rendering feedback. Supervised data expand Helm's 274 factory presets and one initialized state into two million five-second stereo audio-program pairs; post-training uses a weighted sum of eight reward components and needs no differentiable synthesizer. Ablation shows that relative to SFT alone, mixed-reward GRPO improves all six metrics on the Gemini-Woosh, FSD50K, and in-domain Helm inversion test sets, including LanguageBind, which is not used in the reward; for text-driven generation on Gemini, text CLAP, LanguageBind, CE, CU, and PQ all improve.
One inversion model supports both audio-driven and text-driven paths, producing synthesizer programs that users can edit directly. The text path first uses a frozen Stable Audio 3 Medium model to generate a waveform, which the same encoder-decoder inverts; text is not fed directly into the program decoder, and both paths share the encoder, decoder, and renderer. On Gemini-Woosh inversion, AutoSynth reaches an audio CLAP distance of 0.37 versus 0.54 for Synth Permutations and 0.49 for DDSynth-RL; on Gemini text generation its text CLAP is 0.550 and PQ is 7.433, higher than the other two inversion-based cascades, and PQ is not directly optimized.
In subjective listening tests, AutoSynth leads on both alignment and content quality across both tasks and both test sets. Ten sound-creation professionals familiar with synthesizers rated inversion and text-driven generation in an anonymous, blinded study with hidden method names and independently randomized candidate order. The results contain 1,800 valid scores, with 50 ratings per method, test set, and dimension; on inversion, AutoSynth leads both ratings on both test sets even though DDSynth-RL has higher RMS similarity.
Perspective
The work targets short sounds useful for music production and sound design rather than complete songs; the data contain single-note and chord performances rather than long melodies. Results apply to sounds within Helm's expressive range: the advantage is clear on the sound-design-oriented Gemini-Woosh test set, while on FSD50K, which covers broader real-world recordings including environmental sounds and speech, the CLAP gap narrows, which the paper attributes to sounds beyond the synthesizer's expressive range. The approach can extend to other synthesizers but requires redefining the program format, data-augmentation rules, and post-training reward design. Training and evaluation code are open-sourced to support independent reproduction and comparison.
The paper notes that a unified benchmark for complete synthesizer-program generation is still needed, and that differences in synthesizers' expressive capabilities and control structures make it difficult to disentangle the contributions of generation methods and underlying sound engines in cross-system comparisons. Reward components and weights translate human sound-design preferences into training objectives, but program formats vary across synthesizers and legally usable training data remain scarce, making data-augmentation rules and reward design particularly important. Sound design is a creative activity, and neither acoustic-feature fidelity nor semantic distance fully captures creative intent and aesthetic preferences, while timbral aesthetics involve subtle perceptual differences that are difficult to verbalize and vary across listeners, making context-free judgments difficult. In addition, text-driven generation does not achieve the best LanguageBind performance on FSD50K, indicating room to improve synthesizer-based sound matching for broader sound-effect tasks.
