Skip to main content
Back to timeline
arXivSource publication:

PEACE delivers the first joint embedding of DSP effect code and audio, reaching 53.5% R@10 on reverb retrieval versus 28.3% for pretrained AFx-Rep

Related research and updates

Synopsis

The work introduces PEACE, the first joint embedding that maps Faust effect code and processed audio into one space: the audio side reuses AFx-Rep while the code side is either a fine-tuned T5 or BoxGraph, a message-passing graph neural network over the Faust compiler's Box API intermediate representation, trained with SLAP's non-contrastive BYOL objective; across 200K training pairs, 21 Faust effects and chains of length 1-3, the two code encoders roughly tie on mixed-length chain retrieval, T5 leads slightly on single-effect galleries and longer chains, BoxGraph is stronger at parameter-masked chain-topology retrieval (17.4% versus 12.4% R@1) and reaches 53.5% R@10 on a reverb benchmark built from professional plugins unseen in training, above pretrained AFx-Rep's 28.3%.

Source-provided article image: PEACE: Joint Embeddings of DSP Effects Code and Audio
Figure 1 ·

Figure 1 : SLAP architecture. Each modality’s online branch pairs encoder-projector ℰ M \mathcal{E}_{M} (producing y M y_{M} then z M z_{M} ) with predictor 𝒫 M \mathcal{P}_{M} to produce q M q_{M} . The EMA target branch ℰ ¯ M \bar{\mathcal{E}}_{M} produces stop-gradient z ¯ M \bar{z}_{M} with no predictor. The ℒ \mathcal{L} terms show loss gradients flowing through the predictors into the online branches, not the EMA targets.

arXiv

Interpretation

PEACE establishes the first joint embedding space between audio effect code and effected audio, so an audio query can retrieve the Faust program that produced it by cosine similarity, and vice versa. Prior work either relies on natural-language descriptions, regresses parameters, or predicts fixed labels, none of which offers a similarity metric for comparing programs; PEACE makes the code modality a retrieval target that encodes both program structure and parameter values. Trained on 200K training, 8,192 validation and 4,096 test audio-code pairs, with cross-modal retrieval reported as Recall@1/@10 and normalized ranks on galleries of chains of length 1-3, alongside modality gap distance and modality/source separability probes.

The two code encoders trade off: fine-tuned T5 is slightly better on average in single-effect galleries and degrades more slowly on chains beyond the training distribution, while BoxGraph wins on about half the effect types and is clearly stronger at parameter-masked chain-topology retrieval. The paper turns the choice of code encoder from a single design into two comparable designs, and shows that masking parameters lets the embedding express only effect selection and ordering without a supervised classifier. On single-effect R@10, BoxGraph wins 11 of 21 effects while T5 leads on Chorus, Delay, Flanger, Phaser, ReverseEcho and three of five reverbs; on chain retrieval AFx-Rep-BoxGraph reaches 17.4% R@1 and 61.4% R@10 versus T5's 12.4% and 43.2%.

Cross-modal SLAP training improves the audio side's intramodal understanding: on a benchmark of 1,831 room impulse responses from four professional reverb plugins unseen in training, PEACE trained from scratch with BoxGraph reaches 53.5% R@10, far above pretrained AFx-Rep at 28.3%, Fx-Encoder++ at 11.6% and LAION-CLAP at 10.7%. This shows joint embedding training does more than align two modalities: it can also improve the audio representation itself, and pairing frozen AFx-Rep latents with BoxGraph already lifts R@10 from 28.7% to 38.6%. The benchmark uses 1,831 ReverbFX stereo RIRs and 250 jaCappella vocal tracks, with the task of identifying which of 1,831 candidates shares the same RIR; PEACE reports both latents and predictor outputs, and the predictor always outperforms the latent.

Retrieval differences across effect types point to concrete directions for improving the audio encoder and data: time-based reverb and delay perform best, non-linear and dynamics effects such as compressor and distortion are weaker, and panner effects are limited because mid/side spectrograms cannot distinguish left from right. The paper breaks the aggregate metric down across 21 effects and identifies which effects pose structural difficulty for the current audio front end, rather than reporting only an average. In single-effect R@10, ReverbFreeverb reaches 84.5% (BoxGraph) while Gate reaches only 2.7%, Limiter 2.8% and StereoWidth 0.9%; the paper explains that Limiter and Gate may not change audio for some inputs and StereoWidth does not change mono inputs.

Perspective

The result targets software audio effects written in Faust and chains of length 1-3, with training data of randomly sampled parameter presets paired randomly with dry audio, and galleries ranging from 21 single effects to 420 ordered effect pairs. It lets audio queries retrieve effect-chain topology and code queries retrieve audio, and offers a target space for style transfer, dry stem recovery and DSP code generation; the paper also envisions replacing a DNN effect classifier in an existing pipeline with chain retrieval, then handing candidates to CMA-ES parameter search. The intended setting is music production and audio effect analysis, not real-time plugin processing or human perceptual evaluation.

The paper notes a sim-to-real gap in its training pipeline: parameters are randomly sampled rather than human-curated, presets are randomly paired with source audio, and the source audio is not always fully dry; out-of-domain evaluation covers only reverb rather than a broader range of plugin effects, and there is no user study. Variance in within-effect retrieval is large, and which effects would benefit from time-domain encoders or better data remains open. BoxGraph currently uses only the par and seq composition operators, leaving split, merge and recurse to future work, so behavior after extension to modular synthesis (envelopes, LFOs) is not yet known.

Sources