Skip to main content
Back to timeline
arXivSource publication:

Extending mDPO to Audio: Qwen2-Audio Gains 14.0 Points on DCASE 2025 and 27.4 Points on AH Existence

Synopsis

This work ports multimodal Direct Preference Optimization (mDPO) from vision-language models to the audio domain, building preference pairs that contrast intact audio with acoustically perturbed audio (reversal, random noise, time and frequency masking) and training Qwen2-Audio-7B-Instruct, raising DCASE 2025 audio QA accuracy from 56.0% to 70.0% and AH Existence accuracy from 55.6% to 83.0%.

Source-provided article image: Hallucination Reduction for LLM-Based Audio Understanding via Multimodal Direct Preference Optimization
Figure 1 ·

Figure 1: Overview of the mDPO framework for audio hallucination reduction.

arXiv

Interpretation

It introduces an audio-domain mDPO objective that turns the negative audio branch of contrastive decoding into the rejected input of preference learning, correcting acoustic grounding in the weights rather than at inference time. The original mDPO relies on visual distortions such as image cropping, which do not transfer to audio; this work substitutes acoustic perturbations (reversal, Gaussian noise, time masking, frequency masking, no audio) while keeping the standard DPO term, the conditional preference term, and the anchored preference term. On Qwen2-Audio-7B-Instruct with classification accuracy as the metric, ablating the three loss terms shows the full objective lifts accuracy from 56.0% to 71.4%, while removing the anchor term drops it to 64.0%.

It identifies temporal reversal, frequency masking, and random noise as the most effective perturbations and combines them into a Triple Mixed training set. Prior contrastive decoding work mostly applies perturbations at inference time; this work systematically compares perturbations inside preference training and yields a reusable perturbation-selection recipe. On the DCASE 2025 test set, reverse reaches 67.1%, random noise 67.4%, and frequency masking 67.3%, all above no-audio at 66.7% and time masking at 66.6%; Triple Mixed gives the highest overall accuracy at 70.0%.

On DCASE 2025, perturbation-based preference training reaches Complex-subset performance comparable to challenge teams that fine-tune with external data. Those systems lean on manual data curation, annotation reformatting, and iterative pseudo-labeling; this work offers perturbation-constructed preference pairs as an alternative route. On the Complex subset this method scores 80.0% versus DCASE Team 1 at 83.3% and Team 2 at 80.1%; on the Temporal subset it scores 43.8% versus 60.4% and 62.6% for the two teams.

On AH Existence, the reverse perturbation reaches 83.0% accuracy, outperforming inference-time contrastive decoding at 76.7%. Hyperparameters selected on DCASE are transferred directly to AH Existence without retuning, indicating that perturbation-based preference training transfers across datasets. On the AH Existence test set the baseline is 55.6%, with no-audio 81.3%, reverse 83.0%, random noise 81.2%, and frequency masking 79.2%, against contrastive decoding at 76.7%.

Perspective

The results apply to audio question answering with Qwen2-Audio-7B-Instruct as the backbone and classification accuracy as the metric: the Temporal and Complex subsets of DCASE 2025 and the AH Existence subset. The method requires regenerating a preference training set for each perturbation, so perturbation choice and hyperparameter configuration are prerequisites for use. The authors state that future work will extend the method to open-ended generative tasks such as audio captioning and explore adaptive perturbation strategies that dynamically adjust distortion difficulty during training.

The Bioacoustics subset of DCASE 2025 is excluded because its audio data is not publicly accessible, so conclusions do not cover that domain. On the Temporal subset this method scores 43.8%, below the two challenge teams at 60.4% and 62.6%, indicating limited gains for temporal reasoning. Only the AH Existence subset is used, leaving AH Order and AH Attribute unevaluated. In addition, several equations, the anchor margin value, and specific perturbation parameter values appear in the text as symbols or tables, so reproduction requires checking Figure 1 and the full table values.

Sources