Fine-tuning large audio language models with low-level acoustic features and speaker demographics surpasses deep-learning baselines for dysarthric speech detection, with Qwen2-Audio-Instruct reaching state-of-the-art performance
Related research and updatesSynopsis
The work proposes a framework that fine-tunes Large Audio Language Models (LALMs) for dysarthric speech detection on speech recordings combined with textual information comprising low-level acoustic features and speaker demographics; across two LALMs the framework outperforms deep-learning-based baselines, with Qwen2-Audio-Instruct achieving state-of-the-art performance, and an ablation study shows that incorporating acoustic features and speaker demographics during fine-tuning improves LALM performance, while LALMs alone exhibit only chance-level zero-shot performance.
Figure 1: Schematic illustration of the proposed method: a) Conventional fine-tuning of the LALM using raw audio and textual inputs. b) Fine-tuning with the proposed approach, where besides the raw audio and textual inputs, low-level audio features are extracted and incorporated into the textual input together with the task description and metadata.
arXivInterpretation
The framework fine-tunes Large Audio Language Models (LALMs) for dysarthric speech detection, taking speech recordings together with textual information that comprises low-level acoustic features and speaker demographics. Existing automatic dysarthric speech detection approaches predominantly rely on deep learning, and although LALMs have shown strong performance across various tasks, their application to dysarthric speech detection had not yet been established; this work proposes a concrete framework for adapting LALMs to this task. The abstract reports that across two LALMs the framework outperforms deep-learning-based baselines, with Qwen2-Audio-Instruct achieving state-of-the-art performance, so the claim rests on a cross-model comparison.
An ablation study shows that incorporating acoustic features and speaker demographics during fine-tuning improves LALM performance. This attributes the gain to the textual-side low-level acoustic features and demographic information rather than to the LALM alone, giving a testable basis for input design. The evidence comes from an ablation study, a controlled comparison over input components, though the abstract gives no specific values or effect sizes.
LALMs alone exhibit only chance-level zero-shot performance, indicating that an un-fine-tuned LALM cannot directly handle dysarthric speech detection. This contrast highlights the necessity of fine-tuning and feature injection, running against the impression that LALMs perform strongly across many tasks. The evidence is an observation of performance under the zero-shot condition, described in the abstract as chance-level without specific metrics.
Perspective
The result targets the specific task of dysarthric speech detection and applies in settings where speech recordings are available and low-level acoustic features and speaker demographics can be extracted; its intended use is to support, not replace, traditional clinical diagnosis that relies on a speech and language pathologist. For readers who want to bring large audio language models into speech-pathology-related tasks, the work offers a reusable adaptation pattern: fine-tuning as the main mechanism, with low-level acoustic features and demographic information injected as textual-side input.
The abstract does not give dataset size, recording conditions, speaker distribution, evaluation metrics, or specific performance numbers, nor does it specify how the low-level acoustic features are composed or extracted, so the magnitude of the improvement and its stability across populations or recording environments cannot be judged. The exact measure behind the chance-level zero-shot result is likewise not stated in the abstract. In addition, this document is summary-level material lacking the body, figures, and experimental details, so these questions require the original paper to resolve.
