Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Fine-tuning large audio language models with low-level acoustic features and speaker demographics surpasses deep-learning baselines for dysarthric speech detection, with Qwen2-Audio-Instruct reaching state-of-the-art performance

The work proposes a framework that fine-tunes Large Audio Language Models (LALMs) for dysarthric speech detection on speech recordings combined with textual information comprising low-level acoustic features and speaker demographics; across two LALMs the framework outperforms deep-learning-based baselines, with Qwen2-Audio-Instruct achieving state-of-the-art performance, and an ablation study shows that incorporating acoustic features and speaker demographics during fine-tuning improves LALM performance, while LALMs alone exhibit only chance-level zero-shot performance.