Public articles linked to the same research event.
arXiv The work proposes a framework that fine-tunes Large Audio Language Models (LALMs) for dysarthric speech detection on speech recordings combined with textual information comprising low-level acoustic features and speaker demographics; across two LALMs the framework outperforms deep-learning-based baselines, with Qwen2-Audio-Instruct achieving state-of-the-art performance, and an ablation study shows that incorporating acoustic features and speaker demographics during fine-tuning improves LALM performance, while LALMs alone exhibit only chance-level zero-shot performance.
The work proposes a framework that fine-tunes Large Audio Language Models (LALMs) for dysarthric speech detection on speech recordings combined with textual information comprising low-level acoustic features and speaker demographics; across two LALMs the framework outperforms deep-learning-based baselines, with Qwen2-Audio-Instruct achieving state-of-the-art performance, and an ablation study shows that incorporating acoustic features and speaker demographics during fine-tuning improves LALM performance, while LALMs alone exhibit only chance-level zero-shot performance.
The work proposes a framework that fine-tunes Large Audio Language Models (LALMs) for dysarthric speech detection on speech recordings combined with textual information comprising low-level acoustic features and speaker demographics; across two LALMs the framework outperforms deep-learning-based baselines, with Qwen2-Audio-Instruct achieving state-of-the-art performance, and an ablation study shows that incorporating acoustic features and speaker demographics during fine-tuning improves LALM performance, while LALMs alone exhibit only chance-level zero-shot performance.
The work proposes a framework that fine-tunes Large Audio Language Models (LALMs) for dysarthric speech detection on speech recordings combined with textual information comprising low-level acoustic features and speaker demographics; across two LALMs the framework outperforms deep-learning-based baselines, with Qwen2-Audio-Instruct achieving state-of-the-art performance, and an ablation study shows that incorporating acoustic features and speaker demographics during fine-tuning improves LALM performance, while LALMs alone exhibit only chance-level zero-shot performance.