Artificial Intelligence-Based Hypernasality Diagnosis Using CAPS-A-AM Rated Speech Samples in Pediatric Velopharyngeal Dysfunction
Synopsis
In this prospective, single-center study, speech samples from 40 children aged 2 to 17 (including individuals with velopharyngeal dysfunction, conditions associated with VPD, and healthy participants) were collected during speech-language pathologist-guided evaluation with consensus CAPS-A-AM ratings, and mel spectrograms of high vowels /i/ and /u/ from sustained vowels, isolated words, and sentences were used to train logistic regression, an EfficientNet-V2-S attention multiple-instance-learning CNN, and a CNN-XGBoost hybrid for binary hypernasality classification at two CAPS-A-AM thresholds (absent 0 versus any hypernasality 1-4, and absent/borderline 0-1 versus mild-to-severe 2-4); multiple independent modeling approaches detected clinically rated hypernasality, with EfficientNet-V2-S a
Interpretation
The study established a workflow and preliminary algorithm for AI-based hypernasality detection grounded in CAPS-A-AM standardized auditory-perceptual assessment and mel spectrograms. Prior hypernasality evaluation relies on speech-language pathologists' auditory-perceptual judgment; this work combines the standardized CAPS-A-AM framework with a machine learning pipeline, exploring a path toward accessible speech assessment tools when SLP expertise is limited. Prospective single-center design with speech samples collected during SLP-guided evaluation and consensus ratings, providing a standardized reference for model development; sample size was 40 children.
Multiple independent modeling approaches (logistic regression, EfficientNet-V2-S attention MIL CNN, and CNN-XGBoost hybrid) were able to detect clinically rated hypernasality. The study evaluated three methods of differing complexity rather than a single model, indicating that hypernasality detection is not tied to one specific modeling strategy. Three approaches were compared within the same cohort and assessment framework, with results consistently pointing to detectability; however, overlapping confidence intervals precluded demonstration of model superiority.
EfficientNet-V2-S achieved the highest observed performance estimates in this cohort, with F1 scores of 0.786 to 0.900. This provides a concrete performance reference range for mel-spectrogram-based hypernasality classification across binary tasks at two CAPS-A-AM thresholds. Performance estimates come from a single-center cohort of 40 children with overlapping confidence intervals, so the values should be read as preliminary observations rather than established superiority.
Models were tested at two CAPS-A-AM thresholds: absent (0) versus any hypernasality (1-4), and absent/borderline (0-1) versus mild-to-severe hypernasality (2-4). By setting two clinically relevant thresholds, the study examined both whether hypernasality is present and whether it reaches mild-to-severe levels, aligning with clinical triage needs. Thresholds map directly onto the CAPS-A-AM rating system with clear evaluation criteria; however, the limited sample size means performance differences between thresholds require larger-sample validation.
Perspective
The study defines the preliminary feasibility of mel-spectrogram-based AI hypernasality classification in a pediatric velopharyngeal dysfunction cohort, applicable to speech samples collected under speech-language pathologist guidance and referenced to consensus CAPS-A-AM ratings. Its next-step significance lies in expanding sample acquisition and incorporating increasingly diverse speech inputs for training and rating, laying groundwork for assistive assessment tools in settings with limited expert resources.
Readers should still watch: overlapping confidence intervals mean model superiority is not established, and the EfficientNet-V2-S F1 of 0.786-0.900 should be treated as an observed estimate in this cohort; the single-center sample of 40 children limits generalization to broader populations; performance differences of mel spectrograms for high vowels /i/ and /u/ across sustained vowels, isolated words, and sentences warrant further analysis; and the influence of different language backgrounds, recording equipment, and rater consistency on the model remains an open question.
