Skip to main content
Back to timeline
PloS oneSource publication:

Explainable AI for Sentiment Analysis of HMPV Using XLNet and SHAP

Synopsis

The study scraped 15,300 HMPV-related comments from YouTube news channels in 2024-2025, retained 9,758 after preprocessing, generated weak sentiment labels with VADER, compared six transformer models (ELECTRA, RoBERTa, ALBERT, DistilBERT, XLNet, BERT), found XLNet best at 93.50% accuracy, and used SHAP's PartitionExplainer to produce word-level attributions on correctly classified samples, showing words such as "flu" and "fear" driving negative predictions, "save", "help" and "mad" associated with neutral, and "falling", "save", "god" and "america" associated with positive.

Source-provided article image: Explainable AI for sentiment analysis of human metapneumovirus (HMPV) using XLNet.
PubMed

Interpretation

It builds an end-to-end sentiment analysis pipeline for HMPV-related YouTube comments, covering multilingual translation, emoji transliteration, removal of links and special characters, stopword removal and lemmatization, with the dataset and code released publicly. The authors describe it as the first sentiment analysis study on HMPV, since prior HMPV literature centered on epidemiology and clinical diagnosis rather than online public discourse. 15,300 comments were scraped and 9,758 retained after preprocessing; data and training code are hosted in a public GitHub repository, giving a reproducible base.

Across six transformer models evaluated on the same held-out test set, XLNet led with 93.50% accuracy, ahead of ELECTRA at 91.14% and BERT at 91.00%, while RoBERTa, ALBERT and DistilBERT ranged from 89.50% to 89.96%. It puts XLNet's permutation language modeling advantage to an empirical test on informal user-generated text about a specific public health topic rather than only on general benchmarks. Each model used its own tokenizer on the same held-out test set with weighted precision, recall and F1; the confusion matrix shows 708 correct Negative, 629 Neutral and 424 Positive cases, and one-vs-rest AUCs of 0.97 for Negative, 0.97 for Neutral and 0.96 for Positive.

It applies SHAP's PartitionExplainer to the fine-tuned XLNet, using 100 neutral training comments as the background set, summing subword attributions into word-level scores, and explaining only correctly classified samples. It brings explainable AI into HMPV sentiment classification so that which words push a sentiment decision becomes visible, rather than reporting only an accuracy figure. Attributions yield concrete word examples: "flu", "fear" and "shit" show high SHAP values in negative predictions, "save", "help" and "mad" relate to neutral, and "falling", "save", "god" and "america" push toward positive; the authors also manually reviewed 50 high-importance comments per class and report that under 10% were only weakly related to HMPV.

It adds a human qualitative validation step to check whether SHAP-identified important words actually sit within HMPV discourse rather than political discussion, generic profanity or spam. It explicitly quality-checks the explanations themselves within a sentiment classification study, responding to the concern that attributions can be misled by topic drift. For each class, 50 correctly classified comments with high absolute SHAP values were manually reviewed by the authors, who concluded that most high-importance words appeared in comments about infection severity, public health response, personal anxiety or trust in authorities.

Perspective

The framework suits exploratory online discourse monitoring of YouTube news-channel comments, especially where informal text is multilingual and contains emojis and noise; the authors state its public health relevance lies in supporting infodemiological analysis and early detection of discourse trends, with future extension to multi-platform data such as Twitter/X, Facebook and Reddit, demographic metadata where ethically permissible, and triangulation with survey or epidemiological data.

Readers should still watch that sentiment labels are VADER-generated weak labels, with the human check being qualitative and without a formal agreement measure; SHAP explanations cover only correctly classified samples and the authors stress they are approximate contribution estimates rather than causal attributions; the model separates the Positive class less well than Negative and Neutral, with the confusion matrix showing 56 Positive cases predicted Negative and 17 predicted Neutral; mixed sentiments and sarcasm remain hard to identify; and XLNet pretraining and training are computationally costly. In addition, in the text available here figures and tables appear as numbered references, so specific hyperparameter values and per-class precision, recall and specificity numbers could not be read directly; reproduction details should be checked against the original tables and the public repository.

Sources