Skip to main content
Back to timeline
Journal of Medical Internet ResearchSource publication:

What Single-Topic Summaries Miss in Hospital Reviews: Aspect-Level Evaluative Structure Using Generative Pretrained Transformer-Based Sentiment Analysis

Synopsis

Using 5,467 Google Reviews posted in 2024 from all 24 medical centers in Taiwan, this study compared the common LDA dominant-topic assignment with GPT-based aspect-based sentiment analysis (ABSA) on the same corpus, finding that aspect-bearing reviews discussed an average of 2.05 distinct service aspects, that dominant-topic assignment yielded an illustrative 51.2% representational compression, that a soft-assignment LDA baseline reduced count-level compression to 1.7% but left semantic alignment limited (mean set Jaccard=0.33) and carried no aspect-level sentiment polarity, that 11.0% of multiaspect reviews showed cross-aspect mixed sentiment with Technical-Functional Divergence accounting for 61.

Interpretation

It quantified how much dominant-topic assignment compresses multidimensional patient feedback: aspect-bearing reviews discussed an average of 2.05 distinct aspects (95% bootstrap CI 2.02-2.08), yielding an illustrative 51.2% compression estimate (95% bootstrap CI 50.6%-51.9%). Applied LDA research typically assigns each review to a single topic; by contrasting that practice with aspect-level analysis on the same corpus, the study turns a conceptual concern about compression into an estimable quantity. Based on the full corpus of 5,467 reviews with bootstrap confidence intervals; LDA was set at K=7 and aligned to the 7 ABSA categories, forming a controlled but information-asymmetric comparison.

Soft-assignment LDA reduced count-level compression to 1.7%, yet semantic alignment with ABSA aspects remained limited (mean set Jaccard=0.33), and topic assignments carry no aspect-level sentiment polarity. It shows that recovering topic counts and recovering semantic and sentiment information are distinct achievements, offering a finer distinction for choosing analytic tools. A soft-assignment LDA baseline and GPT-ABSA were run in parallel on the same corpus, with set Jaccard quantifying semantic alignment; the absence of sentiment polarity is an observation about the method's properties.

Among multiaspect reviews, 11.0% exhibited cross-aspect mixed sentiment, and Technical-Functional Divergence, praising technical quality while criticizing functional quality, appeared in 61.6% of these cases. It moves mixed sentiment from a review-level phenomenon to sentiment divergence across aspects within the same review and identifies the most common form of that divergence. Mixed sentiment was defined as containing both positive and negative aspect evaluations; GPT-4o achieved accuracy 0.89, weighted F1 0.89, and Cohen κ=0.78 against the consensus gold standard, with interrater reliability of Cohen κ=0.82 between two independent annotators.

Clinical dimensions were more frequently co-mentioned in positive reviews and operational dimensions in negative reviews, but pointwise mutual information indicated these differences were substantially confounded with marginal aspect prevalence rather than reflecting prevalence-independent co-occurrence tendencies. By adding a prevalence-adjusted sensitivity analysis alongside the reported co-mention differences, it guards against reading which aspects patients discuss as associations among aspects. Rating-stratified network analysis compared aspect co-mention patterns using Jaccard similarity, with pointwise mutual information as a prevalence-adjusted sensitivity analysis.

Perspective

The study addresses service-quality analysis of patient review text and applies to settings where multiple service aspects and their sentiment polarities need to be distinguished within a single review, for example organizing feedback across clinical and operational dimensions. Its comparison design aligns the 7 ABSA categories with the K=7 LDA topic labels, so the conclusions apply directly to comparisons under that alignment; for users who want soft assignment to retain topic counts, the results suggest that count recovery and semantic and sentiment recovery should be evaluated separately.

This reading is at the summary scope, with figures and supplementary materials not included, so the specific definitions of aspect categories, the stratification details of the network analysis, and the specific pointwise mutual information values cannot be checked here. The compression estimate is described by the authors as illustrative, and its interpretive boundary is worth noting; the mixed-sentiment proportion and the share of Technical-Functional Divergence come from this corpus and would need re-estimation when extrapolating to other regions, platforms, or years. How much aspect-level analysis changes hospitals' actual handling of feedback remains an open question.

Sources