Skip to main content
Back to timeline
IJIS - Indonesian Journal On Information SystemSource publication:

Across seven algorithms for SPKLU review sentiment on PLN Mobile, MLP wins with 89.16% accuracy but 59.42% macro F1

Synopsis

The study collected user reviews of the EV Charging/SPKLU feature in PLN Mobile from Google Play Store, labeled sentiment using ratings as weak labels, applied algorithm-specific class weighting on the training data to address class imbalance, and compared seven classifiers—Logistic Regression, SVM, Naive Bayes, MLP, RNN, BiLSTM, and DistilBERT—finding that MLP achieved the highest validation macro F1 (67.80%) and was selected for final testing, where it reached 89.16% accuracy and 59.42% macro F1 on 286 test instances, indicating high overall classification performance with uneven performance across sentiment classes.

Source-provided article image: EVALUASI MODEL MACHINE LEARNING, DEEP LEARNING, DAN TRANSFORMER UNTUK ANALISIS SENTIMEN ULASAN SPKLU PADA PLN MOBILE DENGAN PENANGANAN KETIDAKSEIMBANGAN KELAS

Interpretation

The study systematically compares seven classification algorithms spanning machine learning, deep learning, and Transformer families on PLN Mobile EV Charging/SPKLU reviews. Prior PLN Mobile sentiment work has largely focused on a few algorithms such as Naive Bayes, SVM, and Random Forest; this study brings MLP, RNN, BiLSTM, and DistilBERT into a single evaluation pipeline. The abstract explicitly names the seven algorithms and their three groupings and states the data come from Google Play Store, but does not provide per-algorithm comparison figures.

Sentiment labels were assigned via a rating-based weak-labeling strategy, and class weighting tailored to each algorithm was applied on the training data to handle class imbalance. Compared with similar work using SMOTE or lexicon-based labeling, this study chooses class weighting as the imbalance remedy and adjusts it per algorithm. The abstract states that data underwent text preprocessing and rating-based weak labeling and that class weighting was applied to training data; specific weight values are not given in the text.

MLP achieved the highest validation macro F1 (67.80%) and was selected as the final model, yielding 89.16% accuracy and 59.42% macro F1 on 286 test instances. The result provides an empirical record that MLP outperformed deep sequential models and a Transformer under the same data conditions, while reporting the gap between accuracy and macro F1. The abstract reports specific validation and test values and a test size of 286, and notes high accuracy with uneven per-class performance; confusion matrix details and confidence intervals are not provided.

Perspective

The work targets product and operations teams that need to automatically identify sentiment from PLN Mobile EV Charging/SPKLU user reviews, and applies to settings where Google Play Store reviews are the data source, ratings serve as weak labels, and class weighting handles class imbalance. Its pipeline can serve as a reference template for sentiment analysis of similar app reviews, especially when choosing among machine learning, deep learning, and Transformer approaches.

Because the reading scope is incomplete, the body, figures, and per-algorithm comparison results beyond the abstract were not available, so the full validation ranking of the seven algorithms, the class distribution revealed by the confusion matrix, and the specific class-weight values cannot be confirmed. Readers concerned with the class structure behind the macro F1 versus accuracy gap, or wishing to reproduce the pipeline, would still need the original methods and results details.

Sources