Two-stream Transformer fusing PTR-ToF-MS volatiles with targeted non-volatile metabolites grades Baimudan white tea at 95.8% accuracy on 24 held-out samples
Synopsis
Using PTR-ToF-MS headspace volatile fingerprints plus HPLC- and amino-acid-analysis-quantified non-volatile metabolites, this study built a two-stream Transformer that encodes each modality separately and fuses them for four-class grading of Baimudan white tea, reaching 95.8% accuracy (23/24) and a 0.958 macro-F1 on an independent prediction set drawn from 120 samples (30 per grade; 96 training, 24 prediction), with a single Special-grade sample misclassified as Grade I, mean cross-validated accuracy of 0.979±0.026 within the training set, and SHAP/attention analyses linking high grades to floral/sweet volatile ions plus higher amino acids and soluble sugars and lower grades to greener/woody volatile ions and kaempferol-related markers.
Multimodal Transformer architecture for Baimudan white tea quality grading.
PubMedInterpretation
A two-stream Transformer encodes PTR-ToF-MS volatile fingerprints (268 dimensions) and targeted non-volatile metabolites (34 dimensions) separately before fusion, achieving 0.958 accuracy and 0.958 macro-F1 on the independent prediction set (n=24, six per grade), with only one Special-grade sample predicted as Grade I. Prior Baimudan grading work mostly analyzed volatile or non-volatile modalities separately using linear chemometrics or unimodal models; this study preserves modality-specific representations before fusion and reports four-class results on a held-out set. The prediction set was fully held out from training, hyperparameter selection, early stopping, and normalization; within the training set, three repeats of stratified five-fold cross-validation (15 validation runs) gave mean accuracy 0.979±0.026 and macro-F1 0.977±0.028, with normalization parameters fitted only on each fold's training portion to avoid leakage.
SHAP attribution and modality-specific self-attention heatmaps jointly indicate that grading relies on both volatile and non-volatile features: high grades (Special + Grade I) are enriched in floral/sweet-associated ions such as phenylacetaldehyde, γ-nonalactone, ethyl acetate, and linalool plus valine and total soluble sugars, whereas low grades (Grade II + Grade III) are enriched in green/woody ions such as 1-octen-3-ol, cis-3-hexen-1-ol, and (E)-2-hexenal and in kaempferol-related variables. The study couples attention-based fusion with SHAP interpretation to produce a cross-aroma-and-taste list of grade-discriminative drivers rather than reporting classification accuracy alone. Based on the top-20 chemical variables ranked by mean absolute SHAP value, together with a literature-supported sensory/taste annotation table and grade-wise statistics; the authors state these sensory labels are literature-derived descriptors and do not constitute direct sensory validation or causal attribution.
Confidence analysis shows 79% of the 120 samples were predicted with top-1 probability above 0.95, while misclassified samples showed reduced confidence and elevated Shannon entropy, and a reliability diagram indicated top-1 probabilities were broadly aligned with empirical accuracy. This supports probability- or entropy-threshold rejection strategies in an automated grading pipeline rather than a single hard classification output. Based on top-1 probability and Shannon entropy across all 120 samples and a reliability diagram binning samples by confidence; the authors frame this as a potential basis for automated screening or rejection.
A t-SNE projection of fused embeddings shows samples largely forming grade-wise clusters, with the misclassified prediction sample located near the Special–Grade I boundary, suggesting a borderline chemical profile rather than random instability. The visualization links the single misclassification to chemical continuity between adjacent grades, complementing the confusion matrix with representation-space evidence. t-SNE is descriptive (perplexity=20, n_iter=600, exact method, random seed 7); the authors note it does not affect training or evaluation, and formal performance evaluation used only the independent prediction set.
Perspective
The result targets instrument-assisted grading of Baimudan white tea within the four-grade GB/T 22291-2017 system, under the setting of 120 samples from one cultivar (Fuding Dabai), one harvest period (spring 2023), and one supplier's grading; for quality-control and testing contexts seeking to complement or replace sensory panels with PTR-ToF-MS plus targeted non-volatile metabolites, it provides a reproducible two-stream Transformer pipeline, modality-specific attention and SHAP interpretation, and a probability/entropy-based rejection idea. The authors state that broader external validation using independent samples from additional harvest years, producers, geographic origins, and processing batches will be necessary to evaluate generalizability, transferability, and robustness.
The sensory-related interpretation is exploratory: the sensory subscores are supplier-provided archival records from routine commercial grading under GB/T 22291-2017, no independent sensory panel was conducted, so links between chemical markers and aroma/taste attributes should be read as contextual associations rather than causal or definitive sensory markers, and targeted sensory validation would still be required. Galloylated catechins (EGCG/ECG) and caffeine ranked among top features and were higher in the high-grade group in this dataset; the authors suggest this may reflect batch-specific raw-material or processing characteristics and could be sensorially moderated by concurrently elevated amino acids and soluble sugars. In addition, the independent prediction set is small (six per grade) and all samples come from a single harvest period, so performance across years, origins, and batches remains an open question; this reading is full text, but some supporting details sit in supplementary tables and figures, so item-level values still require the supplementary materials.
