AI model reached only 0.56 accuracy on seven-category BITSS infant stool grading, rising to 0.76 when collapsed to four groups
Synopsis
Using 1312 photographs of infant stool in diapers (1100 meeting quality criteria, of which 896 had majority agreement among three pediatric gastroenterologists as the reference set), the study trained 20 convolutional neural network models on Teachable Machine with an 85/15 training and internal-validation split and assessed agreement with specialists on the BITSS: the best seven-category model reached overall accuracy 0.56 with linear and quadratic weighted Kappa of 0.59 and 0.75, while a four-category grouping (constipated, formed, soft, liquid) reached accuracy 0.76 with weighted Kappa of 0.69 and 0.78, whereas agreement among the three specialists showed a Fleiss Kappa of only 0.24 and 24.7% complete agreement.
Interpretation
In a Latin American setting, the AI model achieved moderate to substantial agreement with the pediatric gastroenterologists' reference classification of infant stool in diapers on the BITSS, with quadratic weighted Kappa of 0.75 for seven categories and 0.78 for four categories. Prior image-recognition work on this task came largely from other regions; this study provides internal-validation evidence on local data and reports both inter-specialist agreement and AI-reference agreement. Based on 896 images with majority agreement, 20 independent models (10 for the seven-category and 10 for the grouped version), and linear/quadratic weighted Kappa with 95% confidence intervals; this is internal validation, not external validation.
Collapsing the seven BITSS categories into four functional groups markedly improved model performance, with overall accuracy rising from 0.56 to 0.76, indicating that fewer categories help the algorithm discriminate and remain stable. The study directly compares seven-category and four-category settings on the same image bank, quantifying how category granularity affects AI performance. Derived from the confusion matrices and per-category sensitivity, specificity, precision, and F1 scores of Model 7 and Model 19 on the same reference set; in the four-category version the constipated group had the highest F1 (0.90) and the formed-stool group the lowest (0.45).
Specialists themselves showed substantial variability in BITSS grading, with complete agreement among the three raters in only 24.7% of images and a Fleiss Kappa of 0.24, with disagreements concentrated in adjacent categories, especially BITSS 3 and 4. This places AI performance against a human-reading baseline, suggesting that part of the classification difficulty stems from ambiguity in the scale's intermediate categories rather than from the algorithm alone. Based on independent blinded grading by three pediatric gastroenterologists with majority vote defining the reference category, plus pairwise weighted Kappa values (linear 0.43-0.55, quadratic 0.63-0.76).
Model errors clustered mainly in adjacent categories, and BITSS 4 (formed stool) was the worst-discriminated category in both the seven-category and four-category models. This pattern is consistent with prior observations in ordinal classification of medical images, where errors cluster in adjacent categories, and it points to class imbalance as a possible influence on performance. From the confusion matrices and per-category metrics: in the seven-category model BITSS 4 had sensitivity 0.11 and F1 0.20; in the four-category model the formed-stool group had sensitivity 0.39 and F1 0.45.
Perspective
The study addresses automated classification of photographs of stool in diapers from infants under 12 months, applicable under standardized capture conditions (resolution at least 640 x 480 pixels, uniform lighting, more than 50% of the stool sample visible), and can serve as a research tool to reduce visual grading variability in pediatric gastroenterology. Its value lies in providing internal-validation and preliminary reproducibility evidence for subsequent external validation, and in suggesting that the algorithm is more stable at coarser category granularity (four functional groups), suitable for exploratory screening or grouping applications.
The authors note that the Teachable Machine platform limits access to technical model details, restricting full reproducibility of the training process; the image database did not allow establishing independence by patient or the individual contribution of each infant; and excluding images without consensus and applying minimum photographic quality criteria may have reduced the dataset's representativeness compared with typical clinical scenarios. In addition, class imbalance may skew performance toward the most represented categories, and future versions should increase representation of less frequent categories. Readers may still watch whether the model remains stable on the unanimous-agreement subset, how it performs at different levels of diagnostic certainty, and whether external validation reproduces the higher agreement seen in the four-category version.
