Skip to main content
Back to timeline
Health Science ReportsSource publication:

Machine Learning for Noninvasive Anemia Diagnosis: A Systematic Review Based on CRISP-DM

Synopsis

Following PRISMA, this review searched PubMed, Web of Science, and Scopus (January 1, 2019 to March 27, 2025), included 58 of 1923 records, and used the CRISP-DM phases (problem understanding, data understanding, data preparation, modeling, evaluation, deployment) to organize machine learning approaches to noninvasive anemia detection and hemoglobin estimation, covering algorithms, data sources, acquisition sites, light sources, evaluation metrics, and deployment, alongside a PROBAST appraisal of risk of bias and applicability.

Source-provided article image: Applications of Machine Learning in Noninvasive Anemia Diagnosis: A Systematic Review Based on Cross-Industry Standard Process for Data Mining.
Figure 1 ·

The PRISMA flowchart detailing the number of articles retrieved from each database and produced after screening and eligibility assessment.

PubMed

Interpretation

The review organizes 58 included studies through the six CRISP-DM phases and provides a classification of techniques by machine learning algorithm, data type, and data acquisition site. Earlier reviews focused mainly on deep learning, image data, or smartphones alone; this review also incorporates other machine learning algorithms, other input data types, and multiple capturing devices, and extends the search window to March 2025. Methodologically it follows the PRISMA checklist, with two independent reviewers screening for eligibility (Cohen's κ = 0.94) and disagreements resolved by consensus with a third reviewer; search queries and inclusion/exclusion criteria are listed explicitly.

Included studies are highly dispersed in task, site, and signal: 33 studies performed continuous hemoglobin regression, 22 performed anemic/non-anemic classification, and 3 did both; the eye (26 studies, all imaging, most often the conjunctiva) and the finger (25 studies, mostly PPG) were the most frequent acquisition sites, with the palpable palm third at 7 studies. The review consolidates these scattered contributions into comparable distributions by task, body site, data type, and light source wavelength, noting three noticeable peaks at 660, 850, and 940 nm. Conclusions derive from item-by-item tabulated extraction across the 58 studies (Table 2) and distribution figures by body site and wavelength, constituting a qualitative rather than quantitative synthesis.

Performance metric distributions vary markedly: among regression metrics, MSE showed the greatest dispersion (median 0.655 (g/dL)², IQR 1.343), MAE median 0.737 g/dL (IQR 0.830), and RMSE median 0.685 g/dL (IQR 0.589); among classification metrics, precision was most scattered (median 97%, IQR 11.26%), sensitivity median 95.35%, and accuracy median 96.49%. Beyond listing each study's best result, the review reports cross-study medians and interquartile ranges, letting readers judge how concentrated or dispersed overall performance is. Based on aggregating the best test or validation metrics reported by each study; the authors also note that the 46% high risk of bias in the analysis domain may inflate these near-optimal metrics.

PROBAST appraisal rated 28 studies at low risk of bias, 30 at high risk, and 3 unclear, with high risk arising mainly from the analysis domain (28 studies) due to absent optimism adjustment, unclear outcome determination and timing, and a limited number of outcome events per candidate predictor (EPV). The review maps risk-of-bias domains onto CRISP-DM phases (Figure 11), locating methodological issues within specific data mining stages. Assessed by two independent evaluators using PROBAST's 20 signaling questions across 4 domains (participants, predictors, outcome, analysis); the predictors domain achieved low risk of bias in all papers.

Perspective

The review's conclusions apply to the scope of machine learning research on noninvasively estimating hemoglobin or screening for anemia, drawing on English-language literature from PubMed, Web of Science, and Scopus between January 1, 2019 and March 27, 2025; inclusion required a self-designed and implemented machine learning algorithm, full-text availability, and English language, while invasive methods, animal studies, monitoring of already diagnosed patients, evaluation of ready-to-use commercial models, and protocols, case reports, reviews, meta-analyses, conference proceedings, preprints, and industry reports were excluded. Organized by CRISP-DM phases, it can guide algorithm and acquisition-site selection, light source and hardware design, choice of evaluation metrics, and mobile deployment thinking; the authors note that quantitative synthesis such as meta-analysis would require a separate study design.

Readers should keep in mind that included studies differ widely in task definition, population, device, and metrics, so cross-study medians and interquartile ranges describe distributions rather than enabling direct comparisons of superiority; the authors note that high risk of bias in the analysis domain stems mainly from absent optimism adjustment, so near-optimal metrics may be inflated by overfitting on validation data and need independent external validation; most studies did not specify the underlying etiology, and differential diagnosis between iron deficiency anemia and anemia of chronic disease remains underexplored; calibration metrics are underused, so regression outputs near clinical decision thresholds are best interpreted as intervals rather than precise values; public datasets remain limited (such as Eyes-defy-anemia and UK Biobank), and building and releasing public multimodal datasets is still an open task; moreover, this is a text-level review, and specific numeric distribution details in the figures should be checked against the original figures.

Sources