Tumour region identification guided scoring (TRIGS) and foundation model-based Tumour Infiltrating Lymphocyte scoring are prognostic for pathological complete response/event free survival in the triple negative patients in the PARTNER randomized controlled trial
Synopsis
This study updated an automated tumour infiltrating lymphocyte (TIL) assessment pipeline, proposing TRIGS aligned with clinical scoring guidelines and foundation model-based SAM-TIL, and showed in 166 neoadjuvantly treated patients in the TransNEO cohort that they predicted pathological complete response with odds ratios of 1.95 (95% CI 1.22-3.03, p=0.005) and 2.32 (95% CI 1.43-3.77, p=0.001), and in 277 triple negative and HER2-positive patients in The Cancer Genome Atlas showed overall survival hazard ratios of 0.79 (95% CI 0.63-1.00, p=0.05) and 0.80 (95% CI 0.67-0.97, p=0.02), with correlation to gold standard clinical assessment of 0.59-0.69 and no substantial difference from gold standard assessment in predicting pathological complete response (AUC 0.60-0.
Interpretation
The study proposed and updated an automated TIL assessment pipeline, including tumour region identification guided scoring (TRIGS) aligned with clinical scoring guidelines and foundation model-based lymphocyte detection (SAM-TIL), and compared it with alternative methods (HoVerNet, muTILs) and gold standard clinical assessment. Relative to previously established automated TIL assessment pipelines, this work aligns the scoring pipeline with clinical scoring guidelines and introduces a foundation model for lymphocyte detection, while including multiple alternative methods for comparison. The study evaluated these methods across multiple cohorts, including 166 neoadjuvantly treated patients in the TransNEO cohort, 277 triple negative and HER2-positive patients in The Cancer Genome Atlas, and 285 patients from the PARTNER randomized trial, reporting odds ratios, hazard ratios, correlations, and predictive performance metrics.
TRIGS and SAM-TIL had predictive value for pathological complete response, with odds ratios of 1.95 (95% CI 1.22-3.03, p=0.005) and 2.32 (95% CI 1.43-3.77, p=0.001), respectively. These results provide quantitative evidence that automated TIL scoring methods predict pathological complete response in the neoadjuvant setting, complementing prior literature on TILs as a prognostic biomarker. Based on analysis of 166 neoadjuvantly treated patients in the TransNEO cohort, reporting odds ratios, 95% confidence intervals, and p-values.
TRIGS and SAM-TIL had prognostic value for overall survival, with hazard ratios of 0.79 (95% CI 0.63-1.00, p=0.05) and 0.80 (95% CI 0.67-0.97, p=0.02), respectively. These results extend the prognostic value of automated TIL scoring methods to the overall survival endpoint in 277 triple negative and HER2-positive patients in The Cancer Genome Atlas. Based on analysis of 277 triple negative and HER2-positive patients in The Cancer Genome Atlas, reporting hazard ratios, 95% confidence intervals, and p-values.
Despite differences from gold standard clinical assessment, no substantial difference was observed between the tested methods and gold standard assessment in predicting pathological complete response (AUC 0.60-0.64) or event free survival (Integrated Brier score approximately 0.10), suggesting AI tools that do not follow manual scoring guidelines could be validated and subsequently used. This finding suggests that automated TIL scoring methods may achieve predictive performance similar to gold standard assessment even when not aligned with manual scoring guidelines, offering a reference for validation pathways of AI tools. Based on correlation analysis in 285 patients from the PARTNER randomized trial (0.59-0.69) and comparison of predictive performance metrics (AUC 0.60-0.64, Integrated Brier score approximately 0.10).
Perspective
The results apply to automated assessment of TIL scoring in HER2-positive and triple negative breast cancer patients, particularly for predicting pathological complete response in the neoadjuvant setting and overall survival in triple negative and HER2-positive patients. The study suggests automated methods could be validated and subsequently used, but lymphocyte detections are not interchangeable with gold standard clinical assessment and therefore cannot support pathologist assessment.
Correlation between automated methods and gold standard clinical assessment was 0.59-0.69, indicating differences; the study suggests further validation for independent use or improvement of prognostication is required. Readers may watch for consistency of results across cohorts, robustness of the foundation model under different staining and scanning conditions, and practical integration of automated scoring into clinical workflows.
