PIRTA pairs image-domain retrieval with text-domain augmentation, lifting ischemic-territory accuracy in 3D brain MRI reports by up to 57.2 points across cohorts
Synopsis
The study proposes PIRTA, a retrieval-augmented generation framework in which a 3D ViT encoder pretrained with large-scale MAE self-supervision retrieves clinically similar 3D DWI/ADC volumes in the image domain, and their paired clinician-authored reports ground LLaMA3-8B-Instruct report generation, avoiding explicit image-text alignment; on a 1,831-case multi-institutional internal cohort, a 580-case privacy-preserving external cohort, and the 206-case public ISLES benchmark, PIRTA achieves strong image-domain retrieval (internal mAP@1 of 94.0%) and consistently improves ischemic-territory accuracy, a clinically grounded surrogate for factuality, over direct image-to-text 2D baselines, leading the strongest 2D baseline by +57.2 / +34.5 / +30.6 multi-class accuracy points.
Fig. 1. Conceptual comparison of cross-modal alignment strategies. (a) Conventional methods align image and text encoders in a shared representation space, requiring joint modeling across modalities. (b) PIRTA avoids explicit cross-modal alignment by retrieving similar images in the image domain, and using their paired clinician-authored reports for text-domain augmentation.
· Page 2Interpretation
PIRTA reformulates 3D MRI report generation as image-domain retrieval plus text-domain augmentation, training only an image encoder and reducing the hypothesis space from the joint P(x,y) to the marginal P(x), thereby avoiding explicit cross-modal alignment. Prior report generation relied on supervised learning with extensive annotation or segmentation pipelines requiring joint modeling of image and text encoders; PIRTA instead retrieves similar cases in image space and uses their paired clinician reports as external evidence to drive LLM generation. The paper formalizes the contrast between joint hypothesis complexity C(H)=C(H_image)+C(H_text) and PIRTA's H_PIRTA=H_image, and reports retrieval and generation results on internal and two external cohorts.
Large-scale MAE self-supervised pretraining substantially improves cross-institution retrieval generalization of the 3D DWI/ADC encoder. The paper reports that as pretraining scale grows from random initialization to the 1,831-case internal cohort to 38,532 UK Biobank subjects, retrieval metrics improve most on external cohorts, e.g. ISLES mAP@1 rising from 38.83% to 70.87%. Table 1 reports mAP@1/5/10 and Acc@1/5/10 for the internal cohort, External Cohort 1, and ISLES; Table 2 uses four-class territory classification with the same encoder as a sanity check of representation quality, showing parallel trends.
Using ischemic-territory accuracy as a clinically actionable factuality surrogate, PIRTA outperforms 2D multimodal baselines on the internal cohort, BRMH, and ISLES. The paper argues that n-gram metrics such as BLEU and ROUGE miss clinically important factual errors, motivating ischemic-territory accuracy, which directly informs reperfusion triage. Fig. 4 shows PIRTA leading the strongest 2D baseline by +57.2 / +34.5 / +30.6 multi-class accuracy points, with per-class gaps persisting across territories; the qualitative case in Table 3 shows higher retrieval similarity aligning with closer agreement to the expert reference.
Retrieval similarity can serve as a per-case confidence cue to flag unreliable retrievals for radiologist review. The paper positions the explicit cosine-similarity score as a per-case confidence cue, directly linking retrieval quality to report factuality. The paper reports that higher retrieval similarity consistently corresponds to better factual accuracy, and that failures concentrate at low-similarity, anatomically rare infarct patterns.
Perspective
The work targets radiology report generation for acute ischemic stroke from 3D DWI/ADC, suited to institutions that hold paired image-report databases and can build a retrieval corpus; its conclusions rest on a multi-institutional internal cohort, a privacy-preserving external cohort, and the public ISLES benchmark, with ischemic-territory accuracy as the evaluation axis. For readers, it points to a reusable path: pretrain a 3D encoder on large-scale unlabeled MRI, then ground LLM generation in clinician reports retrieved in the image domain, reducing optimization complexity under data constraints and offering similarity scores as a per-case confidence cue for radiologist review.
The paper itself notes that factual accuracy degrades when rare or atypical infarct patterns are underrepresented in the retrieval database, yielding low similarity and weak grounding; territory accuracy is explicitly framed as a focused factuality measure rather than a complete report-quality metric, and does not directly score lesion size, perfusion mismatch, hemorrhagic transformation, or report fluency. Planned evaluation axes include finer-grained ASPECTS-style per-region scoring, prospective clinical evaluation with radiologist-rated factuality on held-out reports, and hybrid multimodal retrieval that augments image embeddings with clinical metadata such as NIHSS, presentation, and past medical history. Readers should watch whether these directions replicate in broader populations and real clinical workflows.
