Skip to main content
Back to timeline
bioRxivSource publication:

Hierarchical Temporal Transformer for Cancer Grade Prediction and Cross-Cancer Transfer Learning from Pathology Reports

Synopsis

This work presents the Hierarchical Temporal Transformer (HTT), a two-level architecture in which level 1 encodes each report with BiomedBERT adapted by LoRA and level 2 is a temporal transformer that reads a patient's full report sequence using a continuous-time positional encoding built from the measured number of days between visits plus learnable cancer-type embeddings; on a controlled synthetic corpus of sequential radiology reports it reaches validation AUROC 0.942 versus 0.881 for a single-report baseline and transfers to held-out pancreatic cancer at 0.995 versus 0.949, and on 4,786 real pathology reports from TCGA-Reports spanning 14 cancer types it predicts tumor grade for three types withheld entirely from training with AUROC 1.000 on thyroid carcinoma, 0.960 on sarcoma and 0.

Interpretation

HTT uses a two-level structure to address both the temporal evolution of report sequences and cross-cancer generalization: level 1 encodes each report with LoRA-adapted BiomedBERT, and level 2 is a temporal transformer that reads a patient's full report sequence. Relative to existing language models that read each report in isolation, this architecture explicitly models how a patient's disease changes across visits and supplies per-family conditioning through learnable cancer-type embeddings. The architecture is described in full with a clear division between the two levels, and the two capabilities are tested separately on a synthetic corpus and a real corpus.

On a controlled synthetic corpus of sequential radiology reports, HTT reaches validation AUROC 0.942 against 0.881 for a single-report baseline, and transfers to held-out pancreatic cancer at 0.995 against 0.949, while the two models are indistinguishable on a 60-patient test set. This experiment tests temporal modeling and cross-cancer transfer separately, since no fully open corpus contains longitudinal reports for many cancer types. Progression phrases in the synthetic corpus are inserted from templated trajectories, a controlled setting, and the pancreatic transfer result is reported alongside the indistinguishability on the 60-patient test set.

On 4,786 real pathology reports from TCGA-Reports spanning 14 cancer types, HTT predicts tumor grade for three types withheld entirely from training, reaching AUROC 1.000 on thyroid carcinoma, 0.960 on sarcoma and 0.808 on lung squamous cell carcinoma, with mean held-out AUROC 0.923 equal to the in-distribution test AUROC of 0.923. Existing evaluations typically cover only cancer types present in training data, whereas this result extends evaluation to entirely held-out cancer families with no measurable transfer penalty observed. Based on a real pathology report corpus of 4,786 reports across 14 cancer types, with a reported comparison between held-out and in-distribution AUROC.

Ablation on the real corpus shows the transfer is carried by the pre-trained encoder rather than by the temporal components, which contribute 0.39 AUROC points with one report per patient. The ablation attributes the source of transfer to the pre-trained encoder, suggesting grade-related pathological language may be learnable in a cancer-type agnostic way. The ablation is run on the real corpus and explicitly reports the temporal components' contribution in the single-report setting.

Perspective

The result targets predicting tumor grade from pathology reports and transferring across cancer types, applying to real corpora with either multi-visit report sequences or one report per patient, and to the controlled synthetic corpus used to test the two capabilities separately; its significance lies in pointing toward unified cancer NLP systems that require no per-type retraining.

A careful reader would still watch: on the real corpus with one report per patient the temporal components contribute only 0.39 AUROC points, so the role of temporal modeling on real longitudinal data awaits more evidence; progression phrases in the synthetic corpus are inserted from templated trajectories, leaving open how far those conclusions extend to real clinical trajectories; and this load is a summary-level text missing figures and full method details, which may affect a complete grasp of the experimental settings and result details.

Sources