OmniTCR: a foundation model unifying T cell receptor recognition prediction and conditional sequence generation
Synopsis
The study presents OmniTCR, a 113-million-parameter autoregressive foundation model pretrained on 328 million formatted human immune-sequence records that uses sequence-type tokens and complementary component orders to jointly learn from individual TCR chains and partial or complete TCR-pMHC associations, thereby performing both TCR recognition prediction and conditional sequence generation within one model, achieving AUPRCs of 0.7009 for peptide-TCRβ recognition and 0.8235 for TCR-pMHC interaction prediction on unseen epitopes and a mean AUROC of 0.9436 in distinguishing cancer from healthy repertoires across 11 independent pan-cancer cohorts.
Interpretation
OmniTCR unifies TCR recognition prediction and receptor generation in a single autoregressive foundation model rather than modelling them separately as is traditionally done. The text states that recognition prediction and receptor generation are 'traditionally modelled separately,' leaving large TCR sequence collections disconnected from smaller TCR-peptide-MHC datasets; this work connects the two tasks and the two data types using sequence-type tokens and complementary component orders. The methodological basis is a 113-million-parameter model pretrained on 328 million formatted human immune-sequence records, which is evidence at the level of model and pretraining scale.
On unseen epitopes, OmniTCR's recognition prediction outperforms the strongest evaluated comparators. The text reports AUPRCs of 0.7009 for peptide-TCRβ recognition and 0.8235 for TCR-pMHC interaction prediction, exceeding the strongest evaluated comparators by 0.3396 and 0.3451, respectively. The evidence comes from AUPRC metrics on 'unseen epitopes' and the margins over the strongest evaluated comparators, which is evidence at the benchmark-evaluation level.
OmniTCR can distinguish cancer from healthy repertoires and achieves the highest sequence recovery on generation tasks. The text reports a mean AUROC of 0.9436 across 11 independent pan-cancer cohorts and the highest sequence recovery on internal and external generation benchmarks; structural modelling additionally supported the plausibility of selected pMHC-conditioned CDR3β candidates. The evidence includes AUROC across 11 independent cohorts, sequence recovery on internal and external generation benchmarks, and structural modelling support for selected candidates.
Perspective
The work is aimed at computational immunology and receptor design settings, applies to human immune sequence data, and covers individual TCR chains as well as partial or complete TCR-pMHC associations; its recognition prediction results concern unseen epitopes, its repertoire discrimination results concern 11 independent pan-cancer cohorts, and its generation results concern internal and external generation benchmarks. For researchers and designers who want to perform TCR recognition prediction and conditional sequence generation within one framework, this model provides a reusable foundation.
The available text is an abstract-level overview without figures or supplementary material, so the specific composition of the pretraining data, the details of the comparators on each benchmark, how the structural modelling was evaluated, and the extent of experimental validation for generated candidates remain open questions a reader would need to check in the original text.
