Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

OncoNoteBERT trains oncology-specific encoders on 290,000 UK outpatient notes: continued pretraining cuts perplexity to 2.10, while a from-scratch tokenizer shortens sequences by about 23.6%

Using 290,026 outpatient oncology notes from 21,564 patients treated for lung and head-and-neck cancer at a UK specialist oncology centre, the study compared RadBERT and PathologyBERT with two local strategies: OncoNote-RadBERT, produced by continued masked language model pretraining, reached a validation perplexity of 2.10, while OncoNoteBERT, trained from scratch with an oncology-specific WordPiece tokenizer, reached 2.83 but the most efficient tokenisation (subword fertility 1.224, normalised sequence length 0.7636, about 23.6% fewer tokens) and returned clinically acceptable predictions for 12 of 13 masked-token probes versus 7 of 13 for OncoNote-RadBERT; both local models represented the institutional placeholder [REMOVED] as a single learnable token.