Skip to main content
Back to timeline
arXivSource publication:

OncoNoteBERT trains oncology-specific encoders on 290,000 UK outpatient notes: continued pretraining cuts perplexity to 2.10, while a from-scratch tokenizer shortens sequences by about 23.6%

Related research and updates

Synopsis

Using 290,026 outpatient oncology notes from 21,564 patients treated for lung and head-and-neck cancer at a UK specialist oncology centre, the study compared RadBERT and PathologyBERT with two local strategies: OncoNote-RadBERT, produced by continued masked language model pretraining, reached a validation perplexity of 2.10, while OncoNoteBERT, trained from scratch with an oncology-specific WordPiece tokenizer, reached 2.83 but the most efficient tokenisation (subword fertility 1.224, normalised sequence length 0.7636, about 23.6% fewer tokens) and returned clinically acceptable predictions for 12 of 13 masked-token probes versus 7 of 13 for OncoNote-RadBERT; both local models represented the institutional placeholder [REMOVED] as a single learnable token.

Source-provided article image: OncoNoteBERT: A Foundation Representation Model for Natural Language Processing of Real-World Outpatient Oncology Notes
Figure 1 ·

Figure 1: Study workflow for data preparation, encoder development, and intrinsic representation evaluation. All evaluations were performed on the validation split of patient data and did not alter model parameters.

arXiv

Interpretation

Domain-adjacent encoders fit outpatient oncology language poorly in zero-shot evaluation: perplexity 113.04 for RadBERT and 2,035.03 for PathologyBERT. Prior oncology-related BERT models were designed for downstream tasks rather than as general-purpose language encoders; this work evaluates radiology- and pathology-domain encoders directly on a real outpatient oncology corpus at the intrinsic level. Evaluated on the held-out validation set using masked language modelling loss and perplexity, with RadBERT loss 4.728 and PathologyBERT loss 7.618; the authors note that the PathologyBERT comparison combines model and tokenizer fit.

Continued pretraining produced the strongest corpus-level language fit: OncoNote-RadBERT reduced validation loss from 4.728 to 0.7406, corresponding to a perplexity reduction from 113.04 to 2.097. Because OncoNote-RadBERT and RadBERT share the same tokenizer lineage apart from the added [REMOVED] special token, the authors treat this as the clearest evidence of continued pretraining. A comparison within the same tokenizer lineage; the authors note the frequent [REMOVED] token may also have contributed, and each model was trained once, so results are point estimates.

From-scratch training with an oncology-specific tokenizer gave the greatest tokenisation efficiency: OncoNoteBERT showed subword fertility 1.224, continuation-token rate 0.01781, and normalised sequence length 0.7636, requiring about 23.6% fewer tokens on the same sampled documents. This tokenizer represented specialised expressions such as ECOG, T3N2M0, chemoradiotherapy, cisplatin, durvalumab, dysphagia, hypopharynx, mucositis, and pembrolizumab as single tokens, whereas the inherited RadBERT tokenizer fragmented many of them into multiple subword pieces. Tokenizer metrics computed on a 5,000-document sample plus a targeted audit of oncology-related strings; unknown-token rates were negligible for RadBERT, OncoNote-RadBERT, and OncoNoteBERT (about 4e-8 to 6e-8) but higher for PathologyBERT at 7.42e-3.

Corpus-level fit and masked-token behaviour diverged: OncoNoteBERT returned a clinically acceptable top-five prediction for 12 of 13 prespecified fill-mask prompts, versus 7 for OncoNote-RadBERT, 2 for RadBERT, and 1 for PathologyBERT. The authors attribute this divergence partly to tokenizer fragmentation rather than learned semantics alone: a single masked position can yield only one vocabulary token, so expressions split into several WordPieces cannot be produced in full. The probe set comprised only 13 prespecified prompts with prespecified acceptable answers, and the authors frame the results as a qualitative audit rather than a tokenizer-independent measure of clinical knowledge or downstream performance.

Perspective

This work targets teams building reusable representation backbones in governed real-world settings, particularly those handling longitudinal, multipurpose outpatient oncology documentation while preserving institutional de-identification conventions. Its findings apply to the setting of a single UK specialist oncology centre and lung and head-and-neck cancer outpatient notes, with evaluation limited to the intrinsic level: held-out masked language modelling fit, tokenisation efficiency, clinical-term tokenisation, masked-token behaviour, and exploratory representation geometry. The authors position OncoNote-RadBERT and OncoNoteBERT as starting points for subsequent task-specific studies and list downstream directions including treatment and toxicity extraction, treatment-response classification, phenotype identification, tumour-stage extraction, and survival or treatment-decision modelling. Retaining [REMOVED] as a dedicated token lets the models learn the linguistic contexts in which redacted content occurred without access to the underlying identifiable information, offering a transparent way to represent an existing de-identification convention within a data-resident workflow.

Readers should still watch several points: the corpus came from a single UK specialist centre and was restricted to lung and head-and-neck cancer notes, so transferability across institutions, tumour groups, and documentation practices remains to be established; each local model was trained once with one architecture and one prespecified configuration, so reported outcomes are point estimates rather than averages across initialisations or hyperparameter settings; perplexity comparisons involving PathologyBERT and OncoNoteBERT reflect combined model-parameter and tokenizer effects and should not be treated as tokenizer-independent rankings; the fill-mask evaluation comprised only 13 prespecified prompts and did not capture the full range of clinically plausible completions; and the representation analyses were exploratory and do not establish model quality or clinical utility. In addition, [REMOVED] does not distinguish among categories of redacted information, so where category-specific semantics matter, future models could use separate tokens such as [NAME], [DATE], or [LOCATION], subject to de-identification system performance and governance requirements; retaining the token does not constitute a formal privacy guarantee. Future work could score complete candidate strings using tokenizer-specific pseudo-log-likelihood or multiple-mask procedures, and use methods such as centred kernel alignment for cross-model representation comparison.

Sources