Bangla Medical NER Benchmark: Fine-Tuned XLM-RoBERTa Sets New F1 of 0.5959, While Language-Specific BanglaBERT Trails at 0.4937
Synopsis
This work benchmarks three fine-tuned transformer encoders (BanglaBERT, mBERT, XLM-RoBERTa) against GPT-4o mini under zero-shot and few-shot prompting for Bangla medical named entity recognition across the full test set of 3,179 samples, where fine-tuned XLM-RoBERTa reaches an F1 of 0.5959, surpassing the previously reported best of 0.5848, while the language-specific BanglaBERT reaches only 0.4937, and fine-tuned models outperform the optimal prompting configuration by a factor of 3.76.
Figure 1. Overview of the unified experimental framework, showing
arXiv · Page 3Interpretation
Fine-tuned XLM-RoBERTa achieves an F1 of 0.5959 on Bangla medical NER, surpassing the previously reported best result of 0.5848 and establishing a new state of the art. Compared with prior work, this result provides a reproducible baseline on the full test set rather than an estimate on a small subset. A large-scale evaluation across the full test set of 3,179 samples, which the authors describe as statistically robust and reproducible.
The language-specific BanglaBERT consistently underperforms its multilingual counterparts with an F1 of 0.4937, indicating that pretraining domain diversity can outweigh language specificity in highly specialized clinical settings. This comparison directly challenges the intuition that a language-specific model must be better for low-resource language tasks, placing language specificity and domain diversity under the same benchmark. A controlled evaluation of three encoders on the same task and test set, with a clear gap between BanglaBERT and the multilingual models.
Per-entity-type analysis shows that Medicine and Specialist categories are recognized reliably with F1 scores above 0.83, while the Symptom category is the most challenging at an F1 of 0.4367 despite being the most frequent training class. It provides a fine-grained per-entity-type performance profile for this task, revealing that high frequency does not necessarily mean high recognition. Per-entity-type F1 statistics computed over the full test set.
Fine-tuned transformer models outperform the optimal prompting configuration by a factor of 3.76, showing that prompt-only pipelines remain inadequate for structured clinical entity extraction in low-resource language environments. It directly quantifies the comparison between the fine-tuning paradigm and GPT-4o mini zero-shot and few-shot prompting configurations, yielding a multiple-fold gap. A comparison of fine-tuned models against the optimal prompting configuration on the same task, with a 3.76-fold difference.
Perspective
This benchmark targets the specific task of Bangla medical named entity recognition and applies to research and engineering settings involving low-resource languages and clinical entity extraction, providing a reproducible comparison baseline for subsequent model selection and evaluation. Its conclusions pertain to the three evaluated encoders and the GPT-4o mini prompting configurations, as well as to this task's entity-type system.
Readers should still note that the Symptom category has the lowest F1 despite being the most frequent training class, and the reasons are not elaborated in the text; whether the trade-off between language specificity and domain diversity holds in other languages or clinical tasks awaits further evaluation; moreover, the loaded text is abstract-level content and does not include data construction details, annotation guidelines, or statistical testing information, which require consulting the original paper.
