Skip to main content
Back to timeline
arXivSource publication:

MITE instruction-tunes with Python, C++, and Java code formats and ensembles by voting, consistently beating BERT and LLM baselines on six BioNER datasets

Related research and updates

Synopsis

The work proposes MITE, which reformulates biomedical named entity recognition as a structure-to-structure generation task by rendering each training instance in multiple programming-language formats such as Python, C++, and Java to supply structurally diverse supervision without external biomedical knowledge or extra annotations, and at inference aggregates predictions from the different code formats through entity-level voting; on six widely used BioNER datasets MITE consistently outperforms representative BERT-based and LLM-based baselines and shows strong cross-dataset generalization, with ablation and parameter analyses further supporting the effectiveness and robustness of the proposed components.

Source-provided article image: Enhancing Biomedical Named Entity Recognition via Multiple Programming Languages Instruction Tuning and Ensemble Method
Fig. 1 ·

Fig. 1: Examples of performing structured BioNER task with NL and Code instructions. In contrast to prompting with NL instructions, we utilize structured Code instructions to mitigate the output discrepancy between the pretraining and inference stages.

arXiv

Interpretation

MITE recasts BioNER from flat textual sequence labeling under natural-language instructions into a structure-to-structure generation task in which both instructions and entity outputs are carried in code-formatted representations. Existing instruction-tuning approaches typically serialize annotations as flat text, offering limited structural constraints for typed entity extraction; MITE replaces that serialization with code formats. Method description at the abstract level, stating that instructions and entity outputs are represented in code format; template details are not given.

Each training instance is transformed into multiple programming-language formats, including Python, C++, and Java, providing structurally diverse supervision while preserving the same underlying entity semantics. Prior work often alleviates annotation scarcity by introducing external biomedical knowledge, which requires costly resource construction; MITE obtains diverse supervision without external knowledge or additional annotations. Method description at the abstract level, explicitly listing Python, C++, and Java and emphasizing that semantics are preserved.

At inference, an entity-level voting strategy aggregates predictions from the different code formats, reducing language-specific prediction variance and improving robustness. Learning from a single serialized output form restricts structural diversity and reduces model robustness; the voting ensemble targets variance from that source. Method description at the abstract level; the concrete voting implementation and statistics are not reported.

On six widely used BioNER datasets, MITE consistently outperforms representative BERT-based and LLM-based baselines and exhibits strong cross-dataset generalization, with ablation and parameter analyses supporting the effectiveness and robustness of the components. Relative to existing BERT-based and LLM-based baselines, the work reports consistent advantages across multiple datasets and examines component contributions through ablation and parameter analysis. Experimental conclusions at the abstract level; specific dataset names, metric values, and effect sizes are not provided.

Perspective

The work targets biomedical named entity recognition, a structured extraction task, and fits settings where annotation resources are limited but LLM instruction tuning is desired; its method setting renders training instances into code formats such as Python, C++, and Java while keeping entity semantics consistent, and aggregates multi-format predictions by entity-level voting at inference, so it needs no external biomedical knowledge base or additional manual annotation. The conclusions described in the abstract rest on six widely used BioNER datasets and report cross-dataset generalization, indicating a multi-dataset, cross-dataset evaluation setting. For biomedical text-mining researchers and practitioners seeking to lower annotation and knowledge-construction costs, this route offers a directly reusable paradigm for supervision construction and ensemble inference.

The visible text is only an abstract, lacking specific dataset names, metric values, baseline configurations, voting-strategy implementation, and concrete ablation and parameter-analysis results, so the size of the advantage and which entity types or text domains benefit most cannot be judged. How semantic equivalence across programming-language formats is guaranteed, and whether code-format supervision is sensitive to model scale or pretraining corpus, remain open questions worth watching. The abstract also does not state performance on non-biomedical domains or non-entity-extraction tasks, so the transferable scope awaits further verification.

Sources