Skip to main content
Back to timeline
arXivSource publication:

A Scalable Framework for Automated NER Annotation Correction in Low-Resource Languages

Synopsis

The work presents a multi-step framework that first corrects word segmentation with ChatGPT-4o, then fills missing entities via context-free lexicons with a minimum frequency of three, and finally applies frequency-based iterative self-training with a dual threshold on logits probabilities and the 90th percentile of self-attention scores, using a multilingual XLM-RoBERTa-large model to select candidates by F1; on manually revised validation/test splits for Urdu MK-PUCIT, Shahmukhi (Western Punjabi), and Sindhi SiNER, fine-tuning XLM-RoBERTa-large yields test micro-F1 gains of 3.96, 1.40, and 1.44 points, while ChatGPT-4o zero/few-shot NER remains below the supervised model.

Source-provided article image: A Scalable Framework for Automated NER Annotation Correction in Low-Resource Languages
Figure 1 ·

Figure 1: Framework for NER Dataset Correction: An iterative process leveraging word segmentation correction, context-free entity augmentation, and iterative self-training with dual-threshold validation to enhance the annotation quality.

arXiv

Interpretation

Proposes a frequency-based iterative self-training correction pipeline to fill missing annotations in low-resource NER datasets. Unlike prior correction work focused on high-resource corpora such as OntoNotes 5.0 and CoNLL-03 that relies on external resources like Wikipedia and expert annotators, the pipeline targets low-resource languages without manual re-annotation. Evaluated on three manually revised validation/test sets, with test micro-F1 gains of 3.96, 1.40, and 1.44 points; best macro-F1 occurs at iteration 5 for MK-PUCIT and SiNER and iteration 9 for Shahmukhi.

Introduces a dual-threshold mechanism combining a logits probability threshold with a dynamic attention threshold (90th percentile of diagonal self-attention scores) to reduce propagation of erroneous annotations. Confidence filtering is applied jointly to label probabilities and self-attention scores, paired with permutation-based candidate generation and F1-based validation by a multilingual model to choose the final label sequence. Appendix A.2 compares detected entity counts across attention percentiles 0.75-0.95 on a Shahmukhi subset (e.g., 359 at 0.95, 424 at 0.90, 451 at 0.75), motivating the 0.90 percentile choice.

Demonstrates the effect of word segmentation correction on transformer-based NER and shows that causal LLMs can improve the quality of non-English datasets. Highlights that space inconsistencies in Perso-Arabic script lead to different SentencePiece tokenizations (an example sentence changes from 21 to 17 tokens), and uses ChatGPT-4o few-shot prompting to correct word boundaries while preserving annotations. All three datasets improve on validation micro-F1 after word segmentation correction (MK-PUCIT 76.31 to 77.08, Shahmukhi 79.40 to 81.14, SiNER 86.70 to 87.11).

Prepares and releases manually reviewed validation and test sets for the three datasets as unbiased evaluation benchmarks. Evaluation sets are kept independent of the label propagation process, and inter-annotator agreement is computed on 100 reference sentences (Cohen's Kappa between 0.9440 and 0.9765). Validation/test set sizes and entity statistics appear in Table 2, and Appendix A.3 reports entity-count changes before and after manual correction.

Perspective

The framework targets low-resource NER datasets in Perso-Arabic script with PER/LOC/ORG entity types, suited to settings where large noisy datasets need quality improvement without manual re-annotation; the paper notes the three selected languages are topologically related and culturally similar, sharing similar person names, locations, and organizations, so cross-lingual representation may help. The method depends on entity lexicons (minimum frequency of three) and a multilingual XLM-RoBERTa-large validator, and per-iteration dataset versions are released for downstream use.

The paper notes that automated annotation quality depends on the self-trained model and the validation process, that errors in initial pseudo-labels can be reinforced across iterations, and that the fully automated pipeline lacks the contextual understanding human annotators provide, so enhanced datasets may still contain annotation inconsistencies; moreover, increased entity counts alone do not confirm each added label is correct, and improved F1 on clean test sets provides supporting evidence. Readers may further watch the stability of the dual thresholds across languages or entity types, the fact that the attention threshold sensitivity analysis was conducted only on a Shahmukhi subset, and how the pipeline performs on languages that are not topologically related.

Sources