Autogenerating a Domain-Specific Question-Answering Data Set to Enable High-Performing Language Models for Magnetic Materials
Synopsis
This paper presents a method for autogenerating domain-specific question-answering data, builds MagQA with 168,080 magnetic-materials QA pairs, and uses it to finetune BERT-style models; on a manually annotated magnetic-materials test set, a vanilla BERT-base-cased finetuned on MagQA mixed with SQuAD v2 (MagBERT_MagQA_Mixed) achieves the best result with an F1 of 78.43% and an exact-match score of 72.84%, indicating that, given sufficiently large and high-quality domain-specific QA data, domain-specific BERT models need only be finetuned from vanilla BERT without domain-adaptive pretraining.
MagBERT workflow presented in this study. Abbreviations used for visual clarity denote Artificial Intelligence (AI), Bidirectional Encoder Representations from Transformers (BERT), ChemDataExtractor (CDE), Domain-Adaptive Pretraining (DAPT), Question-Answering (QA), our magnetic QA data set (MagQA) developed herein and the Stanford Question-Answering Data set, version 2 (SQuAD v2) that contains QA pairs that have been constructed from generic English language.
PubMedInterpretation
It proposes a pipeline that autogenerates domain-specific QA data sets and uses it to build MagQA, containing 168,080 magnetic-materials QA pairs. Addressing the scarcity and high manual-annotation cost of domain-specific QA data, it offers a scalable autogeneration approach rather than relying on item-by-item human writing. The paper supports this with the data set size (168,080 QA pairs) and the subsequent finetuning experiments, making it method-plus-data-construction evidence.
On magnetic-materials QA, a vanilla BERT-base-cased finetuned on MagQA mixed with SQuAD v2 (MagBERT_MagQA_Mixed) performs best, with an F1 of 78.43% and an exact-match score of 72.84%. Among the six BERT-base-cased models compared, mixing domain data with general QA data outperforms configurations using a single data set, and the best model did not undergo domain-adaptive pretraining. The result is based on a manually annotated magnetic-materials test set and reports both F1 and exact-match metrics, making it an empirical result with an explicit evaluation benchmark.
Domain-adaptive pretraining is not a necessary condition for domain-specific BERT models: a vanilla BERT-base-cased finetuned only on MagQA mixed with SQuAD v2 reaches the best performance. Compared with the three models that first underwent domain-adaptive pretraining on a corpus of 97,308 magnetic papers before finetuning, this finding indicates that when sufficiently large and high-quality domain QA data are available, the domain-adaptive pretraining step can be omitted. The paper derives this from comparing the relative performance of three vanilla models and three domain-adaptively pretrained models, making it controlled-comparison evidence.
Through an ablation study using four sizes of QA data sets, it examines how much data is needed for sufficient domain specification, and shows the approach can be extrapolated to larger architectures such as RoBERTa-base and DeBERTa-base. Beyond a single best configuration, it provides empirical observations on the relationship between data scale and domain specification, and shows the pipeline is not limited to BERT-base. The ablation covers four data sizes, while the extrapolation is an architecture-level feasibility demonstration with weaker evidence strength than the main comparison experiments.
Perspective
The results are aimed at QA in the magnetic-materials domain and apply to settings where sufficiently large, high-quality domain QA data are available and where one wants to build domain-specific BERT-style models at lower compute cost; the method is shown to extrapolate to RoBERTa-base and DeBERTa-base, but the main evaluation remains benchmarked on BERT-base-cased, and the best configuration relies on mixing MagQA with SQuAD v2.
Details of quality control for the autogenerated QA pairs, the size and annotation consistency of the manually annotated test set, and the specific values of the four data sizes in the ablation are not expanded at the abstract level; in addition, how the coverage and domain boundaries of the 97,308-paper domain-adaptive pretraining corpus affect the extrapolation of the conclusions remains a question worth watching.
