Large language model-enabled automated data extraction for concrete materials informatics
Synopsis
This work introduces a modular, large language model (LLM)-powered agent pipeline that automatically extracts and structures composition–process–property attributes of concrete materials from tables and text in scientific publications, achieving F1 scores up to 0.98 across 17 open and proprietary models and, within about one hour, extracting over 10,000 records from 278 papers screened from more than 27,000 publications; after postprocessing this yields the largest open laboratory database for blended cement concrete with nearly 9,000 high-quality records and over 100 attributes, and machine learning analyses indicate that large, diverse, information-rich datasets improve both in-distribution accuracy and out-of-distribution generalization to unseen materials systems.
Fig. 1 | Survey of dataset sizes used in peer-reviewed publications applying
· Page 2Interpretation
A sequential, modular chain of LLM-based extraction and processing agents handles both tables and text and performs cleaning steps such as acronym expansion, header consolidation, unit normalization, naming standardization, and record-level merging. Compared with traditional text-mining approaches such as ChemDataExtractor and MatSciBERT that rely on fixed rules or large human-annotated corpora, the pipeline adapts to heterogeneous reporting formats without domain- or task-specific fine-tuning and extends prior LLM applications, which focused mainly on text extraction, to tables and cross-passage information integration. Benchmarked against 3,015 records manually extracted from 58 randomly selected publications spanning journals and publishers, across five target categories and 17 LLMs; most models reached F1 of 0.90 or higher per category, all 17 exceeded 0.85 overall, 12 exceeded 0.90, 8 surpassed 0.95, and the best (Claude Sonnet 4.5) reached 0.98; GPT-4o achieved overall precision 0.98, recall 0.96, and F1 0.97.
The work constructs the largest openly available laboratory database for blended cement concrete compressive strength, with 8,979 unique high-quality records and over 100 attributes covering fly ash, blast-furnace slag, silica fume, limestone powder, and calcined clay. The previously largest reported dataset (Jiang et al.) contained 5,026 records but was not publicly released; the largest openly available dataset (Imran et al.) had 2,171 records, and the most widely used Yeh dataset only 1,030 records from 17 publications with two supplementary cementitious materials; this database draws on 278 publications, contains 3,121 unique mixtures, spans water-to-binder ratios of 0.09–2.2, and systematically includes oxide compositions, loss on ignition, specific gravity, and Blaine fineness of raw materials. Derived from 10,313 automatically extracted records after removing duplicates, non-target supplementary cementitious materials, and abnormal total masses; more than 120 publications were cross-checked, and together with the 58 manually extracted benchmark publications this covers over 7,500 records, more than 80% of the final database.
Machine learning analyses show that including chemical and physical descriptors of binders yields slight to modest in-distribution accuracy gains, with CaO content and Blaine fineness among the influential features alongside curing age and water-to-binder ratio. Prior concrete strength datasets generally lack binder chemical and physical attributes (surveys indicate more than 90% of publications' datasets lack such descriptors); this work is the first to systematically evaluate their contribution at a scale of nearly 9,000 records. Evaluated with random 80/20 splits over 10 random seeds, comparing XGBoost, random forest, LightGBM, multilayer perceptron, support vector machine, and a linear regression baseline; adding five major oxides and Blaine fineness produced slight to modest improvements across models, and SHAP analysis placed CaO content among the most important features.
Training-data-size and out-of-distribution experiments show that predictive accuracy improves approximately as a power law with training data size, that adding even a small number of test-domain samples markedly improves generalization to ternary and quaternary blends, and that chemical and physical descriptors reduce the number of additional experimental samples needed. The work transfers neural scaling-law observations from deep learning to concrete strength prediction and quantifies how descriptors lower experimental cost, providing actionable evidence for data-driven concrete design. The out-of-distribution task trains only on plain portland cement and binary blends and tests on ternary and quaternary systems; XGBoost with base features exceeded 14 MPa RMSE even with the full training set, and adding test-domain samples improved performance noticeably, with an RMSE of 10 MPa requiring about 50 additional samples when descriptors were used versus more than 100 without them.
Perspective
The pipeline targets full-text articles in XML or HTML, focuses on blended cement systems with mixture proportions reported in mass per unit volume, and extracts binder properties, mixture proportions, curing conditions, specimen dimensions, and compressive strength; the database covers five supplementary cementitious materials: fly ash, blast-furnace slag, silica fume, limestone powder, and calcined clay. The authors note the framework can be extended to other concrete technologies and materials systems and applied continuously to newly published literature for ongoing database expansion and updates.
The pipeline still makes occasional arithmetic and unit-conversion errors, for example interpreting '0.5% of binder mass' as '0.5 × binder mass' or converting '0.56 m2/g' incorrectly to '5.6 m2/kg'; complex multi-level headers and merged cells can misalign values into wrong columns; and very large tables may be truncated by maximum output token limits. On the literature side, challenges include undefined acronyms (e.g., 'FA' may mean fly ash or fine aggregate), missing key information (nearly 30% of publications do not report binder properties in tables, 36% and 16% lack explicit curing conditions and specimen dimensions, 34% do not report mixture proportions in tables, and more than 60% do not report compressive strength in tables, often placing it in figures), and source-publication errors such as inconsistent units, typographical mistakes, and data entry errors. The authors also note that the evaluation is end-to-end and does not isolate error contributions from individual intermediate steps, and that information in figures and supplementary files is not yet covered.
