A corrosion inhibitor dataset built from 5,597 publications
Synopsis
This work assembles a corrosion inhibitor dataset from 5,597 publications using a multi-agent extraction pipeline, releasing a manually verified IE_datasets.xlsx, LLM-extracted json_extracted.zip, and IE_PH_4_materials.xlsx for model construction, with fields covering reference metadata, inhibitor name and composition, anodic/cathodic/mixed type, film mechanism, SMILES, corrosion material name, grade, composition, processing and heat treatment, corrosive medium, medium type, concentration and temperature, and test temperature, test time, inhibitor concentration, test method and inhibition efficiency percentage.
Interpretation
It releases a corrosion inhibitor dataset covering 5,597 publications, with the main table IE_datasets.xlsx described as the manually verified and processed dataset file. Compared with individual experimental reports, it consolidates inhibitor and corrosion-evaluation information scattered across a large literature base into a unified table, and explicitly separates manually verified data from model-extracted data. The description lists a complete field inventory and states that IE_datasets.xlsx is the dataset file after manual verification and processing; no field-completeness or entry-count statistics are provided.
It provides json_extracted.zip containing files extracted by an LLM, and publicly releases the multi-agent extraction pipeline code covering schema handling, extraction, merging, review and evidence verification. By releasing both the extraction pipeline and its outputs, the extraction results can be traced back to specific pipeline stages rather than only presenting a final table. The steps state that the repository contains the multi-agent extraction pipeline and analysis scripts that regenerate every statistic and figure (including corpus ingestion, gold-standard alignment and scoring, evidence localization, schema ablation and the confidence model); document parsing used MinerU, language-model inference used deepseek-v4-flash at temperature 0.1 with a 65,536-token output budget, sequence alignment used RapidFuzz, and package versions are pinned.
It supplies IE_PH_4_materials.xlsx as the subset used for constructing models. Beyond the full dataset, it separately marks a curated version intended for modeling, making it directly usable in downstream modeling work. The description explicitly identifies this file as the dataset used for constructing models; it does not state the selection rules, sample size, or modeling task type.
Perspective
This dataset is intended for literature information consolidation and modeling preparation in corrosion inhibitor research, and suits researchers who need fields such as inhibitor chemical structure, corrosion material, medium conditions, and inhibition efficiency; IE_PH_4_materials.xlsx is positioned for model construction, so its applicable setting is limited to modeling.
The current text is a dataset description and does not give quantitative information such as entry counts, field completeness, extraction accuracy, or the proportion manually verified, nor does it state the selection rules and modeling task for IE_PH_4_materials.xlsx; these are open questions a reader would want to confirm before use.
