CoNN-AL reaches R² 0.967–0.982 in automotive glass run channel design using about 2.3% of labels via stream-based active learning
Synopsis
The study proposes CoNN-AL, which adds stream-based active learning to the Cooperative Neural Network with Denoising Autoencoder (CoNN-DAE) and uses Monte Carlo dropout uncertainty to decide in real time which incoming candidates are worth labeling for data-efficient partial inverse design; on a real-world automotive glass run channel dataset of more than 900,000 unique simulated designs, only 20,000 actively selected labels (about 2.3% of the training pool) yield R-squared of 0.967 to 0.982 across all missing-variable levels, approaching upper-bound models trained on far more data, and at harder levels it reaches R-squared of at least 0.95 with 30 to 40% fewer labels than random sampling.
Interpretation
CoNN-AL layers stream-based active learning onto CoNN-DAE, using Monte Carlo dropout predictive uncertainty to select in real time which incoming candidate samples enter labeling, so the limited labeling budget goes to the most informative designs. Prior partial inverse design relied on large simulation-labeled datasets; this framework moves the labeling decision into the data stream rather than building a dataset first and training afterward. Validated on a real-world automotive glass run channel dataset of more than 900,000 unique simulated designs, an empirical engineering setting.
With only 20,000 actively selected labels, about 2.3% of the training pool, CoNN-AL reaches R-squared of 0.967 to 0.982 across all missing-variable levels, approaching upper-bound models trained on far more data. Shows that a very small labeled fraction can approach high-data baselines, rather than holding only on small-scale or synthetic tasks. Reports an R-squared range across all missing-variable levels and compares against upper-bound models.
At the more difficult missing-variable levels, CoNN-AL reaches R-squared of at least 0.95 with 30 to 40% fewer labels than random sampling, and at the most challenging level it is the only strategy in this study to reach R-squared of 0.98. Locates the active-learning benefit in the harder partial inverse design regimes and provides a direct comparison against random sampling. Uses random sampling as a control and reports label savings plus the unique result at the most challenging level.
The authors publicly release the automotive glass run channel dataset alongside the work to support future data-driven design research. Provides a reusable real-world engineering data resource for partial inverse design and active learning research. The dataset comprises more than 900,000 unique simulated designs from a real automotive glass run channel setting.
Perspective
The result targets partial inverse design settings where labels come from expensive simulation, and where only some design variables are specified while the rest must be inferred to reach a target performance value. Its validation scope is the real-world automotive glass run channel case, with a dataset of more than 900,000 unique simulated designs; the publicly released dataset can be reused in future data-driven design research. For engineering teams training surrogate models under a limited labeling budget, the framework offers an actionable path that moves labeling decisions into the data stream.
A careful reader would still watch how the framework performs on other engineering inverse design problems, which this paper does not test; how stable Monte Carlo dropout uncertainty estimates are across different models and data distributions; and that the label savings and R-squared results come from a single case and setting in this study. This is summary-level material without figures or full experimental detail, so specifics of training configuration, hyperparameters, and statistical variation cannot be judged here.
