CISE constrains self-evolving search with conformal interval rewards, returning only true positives under high-fidelity evaluation across three materials-science tasks
Related research and updatesSynopsis
The work proposes Conformal Interval-Driven Self-Evolution (CISE), which builds candidate-specific reward intervals via conditional conformal inference and iteration-wise online density-ratio estimation, uses conservative interval-based rewards for evolutionary feedback, and returns candidates only when all required property intervals lie entirely within their feasible regions; across three self-evolving search tasks in materials science, all candidates returned by CISE are true positives under high-fidelity evaluation, whereas baselines return more candidates but include false positives.
Figure 1: Overview of CISE. Offline, we fit the conditional proxy-error quantile using 𝒟 fit \mathcal{D}_{\mathrm{fit}} , 𝒟 cal \mathcal{D}_{\mathrm{cal}} provides calibration scores, and 𝒟 src \mathcal{D}_{\mathrm{src}} provides the source reference for density-ratio estimation. During evolution, unlabeled probes are used to estimate the current distribution shift, which is incorporated into conditional conformal inference to construct candidate-specific reward intervals. These intervals guide subsequent evolution and candidates are returned only when the entire interval satisfies the task constraint.
arXivInterpretation
CISE constructs candidate-specific reward intervals using conditional conformal inference, paired with iteration-wise online density-ratio estimation to handle distribution change across iterations. Relative to self-evolving search that directly uses low-cost proxy rewards, this replaces point-estimate rewards with statistically calibrated intervals, so the feedback signal itself carries an uncertainty characterization. The method description and theory provide fixed-iteration coverage results whose validity rests on explicit independence and covariate-shift assumptions.
CISE uses conservative interval-based rewards for evolutionary feedback and returns a candidate only when all required property intervals fall entirely within their respective feasible regions. This changes the selection criterion of self-evolving search from 'high proxy score means selected' to 'the whole interval must be feasible', suppressing false positives caused by proxy rewards assigning high scores to infeasible candidates. The mechanism is evaluated on three self-evolving search tasks in materials science, with the reported result that all candidates returned by CISE are true positives under high-fidelity evaluation.
Across three self-evolving search tasks in materials science, all candidates returned by CISE are true positives, while baselines return more candidates but include false positives. The result shifts emphasis from candidate quantity to candidate precision, indicating the practical value of a smaller, more precise shortlist when downstream validation budgets are limited. Evidence comes from experimental comparison on three tasks; at the abstract level the reported difference is qualitative in terms of true positives versus false positives, without specific sample sizes or numeric metrics.
Perspective
The work targets self-evolving search settings where high-fidelity evaluation is prohibitively expensive and only low-cost proxy rewards are available, especially candidate-discovery tasks that must satisfy multiple property constraints simultaneously, with three materials-science tasks as the evaluation vehicle. Its coverage results hold under explicit independence and covariate-shift assumptions, so applicability depends on how well those assumptions match the iteration-to-iteration distribution change in a target domain. For readers, this means the interval-based screening idea is most worth borrowing when validation budgets are limited and the cost of a false positive exceeds the cost of missing some candidates; it offers a smaller, more precise shortlist rather than larger candidate output.
At the abstract level, the specific sample sizes, candidate counts, high-fidelity evaluation metrics, and effect sizes for the three materials-science tasks are not given, so the scale and statistical robustness of 'all true positives' still need confirmation in the main text. The fixed-iteration coverage results depend on independence and covariate-shift assumptions; how far real iterative search departs from these assumptions, and the stability of online density-ratio estimation, are questions a reader would continue to watch. In addition, conservative interval-based rewards may reduce the number of returned candidates, and the trade-off with search efficiency is not developed in the abstract, which is a direction for further observation.
