Skip to main content
Back to timeline
arXivSource publication:

Paired forecasts across 336 prospective states leave explanation credit unconfirmed, while structured cards cut Tox21 repeat drift by about 60 percent

Synopsis

The work introduces a predictive-credit protocol that compares description, matched explanation, and donor explanation under paired forecasts sharing an intervention, forecaster, and outcome across 336 prospective states (controlled learning, 12 Tox21 endpoints, 24 OpenML tasks); the frozen v5 decision was inconclusive and Tox21 and OpenML returned confirmation_no_go, matched point-accuracy gains remained unconfirmed, yet structured cards reduced Tox21 repeat drift by 64.5 and 59.1 percent and a researcher-authored mechanism positive control lowered point MAE by 2.60 percentage points.

AI-generated editorial illustration: Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts

Interpretation

The paper proposes and runs a predictive-credit protocol in which the same forecaster receives a public description, the current state's matched explanation card, or a donor explanation card from another state for the same intervention and outcome, with paired differences measuring forecast gain and local alignment. Prior explanation evaluation targeted simulatability or faithfulness of model behavior; this work moves the target to research agents' numerical forecasts of executed experiments and places description, matched, and donor contexts under one loss. The protocol froze records in three prospective studies: controlled v5 with 120 states, Tox21 with 72 states, and OpenML with 144 states, totaling 336 prospective states and 3,708 prospective calls; all assigned calls remain in intention-to-treat analysis with deterministic fallbacks.

The frozen decisions did not support predictive credit for natural explanations: v5 was inconclusive, Tox21's preregistered ROC AUC interval-score harm test was unmet (D-M=-.0026, 95 percent interval [-.0174, .0104]), and OpenML's joint formation, point-equivalence, and repeatability rule was unmet. Against prior positive reports that explanations help prediction, this work reports an unconfirmed outcome under preregistered gates and locates that outcome at the tested donor resolution and gate thresholds. Tox21's frozen gate required both contrasts to reach .005 with positive one-sided lower bounds, joint positivity in at least 8 endpoints and 4 actions, positivity at every budget, complete population, and integrity of at least .95 in every stratum; only 5 of 12 endpoints and 3 of 6 actions had both contrasts positive, and contrasts changed sign across label budgets.

Structured prediction cards mainly changed repeat consistency and expressed uncertainty rather than point accuracy: in Tox21, matched and donor cards reduced Pro repeat drift from .01004 to .00357 and .00411 (64.5 and 59.1 percent reductions), the Flash replay reduced drift by 60.6 and 59.2 percent, yet matched-card point MAE rose from .01823 to .02020. This separates output coordination from predictive gain by reporting drift and outcome loss jointly, showing that card-induced agreement can decouple from accuracy. Pro and Flash results come from the same 72 Tox21 states; the Flash replay used 432 calls with integrity 1.0 in every stratum, but the frozen joint status was replication_not_supported, with interval-score equivalence the sole failing gate.

Known signals are taken up by the forecaster: a researcher-authored mechanism positive control lowered point MAE from a description baseline of 6.391 to 3.795 (a 2.596-point reduction versus description, 95 percent interval [2.314, 2.880]) and by 4.355 points versus a paired false mechanism; exact-signal calibration won in 144/144 states with a point-value Spearman correlation of .99996. This supplies a sensitivity ceiling for the protocol: when an explanation carries strong known mechanism information, forecasts do improve, separating unconfirmed natural-explanation credit from a forecaster that ignores explanations entirely. The positive control used 40 new states and 240 valid Flash forecasts, with outcomes computed before prompt freeze and withheld from the forecaster; calibration used 1,440 calls and returned the frozen status assay_sensitive.

Perspective

The protocol targets research-agent benchmarks and scientific forecasting: it can score experimental rationales, compare explanation interfaces, and inform cost-aware experiment ranking. It applies where an intervention has been executed, outcomes are measurable, and matched and donor explanations can be constructed; controlled learning, Tox21 molecular endpoints, and OpenML tabular tasks provide three worked settings. For readers reusing the framework, the frozen records support benchmark reanalysis and method comparison, and the positive control and exact-signal calibration show the protocol detects gains when explanations carry strong known information.

Predictive credit for natural explanations remained unconfirmed at the tested donor resolutions, so it is still open whether gains appear with more strongly aligned donors, finer mechanism text, or different forecasters. OpenML full-card assignment widened nominal 80 percent intervals by about 21 percent, with 49.3 percent coverage versus 51.4 percent for description and content, leaving the tension between interval width and coverage open. Two Flash runs on the same stimuli produced different full-card drift magnitudes (.00108 and .00426), suggesting interface and run-to-run variation worth characterizing further. This summary is based on the paper's full text without all figures and appendix detail, and some numeric intervals are not fully rendered in the text.

Sources