Can offline evals ship a change on their own? On 183 headline experiments the pooled eval-to-online link was too weak, so every change went to an A/B test
Lead
Deciding whether a change can ship on offline evals alone requires that eval gains predict online gains across comparable changes, not that the eval agrees with human judgment; on 183 independent interventions from the Upworthy archive, both sides were measured precisely yet the pooled relationship stayed weak, and the illustrative gate sent every change to an A/B test.
Story
Whether a change can ship on offline evals alone depends on the relationship between eval gains and online gains across comparable changes, not on whether the eval itself is trustworthy. Earlier validation asked only whether judge labels agree with human labels, or whether an eval ranks different systems the way human preference does, and neither kind of evidence concerns the gain a change produces. On 183 independent interventions from the Upworthy Research Archive, both the eval and the online outcome were measured precisely, and the pooled relationship between eval gains and online gains was still weak.
Treating each change as a trial, the relationship that predicts online gains from eval gains is estimated from historical change logs and turned into a ship, kill, or A/B-test decision. Historical eval gains and online gains both carry sampling noise, which flattens the observed correlation, so the reported within-change measurement variances are subtracted before the between-change covariance is recovered. In the illustrative configuration the fitted shipping threshold exceeded the largest possible eval gain, so no change was shipped and every one was routed to an A/B test.
What to watch
A team can next rerun this pipeline on its own change logs within pre-specified classes to see whether eval gains predict online gains stably, and use that to decide which changes can skip an online experiment. Teams that want to keep the guarantee testable can send a small random fraction of would-be ships to A/B tests and use those labels to check whether the relationship still holds in the ship region.
The pooled relationship in the public illustration is weak and the per-class samples hold only 9 to 27 changes, so this demonstration neither establishes nor rules out a class-specific relationship. A mismatch between what the eval targets and what the online outcome measures, a judge that is more accurate on one arm than the other, and the loss of outcome labels in the ship region after deployment can each affect whether the relationship is observable, and each is worth checking before adoption.
