Scoring the scorer: mutation analysis finds the official GPU-kernel check misses one in six witnessed faults, with 78.6% of precision faults escaping
Synopsis
The work adapts mutation analysis into an adequacy metric for graded numerical oracles, injecting 10,303 compilable faults into verified CUDA implementations of 188 KernelBench problems (7,384 with independent kill witnesses), measuring that the official check deterministically misses 16.9% of witnessed faults with misses skewed by family (8.7% arithmetic, 78.6% precision), and using that measurement to explain mechanisms, audit existing patches, and synthesize two-input suites reaching 98.0% detection.
Interpretation
It builds the first at-scale measurement of kernel-benchmark oracle strength: 124 deterministic mutation rules generate 10,303 compilable faults across 188 verified CUDA substrates, 7,384 backed by kill witnesses, and any test protocol is scored by the fraction it detects. Prior work (KernelBench-Verified, The Correctness Illusion, robust-kbench) patched checkers by hand or seeded a handful of bugs, leaving unanswered whether a patch suffices and what it still misses; this turns checker quality from opinion into a comparable number and uses witnesses so only provably-detectable faults enter the denominator. Table 1 gives the pipeline counts: 10,303 generated mutants, 1,594 non-compiling rejected by NVRTC, 279 duplicate compiled images, 177 equivalent, leaving 8,253 distinct and 7,384 witnessed; over 120,000 mutant-input evaluations build the kill matrix; bootstrap resampling over problems shows 16.9% is a population property rather than sampling noise.
The official check deterministically misses one in six provably-detectable faults, and the misses are strongly skewed by family: only 8.7% of arithmetic faults escape, against 78.6% of precision faults, 27.8% of synchronization faults and 22.9% of boundary faults. Earlier claims that checkers are weak rested on anecdotes or hand-seeded bugs; this provides per-family miss rates and shows a checker validated on hand-seeded arithmetic bugs would look excellent while being blind where real GPU bugs live. Table 2 lists witnessed counts and miss rates for six families (arithmetic 3,349 witnessed, 8.7% missed; precision 271 witnessed, 78.6% missed); Appendix A Table 6 gives operator detail such as store-fp16 missing 169/183 (92%) and sync-remove missing 58/198 (29%).
It quantifies the domain's two governing mechanisms: tolerance vacuity, a blind band growing with reduction size (an all-zeros output passes the official check on a softmax problem), and a legitimate-variance ceiling above which harsher inputs reject correct kernels. This explains why the misses exist and turns the intuition 'test with harsher inputs' into a constraint requiring a validity gate, which the text reports finding in no prior work. The text reports specific magnitudes for the largest undetectable relative error at two reduction sizes (Fig. 2b); softmax contributes 28 tolerance-blind survivors where structurally identical log-softmax contributes 9; the reconstructed fuzzing baseline falsely rejects correct kernels 107 times at the top of its magnitude range.
The metric is directly usable for auditing and synthesis: it decomposes KernelBench-Verified's gain into hidden inputs versus tighter tolerance, and set cover over the kill matrix reaches 98.0% detection with two inputs per problem (94.8% held-out) against 83.1% for the official five. KernelBench-Verified's authors state they cannot compute how much of the failure is tolerance versus distributional mismatch, and the kill matrix computes exactly that; suite design also becomes an optimization problem whose gain comes from choosing well rather than testing more. Table 3 gives the budget curve: at budget two, 98.0% on the full pool and 94.8% held-out, versus 83.1%/82.9% for the official five random inputs; a 50/50 id-hash dev/test split guards against overfitting; the knowledge ladder (Table 4) shows the fault taxonomy lifts a fuzz baseline from 21.6% to 61.0%, while white-box access to the concrete faults reaches only 57.3%.
Perspective
The metric is aimed at benchmark maintainers, evaluators of kernel-generation systems, and reinforcement-learning researchers who use the checker as a reward signal, in settings where a GPU-kernel benchmark decides correctness with a graded numerical tolerance. It licenses comparative claims (suite A misses faults suite B catches) and existence claims (this protocol misses these witnessed faults), and it can be used to design suites that need fewer inputs; the authors state it cannot certify a passing kernel correct and never use it that way. The fault model covers 124 rules in six families over LLM-authored, gate-verified CUDA substrates, with results from a single H100 generation; the level-3 architecture study covers 48 architectures and its witness search is shallower than at operator scale, so it is treated as a one-sided bound.
The fault model cannot express faults living in structures it does not generate (tensor-core paths, double-buffered pipelines, some warp-level idioms) or multi-site interactions, so the miss rates are adequacy relative to that fault model rather than absolute correctness. The level-3 witness search is shallower, so 17.3% should be read as a floor; the two problems judged unrefereeable (48_Mamba2ReturnY and 45_UNetSoftmax) rest on comparing the official fp32 forward against its own fp64 evaluation, and readers may watch for later re-checks under different tolerances or references. The realism probe uses only 60 problems and three incorrect kernels, which the authors themselves call a small sample. This reading was of the full text, but figure-level details were taken from the prose and tables rather than verified panel by panel.
