KernelZero's 7B Proposer–Coder co-evolution reaches 75.8 CUDA and 77.2 Triton pass@1 on KernelBench
Synopsis
KernelZero introduces a co-evolution framework in which a Proposer generates Torch modules at the Coder's current capability frontier and the Coder is trained with CA-GRPO so that performance is optimized only after correctness becomes reliable; both models start from Qwen2.5-Coder-7B and reach 75.8/69.6 pass@1 on KernelBench CUDA Level 1/2 and 77.2/72.5 on Triton, surpassing Claude-4.5-Sonnet on CUDA and DeepSeek-V4-Pro on Triton.
Interpretation
The framework splits GPU kernel generation into two models with an explicit interface: a Proposer that generates Torch modules from API sets as high-level specifications, and a Coder that translates those modules into CUDA or Triton kernels, with the two updated in turn using their own GRPO rewards to form an automatic curriculum. Unlike prior routes that treat kernel generation as a single-model task, this separates problem-setting from problem-solving into two roles that can be optimized separately, so training data shifts with the Coder's capability instead of relying on a fixed corpus. The paper gives a formal two-stage mapping, the alternating optimization procedure, and separate reward designs for Proposer and Coder; training details state both are initialized from Qwen2.5-Coder-7B and alternate for 4 rounds of 20 steps each.
A frontier-driven module generation mechanism extracts API sets from modules the Coder fails to solve to construct Proposer inputs, and a Frontier Reward pushes the generated modules' success rate toward roughly one half, so modules are neither trivial nor impossible. This directly addresses the scarcity of high-quality data aligned with the model's current capabilities that the paper identifies, producing capability-aligned training data automatically from correctness feedback without manual annotation. The paper describes API co-occurrence sampling plus a three-stage validation of AST check, runnable check, and numerical stability, and gives the form of the Frontier Reward maximized when the success rate is near one half; the ablation shows that freezing the Proposer and continuing to train the Coder yields lower pass@1 and fast1@1 on both CUDA and Triton than continuing to update the Proposer.
CA-GRPO separates correctness and performance signals through a group-level gate, so that correct kernels receive a normalized speedup reward only after group correctness reaches a threshold. The paper notes that prior RL methods typically treat correctness and performance as a single undifferentiated reward; this work explicitly separates them to ease the trade-off where optimizing only for correctness yields conservative implementations while aggressive performance optimization breaks functional equivalence. The paper compares correctness thresholds on the combined 200 tasks of KernelBench Levels 1 and 2, reporting that an intermediate threshold balances correctness and speedup better, and presents the fastest correct kernel among four training variants on six representative operators.
On KernelBench, KernelZero-7B reports CUDA results of 75.8 pass@1, 98.87 pass@5, and 100 pass@10 on Level 1 and 69.6 pass@1, 93.70 pass@5, and 97 pass@10 on Level 2; Triton results are 77.2 pass@1, 97.41 pass@5, and 99 pass@10 on Level 1 and 72.5 pass@1, 94.96 pass@5, and 98 pass@10 on Level 2. The paper states it surpasses Claude-4.5-Sonnet on CUDA and DeepSeek-V4-Pro on Triton, with CUDA Level-1 pass@1 within 0.6 percentage points of Claude-4.5-Sonnet while pass@5 and pass@10 are higher on both levels. Results come from KernelBench's 250 tasks (Level 1 has 100 single-kernel operators, Level 2 has 100 fusion patterns) under a one-shot prompt, reporting pass@k and fastp@k plus average output tokens; evaluation uses the Keck validation service built on cudaLLM, including single-sample GPU execution, parameter consistency, and correctness-gated performance measurement.
Perspective
This work targets kernel-generation settings that start from a PyTorch reference implementation and require generating CUDA or Triton kernels, and it applies to evaluation setups like KernelBench that measure correctness and speedup relative to PyTorch Eager. It lets capability-aligned training data be produced automatically and lets a smaller model keep improving without repeating the full pipeline, which is directly useful to researchers training and evaluating kernel generation and to engineering teams seeking efficient implementations for particular operator combinations. The paper also notes that generated CUDA kernels do not call external high-performance libraries, so many Level-2 tasks containing convolutions or GEMMs face PyTorch references backed by cuDNN or cuBLAS, which delimits the scope in which the results apply.
Several specific numbers appear as placeholders in the abstract and experiment sections, for example the CUDA Level 1/2 pass@1 and pass@10 in the abstract, and some percentages in the Proposer-freezing and correctness-threshold ablations, so those details can only be taken from the numbers actually given in the table and conclusion. How the correctness threshold in CA-GRPO and the roughly-one-half success-rate target of the Frontier Reward behave across backends and difficulty levels, and how the Keck evaluation service performs at larger task scale, are directions a reader can keep watching. In addition, the paper attributes the Level-2 fast1@1 gap between CUDA and Triton to backend abstractions and whether vendor libraries are called, and how far that explanation extends across more operator types remains an open question.
