Skip to main content
Back to timeline
arXivSource publication:

CorrGRPO swaps covariance for Pearson correlation in the normalization denominator, beating GRPO on code generation, tool calling, and agent security

Synopsis

The work observes that when GRPO trains with multiple rewards it normalizes the summed reward by its within-group standard deviation, whose variance equals the sum of all pairwise reward covariances, so large-scale rewards can dominate the normalization and suppress smaller-scale reward signals; the authors propose CorrGRPO, which replaces pairwise covariances in the denominator with Pearson correlation coefficients while keeping the centered total reward unchanged, and compare it with GRPO and variants such as GDPO on code generation, tool calling, and agent utility versus security across models from 0.5B to 8B parameters, reporting improvements in all three domains and an outward expansion of the empirical Pareto frontier between competing objectives.

AI-generated editorial illustration: CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning

Interpretation

The paper gives a covariance view of multi-reward advantage estimation in GRPO: when GRPO sums reward components and normalizes by the within-group standard deviation, the corresponding variance equals the sum of all pairwise reward covariances, so the denominator carries both individual reward variances and pairwise dependence. Previously the group-relative normalization was typically treated as a single scaling term; this work decomposes that denominator explicitly into entries of the reward covariance matrix and shows that, holding the centered total reward fixed, stronger positive correlations reduce advantage magnitude while weaker or more negative correlations increase it, turning the statistical relationships among rewards into an analyzable normalization mechanism. The claim rests on a variance-of-a-sum identity; the appendix derives the population identity and shows an exact finite-sample equivalence between the sample variance of the total reward and the sum of sample covariances, noting that the equality requires the same samples, means, and divisor for variances and covariances.

The paper argues that covariance normalization entangles reward dependence with reward scale: because a covariance equals the product of the two component standard deviations times their Pearson correlation, large-scale rewards dominate both diagonal and off-diagonal terms, suppressing the influence of smaller-scale rewards on normalization. This offers a quantifiable explanation for why some reward components in multi-reward GRPO appear to have limited influence despite configured weights; the paper illustrates it with a four-trajectory, three-reward example in which the third reward has weak correlations with the others yet its variance alone accounts for a large share of the covariance sum. The argument is supported by the covariance decomposition and a concrete numerical example, and the appendix adds a controlled analysis that fixes the centered total reward and sweeps one reward's standard deviation, showing the GRPO advantage decreasing monotonically with that scale while the CorrGRPO advantage stays constant across the positive-scale sweep.

The paper proposes CorrGRPO: replacing pairwise covariances in the denominator with sample Pearson correlation coefficients while retaining the centered total reward in the numerator, so each non-zero-variance reward contributes equally on the diagonal and off-diagonal entries reflect only the strength and direction of dependence. Unlike variance-based reweighting (as in MO-GRPO), per-reward independent normalization followed by aggregation (as in GDPO), or quantile normalization with Mahalanobis whitening (as in RDPO), CorrGRPO preserves the relative reward weights specified by the training objective and changes only the composition of the normalization denominator. The paper specifies the zero-variance handling convention (zeroing the corresponding rows and columns) and proves the resulting matrix remains positive semidefinite, the correlation sum is nonnegative, and the denominator is strictly positive once the numerical stabilizer is added; the appendix also derives compatibility with DAPO, CISPO, and GDPO and reports a CISPO-integrated comparison on Qwen2.5-Coder-7B-Instruct.

The paper compares CorrGRPO with GRPO and variants such as GDPO on three multi-reward domains, code generation, tool calling, and agent utility versus security, using models from 0.5B to 8B parameters, and reports improvements in all three domains plus an outward expansion of the empirical Pareto frontier. The work tests correlation-based normalization across three settings with clearly different reward structures: coding with seven rewards that can reinforce one another, tool calling with four partial-credit rewards, and agent tasks with utility and security rewards that can trade off. Coding trains on the 2,641-problem LeetCodeDataset split and evaluates on its 228-problem split plus zero-shot transfer to HumanEval, MBPP, and LiveCodeBench v6, reporting average Pass@1 gains over GRPO of 2.09, 0.80, 2.27, and 4.21 percentage points at 0.5B, 1.5B, 3B, and 7B; tool calling trains on 3,920 RLLA-4K examples with 80 evaluation examples and transfers to API-Bank, with higher all-exact scores and API-Bank averages for all three backbones; agent experiments train on 1,584 AgentDojo cases with 411 evaluation cases and transfer to ASB and InjecAgent, reporting improved ASB Joint Accuracy and lower InjecAgent attack success rates.

Perspective

The result targets language-model reinforcement learning pipelines that use group-relative advantage estimation with rewards composed of multiple components, covering code generation, structured tool calling, and joint utility-and-security training for tool-using agents; the paper's implementation conventions (zeroing rows and columns for zero-variance components, clamping the correlation sum to zero, and a numerical stabilizer) and training configurations (for example, batch 64 prompts with group size 8 for coding, batch 128 prompts with group size 4 for tool calling, and batch 16 tasks with group size 4 for agent tasks) provide concrete entry points for reproduction. For a reader, this means that when reward components differ substantially in numerical range and have interpretable dependencies, replacing covariance with correlation in the normalization denominator is a candidate first move; the paper also shows the method can be combined with DAPO, CISPO, and GDPO, so it can be evaluated as a local substitution in an existing pipeline.

Worth watching: when there are many reward components, or when components are strongly negatively correlated or nearly collinear, how an approaching-singular correlation matrix affects advantage scale; the paper addresses this with a numerical stabilizer and a positive-semidefiniteness argument but does not offer a general cross-task conclusion. The reported gains are on benchmarks, and the security experiments state explicitly that improvements on these benchmarks do not establish real-world deployment safety or resistance to unseen attacks, so real-environment utility-security tradeoffs still need application-specific evaluation. Correlation-based normalization also does not correct biased or misspecified rewards, and behavior remains dependent on the chosen objectives and their weights. In the text available here, the main body and appendices provide the method, derivations, training settings, and main result tables, but some appendix example prompts and responses appear as placeholders in the evidence bundle, so verifying reward computation line by line would require the full appendix examples and the code repository.

Sources