Skip to main content
Back to timeline
arXivSource publication:

SLCA-GRPO routes tool-segment and summary-segment advantages separately, raising tool-calling success by up to 10.03 points across three backbones

Synopsis

The work splits tool-calling RL trajectories into a tool segment and a summary segment, argues that broadcasting a single trajectory-level advantage causes cross-segment credit misattribution, and proposes SLCA-GRPO, which normalizes the two segment rewards within the group and routes each only to its own tokens inside a single unified policy, paired with a Schema-Guided LLM Simulator and Hierarchical Rewards; across Qwen2.5-3B/7B-Instruct and Qwen3-8B-Base on Toucan-Test, BFCL V3 and τ²-Bench it reports higher means than matched GRPO, with 7B gaps of +2.53, +1.36 and +9.15 percentage points.

AI-generated editorial illustration: SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL

Interpretation

The paper frames cross-segment credit misattribution as a structural failure mode: standard GRPO broadcasts one trajectory-level advantage to all tokens, so summary-reward variation enters tool-token advantages, and in sign-conflict regimes a failed tool call can be reinforced when followed by a correct summary while a correct tool trajectory can be penalized when followed by an incorrect summary. Earlier temporal credit-assignment methods (VinePPO, GiGPO, SPO) still aggregate before advantage estimation, ToolPO adds local tool rewards but still lets summary-dependent noise reach tool tokens, and RLTR separates planner and summarizer into a pipeline, abandoning a unified backbone; this work locates the problem on the structural execution-versus-articulation axis rather than the temporal axis. The paper presents diagnostic cases plus gradient diagnostics: tool and summary gradient directions remain near-orthogonal while gradient-norm spikes appear, which the authors read as advantage contamination rather than directional conflict (Appendix B).

SLCA blocks the summary-to-tool advantage path before the backward pass: tool-side advantages route only to tool tokens and summary-side advantages only to summary tokens, with each segment reward normalized separately within the group, at no additional rollout cost and without splitting the model. Unlike additive augmentation (ToolPO) or pipeline separation (RLTR), this decouples advantage estimation along a structural axis inside one unified policy; the authors explicitly frame it as a bias-variance trade-off and provide a per-update advantage-isolation proposition, a conditional nuisance-variance-removal theorem and a directional-fidelity corollary. The theoretical results are conditional per-update guarantees, which the authors state are not claims of globally bias-free optimization; a support-control experiment estimates a 5.51-point main effect on τ²-Bench from closing the summary-to-tool support path, while the matched w/o SLCA condition changes normalization and token support together.

The accompanying SGLS combines deterministic schema validation with a frozen cross-family LLM that mocks tool responses, and HierR splits the return into a dense tool process score and a terminal summary preference score, making segment routing practical at scale. SGLS treats the schema as a control plane and only mocks calls that pass validation under a fixed template; the HierR process score is a weighted sum of format, tool-name, argument-key, argument-value and parallelism terms, while the summary score comes from an LLM-as-a-judge. Ablations show removing HierR produces the largest Toucan success drops on all three backbones, while removing SGLS has a smaller in-domain effect but changes BFCL and τ²-Bench; replacing the response mocker with GPT-OSS-120B leaves SLCA positive on all three benchmarks (+2.83/+1.24/+7.16).

Across three backbones and three benchmarks, SLCA-GRPO reports higher means than matched GRPO: on 7B, Toucan strict Success@0.9 is 79.13% (+2.53 pp), BFCL 69.77% (+1.36 pp) and τ²-Bench Pass1 41.02% (+9.15 pp), with corresponding 3B gaps of +2.35/+0.40/+1.09 and 8B gaps of +2.05/+3.35/+10.03 percentage points. In the matched comparison GRPO and SLCA-GRPO share data partitions, SFT initialization, SGLS endpoint, decoding settings, evaluation protocol, rollout group size and HierR rewards, differing only in unified versus segment-locked advantage computation; training traces show SLCA ending with fewer average tool turns and higher success, while standard GRPO stabilizes at longer trajectories without a corresponding success gain. Main results are three-run means with standard deviations; the direction of the SLCA gain persists across reward-protocol sensitivity checks including a unified reward-ratio sweep, an execution-based reward swap and a response-mocker replacement.

Perspective

The result targets tool-calling post-training where a trajectory cleanly separates a tool-call/reasoning segment from a final natural-language answer, with training done in the SGLS simulated environment and models up to 8B. For such readers, the directly reusable change is segment masking plus separately normalized segment advantages: it adds no rollout cost, does not split the model, and can be combined with temporal credit-assignment methods such as VinePPO/GiGPO/SPO. The authors also position SGLS and HierR as the infrastructure that makes routing practical at scale, so anyone reproducing this needs both a schema-validating simulator and segment-specific rewards.

The segmentation assumption is itself a scope condition: SLCA assumes a clear boundary between tool calls and free-form text, and settings such as inline code generation would require learned segmentation. On the reward side, the process score relies on matching to gold trajectories and may under-reward alternative valid tool sequences when multiple API plans solve the same request; Toucan's strict Success@0.9 thresholds that score rather than using an independent task outcome label. Temporal credit within the tool segment is not distinguished, since one scalar is routed to every tool token, so a correct Action 1 and a failing Action 2 cannot be told apart, and a combination with VinePPO/GiGPO/SPO remains untested. On scale, experiments stop at 8B and use SGLS, leaving larger-scale training and direct real-API comparisons as future work. In addition, this reading covers the full main text and appendix, but figures appear here as text and tables, so curve-level details can only be understood through the prose descriptions.

Sources