Skip to main content
Back to timeline
arXivSource publication:

PTH checks show four harness details can reverse the ranking of stale-data RL methods, and after correction TIS matches SAN on verl

Related research and updates

Synopsis

The work introduces PTH (Probe The Harness), a set of checks that makes the experimental harness visible, and uses a comparison between SAN, a behaviour-free method, and truncated importance sampling (TIS) on verl and in a single-GPU trainer to show that four harness details—the PPO ratio taken against the learner's own recomputed probabilities, the data seed not reaching the TIS arm, the replay queue reusing its first batch for 33 updates, and two loss normalisers differing from their description—can reverse the observed ranking; with the harness checked, TIS matches SAN on verl, and in the trainer TIS learns steadily while SAN keeps a margin.

Source-provided article image: Probe the Harness: Setup Checks for Stale-Data RL Comparisons in Language Models
Figure 1 ·

Figure 1: The PPO-clip arm trains without an active clip. verl, Qwen2.5-Math-1.5B, GSM8K, refresh interval 96. (a) Greedy validation accuracy, one line per data seed; the dotted line marks the refresh. The PPO-clip arm (orange) drops after update 80 in every seed; SAN + \text{SAN}^{+} (blue) and TIS (green) end level. (b) For the PPO-clip runs, the per-token KL between sampler and learner grows by four orders of magnitude before the refresh, while the clip fraction stays at zero on every update.

arXiv

Interpretation

Harness details can reverse the ranking of methods trained on stale samples: in the SAN versus TIS comparison, SAN first finished ahead in both stacks. Such methods are judged by comparisons against importance-corrected baselines, and this work treats the harness that defines the comparison as an object of inspection, noting that logged quantities looked consistent with a working setup while the quantity defining the comparison went unchecked. Based on a SAN versus TIS case on two stacks, verl and a single-GPU trainer, with four harness details each linked to its effect on the comparison.

Four harness details are identified with their signatures: the PPO ratio was taken against the learner's own recomputed probabilities, the data seed did not reach the TIS arm, the replay queue reused its first batch for 33 updates, and two loss normalisers differed from their description. These details were not previously treated as checks on the validity of the comparison; the work makes them explicit and gives each a signature. Grounded in the case, where for each detail the logged quantity looked consistent while the quantity defining the comparison went unchecked.

With the harness checked, the conclusion changes: TIS matches SAN on verl, and in the trainer TIS learns steadily while SAN keeps a margin. The same methods yield different ranking conclusions before and after the harness is checked, showing the ranking is sensitive to harness settings. From corrected-harness comparisons on both stacks, together with reference results for TIS and uncorrected GRPO under sampler lag.

The work contributes the PTH checklist and reference results for TIS and uncorrected GRPO under sampler lag. It turns harness visibility into a reusable checking procedure and provides reference results for the lagged-sampling setting. The checklist is built from the signatures and effects of the four details in the case, and the reference results cover TIS and uncorrected GRPO under sampler lag.

Perspective

The work addresses researchers and engineers who train language models on stale samples and need comparisons against importance-corrected baselines, in settings using stacks such as verl or a single-GPU trainer where the comparison involves the PPO ratio, data seed, replay queue, and loss normalisers. The PTH checklist and the signatures of the four details can be used to check whether the harness makes the comparison valid before reporting a ranking; the reference results for TIS and uncorrected GRPO under sampler lag serve as a comparison point in that setting.

The abstract does not give experimental scale, model specifications, training steps, or statistical uncertainty, nor how broadly the four details occur, so whether other harness details can also reverse comparisons remains an open question. The abstract does not provide the specific items and steps of the PTH checklist or the numerical reference results under sampler lag, so readers who need to reproduce or directly apply them would still need the checklist and result details in the original text.

Sources