Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

PTH checks show four harness details can reverse the ranking of stale-data RL methods, and after correction TIS matches SAN on verl

The work introduces PTH (Probe The Harness), a set of checks that makes the experimental harness visible, and uses a comparison between SAN, a behaviour-free method, and truncated importance sampling (TIS) on verl and in a single-GPU trainer to show that four harness details—the PPO ratio taken against the learner's own recomputed probabilities, the data seed not reaching the TIS arm, the replay queue reusing its first batch for 33 updates, and two loss normalisers differing from their description—can reverse the observed ranking; with the harness checked, TIS matches SAN on verl, and in the trainer TIS learns steadily while SAN keeps a margin.