OnePO Adapts Base Models to Medicine with RL Alone, as HuatuoGPT-3-27B Reaches 71.4 on HealthBench Professional
Synopsis
The work proposes One-stage Policy Optimization (OnePO), which treats teacher outputs as transient guidance rather than persistent training targets: a probability floor and gradient rescaling strengthen learning on low-probability teacher tokens, and Teacher Retirement discards a teacher output once the current policy's reward surpasses it. With only 20K medical samples, OnePO reaches 67.2 on HealthBench (Total), outperforming SFT+RL and pure RL under the same teacher, and scales to the open-source HuatuoGPT-3 series, whose 27B variant reaches 71.4 on HealthBench Professional.
Interpretation
The paper names two failure modes of standard mixed-policy RL in domain adaptation: Gradient Starvation, where informative teacher-output tokens receive only weak gradients because the current policy assigns them very low probability, and Teacher-Distribution Anchoring, where teacher outputs keep pulling the policy toward the teacher distribution even after the policy can surpass them. Where mixed-policy RL typically keeps teacher outputs as a persistent training signal, this work reframes them as stage-dependent guidance and isolates the two effects with two controlled pilot studies. The diagnosis rests on two pilots: the AHA-Medicine pilot introduces fictional knowledge solely through teacher outputs under zero data leakage, where pure RL and mixed-policy RL stay near zero while OnePO absorbs the fact faster; the anchoring pilot starts all variants from the same 2K teacher-output cold-start model and varies only whether teacher outputs are dynamically retired.
OnePO combines two mechanisms: Adaptive Objective Evolution, which uses a probability floor and gradient rescaling so low-probability teacher tokens receive near-SFT-level gradients and then automatically reverts to standard GRPO as token probability rises, and Teacher Retirement, which retains a teacher output only when its reward strictly exceeds the best on-policy reward in the same group. Unlike the two-stage SFT+RL pipeline, OnePO performs knowledge absorption and the return to autonomous exploration within a single RL stage without a manual schedule; unlike mixed-policy RL that keeps teacher outputs throughout, teacher signals phase out as the policy strengthens. Ablations show all three components matter: removing the probability floor drops HealthBench from 65.4 to 49.1, removing rescaling drops it to 57.3, and removing Teacher Retirement drops it to 52.7, below pure RL at 59.8; stricter retirement thresholds yield higher ceilings (Max > Mean > Min > None).
In medical domain adaptation with 20K training samples, OnePO consistently outperforms SFT+RL and pure RL under the same teacher source on HealthBench and closed-ended medical benchmarks; with DeepSeek-V3.2 (thinking) as teacher, OnePO scores 67.2/44.5 on HealthBench Total/Hard versus 64.5/40.7 for SFT+RL. The paper reports gains of 2.7 points over SFT+RL and 7.4 points over pure RL on HealthBench Total, and notes that teacher quality matters: DeepSeek-V3.2 (thinking), whose outputs contain an explicit reasoning trace, outperforms GPT-5 Chat. In the controlled comparison all three routes share the same training data, RL algorithm, and hyperparameters, with Qwen3-8B-Base on a single node of 8 H200 GPUs; closed-ended benchmarks use exact match and open-ended evaluation follows the official HealthBench protocol.
Scaling the same paradigm to the HuatuoGPT-3 series yields 9B and 27B variants scoring 68.3 and 71.4 on HealthBench Professional, with the 27B variant reaching 70.1 on HealthBench Total, which the paper reports as surpassing frontier models such as GPT-6 Astra; preliminary cross-domain experiments in writing and law also show OnePO above SFT+RL. The work extends OnePO from controlled 8B experiments to an open-source medical model series and provides initial validation in two non-medical domains, writing with rubric rewards and law with exact-match rewards. The scaled training expands the open-ended subset from 10K to 20K and adds 20K rubric tasks from RubricHub, scales closed-ended medical multiple-choice questions to 30K, and trains for roughly 600 RL steps; cross-domain experiments use Qwen3-4B-Base with 10K samples per domain.
Perspective
The result targets training-pipeline designers who want to adapt a general base model to a specific domain, especially medical settings that offer both verifiable and rubric-based rewards. The method still depends on two external signals, teacher outputs and rewards; Teacher Retirement reduces long-term reliance on imperfect teachers, but weak or biased teachers can still slow early learning or limit domain coverage. The controlled experiments use Qwen3-8B-Base on a single node of 8 H200 GPUs, the scaled models are 9B and 27B, and 100B+ frontier scales are not validated; controlled experiments use one teacher output per prompt. The authors state the model is a research prototype not intended for clinical deployment without rigorous validation by medical professionals.
Readers should still watch how the quality of the two external signals, teacher outputs and rewards, shapes results, and note that model-based grading cannot fully rule out reward misspecification or hacking; the grader validation reduces this concern but does not replace expert evaluation. Cross-domain evidence is currently limited to writing and law, and domains with less mature reward design or hard-to-formalize expert criteria still need validation. Multi-teacher or multi-output settings would alter retirement dynamics, which the paper leaves for future work. The loaded material contains the full text and appendices, but some figures are presented as textual descriptions, so exact curve values would require consulting the original figures.
