Skip to main content
Back to timeline
arXivSource publication:

POISE predicts value baselines from a model's own internal states, beating GRPO and PPO with more stable training on a six-domain verifiable-reward corpus

Synopsis

The work introduces POISE (Policy Optimization with Internal State Value Estimation), which uses a lightweight probe reading signals already computed during the forward pass to predict the value baseline, and a cross-rollout construction that predicts each rollout's value from an independent rollout's internal states to preserve gradient unbiasedness; on Qwen3-4B and OLMo3-7B-Instruct-DPO across a six-domain verifiable-reward corpus, POISE outperforms other RLVR baselines with more stable training, and the probe matches a separate LLM-scale value model, generalizes to various tasks, and remains accurate as the policy scales.

Source-provided article image: Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
Figure 1 ·

Figure 1 : Estimated policy-gradient variance under a fixed total rollout budget. The m = 1 m=1 condition allocates one rollout to each of B B distinct prompts, whereas the m = 8 m=8 condition allocates eight rollouts to each of B / 8 B/8 prompts.

arXiv

Interpretation

POISE turns the internal states of a large reasoning model into a value model: a lightweight probe reads signals already computed during the forward pass to predict the baseline and is trained online alongside the policy. Where GRPO estimates its baseline as the group mean over rollouts from the same prompt and PPO trains a policy-scale critic, POISE reuses representations the model already computes, avoiding the trade-off between baseline accuracy and in-batch prompt diversity as well as the extra cost of a critic. The abstract states the algorithm and its online probe training, and reports advantages over other RLVR baselines plus more stable training on Qwen3-4B and OLMo3-7B-Instruct-DPO across a six-domain verifiable-reward corpus.

To preserve gradient unbiasedness, the authors introduce a cross-rollout construction in which each rollout's value is predicted from an independent rollout's internal states. This is a dedicated design for the unbiasedness problem when internal states serve as value estimates, distinct from predicting a rollout's value from its own states. The abstract explicitly states the construction's purpose as preserving gradient unbiasedness, but gives no ablation or quantitative comparison of the construction itself.

The probe matches a separate LLM-scale value model, generalizes to various tasks, and remains accurate as the policy scales. This indicates the value signal carried in internal states can substitute for a separately trained large value model, keeping baseline quality without a second large model. Supported by the abstract's statements that the probe matches a separate LLM-scale value model, generalizes to various tasks, and remains accurate as the policy scales; no specific numbers or scale points are given.

Perspective

The result targets research and engineering settings that train large reasoning models with verifiable rewards, especially general-reasoning training that mixes prompts from multiple domains within a batch. It lets readers obtain a value baseline without training a policy-scale critic, with the probe updated online alongside the policy; the abstract's validation covers Qwen3-4B and OLMo3-7B-Instruct-DPO and a six-domain verifiable-reward corpus, so the intended scope is multi-domain training under verifiable rewards rather than open-ended settings without automatic verification signals.

The abstract gives no specific benchmark numbers, training-budget comparison, probe size or architecture, or ablation of the cross-rollout construction, and it does not specify the composition of the six-domain corpus or the evaluation metrics; the range of policy scales over which the probe remains accurate is also not detailed. Readers wanting these quantitative details will need the tables and experimental setup in the full text.

Sources