A 9B model matches frontier reasoning on GPQA, AIME and LCB via parallel tempering sampling, with no parameter updates
Synopsis
The work introduces Parallel Power Tempering (PPT), which adapts classical parallel tempering (replica exchange) to sequence-level power-sharpened LLM sampling: multiple replicas at different sharpening levels run in parallel so that low-power replicas explore diverse reasoning trajectories while high-power chains exploit higher-likelihood responses, with swaps carrying promising traces to the sharpest output rung, and a fixed-horizon construction removes a structural truncation bias identified in early-stopped power samplers; across MATH500, GPQA, HumanEval, GSM8K, AIME 24&25 and LiveCodeBench v5 on Qwen3-4B, Qwen3-8B and Qwen3.5-9B, PPT attains best or tied-best accuracy in all model-benchmark settings, outperforming Power Sampling, PowerSMC and GRPO, and on Qwen3.
Interpretation
PPT adapts parallel tempering (replica exchange) to sequence-level power-sharpened LLM sampling: it runs replicas in parallel over a ladder of sharpening powers, where lower-power replicas explore diverse reasoning trajectories and higher-power replicas exploit high-likelihood responses favored by the sharpened target, with adjacent replicas exchanging whole records under a Metropolis-Hastings rule. Prior LLM replica-exchange methods index intermediate targets by prefix length, whereas PPT couples replicas at different sharpening powers over a common horizon at each generation stage, exchanging entire sequence records and explicitly coupling exploration under flatter distributions with concentration under sharper targets over the same completion space. Swap acceptance is computed from cached base-model log-probabilities, so an exchange requires no additional model evaluations; the paper reports that on LCB v5 replica exchange accounts for only a small fraction of total time, while peak memory is higher than the single-chain value.
The authors identify a structural truncation bias in early-stopped power sampling implementations: each suffix proposal regenerates only through the current realized length, so a shorter candidate can satisfy the condition while the reverse probability is zero, and the exact MH rule that should reject the move can accept it, creating uncompensated probability flow toward shorter records that parallel tempering cannot repair. The paper derives an asymptotic bias floor for this one-way truncation and proposes a fixed-horizon construction, using deterministic post-EOS padding so proposals always regenerate to a common endpoint, restoring preservation of the true sharpened target. The paper gives theorem-form total-variation lower bounds and a four-state illustration showing the truncating kernel neither preserves nor converges to the target; on Qwen3-4B and Qwen3-8B, moving from the truncating to the fixed-horizon kernel lowers the hit-capacity rate from 9.2% to 8.0% and from 12.2% to 9.5% while accuracy rises.
Across mathematics, STEM and code reasoning benchmarks, PPT attains best or tied-best accuracy in nearly all model-benchmark settings and outperforms GRPO without parameter updates or external rewards; with Qwen3.5-9B, PPT reaches 85.9 on GPQA, 93.3 on AIME 24&25 and 84.9 on LCB v5, comparable to the reported frontier-model numbers. Single-chain power sharpening is unreliable on strong models, as Power Sampling falls below standard decoding on most Qwen3.5-9B benchmarks and PowerSMC trails standard decoding by a wide margin on AIME 24&25, whereas PPT targets the same powered distribution and never regresses below standard decoding. The main table compares Qwen3-4B and Qwen3-8B across six benchmarks, and an extended table covers Qwen3.5-9B on the three most challenging benchmarks; evaluation uses strict pass@1 with no reward- or verifier-based reranking.
Compute-matched controls attribute the gains to the exchange mechanism rather than to more compute or more chains: increasing a single chain's MCMC steps to match the decode-token budget of multi-replica PPT still leaves PPT more accurate on four benchmarks, and an uncoupled ladder with swaps disabled scores lower than PPT on both Pass@1 and voting accuracy. This separates the explanation of spending more tokens from that of letting replicas exchange records, indicating that exchange changes the quality of each chain's trajectory rather than merely the number of samples. The paper reports that on Qwen3-4B and Qwen3-8B, PPT improves Pass@1 and Vote over the uncoupled ladder on MATH500, AIME 24&25 and GPQA; ablations also show accuracy saturating around four to five replicas and equi-accepting ladders outperforming random, arithmetic and geometric spacing.
Perspective
The result is aimed at practitioners who want to improve reasoning quality at inference time without parameter updates or external rewards, in settings where multiple replicas can run in parallel and higher peak memory is acceptable; the paper's evaluation centers on mathematics, STEM and code reasoning benchmarks under strict pass@1 with no reward- or verifier-based reranking. Methodologically, the fixed-horizon construction and the swap kernels preserve the target at a given horizon, and swap acceptance is computed from cached log-probabilities so the exchange itself adds almost no model forward passes; on ladder design, equi-accepting ladders outperform random, arithmetic and geometric spacing in the experiments, and accuracy saturates after four to five replicas.
The paper reports that Qwen3.5-9B reaches performance comparable to frontier models on GPQA, AIME 24&25 and LCB v5, but the frontier numbers are marked as reported values, so readers should check the comparison protocol themselves; PPT raises accuracy while peak memory is substantially higher than the single chain, and how the memory-throughput trade-off shifts across hardware remains an open question; swap acceptance varies with model and task, and the paper notes that newer, stronger models accept fewer swaps, leaving open whether denser ladders are needed; in addition, the convergence theorem relies on a supported-full-restart sufficient condition, and the authors note the coefficient is a worst-case guarantee that does not certify a small error at a practical refinement budget.
