RL-ARC calibrates answer confidence with reasoning confidence, cutting OOD ECE to 0.16 while holding accuracy
Synopsis
The work proposes RL-ARC, a calibration-aware reinforcement learning framework that, on top of RLVR, elicits both the model's confidence in its reasoning process and its confidence in its answer, using reasoning confidence as an auxiliary signal: reasoning-guided regularization when the prediction is correct and an overconfidence penalty when it is incorrect; across mathematical reasoning and out-of-distribution complex reasoning and factual QA benchmarks, RL-ARC improves calibration metrics over existing calibration-aware training methods while keeping accuracy comparable to RLVR, which optimizes correctness alone.
Figure 1: Comparison of reasoning training frameworks: (a) Standard reasoning training often assigns high confidence to incorrect answers (i.e., overconfidence). (b) Previous calibration-aware training reduces overconfidence to some extent, but still tends to assign high confidence to incorrect answers on OOD benchmarks. (c) Our RL-ARC consistently improves calibration while preserving accuracy gains under OOD settings.
arXivInterpretation
RL-ARC uses reasoning confidence as an auxiliary signal for calibrating answer confidence, applying different objectives depending on prediction correctness: reasoning-guided regularization for correct cases and an overconfidence penalty for incorrect cases. Prior calibration-aware training methods mainly calibrate answer confidence without considering the reliability of the reasoning process itself; RL-ARC has the model account for both the reasoning process and the prediction when estimating confidence. The paper formulates the calibration-aware objective (Eq. 5, Eq. 6) and an ablation (Table 5) shows the two objectives act differently: the overconfidence penalty on incorrect predictions achieves the best Brier Score and ECE, while the regularizer on correct predictions improves accuracy, AUROC, and AURC.
Across ID and OOD benchmarks, RL-ARC improves calibration over several confidence estimation and calibration baselines while achieving accuracy comparable to RLVR, which is trained solely to improve accuracy. Existing post-hoc calibration is competitive on ID but its effectiveness does not consistently transfer to OOD; existing calibration-aware training methods generalize robustly on OOD calibration but often incur noticeable degradation in reasoning performance. For Qwen2.5-7B, OOD average ECE drops from 0.50 for RLVR and 0.20 for RLCR to 0.16 for RL-ARC, with accuracy 47.27 close to RLVR's 47.58; on Qwen3-8B, OOD average ECE is 0.17 with accuracy 51.61.
Under distribution shift, RL-ARC produces confidence better aligned with actual prediction correctness and assigns more samples to low and mid confidence rather than concentrating predictions in high confidence. Previous training approaches still concentrate many predictions in high-confidence regions under OOD, whereas RL-ARC distributes predictions over a broader range of confidence levels and assigns low or mid confidence to uncertain predictions. The paper reports confidence-aware calibration analysis on OOD benchmarks (Figure 3) and confidence distribution analysis (Appendix B.2), plus a control experiment that randomly shuffles RL-ARC's confidence scores while preserving the exact confidence histogram and bin occupancy (Table 7), where RL-ARC still achieves lower ECE and Brier than its shuffled counterpart.
Beyond prediction-level calibration, RL-ARC improves reasoning calibration and reduces shallow reasoning. Existing calibration-aware training methods focus mainly on answer calibration; RL-ARC shows that incorporating reasoning confidence indirectly aligns reasoning confidence with reasoning correctness, without direct supervision for reasoning correctness. The paper reports reasoning calibration and shallow reasoning frequency metrics (Figure 6), where the gain in reasoning calibration is substantially larger than the gain in answer calibration and the reduction in shallow reasoning is the largest; reasoning correctness is determined by LLM-as-a-judge.
Perspective
The result targets language reasoning models trained with RLVR-style methods and applies to settings where outputs are filtered by confidence, such as reliability judgments in high-stakes domains. The paper validates on mathematical reasoning (Big-Math, GSM8K, MATH-500, AMC23, AIME24, AIME25) and on out-of-distribution complex reasoning (StrategyQA, GPQA, HotpotQA) and factual QA (SimpleQA, NQ-Open, TriviaQA), with HotpotQA additionally used as an alternative training distribution; calibration improvements and preserved accuracy are reported across Qwen and Llama families and both training distributions.
Reasoning calibration and shallow reasoning rely on LLM-as-a-judge to assess reasoning correctness, and the paper notes that this can introduce errors, so these two metrics are best read as signals awaiting more accurate evaluation methods. Experiments cover only two training settings and use GRPO throughout, so extension to more diverse training datasets, stronger RL algorithms, and multimodal reasoning models remains an open question. In addition, the hyperparameter analysis shows that increasing the influence of reasoning confidence does not necessarily yield further calibration improvements, and the optimal configurations are predominantly positive-dominant, indicating that balancing the two auxiliary objectives requires per-setting tuning.
