TTCL jointly improves LLM reasoning accuracy and confidence calibration on unlabeled target data, with +40.13% average relative accuracy and +70.80% ECE reduction on base models
Related research and updatesSynopsis
The work proposes Test-Time Calibration Learning (TTCL), a label-free framework that derives self-supervision signals for correctness and calibration from multiple model-generated responses and jointly adapts reasoning accuracy and verbalized confidence on unlabeled target-task data, reporting consistent improvements in both accuracy and calibration on mathematical reasoning and factual question answering.
Figure 1: Overview of our work : TTCL enables label-free calibration learning directly on unlabeled target-task data, consistently improving reasoning accuracy while reducing calibration error across models and tasks, with confidence progressively aligning with accuracy during test-time learning.
arXivInterpretation
TTCL is a label-free framework that jointly adapts reasoning accuracy and verbalized confidence directly on unlabeled target-task data, with self-supervision signals derived from multiple model-generated responses. Prior calibration learning is typically incorporated into reinforcement learning and relies on ground-truth correctness supervision; TTCL moves calibration learning to test time without ground-truth labels, targeting practical deployment settings where labels are unavailable and calibration must adapt to newly encountered target tasks. The abstract states that self-supervision signals are derived from multiple model-generated responses and that theoretical analysis establishes TTCL as a bounded surrogate for the ideal calibration objective; the derivation and implementation details are not given in the abstract.
On base models, TTCL achieves an average relative accuracy improvement of +40.13% and an ECE reduction of +70.80% across eight benchmarks. This quantifies the effect of label-free test-time calibration as an average relative improvement across eight benchmarks spanning mathematical reasoning and factual question answering. The abstract reports average relative numbers across eight benchmarks and does not provide per-benchmark results, model sizes, or sample sizes.
TTCL can further improve both accuracy and calibration for already calibrated models under domain shift, particularly when source-domain calibration transfers poorly to target tasks; in the math-to-factQA setting it achieves an average relative accuracy gain of +20.35% and reduces ECE by +53.83%. This broadens applicability: not only base models benefit, but already calibrated models can improve further under domain shift, especially when source-domain calibration transfers poorly. The abstract gives average relative numbers for the math-to-factQA setting and does not report per-task breakdowns or statistical tests.
Theoretical analysis characterizes TTCL as a bounded surrogate for the ideal calibration objective. This links label-free self-supervised calibration to the ideal objective theoretically, rather than only empirically. The abstract only states this conclusion and does not give the form of the bound, its assumptions, or the proof structure.
Perspective
The work targets deployment settings where ground-truth labels are unavailable at test time and calibration must adapt to newly encountered target tasks; it applies to mathematical reasoning and factual question answering and covers both base models and already calibrated models under domain shift. The abstract states that source code has been released, supporting reproduction and extension on similar tasks. Its value lies in letting models adjust both answer correctness and verbalized confidence on unlabeled target data, thereby supporting identification of uncertain predictions and more reliable decision-making.
The abstract does not specify how the self-supervision signals are constructed, the sampling and aggregation strategy over multiple generated responses, or the assumptions behind the bounded-surrogate result, nor does it give per-benchmark results, model sizes, sample sizes, or statistical tests. ECE reduction is expressed in relative form such as "+70.80%", so the baseline values and computation protocol need confirmation in the main text. The abstract also does not state under which task or data conditions the method may not apply; these are areas for further verification in the original paper.
