Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

TTCL jointly improves LLM reasoning accuracy and confidence calibration on unlabeled target data, with +40.13% average relative accuracy and +70.80% ECE reduction on base models

The work proposes Test-Time Calibration Learning (TTCL), a label-free framework that derives self-supervision signals for correctness and calibration from multiple model-generated responses and jointly adapts reasoning accuracy and verbalized confidence on unlabeled target-task data, reporting consistent improvements in both accuracy and calibration on mathematical reasoning and factual question answering.