Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

CITA trains a comparative inference model with a Bayesian tool-graph simulator and LLM comparison judgments, consistently improving Tool F1 and task success across three tool-use benchmarks and multiple backbone LLMs

The work proposes Comparative Inference for Tool-use Agents (CITA), which trains a Comparative Inference Model (CIM) from paired signals combining observed tool behavior, scalable supervision from a Bayesian tool-graph simulator, and semantic judgments from LLM-based comparison, so that the model estimates how likely a possible next tool invocation is to support final task success under the current context; across three tool-use benchmarks and multiple backbone LLMs, CITA consistently improves Tool F1 and task success, and additional analysis shows CIM learns accurate step-level value estimates for comparative tool choices.