Public articles linked to the same research event.
arXiv The work proposes Comparative Inference for Tool-use Agents (CITA), which trains a Comparative Inference Model (CIM) from paired signals combining observed tool behavior, scalable supervision from a Bayesian tool-graph simulator, and semantic judgments from LLM-based comparison, so that the model estimates how likely a possible next tool invocation is to support final task success under the current context; across three tool-use benchmarks and multiple backbone LLMs, CITA consistently improves Tool F1 and task success, and additional analysis shows CIM learns accurate step-level value estimates for comparative tool choices.
The work proposes Comparative Inference for Tool-use Agents (CITA), which trains a Comparative Inference Model (CIM) from paired signals combining observed tool behavior, scalable supervision from a Bayesian tool-graph simulator, and semantic judgments from LLM-based comparison, so that the model estimates how likely a possible next tool invocation is to support final task success under the current context; across three tool-use benchmarks and multiple backbone LLMs, CITA consistently improves Tool F1 and task success, and additional analysis shows CIM learns accurate step-level value estimates for comparative tool choices.
The work proposes Comparative Inference for Tool-use Agents (CITA), which trains a Comparative Inference Model (CIM) from paired signals combining observed tool behavior, scalable supervision from a Bayesian tool-graph simulator, and semantic judgments from LLM-based comparison, so that the model estimates how likely a possible next tool invocation is to support final task success under the current context; across three tool-use benchmarks and multiple backbone LLMs, CITA consistently improves Tool F1 and task success, and additional analysis shows CIM learns accurate step-level value estimates for comparative tool choices.
The work proposes Comparative Inference for Tool-use Agents (CITA), which trains a Comparative Inference Model (CIM) from paired signals combining observed tool behavior, scalable supervision from a Bayesian tool-graph simulator, and semantic judgments from LLM-based comparison, so that the model estimates how likely a possible next tool invocation is to support final task success under the current context; across three tool-use benchmarks and multiple backbone LLMs, CITA consistently improves Tool F1 and task success, and additional analysis shows CIM learns accurate step-level value estimates for comparative tool choices.