Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

CIPO trains LLM agents with counterfactual branches to adaptively pick tool granularity, raising success and decision compression across three benchmarks and two backbones

The work proposes CIPO: it first mines executable composite skills from successful tool-use trajectories via budget-constrained BPE-style merging, then augments GRPO by branching at the first eligible granularity decision of each base rollout and using the paired outcome difference as a bounded supplementary reward; across TOOLATHLON, TOUCAN, and TRAJECT-Bench with Qwen2.5-7B and Llama3.1-8B, CIPO attains the highest Tool F1 and TSR in all six settings with an average Comp./Raw of 0.845 versus 0.912 for GRPO + Skills, and the gain does not come from simply invoking skills more often.