Public articles linked to the same research event.
arXiv The work proposes CIPO: it first mines executable composite skills from successful tool-use trajectories via budget-constrained BPE-style merging, then augments GRPO by branching at the first eligible granularity decision of each base rollout and using the paired outcome difference as a bounded supplementary reward; across TOOLATHLON, TOUCAN, and TRAJECT-Bench with Qwen2.5-7B and Llama3.1-8B, CIPO attains the highest Tool F1 and TSR in all six settings with an average Comp./Raw of 0.845 versus 0.912 for GRPO + Skills, and the gain does not come from simply invoking skills more often.
The work proposes CIPO: it first mines executable composite skills from successful tool-use trajectories via budget-constrained BPE-style merging, then augments GRPO by branching at the first eligible granularity decision of each base rollout and using the paired outcome difference as a bounded supplementary reward; across TOOLATHLON, TOUCAN, and TRAJECT-Bench with Qwen2.5-7B and Llama3.1-8B, CIPO attains the highest Tool F1 and TSR in all six settings with an average Comp./Raw of 0.845 versus 0.912 for GRPO + Skills, and the gain does not come from simply invoking skills more often.
The work proposes CIPO: it first mines executable composite skills from successful tool-use trajectories via budget-constrained BPE-style merging, then augments GRPO by branching at the first eligible granularity decision of each base rollout and using the paired outcome difference as a bounded supplementary reward; across TOOLATHLON, TOUCAN, and TRAJECT-Bench with Qwen2.5-7B and Llama3.1-8B, CIPO attains the highest Tool F1 and TSR in all six settings with an average Comp./Raw of 0.845 versus 0.912 for GRPO + Skills, and the gain does not come from simply invoking skills more often.
The work proposes CIPO: it first mines executable composite skills from successful tool-use trajectories via budget-constrained BPE-style merging, then augments GRPO by branching at the first eligible granularity decision of each base rollout and using the paired outcome difference as a bounded supplementary reward; across TOOLATHLON, TOUCAN, and TRAJECT-Bench with Qwen2.5-7B and Llama3.1-8B, CIPO attains the highest Tool F1 and TSR in all six settings with an average Comp./Raw of 0.845 versus 0.912 for GRPO + Skills, and the gain does not come from simply invoking skills more often.