TopoPlanner lifts tool dependency graphs into cellular complexes, making loop and merge regions explicit as 2-cells and beating GTool and other baselines on n-F1, l-F1 and ACC across four tool-planning benchmarks
Related research and updatesSynopsis
The authors present TopoPlanner, which lifts a tool dependency graph into a cellular workflow complex, attaches a 2-cell along each fundamental cycle induced by non-tree edges, retrieves a request-relevant boundary-closed subcomplex through cosheaf-consistent cellular retrieval, performs multi-dimensional structural reasoning, and injects the resulting cellular representation into the planner LLM as a soft graph token; on topology-guided extensions of the three TaskBench domains and ToolBench, it outperforms prompt-based and graph-enhanced baselines across Vicuna-7B/13B and CodeLlama-7B/13B, for example raising Multimedia/Vicuna-7B n-F1/l-F1/ACC from GTool's 47.17/19.64/7.07 to 86.80/66.24/54.76.
Figure 1: LLM task planning across workflow topologies.
arXivInterpretation
Loop, merge and reuse dependencies are formulated as explicit regions in cellular workflow complexes, giving a unified multi-dimensional representation of non-linear workflows. Prior planners rely mainly on pairwise adjacency or pooled graph summaries, leaving cyclic and convergent regions implicit in edges; this work uses a spanning tree and non-tree edges to promote each independent cycle into a 2-cell, turning region-level dependencies into explicit objects. Appendix A proves that the spanning tree is contractible and that the fundamental cycles form a basis of first homology; propositions in Appendix B show that a shared 2-cell enables constant-depth communication across a region and that closure prevents consistency by omission.
TopoPlanner is introduced as a pipeline that chains lifting, cosheaf-consistent retrieval, multi-dimensional reasoning and topology-guided planning. Retrieval is cast as closed-subcomplex selection: request relevance acts as the prize and cosheaf inconsistency as the cost, with PCST-style greedy collection under boundary closure; boundary and coface relations then drive cross-dimensional message passing, and request-aware pooling yields a soft graph token injected into the planner LLM. The method is given in full in Section 3, with a training objective combining tool-planning negative log-likelihood and a cosheaf-inconsistency regularizer; ablations show that removing 2-cells, removing retrieval, or removing consistency supervision all fall below the full model across 18 backbone-domain-metric comparisons.
Topology-guided extensions of four benchmarks are constructed, and consistent gains are reported across several local LLM backbones. The extension does not let a model freely invent samples: a target topology (directed loop, merge, loop-merge) is fixed first, executable scripts instantiate candidate workflows, and an LLM generates each request under the fixed topology, with grouped splitting by normalized request and ordered node-link structure to reduce structural leakage. The main tables cover HuggingFace, Multimedia, Daily Life and ToolBench with four local backbones; for example, ToolBench/CodeLlama-13B rises from GTool's 57.89/38.74/35.83 to 66.48/49.06/49.17, and gains are often more pronounced in l-F1 and ACC than in n-F1.
Structure-specific diagnostics link the gains to regional dependency preservation rather than merely to longer outputs. On merge/reuse edge recall, dependency order and length-grouped comparisons, TopoPlanner's advantage over GTool grows for medium and long workflows while length MAE is lower, indicating the improvement is not explained only by producing longer tool sequences. On the 804-example 7B subgroup, merge/reuse recall rises from GTool's 19.04/29.06 to 39.72/41.01; execution-order Dep./Pair metrics improve in all three domains; the length-grouped table shows larger gains for medium and long workflows alongside lower length MAE.
Perspective
The framework targets planning settings that already have a tool dependency graph and a tool inventory: lifting can be precomputed and cached per dataset or tool graph, so at request time only retrieval and encoding run, which suits orchestration tasks with a relatively stable tool set and enumerable dependencies. For local open-source backbones the cellular representation is injected as a soft graph token; for models reachable only through an API, the authors keep lifting, retrieval and reasoning modules locally, use the API model only for textual proposal generation, and let the local topology module perform selection and graph-consistent decoding, explicitly not claiming graph-token injection for those models. Cross-domain transfer experiments indicate that a source-domain checkpoint can transfer to unseen-but-described tools and target graphs when target-domain tool descriptions and the target dependency graph remain available.
Several points remain worth watching. First, evaluation covers four benchmark-style tool-planning datasets and unconstrained real-world deployment is not evaluated; loop semantics are restricted to bounded revision, reuse or finite transformation, with runtime repetition and termination governed by the downstream executor or stopping policy. Second, the lifting mechanism targets loop, merge and loop-merge regions and is selected per dataset, so whether it helps other higher-order workflow shapes remains an open question. Third, in the black-box API setting only topology-guided proposal selection and decoding are available, which is not the same mechanism as local graph-token injection. Fourth, the paired-bootstrap intervals quantify example-level uncertainty for fixed checkpoints, not variation across training runs. Fifth, the controlled missing-edge diagnostic covers missing edges only, not general graph corruption.
