CompoWorld composes 448 reusable services into cross-service tasks, lifting Qwen3.6-35B-A3B by 9.17 points on average across eight benchmarks
Synopsis
CompoWorld introduces compositional environment scaling: coding agents turn MCP tool specifications into executable services with typed states and shared interfaces, a random walk connects services through dependency graphs to generate and verify cross-service tasks, verified trajectories support SFT while a Completion-Focused Rubric Reward guides GRPO-based RL, and with 448 services exposing 10,130 tools plus 3K SFT trajectories and 1K RL tasks, Qwen3.6-35B-A3B improves by 9.17 points on average across eight benchmarks and raises AutomationBench task success from 10.33% to 32.33%.
Interpretation
The framework shifts environment scaling from generating more environments toward scaling the dependencies between them: each service owns a typed state, a tool set, and transition logic, and the composed product environment preserves each service's local dynamics while information passes between services through the agent's observations and subsequent actions. Prior work such as AgentScaler, ScaleEnv, and Agent-World mainly expands single environments or task collections, while AppWorld and Terminal-Universe expose cross-application requirements through benchmarks or workspaces; CompoWorld makes service composition itself a controllable dimension of training-environment generation. The paper provides a POMDP formalization of compositional environments and a product-environment construction, noting that namespaces distinguish otherwise identical tools; this is a conceptual and formal contribution, with the accompanying implementation reported as 448 services and 10,130 tools.
Task generation samples services via a random walk and connects them into a service-level dependency graph, where edges denote information dependencies and revisits denote that an earlier value has gone stale; a task-generation agent then instantiates initial states, goals, and a structured rubric across explore, design, probe, and submit phases, with executable final-state checks verifying that the reference solution is reachable and accepted. Unlike synthesis that fixes a tool-call sequence, this records service-level dependencies rather than a unique call order, allowing other successful trajectories to use different tool sequences. The paper states walk constraints (cover every selected service, avoid consecutive visits to the same service, revisit at least one service) and verification conditions, and notes the verifier does not require exact equality to the reference goal state; difficulty is mainly controlled by walk length, with 5 services composed by default plus 5 distractor services and walk length sampled from 7 to 12.
For training, verified successful trajectories support SFT, and the proposed Completion-Focused Rubric Reward inversely weights each rubric criterion by its pass rate within the current rollout group, so less frequently satisfied criteria receive higher weight, combined with GRPO normalization to form advantages. Uniform rubric averaging gives equal weight to easy criteria already satisfied and hard criteria still unmet, so a high average score can still leave the user's goal unmet; this reward concentrates the learning signal on remaining completion gaps. The paper provides a gradient analysis showing that at the update origin a criterion's gradient weight is controlled by its completion gap, with a positive floor retained; Appendix Figure 6 shows the reweighted run starting lower (about 0.63 versus 0.68), overtaking around step 50, and reaching about 0.76 versus 0.67 at step 140.
Across eight benchmarks, CompoWorld improves on Qwen3.6-35B-A3B by 9.17 points on average; AutomationBench success rises from 10.33% to 32.33% (+22.00), exceeding GPT-5.4 (27.67%) and Claude Opus 4.6 (25.50%) and approaching DeepSeek-V4-Flash (36.33%), and it leads all six compared agent-specialized 35B-A3B models on that benchmark. The paper also reports per-domain partial credit on AutomationBench 1.0.6: the mean rises from 41.94 to 72.68, with HR gaining most (+49.40), followed by Marketing (+32.51) and Sales (+31.45), indicating gains span business functions rather than a single domain. Results come from Tables 1 and 2; some compared models use source-reported results, with unavailable or incompatible entries marked by dashes. The paper also notes that the 32.33% pass rate shows completing every requirement remains difficult, and VitaBench 2.0 improves by only 1.62 points.
Perspective
The framework targets general-agent training where information must flow across services, and applies to service collections that can be modeled with typed states and tool interfaces, such as email, Slack, calendars, and database applications; the service library derives from public MCP specifications, so coverage depends on the availability of those specifications. For tools that cannot be reliably implemented as deterministic local code, the paper uses selective world-model simulation and states that only a small fraction of tools rely on it and only about a small fraction of samples invoke such a tool, limiting simulation error to a small part of the training corpus. The scaling experiment shows that as few as 100 SFT examples already exceed the backbone on four benchmarks, and at 3K, -Banking rises from 10.65 to 17.87 and DeepPlanning from 26.04 to 35.02; under a fixed data budget and the same number of environments, composed-environment training beats single-environment training in all four completed comparisons, for example AutomationBench rising from 12.83 to 33.33. These results are meant for teams seeking to expand the diversity and structural complexity of training data through reusable service composition.
The paper notes that automatically generated tests cannot cover the full business rules, implicit dependencies, and rare boundary cases of real services, and current coding agents cannot fully resolve this. World-model simulation covers only a small fraction of tools and samples, but once inaccurate observations are generated they may still affect the agent's subsequent judgments. RL effects are not uniform: the paper reports RL improving on SFT by 1.2 points on average across eight benchmarks, with a substantial SkillsBench gain but a slight decline on -Banking, and overall improvements primarily driven by SFT. Environment quality is assessed indirectly through downstream training gains rather than by directly validating each environment. In addition, some compared models use source-reported results, with unavailable or incompatible entries marked by dashes, so readers should note this difference in protocol when comparing across models. Sales remains the lowest-scoring domain (58.93), and the paper suggests the dependencies and constraints in these workflows still have room to be strengthened.
