Qwen-Planner-Agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model in tool use, memory, skills, and sub-agent coordination
Synopsis
The work builds Qwen-Planner-Agent within a closed-loop AI-for-AI framework that links data production, model training, and deployment through a shared action-feedback-verification contract: AI for Data builds a human-gated agentic data flywheel, AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning and introduces CARE to reduce reasoning and tool-use costs, and AI drives model-harness co-evolution; the agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination, with further gains on non-mobile agentic benchmarks while largely preserving general capabilities.
Figure 2 : AI-for-AI lifecycle of Qwen-Planner-Agent, comprising three interconnected phases. (i) AI for Data combines AI-assisted task generation with automated trajectory collection and curation, using model feedback to refine subsequent tasks and training-data composition. (ii) AI for Training combines planning-oriented cold-start training with hybrid-environment online reinforcement learning, while CARE adapts reward scheduling and advantage signals to the model’s evolving competence. (iii) AI for Harness equips the planner with a unified Harness that manages Skills, persistent Memory, and execution feedback, while AI-assisted diagnosis guides subsequent model and Harness refinement and feeds new requirements back into data production and training.
arXivInterpretation
It proposes and implements a closed-loop AI-for-AI framework that connects data production, model training, and deployment into one iterable whole through a shared action-feedback-verification contract. Where data, training, and deployment are often separated in AI development pipelines, this work places them under a single contract so that training feedback can flow back to guide subsequent data generation. At the summary level, the framework's three components and their connecting mechanism are described, without independent ablations or quantitative comparisons for each component.
AI for Data builds a human-gated agentic data flywheel in which specialized agents construct tasks, collect interaction trajectories, curate and balance training data, and use training feedback to guide subsequent data generation. Data production itself is delegated to agents and driven by a training-feedback loop, rather than relying on a one-off static dataset. The summary describes the flywheel's composition and feedback path, but does not provide data scale, curation criteria, or the proportion of human gating.
AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning, and introduces Competence-Aware Reward-and-Advantage Engineering (CARE) to reduce reasoning and tool-use costs while preserving task performance. CARE shapes rewards and advantages around reasoning and tool-use cost, distinguishing it from training that targets task success alone. The summary states the method's positioning and goal, but reports no magnitude of cost reduction or quantitative task-performance figures.
AI drives model-harness co-evolution through an execution-evidence-driven loop that orchestrates memory, skills, and tools at runtime and feeds structured action feedback and preserved failure traces back into coordinated model and harness adaptation. Failure traces are treated as reusable evidence for joint improvement of the model and the harness, rather than optimizing model weights alone. The summary describes the loop mechanism and what is fed back, but gives no iteration counts or controlled comparisons for the co-evolution.
Perspective
The work targets the setting of real-world mobile planner agents, suited to mobile scenarios requiring long-horizon task planning, tool invocation, memory, and sub-agent coordination, with MobilePA-Bench as the primary evaluation vehicle. Its framework design also speaks to agent development teams that want to close the loop across data production, training, and deployment; the summary notes improvements on non-mobile agentic benchmarks as well, indicating the path may be used in broader agent development workflows.
What is available here is summary-level information, without specific evaluation numbers, sample sizes, ablations, or statistical details, so the magnitude of the 'best overall performance' and of the reduction in reasoning and tool-use costs cannot be judged from the current text. The precise definition of 'competence-aware' in CARE, how human gating intervenes in the data flywheel, and the iteration mechanism of model-harness co-evolution are open questions that require the original methods section to confirm.
