Meituan's LongCat team splits deep research into planning, parallel section research, and global-to-local editing via ResearchSpec, scoring 55.25, 51.35, and 79.83 on three public benchmarks
Synopsis
Meituan's LongCat team presents LongCat-DeepResearch, which shifts early research iteration from the full report to an executable ResearchSpec: multiple planning agents first search external sources and consolidate a research plan, researchers then investigate and draft citation-bearing sections in parallel independent contexts, and a Global Editor assigns cross-section ownership while Local Editors make targeted revisions, yielding 55.25 on DeepResearchBench, 51.35 on DeepResearchBench II, and 79.83 on ResearchRubrics, and 76.04 on an in-house benchmark, second among four compared systems.
Interpretation
The system moves early research iteration from an ever-expanding full report to a compact ResearchSpec that records each section's scope, research questions, required entities or cases, and provisional source leads, validated before dispatch for ID uniqueness, parent relationships, consecutive ordering, and required fields. Prior work either interleaves reasoning, retrieval, and drafting, or uses an evolving report to guide subsequent retrieval and revision; here the emerging global understanding is encoded as an inspectable, revisable execution interface rather than a growing draft or a persistent knowledge graph. The paper gives the method and validation rules, and a case study shows the merged ResearchSpec decomposing four themes into 19 subsections, with the Critic flagging the omitted Deep Portfolio Theory as HP1 and the Reviser adding it to S2.2.
Research runs per section in parallel: each Researcher receives the original query, the complete ResearchSpec, and one assignment, continues gathering evidence in an independent context, and writes a complete citation-bearing section, after which sections are mechanically assembled in ResearchSpec order. Compared with relying on a shared research history or a compressed summary handed to a separate writer, the same agent carries its local evidence into writing, and complete sections are retained through assembly and editing. Component ablations show the full pipeline averaging 63.91 on DeepResearchBench II and ResearchRubrics, above the one whole-report Researcher at 62.07 and One Writer Reviser at 59.56, with identical question IDs across conditions.
Editing assembles first and coordinates second: a Global Editor reads the complete draft, assigns ownership of repeated material, and specifies changes, while Local Editors apply directives within assigned sections, reducing reliance on repeated full-report rewriting. Unlike regenerating the entire document in one invocation, this separates the global judgment needed for coordination from local text generation and keeps the original sections and assembled draft available for comparison. In the Editor scaling study, enhanced editing raised average readability preference from 50.42 at B1 to 53.96 at B3, and B3 preserved or improved native benchmark scores relative to the unedited draft on DeepResearchBench II and ResearchRubrics; the paper notes the enhanced Editor jointly changes source access, instructions, and iteration count, leaving their individual contributions unresolved.
The same stage interfaces support constructing research tasks and trajectories: questions and task-specific rubrics are grounded in independently licensed review articles or frozen multi-source briefs, and teacher executions record candidate and final ResearchSpecs, tool requests and observations, citation-bearing sections, assembled drafts, and editorial decisions. The paper describes rubric construction, searchability checks, and trajectory filtering as a training-data process distinct from evaluation resources, and keeps construction evidence and rubrics separate from the query given to the answering agent. The paper reports a bounded searchability check for the article route (Keep, Revise, or Drop, passing only when all retained criteria receive Keep), deterministic checks for duplicate rubrics, answer leakage, and excluded-source leakage, and trajectory filtering for role identity, tool-call closure, and intermediate-artifact validity; this data feeds mid-training and post-training of LongCat's general-purpose models.
Perspective
The result targets developers and evaluators of research agents that must produce long, evidence-grounded reports, in settings where public-web retrieval is available and tasks can be decomposed into hierarchically identified sections. ResearchSpec is refined during planning and then held fixed during section research, while researchers may investigate further within their assignments; the same interfaces also support constructing research questions, task-specific rubrics, and trajectories for mid-training and post-training of LongCat's general-purpose models.
The paper states that point estimates do not establish statistical significance, scored coverage varies across systems, and Claude-DeepResearch's ResearchRubrics overall is an externally reported aggregate without verifiable per-task grades. The component, planning-refinement, and Editor studies use previously inspected development subsets, some selected using prior model scores; retrieval windows and realized computation vary across studies, and repeated-generation and repeated-Judge variance are not estimated. The enhanced Editor jointly changes source access, instructions, and iteration count, leaving their individual contributions unresolved, and the case study scores only the final report, so planning or editing gains are not isolated. The paper also notes reports can be lengthy, dense, or repetitive, leaving room to improve readability, and that research and editing can omit useful details.
