ANTMAN replaces static partitioning with a revisable Need Graph: a 16x larger search space raises active coordination only 1.23x versus over 15x for partition-driven baselines
Synopsis
The work introduces ANTMAN, a multi-agent information-seeking framework that treats evolving unresolved information needs, rather than input partitions, as the unit of runtime coordination, maintaining a revisable Need Graph that tracks unresolved requirements, accumulated evidence, prior attempts, and search progress to control worker selection, routing, and task-local recovery; across multi-document QA, controlled long-context scaling, and structured navigation, ANTMAN increases active coordination by only 1.23x under a 16x increase in searchable context, compared with more than 15x for partition-driven baselines, while preserving strong answer quality and retaining most of its performance when execution is delegated to substantially smaller 8B worker models.
Interpretation
ANTMAN makes evolving unresolved information needs the unit of runtime coordination, using a revisable Need Graph that records unresolved requirements, accumulated evidence, prior attempts, and search progress to decide which workers become active, how needs are routed, and when task-local recovery is invoked. Existing long-context multi-agent systems typically assign workers to input partitions, as LongAgent does by splitting input into fixed-size chunks each handled by a member agent, which makes the agent population grow with the number of partitions; ANTMAN shifts the coordination state from how the space is segmented to what the query still requires. The paper provides a formal characterization: each coordination step activates at most one territory-backed worker and each realized need can be routed at most a bounded number of times, so the number of activated workers is bounded by realized task demand and the per-need routing budget rather than directly by the number of available territories; Appendix A states the proposition and its proof.
In controlled information-space scaling, growing the searchable space from 32K to 512K (16x) raises ANTMAN's active coordination by only 1.23x, while LongAgent and CoA grow by 15.29x and 15.33x respectively, with the same gap appearing in model calls and inference cost. This directly tests the design claim that coordination should not grow with information-space size, by separating space size from an approximately fixed information need. The experiment follows the multi-needle setting of Needle-in-a-Haystack PLUS introduced by LongAgent, with 10 questions, Early/Middle/Late evidence positions, and 32K/64K/128K/512K contexts, averaging 30 evaluations per entry; Appendix B reports endpoint growth and complete per-length results.
The same need-conditioned coordination mechanism transfers across distinct information substrates without substrate-specific redesign: it outperforms the strongest non-benchmark-specific baselines by 15.3% on RepoProbe, 18.4% on SWE-QA-Pro, and 27.9% on GAIA, and ANTMAN-H, which uses Qwen3-8B for all non-orchestrator components, retains 89.2%, 93.5%, and 98.3% of full ANTMAN's performance respectively. The paper connects prior work on adaptive information seeking with prior work on adaptive multi-agent coordination through a representation of unresolved needs, and shows this connection does not depend on repository-specific navigation machinery. RepoProbe-Python contains 108 questions across eight large Python repositories, SWE-QA-Pro uses a frozen 80-question subset, and GAIA uses the fixed text-only GAIA-Text-103 subset; within each repository benchmark methods share the same frozen question set and model setting, and all methods on GAIA share the same tool interface.
Ablations show the explicit, evolving need state itself contributes: at 512K, Graph-free Adaptive without an explicit Need Graph scores 64.01 versus 84.03 for full ANTMAN, and on GAIA 50.00 versus 60.00, despite activating comparable numbers of workers. This indicates the benefit goes beyond engaging more workers or generic adaptive replanning and comes from persistently tracking unresolved needs; recovery is particularly consequential in controlled scaling, while adaptive rerouting matters in repository navigation. All variants share the same territories, WorkerCards, models, tools, execution budget, and synthesis procedure, and receive the same maximum budget of 100 LLM calls per query; Appendix D reports complete comparisons of answer quality, active coordination, cost, and call counts.
Perspective
The result targets information-seeking tasks that require locating and integrating evidence from an information space far larger than a single context, such as multi-document QA, long-context corpora, code repository navigation, and tool-based web search; it assumes the space can be deterministically partitioned into territories with lightweight WorkerCards and that each unresolved need can be searched locally and boundedly by a single worker. For practitioners, this means the searchable space can be enlarged without expanding the agent population with the corpus, and that most performance is retained when the orchestrator stays strong while execution is delegated to 8B-class worker models. The paper positions the mechanism as reusable across substrates: the controlled long-context and repository evaluations share the same local information-access implementation, with the repository setting additionally supporting structural navigation over symbols, imports, references, inheritance, and call relations, while GAIA uses web search, webpage text extraction, and isolated Python execution, with all methods sharing the same tool interface.
In controlled space scaling, ANTMAN's cost growth is 6.90x, higher than its 1.23x coordination growth and 1.21x call growth, so cost and coordination do not move in lockstep and this difference deserves further observation. The demand-scaling experiment shows that at one required evidence unit, 4.1 of 5.0 worker activities are unrelated to the required evidence (82% of observed activity), indicating exploratory overhead at low demand, which the authors explain through trajectory decomposition rather than treat as a failure. On GAIA most of the gap comes from Level 3, and ANTMAN-H approaches full ANTMAN on less demanding tasks, so the cost of substituting smaller models concentrates on harder cases. In addition, GAIA uses a single territory-backed worker, so rerouting has no effect in that setting. The paper's outlook proposes extending to richer information substrates, developing more transferable representations of information needs, and refining policies for revising and routing needs, all of which remain open questions.
