Skip to main content
Back to timeline
arXivSource publication:

DISCO splits long context into parallel Worker grounding plus Driver reasoning, holding 78.4% on 1M-token RULER-QA while cutting inference cost by over 80%

Synopsis

The work proposes DISCO, which separates long-context processing into local evidence grounding executed in parallel by lightweight Worker LLMs and global reasoning handled by a central Driver LLM whose dynamic DAG planning is optimized with GRPO reinforcement learning; with Qwen3-8B it holds 78.4% accuracy on RULER-QA at 1M tokens (where the RAG baseline collapses to 10.9%), improves up to 9.8 points over full-context baselines on LongBench v2, and matches Gemini-3-Pro-Preview while reducing inference cost by over 80%.

AI-generated editorial illustration: DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation

Interpretation

The paper attributes long-context failure to a structural entanglement of grounding and reasoning, and proposes fully decoupling them: Workers perform only local atomic extraction, while the Driver reasons over refined evidence alone. Unlike RAG, which delegates grounding to shallow embedding or BM25 retrieval, and unlike sequential or single-pass agent workflows such as Chain-of-Agents and LLMMapReduce, DISCO uses generative LLMs for grounding and structurally isolates it from reasoning. The paper provides a factorization (Equations 1-6) and system design, and on LongBench v2 replaces the generative Worker with a same-scale Qwen3-Embedding-4B retriever as a control: 47.3 versus 39.6, indicating generative grounding outperforms static vector retrieval.

DISCO keeps grounding stable at the million-token scale: on RULER-QA at 1M tokens Qwen3-8B reaches 78.4% while the RAG baseline collapses to 10.9%, with 256K/512K/1M at 77.6/77.5/78.4, showing almost no degradation with length. The paper reports grounding performance as invariant to total context size and about 6% above the strongest baseline, LLMMapReduce. Results come from the RULER-QA synthetic multi-hop benchmark at 256K, 512K, and 1M tokens, compared against five baseline families: naive full context, RAG, CoA, LLMMapReduce, and RLM.

Decoupling plus iterative replanning unlocks reasoning: on LongBench v2, DISCO with Qwen3-14B gains 9.8 points over the full-context baseline (48.7% versus 38.9%), and Thinking Mode gains are amplified (8B from +0.9 to +5.5, 14B from +5.6 to +13.5). The paper explains this by Workers having already distilled compact evidence, so the Driver's thinking budget is spent entirely on reasoning rather than searching noisy context. Comparisons are against Full Long Context and LLMMapReduce (33.7%), with a Thinking Mode on/off contrast.

Cost is decoupled from scale: with Gemini-3-Pro as Driver, DISCO retains reasoning parity with the full-context baseline (71.3% versus 72.2%) while cutting evaluation cost from 126.7 to 20.4, a roughly 6.2-fold reduction described in the paper as over 80% lower. The paper offloads token-intensive scanning to lightweight Workers, separating the cost of reading from the cost of thinking. Based on the Gemini-3-Pro cost comparison and latency curves from 256K to 1M under a fixed iso-compute budget of 8 H100 GPUs, where DISCO's latency stays nearly flat while agentic baselines grow steeply and linearly.

Perspective

The result targets reasoning services that must answer multi-hop questions over million-token contexts, and is meant for engineering and research teams that would replace a single long-window model with a lightweight-Worker plus central-Driver combination. It makes inference-time scaling operational, scaling Map actions with context length and Wide-dependency actions with reasoning complexity, and it offers tuning guidance on Worker scale, reward design, and fault tolerance: going from 4B to 14B Workers yields only 2.9% accuracy (50.9% to 53.8%), whereas going from 8B to 14B Drivers yields 39.8% to 50.9%, suggesting compute is better spent on the Driver.

Several open questions remain. The Fact Dropout experiment shows a phase transition: beyond roughly 20% discarded narrow-action outputs, average execution stages jump from 4 to 7, and accuracy falls to 33% at 40% dropout, so where real deployments sit relative to that tipping point is still to be observed. Training data is 4.5k samples (16K-512K tokens) synthesized from MuSiQue via Recursive Context Augmentation, and how that synthetic long-context distribution transfers to real ultra-long documents is not answered in the text. The RLM baseline performs poorly with Qwen3-8B/14B, which the paper attributes to smaller models lacking the coding and planning ability needed for REPL environments, so comparisons against it should be read with that premise in mind. Finally, this reading covered the full text, but figures and tables (for example Figures 2-6) appear as references in the text, so some values are available only through prose paraphrase and fine-grained curves still require the original figures.

Sources