Skip to main content
Back to timeline
arXivSource publication:

RECAST trains a lightweight router to unify retrieval and computation as evidence construction, beating retrieval and program-generation baselines across six benchmarks

Synopsis

RECAST casts evidence construction as a sequential decision process over retrieval and computation operations: a lightweight RouterLM iteratively selects lexical, semantic, or relational primitives, or requests customized operations that a frozen CompilerLM translates into executable code, then passes accepted evidence to a frozen AnswerLM; RouterLM is trained with SFT followed by GRPO, achieving the strongest average success rate across six heterogeneous benchmarks and zero-shot generalization on three held-out benchmarks.

Source-provided article image: RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing
Figure 1 ·

Figure 1: Overview of RECAST. At each round, RouterLM uses the task, compact source profile, and evidence and feedback from earlier rounds to either select and formulate a primitive or synthesized operation or pass the evidence to AnswerLM when it considers the evidence sufficient.

arXiv

Interpretation

The paper broadens context construction from retrieving existing content to actively deriving evidence, treating retrieval and computation as complementary evidence operations that RouterLM selects each round among CALL_PRIMITIVE, SYNTHESIZE, and ACCEPT_CONTEXT. Unlike fixed similarity-based retrieval and retrieval-centric iterative or agentic methods, RECAST allows evidence to be derived through filtering, aggregation, arithmetic, and customized code across multiple source items rather than merely located in a single item. The framework is trained and evaluated on six benchmark families with heterogeneous source representations, 100 unseen questions per family, with multi-round methods capped at six rounds and 4,000 characters of retained evidence, and all methods sharing one AnswerLM and one frozen judge.

Only RouterLM is trained: SFT establishes valid multi-round evidence-construction behavior and GRPO improves the routing policy from downstream task outcomes, with a reward dominated by final-answer correctness and smaller signals for token-level F1 and structural validity. The paper reports that SFT and GRPO are complementary and that their combination exceeds either stage alone, enabling a smaller trainable Qwen3.5-9B router to outperform a large training-free Gemini 3.5 Flash router. Training uses 3,368 SFT trajectories and 608 GRPO instances with LoRA fine-tuning; GRPO applies mixed-outcome filtering that keeps only groups containing both successful and unsuccessful trajectories, and ablations show that removing this filter or using correctness-only reward lowers the average.

Primitive and synthesized operations provide complementary capabilities, and combining them yields the strongest training-free average, while the learned policy shows distinct operation-use patterns across benchmarks. The paper reports operation-space ablations and operation-usage statistics showing that neither operation type is uniformly superior, that synthesis is used on about 27.3% of questions, and that routing length varies from roughly 2.16 to 4.22 rounds. Ablations compare primitive-only, synthesis-only, and combined configurations on the same test set; operation-usage statistics are reported over the 600-question in-domain test set with primitive calls, synthesis calls, synthesis usage, and average routing rounds.

The learned routing policy transfers: it generalizes zero-shot to three held-out benchmarks excluded from training and checkpoint selection, and retains strong performance when CompilerLM is replaced. The paper reports higher average success than the strongest baseline on the held-out benchmarks, and shows that swapping the default compiler for the lighter Gemini 3.5 Flash Lite or Qwen3.5-9B keeps in-domain and held-out performance above the strongest evaluated baselines. Held-out evaluation covers 2WikiMultiHopQA, TAT-QA, and WikiTableQuestions with 100 examples each; compiler-transfer experiments keep the trained RouterLM fixed and replace only models not used as compilers during training.

Perspective

The work targets tasks whose evidence must be derived from heterogeneous sources, including dataframes, financial reports combining text and tables, hierarchical tables, Wikipedia passage collections, and user profiles; the method operates under a cap of six rounds and 4,000 characters of retained evidence per question, with read-only SQLite relational queries and synthesized programs executed in an isolated environment. For engineering practice that wants iterative evidence construction handled by a smaller trainable model while larger models support compilation and final answering, this setting is directly informative; the paper also states that code, data, and the trained RouterLM checkpoint will be released under an open-source license upon acceptance.

Several result tables in the main text and appendices lack numeric values in the provided text, so the summary can only report the directional finding that RECAST attains the strongest average success rate and outperforms the strongest baseline, without verifying specific percentage points; readers needing exact numbers and per-benchmark comparisons should consult the original tables. In addition, training data relies on execution-derived outcomes, and inference-time iterative calls plus occasional code synthesis add latency and compute cost, with the paper listing more efficient trajectory construction and cost-aware routing as future directions; the failure analysis also notes that the router may locate the right topic or an intermediate entity without obtaining the required relation or constraint, or may misinterpret the requested computation, which are open questions for improving evidence requests and sufficiency assessment.

Sources