Skip to main content
Back to timeline
arXivSource publication:

FlyBy lets a 4B small reasoning model query stronger models at knowledge bottlenecks, beating Qwen3-14B on 1,158 hard problems at 2.7x lower serving cost

Synopsis

Through counterfactual interventions at intermediate reasoning states across two model families and multiple scales, the work finds that self-refinement largely consolidates probability mass onto solutions already reachable from the current state rather than making new ones reachable, distinguishes execution bottlenecks from knowledge bottlenecks, and introduces FlyBy, a selective querying framework that trains 4B and 8B models to reason first, diagnose what remains unresolved, and query stronger models at knowledge bottlenecks; on 1,158 hard problems across six benchmarks FlyBy-4B reaches 45.96% pass@8, surpassing Qwen3-14B's 41.64% at 2.7x lower serving cost, and FlyBy-8B raises pass@8 to 51.81%.

AI-generated editorial illustration: Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge

Interpretation

The study separates failures of small reasoning models into two operational regimes: execution bottlenecks, where the correct path remains reachable through the model's own reasoning and reflection can recover it, and knowledge bottlenecks, where further reasoning is insufficient while relevant external information makes a correct path reachable. Prior work on self-refinement and test-time computation often attributed failures to models not recognizing uncertainty or not allocating reasoning effort well; this work operationalizes the boundary through paired counterfactual interventions at intermediate reasoning states, using whether state value is zero to distinguish the two bottlenecks. Across eight reasoning models (Qwen3-0.6B/1.7B/4B/8B/14B and Gemma4-E2B/E4B/12B), 93 competition problems from AIME25/26 and HMMT-Feb2026 plus GPQA-Diamond were sampled with 16 rollouts of up to 32K tokens each; 13.4K endogenous epistemic verbalizations were identified, yielding 215K counterfactual continuations and 57K information-conditioned continuations, 272K in total.

Self-refinement is not absent in small models but rarely converts into progress: epistemic verbalizations remain common in smaller models while their causal effect on state value grows substantially with scale, and reflection mainly raises value in uncertain and execution-like states while reducing answer entropy. This revises the natural hypothesis that small models simply fail to notice their own uncertainty, locating the limitation in converting reflection into progress rather than in expressing uncertainty. Based on endogenous frequency counts using a fixed nine-expression epistemic verbalization lexicon and paired interventions that uniformly sample four verbalization occurrences per rollout, comparing continuation after the expression against decoding with it banned, reporting changes in value and answer entropy.

Relevant external information fills the knowledge gap, but small models are limited on both sides: they enter knowledge-like states more often and are less able to incorporate relevant information into ongoing reasoning once provided. Contrary to the intuition that weaker parametric knowledge leaves smaller models more headroom to benefit from external information, information utilization generally improves with model scale; knowledge-like states are also more prevalent in scientific than in mathematical reasoning. At relative positions 0.25/0.5/0.75 of incorrect traces, three interventions were compared: an epistemic cue, an oracle information cue generated by DeepSeek-V4-Pro without revealing the gold answer, and a random cue from unrelated problems; random cues gave little benefit, epistemic cues modest improvement, and relevant information substantially larger gains, with the increasing scale trend preserved when restricted to problems shared with the next smaller model.

FlyBy exposes external models as a multi-depth query action inside the reasoning trajectory, so the model reasons first, diagnoses what remains unresolved, and then decides whether to query, what to ask, and how much external computation to spend, outperforming larger models on hard problems at lower cost. Unlike approaches that directly optimize tool use, retrieve documents, or query a stronger model once upfront, this method first diagnoses which bottleneck the current state represents and then acquires targeted information; SFT bootstraps the query action space with only 95 rescue trajectories, followed by cost-aware reinforcement learning to calibrate query timing and depth. On 1,158 hard problems across six benchmarks, FlyBy-4B improves over base Qwen3-4B from 21.2% to 46.0% pass@8, surpassing Qwen3-14B's 41.64% at 2.7x lower serving cost, with pass@1 of 16.85% versus Qwen3-8B's 15.31%; FlyBy-8B raises pass@8 to 51.81%; on problems unsolved by Qwen3-4B in 16 no-tool rollouts FlyBy-4B reaches 28.7% pass@8; queries average only 105 output tokens while Search-R1 returns 2,015 characters on average; on 50 SuperGPQA problems querying raises state value by 5.30 percentage points on average while additional reasoning yields essentially no improvement.

Perspective

The result targets deployment settings where a small reasoning model serves as a cognitive core with access to stronger external models, and applies to hard reasoning in mathematics, science, general knowledge, and medicine; the method depends on external API availability and controls answer leakage during training and evaluation by preventing external models from seeing the original problem, rejecting queries with excessive n-gram overlap with the problem, and redacting answer spans from tool observations. For teams seeking higher hard-problem coverage at lower serving cost, the framework offers a trainable template of reasoning first and querying in a targeted way; for settings concerned with privacy and latency, the paper lists privacy-aware selective querying as future work.

The knowledge-bottleneck label is operational: the paper states explicitly that zero observed success is a finite-sample criterion rather than literal unreachability, and under larger continuation budgets some states do show successful continuations, though those states retain a large intervention gap. External models can themselves produce incorrect or hallucinated information that may propagate into subsequent reasoning, which the paper lists as a limitation. The querying policy loses some performance after backend substitution (43.21% and 43.96% versus 45.96%), suggesting some backend-specific influence. In addition, the loaded text includes the abstract, body, and appendices, but figures appear as textual descriptions, so specific figure details would need checking against the original.

Sources