Q&D Trains an 8B Questioner on the Consequences of Its Questions: Required-Evidence Coverage on MuSiQue Rises from 78% to 90%, and Retail Task Success from 13% to 34%
Synopsis
The work adds a content axis to agent proactivity—horizontal proactivity pursues unstated information the current context already identifies, vertical proactivity pursues needs only earlier evidence reveals—scores runs against a need graph recovered mechanically from multi-hop benchmarks' own decompositions with no model judge, and proposes Q&D (questioner and drafter), which forks a recorded run at one step, samples eight candidate questions, continues each to the end, and prefers the question whose continuation retrieves more required evidence, so no reward model or judge is needed; on held-out splits of MuSiQue, StrategyQA and 2WikiMultiHopQA, at equal retrieval spend the trained 8B questioner recovers more required evidence than the same model prompted (89.5% versus 78.
Interpretation
It defines and measures a content axis of proactivity: horizontal proactivity pursues unstated information the current context already identifies, vertical proactivity pursues needs that only earlier evidence makes nameable, and the two are read by breadth, depth-weighted recall and deepest need resolved on a need graph. Prior work on proactive agents mainly studies whether and when an agent acts on its own, while agentic search methods either follow the chain the question spells out or are trained only on final-answer correctness, so what each question pursues is neither measured nor rewarded; this work turns that content axis into a quantity scored from a transcript. Need graphs are recovered mechanically from structure the benchmarks ship, with no person and no model writing a node or an edge: MuSiQue 800 graphs, 2,660 nodes, 1,860 edges; StrategyQA 2,290 graphs, 6,720 nodes, 4,483 edges; 2WikiMultiHopQA 12,576 graphs, 31,120 nodes, 11,956 edges; scoring uses no model judge and comparisons are made at equal question counts.
It proposes Q&D: the agent is split into a questioner and a frozen drafter, a recorded run is forked at one step, eight candidate questions are sampled and each continued to the end, and pairs are ordered by consequence (whether the run answered the task, which reached complete evidence sooner, how much evidence each turn added), so no reward model or judge is needed. Because the drafter is a fixed function of the evidence, every change in the state is caused by a question, so each question can be credited with what followed it; answers decide only 6% of training pairs, so a question that reaches an unstated need early wins even when final answers agree. Training runs in three stages: imitation of good decisions, direct preference optimization on question pairs, then direct preference optimization on question pairs together with contrasts ranking asking above stopping at unfinished states; candidate questions and rollouts come from language models, and two human raters checked the label rule on one suite.
On held-out splits of three multi-hop question-answering benchmarks, at equal retrieval spend the trained 8B questioner recovers more required evidence than the same model prompted, and much of the extra evidence is deep, which is vertical proactivity. The gain comes from what it asks, not from asking more or longer: replacing its questions with as many drawn at random from other tasks of the same benchmark loses 34.7 and 63.3 points of coverage, hiding the evidence costs 8.4 and 8.7 points, and letting the prompted model spend as many question tokens still leaves it 14.9 and 4.7 points behind. MuSiQue coverage 78.3% versus 89.5%, depth-weighted recall 77.9% versus 90.5%, deepest need 1.53 versus 1.75; StrategyQA 78.4% versus 85.5%; 2WikiMultiHopQA 88.2% versus 93.1%; each training seed alone is decided on all three suites, and all four training runs agree within a standard deviation of 0.8 points.
The trained 8B questioner leads a prompted model larger in the same role on two of three benchmarks and, without further training, transfers to a customer-service agent with a simulated customer, where retail task success rises sharply and customer follow-up turns fall. The lead follows depth: it is largest on MuSiQue, whose chains run deepest, and reversed on 2WikiMultiHopQA, whose needs lie at most one step deep, suggesting learned proactivity pays most where evidence must be followed step by step; the transfer shows the behavior is general rather than tied to question answering. At equal spend it leads GPT-OSS-120B by 7.0 and 4.3 points on MuSiQue and StrategyQA and trails by 3.5 on 2WikiMultiHopQA; retail success rises from 13% and 12% to 34% and 32%, better on ten of twenty-five tasks and worse on none, and it beats GPT-OSS-120B by 17.0 and 12.7 points with 1.6 and 1.8 fewer follow-up turns.
Perspective
The result is aimed at tool-using agents that must fill in unstated information over several steps, especially where evidence has to be followed step by step: retrieval-style questioning in multi-hop question answering, and retail customer-service dialogue with a simulated customer. It makes the content of proactivity scorable from a transcript and learnable by preference training, and the questioner never reads a need graph, so it can run where none exists. The setting is a fixed retrieval pool, a frozen drafter and answerer, and comparisons at equal retrieval calls; in retail the questioner is told its questions reach the store's records and never the customer.
The two reported seeds share one supervised checkpoint and one pair export, so intervals are over tasks alone and do not estimate variation from that checkpoint and export; no checkpoint met the pre-committed selection rule as written, and the reported configuration was kept on a second ground fixed in advance. Training labels and the primary quantity share the need graphs, answers decide only 6% of training pairs, and the extra evidence does not yet move answers detectably: on MuSiQue it finds the answer-bearing evidence 9.7 points more often, which at the prompted model's answer rates implies about 4.3 points, below the 4.7 the sample can detect and consistent with the 2.4 observed; on StrategyQA it stops before finding that evidence by choice in 92% and 86% of the runs that end without it. The released retrieval pools are small enough that most prerequisite edges do not gate retrieval, so the vertical result concerns depth in a decomposition; on MuSiQue and StrategyQA each task has one line of inquiry, so breadth records only whether a run got past its first need. Candidate questions and rollouts come from language models, two human raters checked the label rule on one suite, and every user-facing number comes from a simulated customer. Q&D is one round of off-policy training, not compared with on-policy reinforcement learning or iterated on its own rollouts, and the length control was not run held-out. If only a fast parse without figures or tables is available, the specific values and intervals in Table 2, Table 3 and the appendices cannot be checked, which affects how the size of the gains and the stopping behavior can be judged.
