PlanPool externalizes clarification as a mutable question pool, improving ambiguity coverage and cutting silent failures across three benchmarks
Related research and updatesSynopsis
The work argues that in agentic Text-to-SQL a correct execution result does not necessarily mean the underlying underspecification has been adequately resolved, since the agent may silently make unverified assumptions that happen to match the intended answer; it attributes this partly to premature clarification termination and introduces PlanPool, which externalizes the clarification plan as a mutable question pool where every planned question must be explicitly asked or dropped before submission and newly discovered ambiguities can be added during interaction, consistently improving ambiguity coverage and reducing silent failures across three benchmarks derived from BIRD-Interact and Spider while maintaining competitive execution accuracy.
(a)
arXivInterpretation
The paper proposes and examines a distinction: in agentic Text-to-SQL, a correct execution result does not necessarily imply that the underlying underspecification has been adequately resolved, because the agent may silently make unverified assumptions that happen to match the intended answer. Prior evaluation of such systems centers on execution correctness; this work separates correctness from whether underspecification is actually resolved and shows the two can diverge. Based on the behavioral observation reported in the abstract: a correct execution result does not necessarily imply adequate resolution, and this behavior is driven in part by premature clarification termination.
The authors attribute this failure mode in part to premature clarification termination, and report that forcing the agent to ask more questions improves execution accuracy, yet ambiguities are concentrated in earlier interactions, making brute-force questioning inefficient. It locates the source of failure in when clarification stops, while showing that simply increasing the number of questions is not an efficient fix. The abstract reports the relationship between forced additional questioning and execution accuracy, along with the observation that ambiguities are unevenly distributed across interactions.
The paper reports that even when explicitly prompted to plan its clarification process, the agent frequently abandons questions it has already identified as relevant. This indicates that identifying missing information and reliably maintaining and resolving it are distinct capabilities, and prompt-based planning does not close that gap. The abstract reports the phenomenon of abandoning already-identified relevant questions even under explicit planning prompts.
The authors introduce PlanPool, which externalizes the clarification plan as a mutable question pool: every planned question must be explicitly asked or dropped before submission, and newly discovered ambiguities can be added during interaction; across three benchmarks derived from BIRD-Interact and Spider it consistently improves ambiguity coverage and reduces silent failures over unconstrained and prompt-based alternatives while maintaining competitive execution accuracy. It replaces relying on prompts for the agent to maintain its own clarification plan with an externalized, mutable question pool that enforces pre-submission disposition of each question. Based on the three-benchmark comparison reported in the abstract: consistently improved ambiguity coverage, reduced silent failures, and competitive execution accuracy.
Perspective
The result targets agentic Text-to-SQL settings where the agent interacts with users to clarify underspecified queries before generating SQL, and applies to evaluation setups focused on ambiguity coverage and silent failures; its evidence comes from three benchmarks derived from BIRD-Interact and Spider, so the scope of the conclusions is bounded by the interactive query tasks those benchmarks represent. For developers seeking to reduce the risk of answers that look correct but rest on unverified assumptions, PlanPool offers a way to externalize the clarification plan and enforce item-by-item disposition; for researchers, it shifts the evaluation focus from execution correctness alone to whether underspecification is actually resolved.
The abstract does not give concrete metric values, sample sizes, or statistical details for each benchmark, so the magnitude and stability of the improvements still require the full text; the specific distributional form of the observation that ambiguities concentrate in earlier interactions, and how silent failures are determined, are not elaborated in the abstract; how PlanPool behaves as the question pool grows or when user interaction turns are limited remains an open question for further observation.
