Public articles linked to the same research event.
arXiv The work argues that in agentic Text-to-SQL a correct execution result does not necessarily mean the underlying underspecification has been adequately resolved, since the agent may silently make unverified assumptions that happen to match the intended answer; it attributes this partly to premature clarification termination and introduces PlanPool, which externalizes the clarification plan as a mutable question pool where every planned question must be explicitly asked or dropped before submission and newly discovered ambiguities can be added during interaction, consistently improving ambiguity coverage and reducing silent failures across three benchmarks derived from BIRD-Interact and Spider while maintaining competitive execution accuracy.
The work argues that in agentic Text-to-SQL a correct execution result does not necessarily mean the underlying underspecification has been adequately resolved, since the agent may silently make unverified assumptions that happen to match the intended answer; it attributes this partly to premature clarification termination and introduces PlanPool, which externalizes the clarification plan as a mutable question pool where every planned question must be explicitly asked or dropped before submission and newly discovered ambiguities can be added during interaction, consistently improving ambiguity coverage and reducing silent failures across three benchmarks derived from BIRD-Interact and Spider while maintaining competitive execution accuracy.
The work argues that in agentic Text-to-SQL a correct execution result does not necessarily mean the underlying underspecification has been adequately resolved, since the agent may silently make unverified assumptions that happen to match the intended answer; it attributes this partly to premature clarification termination and introduces PlanPool, which externalizes the clarification plan as a mutable question pool where every planned question must be explicitly asked or dropped before submission and newly discovered ambiguities can be added during interaction, consistently improving ambiguity coverage and reducing silent failures across three benchmarks derived from BIRD-Interact and Spider while maintaining competitive execution accuracy.
The work argues that in agentic Text-to-SQL a correct execution result does not necessarily mean the underlying underspecification has been adequately resolved, since the agent may silently make unverified assumptions that happen to match the intended answer; it attributes this partly to premature clarification termination and introduces PlanPool, which externalizes the clarification plan as a mutable question pool where every planned question must be explicitly asked or dropped before submission and newly discovered ambiguities can be added during interaction, consistently improving ambiguity coverage and reducing silent failures across three benchmarks derived from BIRD-Interact and Spider while maintaining competitive execution accuracy.