Skip to main content
Back to timeline
arXivSource publication:

CONTRA screens clarification questions by behavioral divergence, reaching top F1 across four coding agents and beating the best baseline macro-average by 13.88 points

Related research and updates

Synopsis

The authors propose CONTRA, a training-free method that broadly discovers candidate clarification questions, filters them semantically and qualifies them by checking for stable behavioral differences in programs generated under two plausible answers, then uses interaction history to ask or stop; on ClarifyCodeBench it achieves the highest F1 with all four coding agents, exceeding the best baseline macro-average F1 by 13.88 percentage points, and with the same LLM and evaluation protocol it achieves higher clarification recall and F1 than the Claude Code and OpenHands harnesses, and is implemented as a Claude Code plugin.

Source-provided article image: CONTRA: Discovering and Qualifying Behavior-Changing Questions for Selective Clarification in LLM Code Generation
Figure 1 ·

Figure 1: Overview of Contra for selective clarification.

arXiv

Interpretation

CONTRA splits clarification-question identification into broad discovery plus semantic and execution-based qualification, keeping only questions that change program behavior. Existing methods struggle to find key clarification questions while avoiding unnecessary ones; CONTRA grounds selection in generating programs under two plausible answers and checking for stable behavioral differences on shared inputs, rather than relying on semantic judgment alone. The abstract describes the pipeline: generate candidate questions, filter out those unrelated to required behavior or already resolved by the requirement, run conditional generation and behavioral-difference checks on the remainder, then use interaction history to select or stop; no per-stage sample sizes or ablations are given in the text.

On ClarifyCodeBench, CONTRA achieves the highest F1 with all four coding agents, exceeding the best baseline macro-average F1 by 13.88 percentage points. This extends the benefit of selective clarification from a single agent to consistent performance across four coding agents and quantifies the gap to the best baseline. Evidence comes from ClarifyCodeBench experiments reporting F1 and a 13.88-point macro-average gap; the text does not list absolute per-agent F1 values, variance, or significance tests.

With the same LLM and evaluation protocol, CONTRA achieves higher clarification recall and F1 than the coding harnesses Claude Code and OpenHands. This comparison places CONTRA against real coding harnesses under identical model and evaluation conditions, not only against abstract baselines. The abstract states 'same LLM and evaluation protocol' and reports higher recall and F1; specific numbers are not provided in the text.

CONTRA is implemented as a Claude Code plugin that integrates selective clarification into everyday development. The method moves beyond benchmark evaluation to an engineering artifact that can be plugged into a development workflow. The abstract states the plugin implementation; the text provides no usage data, deployment scale, or user-study results.

Perspective

The work targets development settings where coding agents generate code and requirements are underspecified, aiming to avoid behavioral mismatches early with as few questions as possible. Because the method is training-free, it suits teams that cannot or prefer not to fine-tune and want quick integration; the plugin form points to use inside everyday development workflows. Evaluation scope is four coding agents on ClarifyCodeBench plus comparisons with Claude Code and OpenHands under the same LLM and evaluation protocol.

The text is abstract-level and does not list absolute F1 or recall values, variance, or significance tests per agent, so the stability of the 13.88-point gap across agents is hard to judge. The behavioral-difference check depends on how shared inputs are constructed and how the two plausible answers are chosen; these details are not expanded in the text and may affect qualification coverage. How interaction history sets the threshold for continuing to ask versus stopping is also unspecified. In addition, the plugin's effectiveness in real development workflows and whether the interruption cost to developers is actually measured remain open questions.

Sources