Public articles linked to the same research event.
arXiv The authors propose CONTRA, a training-free method that broadly discovers candidate clarification questions, filters them semantically and qualifies them by checking for stable behavioral differences in programs generated under two plausible answers, then uses interaction history to ask or stop; on ClarifyCodeBench it achieves the highest F1 with all four coding agents, exceeding the best baseline macro-average F1 by 13.88 percentage points, and with the same LLM and evaluation protocol it achieves higher clarification recall and F1 than the Claude Code and OpenHands harnesses, and is implemented as a Claude Code plugin.
The authors propose CONTRA, a training-free method that broadly discovers candidate clarification questions, filters them semantically and qualifies them by checking for stable behavioral differences in programs generated under two plausible answers, then uses interaction history to ask or stop; on ClarifyCodeBench it achieves the highest F1 with all four coding agents, exceeding the best baseline macro-average F1 by 13.88 percentage points, and with the same LLM and evaluation protocol it achieves higher clarification recall and F1 than the Claude Code and OpenHands harnesses, and is implemented as a Claude Code plugin.
The authors propose CONTRA, a training-free method that broadly discovers candidate clarification questions, filters them semantically and qualifies them by checking for stable behavioral differences in programs generated under two plausible answers, then uses interaction history to ask or stop; on ClarifyCodeBench it achieves the highest F1 with all four coding agents, exceeding the best baseline macro-average F1 by 13.88 percentage points, and with the same LLM and evaluation protocol it achieves higher clarification recall and F1 than the Claude Code and OpenHands harnesses, and is implemented as a Claude Code plugin.
The authors propose CONTRA, a training-free method that broadly discovers candidate clarification questions, filters them semantically and qualifies them by checking for stable behavioral differences in programs generated under two plausible answers, then uses interaction history to ask or stop; on ClarifyCodeBench it achieves the highest F1 with all four coding agents, exceeding the best baseline macro-average F1 by 13.88 percentage points, and with the same LLM and evaluation protocol it achieves higher clarification recall and F1 than the Claude Code and OpenHands harnesses, and is implemented as a Claude Code plugin.