Skip to main content
Back to timeline
arXivSource publication:

IntentFlux turns "users changing their minds" into a measurable failure mode: mean task score falls from 0.476 to 0.384 in a 627-case calibration, and StateForge lifts General-Test mean score from 0.367 to 0.467

Synopsis

The work introduces IntentFlux, an executable benchmark that rewrites verifiable tasks into multi-turn dialogues with controlled intent changes (Variant replacements and Decoy withdrawals) while preserving the original graders, observes in a 627-case calibration that mean task score falls from 0.476 to 0.384 as superseded and withdrawn information accumulates, finds that across eight models the fully-correct rate is lower when the same final task must be recovered from an evolving dialogue than when given directly in a single turn, and further introduces StateForge, which explicitly maintains the active requirement state before generation and raises General-Test mean score from 0.367 to 0.

AI-generated editorial illustration: When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents

Interpretation

IntentFlux defines intent drift as the failure mode in which superseded parts of the user's intent still influence the final answer or tool action, and makes it executable and scorable by converting verifiable source tasks in code, math, SQL, data-to-text, summarization, and tool use into controlled multi-turn dialogues while preserving each source task's original grader. Earlier multi-turn evaluations, such as sharded-instruction settings and evolving-intent evaluations, measure whether an agent ultimately follows an evolving task but do not separate failures caused by obsolete intent continuing to influence the answer from other multi-turn errors; IntentFlux uses Variants (plausible alternative values later replaced) and Decoys (plausible constraints later withdrawn, retained only when obeying them changes the source grader's verdict) so that stale-intent use shows up directly in task success. A 627-task General calibration pool, a fixed 125-task General-Test, a 100-task Interactive-Test, and a separate 502-task training pool; human annotators verify the variant, decoy, and semantic-dependency pools, and a downstream audit recovers the annotated final intent for 95% of General-Test.

Under controlled stale-information load, performance declines monotonically: mean score falls from 0.476 in the easy stratum to 0.384 in the hard stratum, and full-credit CAR falls from 0.381 to 0.298; all eight tested models show lower full-credit rates under the evolving-dialogue condition than under single-turn presentation, at every difficulty level. The result turns "users change their minds" from a qualitative worry into a reproducible difficulty response, using a paired design (same source case, final intent, grader, and tested model) to report a clean-drift intent-drift gap rather than only an average multi-turn score. Calibration uses gpt-4.1 for construction, gemini-3.5-flash as the tested model, and deepseek-v4-flash as user simulator and rubric judge; easy/medium/hard use variant/decoy budgets of 1/3/5 with realized means of 0.95/0.61, 1.71/1.67, and 2.24/2.50, and mean dialogue length grows from 13.2 to 37.1 turns; key comparisons use case-paired bootstrap confidence intervals.

The loss is not explained by dialogue length alone: on General-Test, a turn-matched no-drift control scores 0.734, close to the 0.778 single-turn condition, whereas the corresponding drift trajectories score 0.469, a 0.265 gap; in the live tool-use Interactive-Test, withholding an explicit restatement of the final intent lowers partial-credit rubric score relative to the clean condition. This rules out generic long-dialogue degradation as the explanation and locates the problem in recovering the current intent from an evolving history; the interactive study shows the effect extends to real tool actions rather than only static final answers. The turn-matched control and drift conditions have identical turn counts and comparable user-message length (13.2k versus 14.6k mean characters); the interactive study uses Qwen3.5-122B as the tested agent, and all three non-oracle conditions share the same CAR of 0.190, so this evidence concerns partial-credit rubric evaluation.

History compression is not sufficient: rolling-summary Deep Agents scores 0.354 and two-phase-compaction OpenHarness scores 0.392, neither reliably above the 0.367 of bare multi-turn execution; StateForge folds the edit history into an explicit active state before generation and raises General-Test mean score to 0.467, while supplying the ground-truth final state raises it further to 0.551, still below the 0.778 clean single-turn score. The authors separate what history management asks (what to retain) from what state folding additionally asks (what is still valid): conclusions derived from a removed intent item can persist even after the item is gone, so an explicit representation of active requirements and their dependencies is needed; the ground-truth-state result shows state-estimation error is material but explains only part of the remaining gap. The harness comparison runs on the same frozen medium-difficulty General-Test trajectories under the same evaluator, temperature, and output budget; in the component-matched ablation, paired differences between each recap baseline and the full harness have 95% confidence intervals that include zero, while the oracle condition shows a positive paired difference from the full harness.

Perspective

The work targets tasks with executable or decomposable intent structure: the General track covers code, math, SQL, data-to-text, summarization, and tool use, and the interactive track uses VitaBench delivery, in-store, OTA, and cross-domain environments. It applies to agent deployments that must recover the current intent from an evolving dialogue before acting, including a modular setup where a smaller tracker supplies the active state at inference time and a setup where state folding is internalized into the base agent. For readers wanting to reuse the results directly, the transferable objects are the IntentFlux construction (Variants/Decoys, preserved original graders, no final-intent restatement) and the StateForge interface of an explicit active state before generation; the tracker-scale comparison finds a 9B tracker statistically indistinguishable from the 122B reference, OPD improves a 2B tracker from 0.156 to 0.399, and thinking distillation raises a standalone 35B agent from 0.296 to 0.461 on the drift condition.

The calibration couples variant and decoy budgets, so it supports a joint stale-information load effect rather than separate causal effects of variants, decoys, or length; the turn-matched control rules out turn count alone, but message ordering and edit wording remain properties of the drift construction. The harness comparison and Interactive-Test each use a single tested model, the former on replayed trajectories. The component-matched ablation does not establish the explicit state representation as the sole source of the system-level gain, and the scaling behavior of both transfer paths remains open. In the interactive study, restating the complete final intent during closing turns removes the measured partial-credit gap, suggesting a benchmark can inadvertently supply the state it intends to test. In addition, several table values are elided in the prose of this evidence bundle, so exact per-model and per-family scores still require the original tables.

Sources