ControlScope compares three revision permissions from the same execution state: wider rewriting solves 2–3 more filesystem tasks but can also interrupt viable continuations
Synopsis
ControlScope compares three nested revision permissions from the same public execution state—continuing the current program (Keep), editing only the next tool call's data arguments (Arg), and replacing the unfinished workflow (Full)—evaluating one-time and repeated reviews across filesystem tasks, ALFWorld, and AppWorld, and finds that Full completes 15–16 of 20 filesystem tasks versus 13 for Keep across two source programs and three reasoning-reviewer draws, 10–13 versus 13 across four fast draws, scores 85/86/87 on 87 ALFWorld tasks and 134/134/127 on 134 tasks, shows small net differences on 585 AppWorld V1 official-test task instances, while frozen replays expose viable agent-written replacements interrupted by later revision.
Interpretation
The work separates available repairs from the repairs an agent actually selects: Keep, Arg, and Full are nested permission sets, and Full includes every smaller-scope operation, so losses from the larger set reflect the deployed selection and execution policy rather than the permission itself. Prior work such as GraSP compared typed local repair with global replanning, and other studies addressed the tension between frequent plan changes and consistent execution; this paper's addition is to make scope a nested operation set and to log both granted scope and realized scope from the same public execution state. Across 20 filesystem tasks, two source programs per task, and three reasoning-reviewer draws, Full completes 15–16 tasks versus 13 for the shared Keep; four fast draws range from 10–13 versus 13, showing that the source program and reviewer draw shape which tasks benefit.
Wider revision scope repairs concrete program defects: batch reads and a parser fix. The paper gives mechanism-level evidence rather than totals alone: a retained batch-read replacement reaches the goal after 94 total calls while the original loop exhausts the 100-call budget, with both branches sharing the first 88 calls; a second source completes at 28 total calls from an identical 26-call prefix. In the student-record task the 19 selected students are independently checked against all original records; the parser repair produces correct results for all 150 students; across 25/75/150-student subsets both reviewer modes score 5/6, 5/6, and 6/6, repeated on three new subsets with two new source programs each.
Later reviews can interrupt a continuation that was already viable, and both tested retention rules yield no panel-level gain. Frozen replays separate whether a candidate is viable from whether it is allowed to finish: in the file-time classification task, 49 of 71 actual replacements in the first source satisfy the goal when retained through the block endpoint or the original call budget, and 46 of 62 do so in the second source. Across every replacement-producing run in the P2 automatic-tool fast panel and the first P2 reasoning draw, sustained revision succeeds on 5/10 and 9/13 runs while retaining only the first revision yields 4/10 and 7/13, with zero paired gains and three losses; five-call protection saves 19.4% of logged output and loses one success across 20 fresh source runs.
Review timing and valid-operation return rates jointly shape task success and model work, not scope alone. The paper folds review scheduling, invalid-review fallbacks, and token cost into the same evaluation: on the 87-task reasoning panel Full gains two terminal successes over Keep with 29.7 times its output tokens, and Arg/Full record 51/16 invalid review replies that remain in the scored outcomes. On the 87-task reasoning panel initial-plan successes are 79/81/85 and become 85/86/87 after ordinary recovery; Full uses 1,124 actions versus 1,336 for Keep and invokes ordinary repair twice versus 17 times; under the same fast protocol the 134-task cohort ends 134/134/127, the opposite direction.
Perspective
The work targets controlling an unfinished program inside an otherwise functioning agent loop: every condition retains ordinary planning, and in ALFWorld every condition also retains the same ordinary failure-recovery procedure, so the comparison concerns control over an already-generated program rather than overall agent capability. It applies to filesystem-style MCP tasks, ALFWorld household plans, and the AppWorld official-test panel, each with its stated execution contract, call budgets, and reviewer configuration. For designers, the directly usable conclusion is to expose both data-argument editing and workflow replacement and to include review timing and valid-operation return rates in evaluation; the batch-read and parser-repair mechanisms show where wider permission pays off.
Results are sensitive to the source program and reviewer draw: two source programs for the same task can diverge, and four fast draws range from three fewer successes to equal. The two fixed ALFWorld cohorts point in opposite directions, with Full gaining two tasks on the 87-task cohort and losing seven on the 134-task cohort, so a single cohort's direction is not a general conclusion. In AppWorld V1 each Full replacement also renews a 250-request block allowance, extra capacity that favors Full, so its observed difference combines revision scope with that capacity. At the fixed early local boundary all 216 paired Arg and Full draws have identical terminal outcomes, so that paired equality does not extend to sustained policies. The interruption evidence remains case-level: neither tested retention rule yields a panel-level gain. Several outcomes also involve service errors and context limits, for example all 12 service errors arise from a single oversized context, and excluding that state preserves the zero observed Arg–Full difference.
