SAKIKO audits tool-call interventions across seven LLMs: behavioral movement is not repair, and only Qwen3-8B earns a formal ADMIT
Synopsis
The work presents SAKIKO, an auditing framework that decomposes internal activation interventions on pre-execution tool decisions into directional error discovery, router-conditioned intervention, destination-resolved verification, and prospectively frozen statistical licensing; across seven LLMs on When2Call and MetaTool, channel-keyed interventions induce direction-specific net gains in five models, yet destination auditing shows behavioral movement is not repair—an intervention with +55 net gain corrupts over half of the baseline-correct decisions it touches, Qwen3-4B and Gemma-2-9B are formally declined due to finite-sample uncertainty, and only Qwen3-8B passes all ten conditions for an ADMIT.
Interpretation
The paper formalizes pre-execution tool-decision errors as ordered directional error channels within a K-way action space, and shows that leaving an error state does not guarantee reaching the correct target. Prior tool activation-steering work evaluates interventions in binary action spaces (call/no-call) or pairwise tool swaps, where vacating an incorrect state coincides with arriving at the alternative; this work extends evaluation to four action modes (direct answer, API call, clarification request, task declination) and defines five mutually exclusive outcome classes: source retained, gold arrival, other wrong, correct retained, and broken. Decisions are read out on the four-action When2Call benchmark via teacher-forced scoring with a 2,556/548/548 train/validation/test split; on Phi-3.5-mini, 153 of 200 routed errors vacate the source state, but only 107 reach the gold target while 46 spill into alternative error classes.
SAKIKO splits repair into five stages that can fail independently, and shows that aggregate net gain simultaneously conceals lateral error redistribution and collateral damage to baseline-correct decisions. Prior work measures intervention success by aggregate accuracy or binary state exits; this work introduces destination-resolved verification and two collateral denominators (population E1 and exposure-conditional E2), showing that the same net gain can correspond to entirely different destination compositions. On Phi-3.5-mini, an intervention worth +55 net gain corrupts 52 of the 93 baseline-correct decisions its Router fires on; on Qwen3-8B, internal activation steering and a score-space baseline reach near-identical net gains (38 versus 37 gold arrivals), yet internal steering misdirects only 14 of 52 exits into alternative errors while logit shifting diverts 21 of 58 exits.
A pre-registered ten-condition statistical licensing gate separates point estimates from statistically robust repair, and two sealed settings are formally declined because their confidence intervals fall short. Prior intervention studies generally report favorable point estimates without an accompanying uncertainty criterion; this work requires confidence intervals rather than isolated point estimates to satisfy prospective thresholds, and adjudicates against 59 budget-matched random directions via an add-one Monte Carlo test. Qwen3-8B clears all ten conditions for an ADMIT (target-hit 0.7308, target gain 0.2759, zero breaks); Qwen3-4B and Gemma-2-9B receive DECLINE verdicts because their target-gain intervals cross zero and target-hit intervals breach thresholds; none of the 59 random directions matches the calibrated direction's target gain.
Linear decodability does not imply steerability, and router gating limits exposure but does not by itself establish preservation among the decisions it touches. The work shows that probe accuracy plateaus across intermediate layers while downstream steering efficacy drops sharply; optimal injection site, dosage, and intervention success diverge from probe accuracy, and the effect of gating on net gain is setting-dependent. On Qwen2.5-7B, linear probe decodability plateaus at roughly 0.9 across layers 16 to 22, while validation net gain jumps from about 0 at layer 16 to about 0.6 at layer 20; removing gating flips net gains negative on Phi-3.5-mini and MetaTool, and on Qwen3-4B raises breaks from 1 to 6 while holding arrivals at 37.
Perspective
The framework targets researchers and evaluators performing inference-time activation interventions on frozen weights, in settings with defined gold labels, a baseline-error population, and an exposed reference partition. Formal licensing outcomes rest exclusively on When2Call; MetaTool supplies binary historical evidence, and ACEBench was retired before intervention for insufficient channel support. The single admitted configuration is gradient activation steering on Qwen3-8B's cannot_answer to tool_call channel, and that license certifies only categorical action-mode selection, not downstream execution validity, and is not a deployment-safety claim.
Destination-resolved results depend on When2Call labels that are synthetically generated by the benchmark's authors without post-hoc re-annotation, so unmodeled systematic labeling biases could shift destination distributions without triggering gate violations. Exposure-conditional preservation is not certified in any setting: Qwen3-8B observes zero breaks over only six router-exposed correct decisions, giving a one-sided upper bound of 1.0, whereas certifying a 0.05 bound requires at least 59 exposed correct instances. The declines for Qwen3-4B and Gemma-2-9B are driven entirely by bootstrap interval conditions rather than point-estimate failures, and the pre-intervention halt of Qwen3.5-9B was split-sensitive, clearing eligibility in roughly 90% of retrospective split reallocations. No stable ordering is established between activation intervention and the score-space comparator, and paired tests after multiplicity correction remain non-significant. In addition, this evidence bundle is a full-text parse, so if figure or table details are not fully rendered, verification of individual numbers may be affected.
