SafeActBench traces the evidence-to-action chain across 656 cases, finding strong static judgment but weaker interactive execution in tool-using agents
Synopsis
The authors introduce SafeActBench, 656 cases across six operational domains and five protocols, using a provenance-bound Evidence Ledger and a deterministic trajectory evaluator to check whether tool-using agents establish required evidence before acting; across ten model–harness configurations, strong static action assessment coexists with much weaker interactive execution, failures often begin with incomplete investigation or premature action, single-action execution is usually reliable once evidence is established, and multi-action workflows additionally expose unresolved prerequisites and incomplete execution.
Interpretation
The paper formulates consequential tool use as an observable evidence-to-action chain, requiring each state-changing action to be supported by evidence established before execution for the correct entity and state, and distinguishing information-gathering calls from consequential actions. Prior interactive evaluations focus on task completion, policy adherence, or overtrust in observations; this work makes whether pre-action evidence is established and bound to the action's entity and state an explicit evaluation target. The framework is instantiated by 656 SafeActBench cases; all author reference solutions pass the deterministic evaluator, and the specifications and evaluator were checked with human audits and negative controls, with scoring based only on observable trajectories and no LLM judge.
Across ten model–harness configurations, static action assessment and interactive execution diverge sharply: several configurations exceed 96% on Legacy yet differ widely on V1–V3. The gap persists when case identity is held fixed: on the same V1 cases, three configurations reach at least 95% static accuracy while interactive ECS is at most 52%. The main evaluation runs each case–configuration pair three times with 100% coverage and no infrastructure failures; paired comparisons use case-paired bootstrap confidence intervals.
Failures often arise before execution: incomplete investigation (BSR) ranges from 21.7% to 62.9% on V0, and premature action (PAR) ranges from 37.0% to 66.9% among V1 episodes with an action attempt, while conditional action success (CAS) after evidence completion is 93.2%–100% for nine of ten configurations. This shows aggregate success rates mask upstream differences in investigation, evidence binding, and action timing rather than reflecting execution capability alone. Diagnostics are computed by the deterministic evaluator on observable trajectories, with each episode counted at most once and rates computed within each repetition then averaged with equal weight.
Controlled interventions show that withholding a decisive record reduces action probability by 37.2–45.2 percentage points, yet agents still act in 46.5%–53.5% of completed Withheld episodes; a requester's claim that the missing record was checked reduces actions to 6/43, 10/43, and 6/43, larger than presenting the evidence package alone. Agents respond more to a requester's conflicting claim than to absent evidence alone, and additional retrieval and action restraint need not move together. Interventions use a frozen cohort of 43 V1 cases with paired case-bootstrap confidence intervals; the authors state these measure behavioral responses, not the agent's internal interpretation.
Perspective
The work targets researchers and evaluation designers studying evidence-to-action behavior in tool-using agents, and applies to stateful synthetic task environments with observable trajectories across six domains: customer and policy operations, engineering and infrastructure operations, legal and financial operations, research assistance, smart-home control, and healthcare operations. Its Evidence Ledger and deterministic evaluator can be used to localize failure points before deployment and support extending interventions to multi-action protocols with evidence varied at each action checkpoint. The authors state these domains are controlled settings and results should not be read as certifications of safety for deployed systems in the corresponding real-world domains.
The analysis captures observable support rather than internal reliance, so a gap remains between evidence being established and the model actually relying on it; the authors list extending interventions to multi-action protocols with evidence varied at each action checkpoint as a natural next step. Controlled interventions use a frozen cohort of 43 V1 cases, with smaller matched samples for Withheld and Contradicted conditions, and non-action is not counted as proof of correct refusal. In harness comparisons, native and Inspect runs were collected at different times with different timeout and recovery histories, limiting attribution to the harness alone. In the supplementary study, SCGR-Select effects vary in direction across configurations and protocols, and some comparisons use different case sets and coverage, so they should be read within their respective case sets. The loaded material is full text, with figure values presented in text and tables; verifying curve details would still require the original figures.
