HarnessSecurity-Bench tests six coding agents: enabling auto-approve raises attack success from 29.2% to 95.6%
Related research and updatesSynopsis
The work builds a ten-mechanism taxonomy across 40 coding agent harnesses and rates 400 harness–mechanism cells (205 confirmed implemented, 83 confirmed absent, 112 unresolved), then introduces HarnessSecurity-Bench, which runs 2,500 trials under GLM-5.2 on 23 tasks across five attack surfaces with separate deterministic oracles for six harnesses, finding that enabling auto-approve raises attack success from 29.2% to 95.6%, that network isolation and read-only mode cut attack success from 57.1% and 8.3% to 0.8% at utility losses of 24.5 and 34.0 percentage points, and that command allowlisting and denylisting reduce attack effects with a small utility loss and a utility gain respectively.
Figure 1: Attack surfaces across the coding agent harness loop. The center shows five stages; panels A–F illustrate paths from adversarial resources or input to changed decisions or execution. Numbered markers identify stages that each surface can affect. HSB evaluates attacks on A–E.
arXiv · Page 3Interpretation
The study derives a ten-mechanism taxonomy (auto-approve, network isolation, command allowlisting, command denylisting, MCP permissions, prompt-injection filtering, path restriction, read-only mode, audit logging, project trust) and confirms 205 implemented, 83 absent, and 112 unresolved cells out of 400 harness–mechanism cells. Prior work was largely architectural surveys or empirical comparisons covering only a few harnesses; this work applies one taxonomy and an auditable rating protocol across 27 open-source and 13 closed-source products. Three researchers independently labeled 44 security labels that were consolidated into ten mechanisms; RQ2 used three LLM raters and two human researchers rating independently, with a majority on 382 of 400 cells and the remaining 18 adjudicated by three human researchers; ordinal Krippendorff's α was 0.740 across all five raters and 0.911 between the two human raters.
About half of confirmed implementations are opt-in (109 of 205 implemented cells rated O), and evidence gaps are larger for closed-source harnesses (69 of 130 closed-source cells unresolved, 53.1%, versus 43 of 270 open-source cells, 15.9%). It separates whether a mechanism exists from whether it is active by default and quantifies the external verifiability gap for closed-source products, for example all 13 closed-source prompt-injection-filtering cells being unresolved. Cell-by-cell ratings grounded in official documentation, CLI materials, and source code, using an Evidence-of-Thought method that requires each non-unknown rating to cite a source path and line range with recorded SHA-256 hashes.
Across 2,500 trials, enabling auto-approve raised utility from 77.1% to 95.4% and attack success from 29.2% to 95.6%; network isolation cut attack success from 57.1% to 0.8% at a 24.5 percentage-point utility loss, and read-only mode cut attack success from 8.3% to 0.8% at a 34.0 percentage-point utility loss. It provides the first paired ON/OFF comparison of native security mechanism settings under the same execution environment, measuring task utility and attack effects with separate deterministic oracles. Six harnesses, nine mechanism benchmarks, 10 trials per task–harness–setting combination, totaling 81,155 tool calls, over 2.2 billion tokens, and 589.32 hours of cumulative execution time, with trials run in isolated Docker containers.
Command allowlisting cut attack success by 48.1 percentage points at a 1.5 percentage-point utility loss, while command denylisting, path restriction, and MCP permissions cut attack success by 40.4, 48.1, and 28.1 percentage points with utility gains of 5.2, 0.3, and 0.5 percentage points; a case shows an allowed interpreter still executed an unauthorized checker despite the command restriction. It shows that restricting capabilities shared with legitimate work lowers both attack success and utility, that controls targeting harmful operations more precisely reduce attack effects at smaller utility cost, and that alternative execution paths leave residual attack surface. Paired ON/OFF comparisons plus task-level cases (in Dependency Lock Repair, Claude Code utility fell from 96% to 76% while attack success fell from 100% to zero; in CLI Arg Parser, Qwen Code under CAL ON still invoked the rejected checker through an allowed Python interpreter, passing all six utility checks and all three attack-effect checks).
Perspective
The results target developers, security assessors, and harness providers using coding agent harnesses, under a controlled setting of isolated Docker containers, GLM-5.2 as the base model, and 10 trials per task–harness–setting combination; the taxonomy and rating matrix cover 40 harnesses (27 open-source, 13 closed-source), while the runtime evaluation covers six open-source harnesses (Claude Code, Codex CLI, Gemini CLI, gptme, Qwen Code, GitHub Copilot) and nine mechanisms. Directly reusable artifacts are the task packages, utility checks, and oracle design for comparing attack effects, task utility, and token and time costs when mechanisms are enabled versus disabled on one's own workload.
RQ3 uses a single base model, GLM-5.2, and differences in tool use and instruction following across models may affect absolute attack success rates and the magnitude of ON/OFF differences; approval decisions in auto-approve-OFF trials come from a fixed LLM reviewer, so model bias and residual nondeterminism may affect individual decisions; Docker containers, automated execution, and simulated services do not fully reproduce real deployments; closed-source harnesses are excluded from the runtime evaluation because their interaction modes differ, particularly GUI-based IDE workflows; prompt-injection filtering supported a paired comparison only for gptme, so the remaining single-setting results cannot establish an ON/OFF effect; and 112 RQ2 cells remain unresolved, including all closed-source prompt-injection-filtering cells, so the actual coverage of those mechanisms remains an open question.
