Skip to main content
Back to timeline
arXivSource publication:

KaliBench's 8,504 query-command pairs show LLMs pick the right Kali tool but stumble on arguments

Synopsis

Researchers built KaliBench, a fine-grained benchmark for schema-free natural-language-to-CLI cybersecurity tool use on Kali Linux, containing 8,504 query-command pairs verified by LLM checking, sandboxed execution, and human review across 1,642 tools, 23 capability dimensions, and five security phases; evaluating 24 open-weight model configurations under unrestricted, restricted, and hinted settings, they found tool selection rises from 72.0% to 95.2% on average once candidates are constrained while optional-argument F1 barely moves, and only hints lift optional-argument F1 from 45.1% to 87.8% and Exact Correct from 22.3% to 73.

AI-generated editorial illustration: KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

Interpretation

KaliBench decomposes natural-language security requests into tool selection, optional arguments (alias-aware flag-value pairs), and positional arguments, and scores each component with Tool Accuracy, Optional-Argument F1, Positional-Argument F1, Total Score, and Exact Correct. Earlier cybersecurity benchmarks mostly test knowledge through multiple-choice or structured questions (SecEval, CyberMetric, CyberBench, SECURE, CS-Eval, SecBench, CTI-Bench) or test agents end-to-end on CTF-style tasks (CyBench, NYU-CTF Bench, DefenderBench), while general function-calling benchmarks (BFCL, ToolSandbox, StableToolBench, HammerBench) assume tools are explicitly defined by JSON schemas; KaliBench targets schema-free real CLI settings where models must infer both the tool and the syntax. Data come from official Kali Linux tool manuscripts (structured Markdown for 2,372 tools); Qwen3-Max generated 27.7K candidates, which passed LLM verification (53.8% filtered in the first round), Docker sandbox execution (a further 11.95% filtered), and human-in-the-loop review (another 4.9% removed), leaving 8,504 pairs, or 30.7% of the original generation.

Across three settings of increasing tool information, tool selection and argument construction behave as distinct challenges: the restricted setting mainly fixes which tool to use, while the hinted setting mainly fixes how the command is written. This decomposition lets failures be attributed to specific components, whereas end-to-end agentic evaluation entangles planning, environment interaction, observation interpretation, tool invocation, and error recovery. Averaged over 24 open-weight model configurations, moving from unrestricted to restricted raises Tool Accuracy from 72.0% to 95.2% but Exact Correct only from 22.3% to 28.3%; moving from restricted to hinted raises Optional F1 from 45.1% to 87.8%, Positional F1 from 61.1% to 90.1%, Total Score from 59.4% to 90.8%, and Exact Correct from 22.3% to 73.1%. The delta analysis in Appendix H shows the restricted-over-unrestricted Exact gain averages only 5.9 percentage points, while the hinted-over-restricted Exact gain averages 44.8 percentage points.

Runtime-free verifiable rewards support training: an 8B model built on the RedSage-Ins backbone reaches mean Total Scores of 77.4, 76.9, and 79.2 after SFT, GRPO, and SFT+GRPO respectively, with SFT+GRPO ranking third in mean Total Score among all open-weight models and only 1.0 point below DeepSeek-V3.2†. The reward is derived from component-level metrics (tool name, positional arguments, alias-aware optional arguments) plus an exact-match bonus, and because ground-truth commands were validated by sandboxed execution during data construction, rewards can be computed without executing model outputs, enabling RL without runtime. Training used LoRA on a single H200; SFT took 32 minutes, while GRPO took 28 hours from the backbone and 17 hours from the SFT checkpoint. The paper reports that SFT and RLVR improve the 8B model's mean Total Score across the three evaluation modes by 7.5 percentage points.

The benchmark remains unsaturated for stronger systems, and query wording affects performance. The paper additionally evaluates proprietary models and a scaffolded system, and systematically examines paraphrase, informal, and messy query variants, which is uncommon among comparable CLI benchmarks. On the 5,000 unrestricted-setting queries, GPT-5.6-Sol reaches 61.68% full-set Exact Correct (answering 4,988 queries), Codex (GPT-5.5) 51.68% (4,633 answered), and Claude Opus 5 44.02% (3,675 answered; 59.89% over answered queries), all above the strongest open-weight model GLM-5.2 at 41.32%. For query variants, paraphrases have little effect and even slightly help in unrestricted and restricted settings, informal wording causes a small consistent drop, and messy queries cause the largest drop (still 1.07 percentage points in hinted mode).

Perspective

The benchmark targets single-turn, documentation-grounded natural-language-to-CLI translation and is meant for security settings where local deployment matters because queries, target information, and operational data may be sensitive; its component-level metrics and runtime-free rewards provide reusable evaluation and training interfaces, and the authors point to multi-step workflows, retrieval-augmented tool use, and environment-aware evaluation as future extensions.

Readers should still watch that ground-truth commands come from the Kali tool manuscripts at construction time, so upstream documentation changes through additions, deprecations, renamings, or reinterpretations may make some canonicalized commands outdated; the manuals themselves occasionally conflict (for example, the regtree manual lists -V with both --no-values and --version semantics, and some Impacket tools document atypical short options such as -aesKey), making fully deterministic canonicalization inherently approximate; alias-aware scoring relies on semi-automatically extracted alias groups and cannot capture all historical or community-adopted variants; and the LLM-as-Verifier pipeline, though human spot checks indicate limited residual errors, is not perfect. In addition, coverage differences among proprietary systems (Claude Opus 5 answered only 73.50% of the test set) suggest provider-side safety restrictions affect availability, and how that shapes real deployment remains to be observed.

Sources