Cybersecurity evaluation of LLMs has long been stuck on two extremes: CTF-style knowledge questions that ask "does the model know about it" and end-to-end agentic run-throughs that ask "can the model finish the loop". The capability sitting in the middle that actually decides whether a security engineer can put AI to work — translating a natural-language request into a single nmap, sqlmap, or burp command that actually executes correctly on the terminal — has not been measured directly. The reason is not subtle: command lines are intolerant of argument order, flag aliases, and key-value bindings, and grading has to escape the model's self-description and align with the real tool's behavior.

A team from Khalifa University and the University of Western Australia closed that gap. Posted on arXiv on October 1 (2610.02206, NeurIPS 2026 Evaluations and Datasets Track), KaliBench is a fine-grained cybersecurity tool-use benchmark targeting Kali Linux. The data is grounded in official tool documentation, goes through LLM verification plus sandboxed terminal execution plus human review plus semantic deduplication, and ends with 8,504 query-command pairs (3,504 for training, 5,000 for evaluation) covering 1,642 tools, 23 capability dimensions, and 5 security phases (reconnaissance & initial access, vulnerability analysis, exploitation & payload delivery, post-exploitation & lateral movement, defensive analysis & reporting).

The hard part is not picking the tool, it is writing the right arguments

KaliBench splits evaluation into four pieces: tool selection, optional-argument F1 (alias-aware matching), positional-argument F1 (multiset matching), and strict "exact-command correctness" (which permits documented aliases and reordered optional flags but preserves positional argument order). Under the unrestricted setting with no tool hints, the strongest of 24 open-weight models is GLM-5.2 (753B) at 41.3%; DeepSeek-V3.2 (685B total / 37B active MoE) reaches 33.0%, Qwen3-Coder-Next (80B / 3B MoE) 26.2%, and the median band sits mostly in the 20-30% range — no model exceeds 42% exact-command accuracy. The restricted setting that provides 20 candidate tools pushes the average up to 28.3%; the hinted setting that hands the model the target tool plus its documentation jumps straight to 73.1%, a spread of close to 50 percentage points.

That curve says one thing: tool selection is not the bottleneck. RedSage-K (the paper names its self-trained model Kali-SFT+GRPO; the project page calls it RedSage-K, the name used here) lands at 77.9% tool selection, in the same band as DeepSeek-V3.2's 75.8%; the real gap lives in optional-argument F1 and positional-argument F1 — the 8B model still falls behind on optional-parameter alias recall, key-value bindings, and argument order preservation. For engineering teams this means KaliBench delivers not a binary answer to "can LLMs do cybersecurity" but a quantifiable diagnostic of which exact row of the tool-call interface each model sits in.

Training signal from ground-truth commands, not from running models

The other value of KaliBench is "runtime-free verifiable rewards". The dataset construction stage has already verified command correctness for every sample; during training, deterministic scoring can return rewards for tool selection, optional arguments, positional arguments, command-level exact correctness, and output format — all without executing model-generated commands inside the training loop. This fits the trend of verifiable-reward RL (RLVR) becoming a popular training paradigm over the past year, but KaliBench pushes that design into a sub-area — cybersecurity tool calling — that RLVR has not yet penetrated.

The authors take RedSage-Ins 8B as the base, do supervised fine-tuning first, then GRPO reinforcement learning, to obtain three RedSage-K variants (SFT only, GRPO only, SFT+GRPO). On the average total score across the three evaluation settings, the 8B baseline 71.7 → SFT only 77.4 → SFT+GRPO 79.2, while DeepSeek-V3.2 at 685B / 37B active sits at 80.2. A 1.0-point gap with a roughly 18× parameter-count spread is effectively a tie. The open-weight RedSage-K (SFT-GRPO) is released on Hugging Face (RISys-Lab/RedSage-K-SFT-GRPO), the training and evaluation code is on GitHub (RISys-Lab/KaliBench), and the data itself is published in the HF collection.

Where the proprietary systems push the ceiling

Paper Table 2 supplies the closed-source comparison: GPT-5.6-Sol answers 4,988 of the 5,000 evaluation prompts with 61.83% exact-command accuracy on answered (61.68% over all 5,000); Codex CLI (GPT-5.5, xhigh reasoning) answers 4,633 with 55.77% on answered (51.68% over all); Claude Opus 5 answers 3,675 with 59.89% on answered (44.02% over all). Note that the all-5,000 number counts refusal-to-answer and non-generation as failures — in a cybersecurity domain that may trigger safety guardrails, whether the model dares to answer is itself part of the score. Read together with the open-weight block, KaliBench's picture is: the closed-source lane clears 60%+ exact-command accuracy, the open-weight lane is stuck in the 30-40% band, but targeted training lets an 8B open-weight model close the gap to 1.0 percentage point.

For teams building cybersecurity tooling, KaliBench turns "autonomous tool reasoning" from a slogan into a quantifiable, trainable, comparable experimental surface — an 8B model scoring 79.2 versus DeepSeek-V3.2 scoring 80.2, a 1-point gap, is closer to actual engineering reality than either "LLMs can't do this" or "LLMs can do this".

References: arXiv:2610.02206 (risys-lab.github.io/KaliBench); GitHub repository RISys-Lab/KaliBench; Hugging Face collection RISys-Lab/kalibench-datasets-and-models; GPT-5.6-Sol, Codex CLI, and Claude Opus 5 numbers from paper Table 2.