Trident cuts DRL cyber-defense performance by an average of 522% with a single trainable 7B planner, autonomously surfacing decoy avoidance and adaptive state prioritization
Synopsis
The work introduces Trident, an agentic LLM red-teaming framework composed of a dynamic benchmark with isolated sandbox servers spanning CybORG CAGE 4 and CyberWheel, a dataset of over 13,000 high-fidelity red-blue interaction trajectories for RLVR, and a "Code-as-Policy" RLVR agentic architecture (Trident Agentic) that reformulates red-agent training as a contextual bandit via a Log Summarizer-Planner-Coder design, where a trainable Planner generates complete attack strategies from compressed execution logs and a frozen Coder translates them into executable Python policies deployed against live DRL defenders; empirical evaluation shows that with a single trainable 7B planner, Trident reduces blue-agent defensive performance by an average of 522% compared to static red-agent baselines whil
Figure 1: Overview of the Trident Agentic framework. A Log Summarizer observes and summarizes sanitized logs, enabling a trainable Planner ( π θ \pi_{\theta} ) to formulate abstract tactics. A frozen Coder translates the tactic into an executable Python policy deployed against a live DRL defender. Execution logs provide a verifiable reward signal to update π θ \pi_{\theta} via GRPO.
arXivInterpretation
Trident reformulates red-agent training as a contextual bandit through a tripartite Log Summarizer-Planner-Coder architecture: a trainable Planner generates complete attack strategies from compressed execution logs, and a frozen Coder translates them into executable Python policies deployed against live DRL defenders. Previously, DRL-based autonomous cyber defense was evaluated almost exclusively against static, heuristic red agents, leaving robustness against adaptive threats critically understudied; RLVR had improved LLM reasoning but its integration into cybersecurity remained elusive due to the absence of suitable benchmark environments and interaction datasets. The architecture is explicitly described as a "Code-as-Policy" RLVR agentic architecture with the three named components; the evaluation uses a single trainable 7B planner.
Trident ships with a dynamic benchmark and dataset: the benchmark includes isolated sandbox servers spanning CybORG CAGE 4 and CyberWheel, and the dataset comprises over 13,000 high-fidelity red-blue interaction trajectories for RLVR. The paper attributes the difficulty of bringing RLVR into cybersecurity to the absence of suitable benchmark environments and interaction datasets, and these two resources are proposed directly against that gap. The dataset scale ("over 13,000 high-fidelity red-blue interaction trajectories") and the two named environments are stated in the source text.
Empirical evaluation reveals a fundamental brittleness in existing defenses: compared with static red-agent baselines, Trident reduces blue-agent defensive performance by an average of 522%. The result shifts evaluation from static heuristic red agents to an adaptive LLM red agent, exposing defensive brittleness that static evaluation does not surface. The source reports the quantitative "average of 522%" comparison and emphasizes that the effect comes from a single trainable 7B planner.
Trident autonomously discovers emergent behaviors such as decoy avoidance and adaptive state prioritization, which static heuristics entirely fail to uncover. These behaviors are not hand-coded rules but arise autonomously as the trainable Planner operates against live DRL defenders, occupying a behavior space that static red-agent baselines cannot reach. The source uses the phrase "autonomously discovering emergent behaviors" and names decoy avoidance and adaptive state prioritization.
Perspective
The work targets autonomous cyber defense systems whose defense core is DRL, evaluated on isolated sandbox servers in CybORG CAGE 4 and CyberWheel, with an agentic LLM red agent (trainable Planner plus frozen Coder) against live DRL defenders. It makes adaptive red-team evaluation and RLVR training actionable: researchers can use the dynamic benchmark and the over 13,000 red-blue interaction trajectories to train and compare red-agent policies, and defender developers can re-examine robustness against adaptive rather than static heuristic opponents.
The reading scope here is the summary, without figures, statistical tests, or experimental configuration details, so the exact measurement basis for the 522% figure, the baseline set, differences across environments, and how emergent behaviors are identified remain open questions. In addition, the result comes from two sandbox environments (CybORG CAGE 4 and CyberWheel) and a single 7B planner setting, so generalization to other defense architectures, other network environments, and other planner scales is still to be observed; the long-term stability of RLVR in cybersecurity and behavior under co-evolution of red and blue also warrant continued tracking.
