Skip to main content
Back to timeline
arXivSource publication:

CVE2AP uses LLMs to auto-generate PDDL attack paths from CVE descriptions, reaching 86.9% syntax and 93.1% semantic correctness with GPT-5.5

Related research and updates

Synopsis

The work proposes CVE2AP, which combines structured prompting with planner-driven error feedback so an LLM can automatically generate PDDL-encoded attack paths from natural-language CVE descriptions; evaluated on 21 CVEs, five models and 12 configurations, the best model GPT-5.5 reaches 86.9% syntax correctness, 78.6% solvability, 53.8% embedding similarity and 93.1% LLM-as-expert semantic correctness, with error feedback giving the most consistent gains and GPT-5.5 the best quality-cost trade-off.

Source-provided article image: CVE2AP: Automated Generation of PDDL-Encoded Attack Paths via Large Language Models
Fig. 2 ·

Fig. 2 : The overview of the CVE2AP.

arXiv

Interpretation

CVE2AP automatically converts natural-language CVE descriptions into PDDL-encoded attack paths while targeting syntax correctness, planning solvability and cybersecurity semantic fidelity. Existing PDDL attack-path modeling largely relies on expert manual construction, while general LLM-based PDDL domain generation does not address attack-path-specific requirements such as threat-model compliance, exploitation-mechanism fidelity and attack-stage completeness. The approach is model-agnostic and can be instantiated with any LLM supporting chat or completion interfaces; it is evaluated on 21 CVEs with the S3Eval three-layer pipeline covering syntax, solvability and semantics.

An error-feedback mechanism driven by a domain-independent planner iteratively refines outputs using planner-reported syntax and solvability errors, with up to three retries in the same conversation. Compared with prompt constraints alone, it brings externally verifiable planner error signals into the generation loop without human intervention. Error feedback improves syntax by up to 49 percentage points, solvability by up to 41 points and LLM-as-expert semantics by up to 12 points, while embedding similarity slightly degrades; planner checking completes within milliseconds per domain.

The study characterizes quality-cost profiles across models and configurations: GPT-5.5 gives the highest quality and best quality-cost trade-off, and locally hosted Qwen3 models improve consistently with scale. Prior work did not jointly report syntax, solvability, semantics, token consumption and generation time across multiple models and configurations for attack-path generation. GPT-5.5 reaches 86.9% syntax, 78.6% solvability, 53.8% embedding similarity and 93.1% expert semantics; Qwen3-32B reaches 39.7% syntax, about one-third solvable and over 20% semantic; GPT-5.5 has a median generation time of about 85 s, while Deepseek-V4-Pro takes about 65 s but about 53,000 tokens per sample.

Configuration analysis shows error feedback is the most consistently effective strategy, whereas few-shot examples and structured output templates yield model-dependent benefits with added cost. It quantifies configuration dimensions as independent variables affecting both quality and cost, rather than reporting only a single best setting. Few-shot examples monotonically increase time and tokens, with Deepseek-V4-Pro requiring more than twice the tokens in 1-shot and 2-shot than 0-shot; the output template improves solvability by only about 1-5 percentage points and can add prompt complexity for weaker local models.

Perspective

The results target generating single-vulnerability attack paths from one CVE description, suited to cybersecurity analysis workflows that need planner-executable and semantically faithful representations. The approach is model-agnostic, instantiable on online commercial models or locally hosted models, and supports combinations of 0/1/2-shot prompting, error feedback on or off, and output template on or off. For teams seeking to reduce expert manual modeling effort and quickly turn threat intelligence into formal representations, this pipeline provides a reproducible starting point; the error-feedback mechanism is especially relevant where a domain-independent planner is already available.

Semantic correctness is currently judged by embedding similarity and LLM-as-expert, where embedding similarity slightly degrades under error feedback, and existing semantic evaluation provides quality judgments rather than locatable correction feedback, so semantic feedback remains an open direction. Evaluation covers 21 CVEs and a specific model set, leaving generalization across vulnerability types and larger scales to be observed; GPT-5.5 is subject to content safety filtering, so a small fraction of generations is non-deterministically blocked and retried. Multi-vulnerability and chained-exploitation paths are outside the current scope.

Sources