CodeScan audits code-generation LLMs with black-box, vulnerability-oriented scanning, reporting 97%+ poisoning detection accuracy across 117 models
Synopsis
The work presents CodeScan, a black-box, vulnerability-specific scanning framework for auditing poisoning and backdoor attacks in code-generation LLMs: it identifies attack targets by analyzing structural similarities across multiple generations conditioned on different clean prompts, combining iterative divergence analysis with abstract syntax tree (AST)-based normalization to abstract away surface-level variation and unify semantically equivalent code, then applies LLM-based vulnerability analysis to determine whether the extracted structures contain security vulnerabilities and flags the model as compromised when such a structure is found; evaluated against four representative attacks under both backdoor and poisoning settings across three real-world vulnerability classes, experiments o
Figure 1 . Example of Attacks to Code Generation LLMs
arXivInterpretation
It introduces CodeScan, described as the first black-box, vulnerability-specific scanning framework for code-generation LLMs, under the assumption that the defender specifies the target vulnerability classes and provides corresponding task-relevant prompts. Existing scanning approaches rely on token-level generation consistency to invert attack targets, which the paper states is ineffective for source code where identical semantics can appear in diverse syntactic forms; CodeScan instead targets vulnerability classes under a black-box setting. This claim comes from the abstract's framing of the method and of prior methods' limitations; it is a design-level statement, and the abstract does not provide item-by-item mechanistic comparisons with prior methods.
CodeScan identifies attack targets by analyzing structural similarities across multiple generations conditioned on different clean prompts, combining iterative divergence analysis with AST-based normalization to abstract away surface-level variation and unify semantically equivalent code, thereby isolating structures that recur consistently across generations. It moves the detection signal from token-level consistency, which is sensitive to syntactic diversity, to a structural level after AST normalization, so that semantically equivalent but differently written code can be merged and compared. The abstract names the method's components (iterative divergence analysis, AST normalization, cross-generation structural recurrence) but does not provide per-component ablation results or specific thresholds.
After extracting structures, CodeScan applies LLM-based vulnerability analysis to determine whether those structures contain security vulnerabilities, and flags the model as compromised when such a structure is found. It chains structure extraction and vulnerability judgment into a single decision pipeline, so the scan output lands on the actionable conclusion of whether the model is compromised rather than only surfacing suspicious generated fragments. The abstract states the existence of this judgment step and its trigger condition, but does not give the accuracy of the vulnerability-analysis step itself or details of human verification.
Evaluated against four representative attacks under both backdoor and poisoning settings across three real-world vulnerability classes, experiments on 117 models spanning three architectures and multiple model sizes report 97%+ detection accuracy with substantially lower false positives than prior methods. Relative to prior methods, the work reports both high detection accuracy and lower false positives, with broad coverage in model count, architectures, and sizes. Evidence comes from the experimental scale described in the abstract (117 models, three architectures, four attacks, two settings, three vulnerability classes) and the quantitative results (97%+ accuracy, substantially lower false positives); the abstract does not provide per-item numeric tables, confidence intervals, or statistical tests.
Perspective
The framework targets a setting in which the defender already knows which vulnerability classes to audit and can provide task-relevant prompts, so it fits code-generation model evaluation scenarios with a defined security audit goal, such as black-box checks on specific vulnerability classes before model selection or deployment. Its output is a flag of whether the model is compromised, which can trigger further human review or deeper forensics. For users who need to cover unknown vulnerability classes, or who cannot provide task-relevant prompts, the abstract does not describe how the method applies.
The abstract does not break down the 97%+ accuracy across different attacks, vulnerability classes, or model architectures, nor does it state the concrete false-positive rate or how it is computed, making it hard to judge how stable that figure is across conditions. The abstract also does not describe the reliability of the LLM-based vulnerability-analysis step, whether it was human-verified, or how sensitive detection is to prompt wording. In addition, the available text is only the abstract and browsing context, without the body, figures, or experimental details, so the description of internal mechanisms and experimental conditions is limited to what the abstract states; concrete implementation and per-item results still require the original paper.
