King et al. build the 4,047-session paired ACE corpus and find kernel-level syscall evidence is discriminative on its own for LLM agent attacks, with cross-layer composition generally beating either single-layer view
Synopsis
The work pairs application-level agent telemetry (served tool manifest, user prompt, model messages) with kernel-level syscall traces to present the first paired-evidence characterization of kernel-level versus application-layer signal for agent security, and introduces Agent Cross-Layer Evidence (ACE), a paired-session corpus of 4,047 sessions and 17 threat models spanning six delivery-vector families and 14 of the 25 OWASP LLM and agentic threat categories organized into 12 attack mechanics; across four detector families it finds kernel evidence is discriminative on its own and that composing it with application-layer evidence generally outperforms either single-layer view, and it further demonstrates generalization to unseen attack families and transfer to an alternate agent runtime.
Figure 1: Cross-View detection on ACE. Each session is configured, captured in-container with a strace kernel-trace sensor, emits three paired artifacts, and is scored across the App-View / Kernel-View / Cross-View by models from the three groups shown (ML/DL, LLM, fine-tuned SLM), which comprise the four detector families of § 3 .
arXivInterpretation
It introduces and releases ACE (Agent Cross-Layer Evidence), a paired-session corpus that pairs application-level agent telemetry with kernel-level syscall traces, at a scale of 4,047 sessions and 17 threat models. Existing agent-security benchmarks and defenses operate almost exclusively at the application telemetry layer, namely the served tool manifest, the user prompt, and the model's messages; this work provides the first paired-evidence characterization of kernel-level versus application-layer signal. The corpus spans six delivery-vector families and 14 of the 25 OWASP LLM and agentic threat categories, is organized into 12 attack mechanics, and includes per-mechanic characterization of where the most discriminative evidence lies.
Across four distinct detector families, kernel evidence is discriminative on its own, and composing it with application-layer evidence generally outperforms either single-layer view. This indicates the two layers carry complementary signals that single-layer analyses can miss, giving empirical support to the value of cross-layer evidence for agent security. The finding comes from comparisons across four distinct detector families rather than a single model or detector.
It demonstrates generalization to unseen attack families and transfer to an alternate agent runtime. It extends the value of cross-layer evidence beyond seen threats to unseen attack families and a different runtime environment. Generalization and transfer are reported as demonstrations in the paper; specific figures are not given in the abstract.
Perspective
The results target settings where LLM agents are deployed into infrastructure granting broad host authority and where kernel-level syscall traces can be collected; for such settings the work offers a reusable paired corpus and per-mechanic guidance on which layer holds the evidence, letting follow-up work choose the evidence layer by attack mechanic and treat cross-layer composition as a default detector design option. For deployments that can only obtain application telemetry and cannot collect syscall traces, the conclusions do not directly apply.
The abstract does not give per-detector-family metric values, the size of the gain from composition over a single layer, or the rationale for selecting the 17 threat models and 12 attack mechanics; the scale and conditions of the generalization and transfer experiments are likewise not expanded in the abstract. In addition, the loaded text is the arXiv abstract page and metadata without body figures or tables, so per-mechanic evidence locations and the four-detector comparisons can only be described at the abstract level and would need the original tables for confirmation.
