When Agents Look Like Beacons: NIDS Evasion by Model Context Protocol Traffic
Synopsis
In a controlled Docker-based testbed, this study evaluated Suricata signature matching and RITA behavioral beacon scoring against eleven mathematically defined traffic profiles (including MCP task-driven, orchestrated, burst, jittered, and User-Agent-spoofed variants) across three TLS conditions (Opaque TLS, TLS-Inspected, and Cleartext), finding that MCP JSON-RPC traffic produced near-zero alerts under the ET Open ruleset and a consistent 0.0 RITA beacon score, with jitter injection and User-Agent spoofing leaving sensor output unchanged, and proposing Agent-Native ALPN and schema-aware stateful inspection rules as remedies.
Fig. 1: RITA beacon scores across representative traffic profiles. The horizontal dashed line at 0.85 marks the CISA operational detection threshold, shown as a reference point from a calibration run using a known C2 packet trace [ 7 ] . All MCP profiles (P4–P10) score 0.0. The synthetic C2 profiles (P2 and P3) also score 0.0 due to a RITA boundary condition explained in Section V-C .
arXivInterpretation
Under Opaque TLS, sensors observed only TLS handshakes, IP addresses, TCP ports, and flow-level metrics, and the JA3 fingerprint Zeek extracted from the ClientHello was cf712d3b01ab6bfd22ca983b4748ab2f, which returned zero malicious associations against the Abuse.ch SSLBL threat intelligence database; under TLS-Inspected conditions, mitmproxy terminated and re-encrypted the session, exposing the full MCP JSON-RPC payload structure including method names such as initialize, tools/list, and tools/call and nested parameters containing terminal commands, SQL queries, and local file paths. Prior discussion of MCP centered on protocol design and threat taxonomies; this provides a measured contrast of what sensors actually observe under different TLS visibility conditions. Controlled measurement across three TLS conditions in a containerized testbed, with the JA3 fingerprint cross-referenced against a threat intelligence database.
Suricata produced zero actionable alerts across all MCP profiles under the ET Open ruleset (44,236 active rules) for both Opaque TLS and TLS-Inspected conditions; the sole exception was a low-volume match in P6 and P7 under Cleartext averaging 0.25 alerts per run, which manual PCAP inspection identified as a generic HTTP protocol anomaly caused by a missing request header in the traffic generator rather than a C2 or malware detection, and no C2-category or data exfiltration rule ID fired in any run. Turns the question of whether MCP triggers false positives or false negatives in signature-based IDS from speculation into measured counts under a full, unmodified ruleset. Each profile-condition combination was repeated N=5 times with 60 seconds of sustained traffic per run, and Appendix Table II reports mean alert counts per profile.
RITA assigned a beacon score of 0.0 to all eleven profiles across all TLS conditions, including the highly periodic MCP Orchestrated profile P5 and the synthetic C2 profiles P2 and P3; for P2 and P3, each 60-second run produced fewer than 12 unique source-destination connection records in Zeek's conn.log, below the minimum connection count RITA requires for stable quartiles, so it reports 0.0 rather than an indeterminate value, a known boundary condition of RITA's statistical model. Surfaces the boundary behavior of behavioral beacon scoring at low connection counts and separates it from a genuine detection gap. Based on connection counts from Zeek conn.log and RITA v5.x scoring output, with the 0.85 reference line in Figure 1 drawn from a separate calibration run using a published C2 packet trace rather than an observed score from any testbed profile.
For MCP profiles P4 through P10, the 0.0 scores reflect a genuine detection gap: lognormal inter-arrival times driven by LLM inference latency produce high Bowley skewness and high MAD values, and because RITA's normalization model was calibrated on low-skewness, low-MAD periodic malware, high skewness and high variance push the normalized output toward 0.0 rather than 1.0; since baseline MCP profiles already produced 0.0 scores and zero alerts, P7 jitter injection and P9a User-Agent spoofing produced no measurable change in sensor output. Shows that MCP's default temporal shape already falls outside the parameter space where RITA's temporal heuristics detect anomalies, making active evasion unnecessary against the tested configuration. Based on the conceptual density comparison of inter-arrival times in Figure 2 and on sensor output comparisons between P7, P9a, and baseline profiles.
Perspective
The results apply within the containerized testbed the authors built: Suricata v8.x with the full, unmodified ET Open ruleset, Zeek v6.x with community JA3/JA4 scripts, RITA v5.x, and the three visibility conditions of Opaque TLS, TLS-Inspected, and Cleartext. They are aimed at enterprise network defenders, MCP specification maintainers, and IDS vendors, to show under which deployment conditions MCP remote tool calls do not trigger the sensors tested; the proposed Agent-Native ALPN and schema-aware stateful inspection rules are next steps that can be pursued at the gateway and ruleset level.
A careful reader would still watch: N=5 repetitions limits the ability to estimate continuous effect sizes, which the authors themselves frame as an initial characterization of the detection gap; RITA reporting 0.0 below a minimum connection count means periodic profiles' scores should be read alongside connection counts; the 0.85 reference line in Figure 1 comes from a separate calibration run rather than a testbed observation; and validation across additional IDS configurations, rule sets, and real-world traffic volumes remains to be done.
