Skip to main content
Back to timeline
arXivSource publication:

SwarmReconGuard benchmark detects 100% of known attack policies but only 3% of unseen ones

Synopsis

The authors formalize Distributed Collective Reconnaissance (DCR), where each identity stays valid and low-rate while the population collectively acquires system knowledge under hidden coordination, and build SwarmReconGuard, a Docker-isolated reproducible benchmark observing only service-boundary telemetry; across 11 behaviors, 10–10,000 virtual identities, 440 runs and 3,666,300 requests they compare semantic, Gaussian, conditional, graph, kernel, hybrid and CUSUM detectors, finding Gaussian likelihood-ratio detection reaches 100% detection with 0% observed false positives on known attacks but only 3% on unseen policies, CUSUM reaches 36.1% overall detection, and hybrid CUSUM reaches 85.7% at 10,000 identities.

Source-provided article image: SwarmReconGuard: Black-Box Detection of Distributed Collective Reconnaissance by Individually Benign-Looking Agent Populations
Figure 1 ·

Figure 1: SwarmReconGuard architecture. Attacker-side policy and coordination are hidden. The defender receives only API-boundary telemetry, from which window-level semantic, graph, distributional, and sequential evidence is derived.

arXiv

Interpretation

The paper defines DCR as a threat whose objective differs from DDoS: rather than degrading availability, the attacker keeps every identity low-rate, valid and benign-looking while the population cumulatively acquires endpoint behavior, identifier structure and diagnostic metadata. Prior intrusion-detection work centers on flow-, host- or identity-level classification, DDoS detection and botnet signatures, and does not directly answer whether a service can detect coordinated information acquisition when every identity stays within ordinary request budgets and the defender cannot observe internal coordination. The threat model assumes boundary telemetry only and states an observability limit: when adversarial and benign observable distributions are identical, no boundary-only detector can outperform chance.

SwarmReconGuard provides a reproducible Docker-isolated benchmark with four benign controls, seven attack policies, four population scales, independent training/calibration/test phases and exact telemetry accounting. Unlike work reporting only in-distribution accuracy, the benchmark explicitly measures policy generalization and includes hard benign controls such as legitimate diagnostic access, benign exploration and benign bulk access. The final study contains 440/440 valid runs and exactly 3,666,300 processed events for 3,666,300 test requests with zero telemetry failures; training and calibration add 3,066,990 measured reference requests, and thresholds are calibrated on 299 independent benign streams.

Detector comparison shows a marked generalization gap between known and unseen policies: Gaussian LLR detects 100% of the two known attack policies at 0% observed false positives but only 3.0% of the five unseen policies, while the heuristic falls from 93.75% to 7.5%. This negative result shows that near-perfect separation on known attacks does not establish robust DCR defense, since an adversary can change coordination, diagnostic usage, timing or attacker fraction while preserving the collective information-acquisition objective. Results come from detection rate, FPR, precision and F1 at independently calibrated operating points, supported by offline leave-attack-out folds where hybrid ranking reaches pooled AUROC 0.841 and macro AUROC 0.846 but calibrated detection is 39.3% at 4.46% FPR.

Sequential evidence accumulation is the most promising direction in these results: CUSUM achieves the highest overall detection (36.1%) and F1 (0.527) at 1.25% FPR and 98.1% precision, and hybrid CUSUM reaches 85.7% detection with 0% observed false positives at 10,000 identities. This suggests DCR is better modeled as weak evidence accumulated over time than as a single anomalous window, and that “more agents means easier detection” is not a universal law since detectability depends on the interaction of policy, windowing, cumulative structure and calibration. Scale analysis shows hybrid CUSUM is ineffective at 10 and 100 identities, reaches 31.4% at 1,000 and 85.7% at 10,000, while CUSUM detects 52.9%, 34.3% and 57.1% at 100, 1,000 and 10,000 identities respectively.

Perspective

The benchmark targets the service-boundary telemetry setting: the defender owns the gateway and target service and sees complete synthetic boundary telemetry, but not prompts, shared memory, coordination messages, model activations, true swarm membership or scenario labels. It is suited to quantifying the residual observable signal and the cost of relying on behavioral detection alone, and to informing API defense architectures that combine detection with contextual authorization, task-scoped capabilities, selective challenges, response minimization or service-level exposure budgets. The authors state that successful DCR defense must exploit an observable distributional difference, introduce new differentiating context, reduce disclosure, or accept residual leakage. The stated research target is high unseen-policy recall at sub-1% FPR, low exposure-at-alarm-or-end, and negligible incremental SLA impact under a service operating below its saturation knee.

Results rest on one synthetic municipal-style API and scripted policies; the authors state that scripted policies provide experimental control at population scale and do not claim all 10,000 entities are LLM agents, so transfer from scripted swarms to real constrained LLM agents remains to be validated. The matched-swarm construction approximates marginal matching rather than proving equality of complete individual distributions. The laboratory supplies trial boundaries, whereas production detection would require continuous segmentation, decay, sessionization or a rolling-horizon alarm policy. The 299-stream rank calibration targets a stream-level false-alarm probability under exchangeability and comparable horizons; workload, scale or seasonal shifts can violate this assumption, so observed FPR and Wilson intervals should supersede nominal claims. Semantic coverage counts unique synthetic resource keys and is not an information-theoretic estimate of secrets learned by a real adversary. The target service saturates at 1,000–10,000 identities, so strict SLA-preservation claims require a separate capacity study. Handcrafted graph evidence does not improve the current hybrid and should be regarded as a negative ablation result. Detectors are observational and do not throttle, redact, challenge or alter responses, so the study measures detection and exposure rather than prevention benefit or user-friction cost.

Sources