Diffusion-synthesized surges expose fixed-threshold alert failure in disaster cellular monitoring, and hard-sample adaptation restores usable alerts
Related research and updatesSynopsis
Treating Internet activity in CDR grids as a proxy for hidden signaling overload, the study trains a lightweight CNN on stylized injections, stress-tests it with diffusion-synthesized surges that preserve normal traffic structure, and finds that at the default threshold F1 collapses from about 99.94% at baseline to 0% while ROC-AUC stays at 0.886, showing ranking survives but the threshold does not; hard-sample adaptation across three random seeds raises F1 to 85.67% and ROC-AUC to 0.99996, yielding a reusable pre-deployment stress test.
Fig. 1: Emergency overload monitoring framework. Online (top): 9 × 9 9\times 9 snapshots scored by a compact CNN and thresholded at fixed τ = 0.5 \tau=0.5 . Offline (bottom): diffusion stress testing and hard-sample adaptation; τ ∗ \tau^{*} is selected offline on stress-test grids and is not the online cutoff. Main detector: Internet-only; ablation: three-channel input.
arXivInterpretation
It formulates emergency signaling-overload monitoring with Internet CDR activity as an indirect proxy and builds labeled synthetic surge grids on the Milano CDR corpus. Unlike prior CDR-grid studies that monitor the channel directly reflecting the event (for example SMS counts for SMS flooding), the target here is control-plane signaling stress while only coarse CDR aggregates are observable. Built on the public Milano corpus (2013-2014, 3G/4G-era metropolitan traffic) following an established CDR-grid preprocessing pipeline, with normal train/test of 614/167, stress train/test of 167/168, and 40 overload-affected cells; the authors state labels come from the synthetic surge construction, not measured RRC.
Diffusion-synthesized surges expose fixed-threshold fragility: at the default cutoff of 0.5 the Internet-only F1 is 0.00% while ROC-AUC remains 0.886, and the post-hoc F1-maximizing threshold shifts to about 0.817, an operating-point drift of 0.317. Matched-condition training and testing on the same stylized injections hides this drift, since baseline F1 reaches 99.94% and looks near-perfect. Mean plus or minus standard deviation over three random seeds; the authors note the F1-maximizing threshold is selected post hoc on the same stress-test grids (oracle) rather than being an independently validated operating point.
Hard-sample adaptation improves both ranking and thresholded performance: F1 rises from 0% to 85.67%, ROC-AUC to 0.99996, and the post-hoc threshold drift falls to 0.167. Threshold-only recalibration without retraining reaches 67.07% F1 while AUC stays at 0.886, so adaptation adds ranking-level gains beyond recalibration alone. Three seeds; post-adaptation F1 standard deviation is 14.37, which the authors cite in stating that F1 is not a certified operating point.
Between the two input-architecture pairs, Internet-only with a compact CNN dominates on every stress-relevant metric (stress AUC 0.886 versus 0.264, post-adaptation F1 85.67% versus 66.80%, drift 0.317 versus 0.450). The authors note the comparison changes channels and backbone together, so a pure channel effect is not isolated. Three seeds; the three-channel model reaches 100% F1 on stress-only samples after adaptation but only 66.80% on the mixed normal-plus-stress set.
Perspective
The result targets disaster-response settings where only coarse CDR aggregates are available and local compute is limited, in which a single Internet channel already yields a clearer overload signature. It supports a two-step emergency workflow: detect drift under stress-test conditions, then retrain or recalibrate thresholds before the next incident window. The diffusion stress-test protocol can serve as a reusable pre-deployment testbed for evaluating monitors under disaster-like surge shift.
Internet CDR is only an indirect proxy, and the text provides no paired RRC/NGAP traces to validate its correspondence to real signaling load; labels and surges are synthetic and cover only one Milano window; post-adaptation F1 standard deviation reaches 14.37 across three seeds, so thresholded performance varies noticeably by seed; reported optimal thresholds and drift are post-hoc values on the evaluation grids with no held-out calibration split; the Internet-only versus three-channel comparison changes the backbone as well, so the channel effect is not isolated; spatial realism is checked only coarsely via z-score, MMD, and correlation. These are scope and open questions, and readers should add a held-out calibration split and real disaster-period traces before using the workflow for an actual alert rule.
