Four network-traffic generators leak their training data: membership inference reaches 88% TPR and network identifiers can be fully recovered
Synopsis
The study introduces a privacy measurement suite for synthetic network traffic spanning membership inference, data extraction, and network-specific attacks on identifiers, attributes, and topology, and evaluates four representative generators (NetShare, NetDiffusion, TrafficLLM, NetSSM) across five datasets, finding that even with black-box-only attackers and minimal training, membership inference reaches up to 0.88 TPR at FPR≤0.01 and up to 100% of network identifiers can be recovered, while anonymization and DP noise each cover only part of the risk at a utility cost.
Figure 1: Pipeline of splitting each dataset into non-training, training and auxiliary data.
· Page 4Interpretation
The paper proposes and operationalizes a multi-dimensional privacy measurement suite for synthetic network traffic: membership inference at the PCAP level (reporting TPR@FPR≤0.01 and AUC-ROC), data extraction (a generated sequence counts as extractable if it contains at least 10 consecutive native representation units matching training data verbatim), and network-specific attacks subdivided into network identifiers (output coverage and output confidence), sensitive network properties (TTL, IP ID, ToS, window size, control flags, TCP data offset, flow size, packet size, measured by normalized EMD), and network topology (node overlap, edge overlap, node degree distribution EMD). Prior privacy evaluation of synthetic network traffic relied mainly on membership inference as the sole measure, even though traffic carries sensitive information beyond membership; this work makes those network-semantic risks explicitly computable so dataset owners can see what is exposed, whether it is sensitive, and to what degree. The metrics are given as explicit formulas (coverage and confidence), topology is restricted to the top 100 communication flows ranked by total bytes with nodes as IP or MAC addresses, and properties use normalized EMD; attacks are run only where model architecture and outputs permit, with a table marking which attacks were conducted per model.
Under the least favorable setting for the attacker (black-box access only, and models other than NetDiffusion trained for just one epoch), leakage remains substantial: MIA achieves AUC≥0.80 on four of five datasets and TPR@FPR≤0.01 up to 0.88, with NetSSM most vulnerable, TrafficLLM next, and NetDiffusion considerably less so. This shows privacy risk is not driven by overfitting alone: even with minimal training and black-box access, byte-level (NetSSM) and hexadecimal (TrafficLLM) encodings leak more membership information than bit-level image encoding (NetDiffusion), which the authors attribute to encoding granularity and native sequence modeling capability. The finding rests on a cross-comparison of four generators and five datasets reporting both AUC and TPR at a stringent FPR; the authors also note MIA effectiveness varies by dataset and offer similarity between training and non-training samples as a possible explanation, framed as a hypothesis.
Data extraction is more effective against autoregressive models: NetSSM and TrafficLLM show extractable rates ≥0.4 on four datasets (all but CIC) and above 0.8 on some, while NetDiffusion shows no extractability; extraction concentrates at positions 1–9 and 27–33, corresponding to identifier fields such as MAC and IP addresses. The attack succeeds using only traffic-type labels as prompts, with no training traffic supplied, showing attackers can recover training samples at high rates without knowledge of the training data; it also reveals that categorical/identifier fields are memorized more readily than continuous-valued fields, giving position-level evidence about memorization in the mixed-type structure of network headers. Extractability is defined as at least 10 consecutive native representation units matching training data verbatim; the position analysis is shown as a byte-position curve for NetSSM with the TrafficLLM curve in the appendix, and the authors map high-extraction positions to Ethernet-layer fields.
Most generators leak network identifiers and sensitive network properties, and NetSSM additionally leaks network topology: source-IP confidence exceeds 0.8 on four of five datasets and coverage exceeds 0.35 on some; without mitigation, normalized EMD for TTL, ID, flow size, and ToS can be 0.10 or lower; NetSSM reaches 1.0 node and edge overlap for MAC-based topologies, indicating complete topology reconstruction. The paper further finds that frequently occurring identifiers (source IPs appearing more than 1,000 times in training) are memorized at substantially higher rates, and that overlapping topology nodes are predominantly high-degree hubs, meaning critical servers and core routers are precisely the entities most exposed. Identifier findings use two complementary metrics (coverage and confidence); property leakage uses per-property normalized EMD comparisons across four models in a table; topology findings use graph overlap metrics and node degree distribution EMD, with the authors explaining higher MAC-level leakage by MAC address spaces typically being smaller than IP address spaces.
Perspective
The study targets multi-flow PCAP-level synthetic network traffic, evaluating four representative generators (NetShare, NetDiffusion, TrafficLLM, NetSSM) on five datasets (IoT, SR, VNAT, USTC, CIC), with black-box access plus minimal training as the conservative baseline and additional conditions covering white-box access, training epochs, and data scale. It applies to data owners and network operators deciding whether and at what granularity synthetic traffic can be released: fidelity can be prioritized for internal use in trusted environments, while external sharing requires stronger protection. Mitigation evaluation centers on NetSSM and covers three anonymization strategies (complete anonymization, pseudonymization, prefix-preserving) plus DP noise at ε=0.1 and ε=1, distinguishing multi-flow from single-flow generation settings.
Several questions remain worth watching. First, MIA effectiveness varies across datasets, and the authors propose similarity between training and non-training samples as a possible cause, which is a hypothesis rather than a verified mechanism. Second, property leakage is measured by normalized EMD, which the authors adopt because ground truth for fingerprinting attacks is unavailable, so the metric reflects exploitability rather than actual attack success. Third, mitigation evaluation is limited to NetSSM, so whether other models show the same ordering of privacy-utility tradeoffs is unclear. Fourth, network property inference effectiveness sometimes rises after anonymization, indicating different attack types do not necessarily move together and a single metric can give a partial picture. Fifth, downstream task labels are unavailable for multi-flow traffic, so task accuracy is evaluated only in the single-flow setting and the utility cost in multi-flow scenarios is not quantified.
