Skip to main content
Back to timeline
arXivSource publication:

A five-condition review of 13 membership inference attacks finds none simultaneously non-overfitted, competitive, reliable, and computationally feasible

Synopsis

The authors propose an evaluation framework with five necessary conditions (C0 sensitive disclosure potential, C1 non-overfitted model, C2 competitive model, C3 reliable membership inference, C4 computational feasibility) and use it to review 13 representative black-box membership inference attacks across 61 attack-dataset pairs, concluding that no attack satisfies C1 through C4 simultaneously and that under these realistic conditions membership inference attacks represent weak privacy threats.

Source-provided article image: A Critical Review on the Effectiveness and Privacy Threats of Membership Inference Attacks
Fig. 1

Fig. 1: Overview of the proposed evaluation framework for MIAs. Condition C0 relates to the MIA disclosure potential, which critically depends on the data set used to train the target model. Conditions C1–C4 characterize the effectiveness of the MIA itself, which should reliably attack non-overfitted and competitive models at a reasonable computational cost.

· Page 5

Interpretation

The paper proposes a five-condition evaluation framework that decomposes whether a membership inference attack constitutes a genuine privacy threat into a data-level precondition C0 (training data must be an exhaustive sample of a population, confidential attribute values must be unique, and the information assumed known to the attacker must be plausible) and attack-level conditions C1 to C4 (non-overfitted, competitive, reliable inference, computationally feasible). Prior work discussed the link between overfitting and attack success, or separately noted high false positive rates and unrealistic priors, but did not integrate these into a single framework of conditions that must hold simultaneously. The thresholds are deliberately permissive: C1 allows a train-test accuracy gap g ≤ 10%, C2 allows test accuracy within 5% of the state of the art, C3 requires weighted precision ≥ 95% evaluated at a membership prior p = 10%, and C4 aggregates three factors (number of additional models M, inference model complexity I, queries per sample Q) into low, moderate, or high overall cost.

At the C0 level, the datasets used by the surveyed attacks (Adult, Purchase-100, Texas-100, UCI Credit, UCI Hepatitis, UCI Cancer, Locations, MNIST, CIFAR-10, CIFAR-100, CINIC-10, GTSRB, ImageNet-1K, LFW, Newsgroups, RCV1X) are none of them exhaustive samples of their corresponding populations, so non-trivial membership disclosure is possible in principle but unequivocal attribute disclosure is not; moreover, none of the reviewed attacks documents or even considers uniqueness of confidential attribute values. The paper systematically transfers the disclosure-risk concepts long established in database privacy (identity disclosure and attribute disclosure) onto the evaluation of membership inference attacks, pointing out a clash between membership disclosure and attribute disclosure. The argument rests on a dataset-by-dataset analysis of the properties above and on logical derivation of exhaustivity and uniqueness as necessary conditions, rather than on new experimental measurement.

At the C1 and C2 levels, 24 of the 61 attack-dataset pairs (39.34%) lack sufficient information to assess overfitting; among the 37 assessable pairs, 27 (72.97%) use overfitted target models with train-test gaps exceeding 10% and only 10 (27.03%) satisfy C1; among the 50 pairs where competitiveness is assessable, only 12 (24.00%) use models within 5% of the state of the art; only 7 pairs (18.92%) satisfy C1 and C2 simultaneously. This provides a quantitative picture of target-model selection in the membership inference literature, indicating that much of the reported attack success comes from overfitted or underperforming target models rather than realistic deployment scenarios. The counts come from the paper's Table 1, which tabulates train accuracy, test accuracy, generalization gap, state-of-the-art accuracy, and the gap to it for each attack-dataset pair, with some values marked NA where the source lacked them.

At the C3 and C4 levels, only 16 pairs (26.2%) maintain reliable inference under realistic priors; LiRA reaches ≥99.95% precision with FPR ≤0.001% on CIFAR-10, CIFAR-100, and ImageNet-1K, and RMIA reports 100% precision and 0% FPR on all datasets, but 8 of the 13 attacks (about 62%) incur high computational cost, for example LiRA requiring 32 to 256 shadow models; combining C1 through C4, no attack satisfies all four conditions, and only [50] and [73] on MNIST each meet three of the four. The paper places reliability (low false positive rate plus high precision under realistic priors) and computational feasibility side by side in one evaluation, showing that recent attacks have improved reliability while scalability remains a critical constraint. Reliability is judged from the paper's Table 2 using precision, true positive rate, false positive rate, and prior settings, with missing metrics derived where possible via standard statistical relationships; computational cost is judged from Table 3 by the number of additional models, inference model complexity, and query counts.

Perspective

The framework addresses classification deep neural networks in predictive machine learning, the setting for which membership inference attacks were designed and on which the surveyed works focus; it evaluates black-box attacks, since white-box access is often infeasible in real-world scenarios and prior work shows the optimal attack strategy is primarily based on the model's loss. Its purpose is to judge whether a membership inference attack constitutes a genuine privacy threat under realistic conditions and, on that basis, whether utility-hampering defenses such as differential privacy are warranted. For practitioners releasing trained models, the paper recommends first checking whether the training data satisfy the necessary conditions for unequivocal disclosure of sensitive information; if they do not, refraining from privacy protection techniques that may cause unwarranted utility loss; and if they do, resorting to mechanisms that target problematic data features (such as sampling or partial synthesis) while ensuring the model is not overfitted. The authors list the analysis of privacy attacks proposed for generative models, including large language models and diffusion models, as future work.

The C1 and C2 thresholds (10% generalization gap, 5% deviation from the state of the art) are set by the authors, who acknowledge the specific thresholds can be debated and emphasize that their choices are deliberately permissive. The C3 reliability judgment depends on weighted precision ≥ 95% and a membership prior p = 10%, while for some attacks the precision, false positive rate, or true positive rate is missing from the source, so the paper can only derive them via standard statistical relationships or mark them NA; the reliability conclusions for those pairs therefore depend on whether those derivations hold. The paper also notes that some attacks report very few true positives with large standard deviations (for example, 2.8 ± 2.4 on CIFAR-10 and 3.0 ± 1.26 on CIFAR-100 for [73]), and cites observations that the same target sample does not consistently yield a positive inference across different initializations and that reproducing the offline LiRA attack with the authors' own code gave lower accuracy than reported, all of which affect judgments about individual attacks. The paper only raises generative models as an outlook without stating conditions, so its conclusions do not currently cover large language models or diffusion models.

Sources