Skip to main content
Back to timeline
arXivSource publication:

MemGuard-Alpha audit finds MIA contamination scores drop to 0.487–0.537 discrimination at a fixed date, and no filtering variant beats the unfiltered ensemble under symmetric transaction costs

Synopsis

The authors introduce MemGuard-Alpha, comprising a composite contamination score combining five membership inference attack (MIA) methods with a temporal proximity feature, plus Cross-Model Memorization Disagreement exploiting variation in training cutoffs across models, and audit both across seven LLMs (124M–7B), 50 S&P 100 constituents, 42,800 prompts and 299,600 prompt-model MIA scores spanning 2019–2024, yielding three negative results: the temporal proximity feature recovers the in-sample label perfectly (ROC-AUC 1.000) because it is a monotone transform of the defining variable, the discriminative power of MIA scores is largely attributable to scale differences between models (raw scores reach AUC up to 0.99 at a fixed date but fall to 0.487–0.

Source-provided article image: MemGuard-Alpha: Limits of Membership Inference for Detecting and Filtering Memorization-Contaminated Signals in LLM-Based Financial Forecasting
arXiv · Page 8

Interpretation

The authors build MemGuard-Alpha, combining five MIA methods with a temporal proximity feature into a composite contamination score, and propose Cross-Model Memorization Disagreement, which exploits variation in training cutoffs across models to characterize memorization. MIA had been proposed as a diagnostic, but it had not been established whether MIA scores are informative about memorization in this setting, or whether signal-level filtering built on them helps once realistic costs are applied; this work makes both the explicit object of an audit. Audited across seven LLMs (124M–7B), 50 S&P 100 constituents, 42,800 prompts, and 299,600 prompt-model MIA scores spanning 2019–2024.

Where in-sample status is defined by a training cutoff, a temporal proximity feature recovers that label perfectly (ROC-AUC 1.000) because it is a monotone transform of the defining variable; any composite score containing such a feature reports separation that is arithmetic rather than detection. This explains why a composite score containing that feature shows apparently perfect separation, attributing the apparent detection power to the label definition itself. Reported as perfect recovery at ROC-AUC 1.000, with the mechanism identified as a monotone transform of the defining variable.

The discriminative power of the MIA scores is largely attributable to scale differences between models: asked at a fixed date which models had that date in training, raw scores appear highly informative (AUC up to 0.99), but under three independent within-model normalizations discrimination falls to 0.487–0.537. This separates the high AUC of raw MIA scores from the near-chance levels after within-model normalization, indicating that the former largely reflects scale differences. AUC of 0.487–0.537 under three independent within-model normalizations, versus raw scores reaching AUC up to 0.99.

With transaction costs applied symmetrically, no filtering variant improves risk-adjusted performance over the unfiltered ensemble, none attains significant Fama-French five-factor alpha, and excluding the single weakest model outperforms every contamination-based filter. This directly tests whether MIA-based signal-level filtering helps once realistic costs are applied, answers in the negative, and points to a simpler alternative that performs better. Evaluated with symmetric transaction costs, comparing the unfiltered ensemble against each filtering variant and testing Fama-French five-factor alpha significance.

Perspective

The work targets research and engineering settings that use LLMs to generate financial alpha signals and define in-sample status by a training cutoff, and is meant for readers who need to judge whether MIA scores can serve as a memorization diagnostic and whether signal-level filtering built on them is effective under symmetric transaction costs; the authors release all artifacts needed to reproduce the results, enabling others to re-check and extend them in the same setting.

Readers should still watch how far these conclusions hold when in-sample status is not defined by a training cutoff, or under different cost models and model families; how Cross-Model Memorization Disagreement behaves under other model combinations; and how robust the observation that excluding the single weakest model beats contamination-based filtering is across broader stock universes and longer horizons. The loaded text is an abstract-level presentation without detailed tables or implementation specifics, so reproducing exact numbers still requires consulting the released artifacts.

Sources