MemLeak measures 70–100% cross-user leakage in multi-tenant agent memory and restores contamination to 1.00/5 with hard ownership gating
Related research and updatesSynopsis
The work formalizes cross-user admissibility failure in multi-tenant AI agent memory and runs six experiments plus follow-up ablations under sparse TF-IDF and production-faithful MiniLM retrieval, finding 70–100% incidental leakage under pooled same-team retrieval, 90–100% top-k placement for adversarially crafted memories, end-to-end response contamination up to 5.00/5 with contaminated answers often rated as helpful as clean ones, and that among three mitigations only hard post-retrieval ownership gating consistently restores the 1.00/5 clean baseline across two generation models at roughly 1.4 ms latency per query.
Figure 1: MemLeak architecture. Alice’s memories and Bob’s query co-exist in a shared vector store; cosine top- k k retrieval pulls Alice’s memories into Bob’s context without an ownership check (dashed red path), constituting a cross-user admissibility failure (§ 3 ). Dashed bordered boxes mark the three mitigation intercept points: M1 (metadata filtering at the store), M2 (ownership-aware embeddings), and M3 (post-retrieval ownership gating).
arXivInterpretation
A shared embedding space is itself a cross-user leakage surface: a user's query can retrieve semantically adjacent memories owned by another user through ordinary cosine-similarity retrieval, with no exploit, prompt injection, or code change. Prior memory-trustworthiness work (e.g., Tan et al., Wu et al.) treated memory as a single-user trust boundary, addressing when a user's own memories should influence their own task; this work extends the question to the multi-user case, i.e., what happens when Alice's memories reach Bob's session. The E1 existence proof shows systematic leakage under pooled retrieval in both configurations (Config A 100%, Config B 70%, 95% Wilson interval [40%,89%]), while partitioned retrieval yields 0% in both, providing the contrast.
Leakage severity does not track intuitive risk signals such as organizational distance or global embedding-centroid proximity, so similarity is a poor proxy for admissibility. The work formalizes cross-user admissibility failure and defines leakage metrics (LR, MCR, score gap, XLR, an exploratory Resonance Index), reporting 95% Wilson confidence intervals throughout. In E2, T1 and T2 show nominally higher leakage (90%) than T0 (70%) despite T0 having the highest centroid similarity (0.570), but the Wilson intervals for 70% and 90% overlap substantially, and the authors explicitly treat this fine-grained ordering as suggestive rather than established; the robust result is the cross-company control T3 at 10% leakage (interval [2%,40%]), clearly separated.
An adversary who knows only the victim's domain and role, without query-level or model-internal access, can achieve near-oracle retrieval advantage with crafted memories. E3b makes attacker knowledge explicit and compares against baselines of increasing query knowledge (organic, domain keyword bag, per-query keywords, crafted bridging, verbatim-copy oracle), separating the marginal contribution of crafting from raw vocabulary overlap. Under Config B, placement rises monotonically with assumed knowledge: 10% (organic) to 40% (keyword bag) to 80% (per-query keywords) to 90% (crafted bridging) to 100% (verbatim-copy oracle); a bare domain keyword bag already reaches 40%, showing part of the advantage comes from shared vocabulary alone.
Leaked memories materially contaminate downstream responses, and contaminated responses are often rated as helpful as clean ones, so standard response-quality monitoring cannot detect this failure. The work evaluates downstream contamination under both a static injection path and a full retrieve-then-generate path with Gemini 2.5 Flash and Claude Sonnet 4.5, and validates the LLM judge against human annotation in a pilot. In E4, contamination rises from the 1.00 clean baseline to 2.67–5.00; the retrieval path with Gemini 2.5 Flash reaches the maximum 5.00/5, and Claude Sonnet 4.5 reaches 4.67/5 on both paths; the E0 human-agreement pilot found MAE 0.53 on clean controls and helpfulness and MAE 0.79 on contamination, with the single disagreement directionally conservative, so contamination numbers are treated as a lower bound.
Perspective
The work targets pooled vector stores or stores whose scope is intentionally widened (under-scoped metadata filters, team- or role-scoped retrieval, shared project namespaces, and scoped HR/finance/security workflows), not the trivial case of a fully absent access-control layer; strict user partitioning serves as the ACL baseline and drives incidental leakage to 0% in their experiments. The results apply to multi-tenant personal agent memory that uses embedding-similarity retrieval with ownership expressed as soft metadata, and they advise system designers to evaluate retrieval architecture and generation model selection jointly. Directions this enables include validating at larger samples and against real enterprise memory distributions, evaluating larger and more diverse enterprise embedding models, running cross-model judge-generator pairings and independent human adjudication, and adopting hard post-retrieval ownership gating as a deployable low-cost default defense.
Fine-grained orderings (such as T1/T2 above T0 in E2, or the Resonance Index ordering across T0–T2 in E5) rest on a small number of queries per tier with overlapping Wilson intervals and should be read as suggestive rather than established; memories and queries are fixed hand-authored fixtures whose representativeness for real enterprise memory distributions, vocabulary overlap patterns, and query phrasing is unverified. E4/M3 contamination and helpfulness are scored by the same model family used as generator, the human-agreement pilot is small and weaker on contamination, and cross-model judge-generator pairings remain future work. Mitigations are proof-of-concept implementations, most experiments use in-memory cosine search, and only the latency benchmark uses production pgvector/HNSW infrastructure at a single over-fetch factor; the namespace-offset variant is not viable as tested and would require full user-conditioned encoder training to evaluate properly. E6's negative result (100% bit-error rate at bits 5–8) indicates that a channel with controllable throughput would require structured error-correcting encoding, which the paper does not claim to have demonstrated. The work also studies only embedding-similarity retrieval and does not empirically evaluate alternatives such as LLM-based memory selection.
