Skip to main content
Back to timeline
arXivSource publication:

AMU adds lineage gating to shared agent memory: column-level authorization zeroes cross-department leakage while keeping most memory reuse

Synopsis

The authors introduce the Analytical Memory Unit (AMU), a memory schema that attaches a full derivation (lineage) graph to every cached result and gates retrieval so a hit is served only when the requester is authorized for every column touched; under complete lineage recording they prove a safety guarantee by construction, and six experiments show lineage-gated retrieval removes the cross-department leakage of content-gated memory at a modest reuse cost, with zero leaks over 9 round-trips in a real-agent integration.

Source-provided article image: Lineage-Aware Memory Governance: A Derivation-Gated Framework for Privacy-Preserving Column-Level Access Control in Enterprise AI Agents
Figure 3 ·

Figure 3: Synthetic vs. TPC-H schema comparison (30 seeds each). TPC-H’s denser sensitive-column structure yields a higher naive leak rate; lineage-aware gating eliminates leakage in both configurations.

arXiv

Interpretation

The paper proposes the AMU schema and a lineage-gated retrieval policy: each cached result carries a derivation graph, and retrieval performs a single set-difference check, returning a hit only when the requester's permissions cover every column the result touched, otherwise skipping that candidate. Existing agent-memory systems (MemGPT, Zep, A-MEM, Governed Memory, SSGM, Oracle AI Agent Memory and others) authorize a memory entry as a whole by content, ownership, or role and do not record which columns produced a result; AMU brings column-level derivation tracking to agent memory entries rather than raw database records. Table I compares six governance dimensions across representative systems, showing no prior system combines column lineage with a retrieval gate; the mechanism itself is presented through algorithms and formal definitions.

Under complete lineage recording (Assumption 1), the authors give a safety guarantee by construction: gated retrieval will not return a result whose derivation touches a sensitive column outside the requester's permissions, with corollaries for unsafe store contents and policy updates plus complexity analysis. This turns whether a result crossed a column-permission boundary from an audit question into a property the system can check mechanically; the authors explicitly frame it as a conditional design guarantee rather than an empirical finding. The proof proceeds by exhaustive case analysis over the algorithm's two return paths; the authors also state the guarantee depends on lineage completeness and measure degradation when completeness falls short in Experiment 3.

Six experiments characterize the governance-efficiency tradeoff: naive content-gated shared memory leaks structurally across departments, lineage gating removes that leakage at a reuse-rate cost of a few percentage points (roughly 13 additional query executions per 100 requests), the gate check itself is sub-microsecond, and worst-case retrieval latency scales linearly with store size. The paper attributes the leak rate to the theorem's guarantee and locates the genuine empirical content in the naive baseline's leak rate, the reuse cost, and degradation under incomplete lineage, replicating the tradeoff on both a synthetic 5-table and the TPC-H 8-table schema. Experiments use a pure-Python discrete-event simulator with 30 seeds, bootstrap 95% confidence intervals, and Mann-Whitney tests; runtime is measured with time.perf_counter() over 2,000 iterations on a single hardware configuration.

A real-agent proof-of-concept uses sqlglot to extract lineage automatically from the executed SQL abstract syntax tree instead of agent self-reporting, yielding zero sensitive-column leaks across 9 cross-department round-trips on the Northwind database, with 2 gate blocks coinciding exactly with 2 automatically detected metric-definition conflicts. This offers a practical route to satisfying Assumption 1 without agent cooperation, which the authors position as a feasibility demonstration rather than evidence of production viability. A small 9-round-trip proof-of-concept; the authors list three static-parse blind spots (views, stored procedures, deep aliasing) and note that the fail-closed design limits damage to over-blocking.

Perspective

The mechanism targets department-level enterprise agents sharing an analytical memory store: requesters have a legitimate business need for a metric, but not all derivation paths fall within their column permissions. It complements rather than replaces source-layer access control (such as Apache Ranger or Databricks Unity Catalog) and applies to named, logged columns under sufficiently complete lineage. The authors recommend automatic lineage extraction (interposing at the database driver layer on executed SQL) over agent self-reporting, and a fail-closed posture: treat an AMU as maximally sensitive when extraction fails, and default retrieval to denial when a permission set is unavailable. For teams wanting to turn whether a result crossed a column-permission boundary into a mechanically checkable property, and for organizations preparing explainability and risk-management material for high-risk AI systems, this layer offers a reusable design pattern.

The safety guarantee holds only when lineage recording is complete; Experiment 3 shows protection degrades as completeness falls, and TPC-H needs higher completeness because its sensitive columns span more join steps. The mechanism does not cover T4 inference attacks: a derived feature that encodes sensitive information without naming its source column (for example, a risk score computed from income) slips past the gate, which the authors call the most consequential gap for production use. Cross-metric correlation and departmental collusion fall outside what a single-retrieval column gate can address. Automatic extraction depends on SQLGlot's static AST parse of SQL text, so unresolved view definitions, stored procedures and user-defined functions, and deeply nested aliasing with SELECT * expansion can leave lineage incomplete. Runtime benchmarks were measured on a single hardware and OS configuration, so absolute constants may shift across platforms. The 43-pair fuzzy conflict-detection dataset is still small, and the proposed D4 composite detector has not yet been evaluated. The real-agent integration spans 9 round-trips and does not yet demonstrate production-scale viability.

Sources