ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions
Synopsis
The work introduces privacy exposure displacement and the ASLEval framework, which pre-registers a hidden target set, an authorization relation, and declared visible exits before a session, then uses a target-blind probe agent and post-hoc target-grounded adjudication to compare local evaluation proxies with full-session visible exposure; across several enterprise-style environments and two independently implemented runtimes it observes three recurring patterns: an expected-outlet-only view misses 46.
Figure 1: Three manifestations of privacy exposure displacement. C1: the expected outlet can miss exposure visible elsewhere in the session. C2: attacker-reported candidates differ from evaluator-confirmed exposed targets through both omissions and false discoveries. C3: exposure is associated with schema-aligned request/probe paths; reducing model-visible return content changes this path, but is not a cost-free defense.
arXivInterpretation
It identifies privacy exposure displacement as a measurement failure and organizes it into outlet displacement (C1, where exposure is counted), recognition displacement (C2, what is recognized), and path displacement (C3, how exposure is mediated). Existing agent privacy and safety evaluations often inspect one designated action, final response, file-sharing event, or red-team report; the paper argues such local proxies may represent only part of the requester-visible session boundary and supplies a unifying taxonomy. The concept and taxonomy come from the paper's problem framing and Figure 1, and are then instantiated in three empirical findings.
It introduces the ASLEval protocol: before execution the evaluator fixes a hidden target set, an authorization relation, an expected outlet, and the declared visible session boundary; a target-blind probe agent issues only interface-valid requests; a post-hoc target-grounded adjudicator maps every declared observation to the pre-specified target set; only requester-visible and external exits count as exposure, while internal traces are diagnostic only. Relative to Privacy in Action, AgentDAM, AgentLeak, and CIPL, ASLEval asks whether a local proxy represents unauthorized exposure of a fixed hidden target set across the declared visible boundary, differing in target definition, counted surface, and measurement question. The protocol is given formally (authorization relation, observation set, visible/internal channel partition, exposure coverage rate, and outlet-loss fraction) and includes an evaluator-only negative control validating authorization-aware metric semantics.
Across runtimes, models, and environments it reports three measurement regularities: outlet-local protection can improve the audited channel while leaving session exposure nearly unchanged; attacker self-reports are noisy predictions rather than exposure ground truth; and schema alignment is associated with visible exposure, with internal evidence usually preceding it at the request/probe level. These regularities come from common-log comparisons under one evaluation contract rather than end-to-end reproductions of prior systems, giving a direct contrast between local views and the visible-exit union. C1 is replicated across two independently implemented runtimes (Agent-R and Agent-F) and an expanded item set; C2 remains sensitive to difficult candidate matches, with human validation showing it is the hardest adjudication slice (primary strict F1 0.762); C3 supports request/probe-level temporal order rather than a nested causal chain.
It provides evidence on the privacy-utility trade-off: minimizing model-visible returns reduces overlapping visible entities to 104.8 (a 45.7% decrease) and lowers the overlap rate from 0.920 to 0.445, but increases tool-call attempts from 304.0 to 775.8; on a 15-task normal slice, unauthorized visible exposure falls from 0.460 to 0 while deterministic task success falls from 0.800 to 0. It frames return minimization as a mechanism intervention that reveals a severe privacy-utility trade-off rather than a deployable defense, and argues tool-interface proposals should report a privacy-utility frontier rather than leakage reduction alone. Based on fixed-request replay and a separate normal-task slice, with backend success remaining 1.0 in both conditions, indicating insufficient model-visible information rather than tool failure.
Perspective
The work is intended for defensive measurement and benchmark analysis, applicable to session-level privacy exposure evaluation over synthetic enterprise-style records in isolated environments, for benchmark designers, agent-runtime developers, and tool-interface designers; its conclusions apply to settings with pre-declared visible exits and a pre-registered hidden target set plus authorization relation, and it argues for reporting local views alongside the visible-exit union and presenting privacy together with task utility.
A careful reader would still watch: the C2 recognition result is positioned by the authors as a robust direction rather than a calibrated estimate of every candidate match, so candidate-match sensitivity analyses are worth following; C3 establishes request/probe-level temporal association rather than a parent-child call-graph causal chain, and schema alignment covaries with entity type and source table; the normal-task slice evaluates a deliberately coarse minimized-output intervention, leaving optimized selective-return policies uncharacterized; and although the full text was loaded, the supplement's run matrices, seeds, and some conditions are not in the body, so those details remain open questions.
