TGCC uses cross-layer trust gating so a compromised LLM agent self-revokes within steps instead of escalating
Related research and updatesSynopsis
The work turns a descriptive layered trust model into an operational controller: a cross-layer synergy operator propagates prerequisite-layer deficits into dependent layers, a no-regret online estimator grounds per-layer weights in observed failures, and Trust-Gated Capability Control issues short-lived, revocable grants only when composite trust and the relevant prerequisite layers clear capability-specific thresholds; the authors prove and numerically confirm that a stealthy compromise inflating behavioral trust while degrading a prerequisite layer cannot escalate privilege and instead self-revokes within a few interactions.
Fig. 1: (a) Steady-state trust (bars) vs. analytic fixed point T ⋆ T^{\star} (diamonds), MAE 0.007 0.007 (Proposition 1 ). (b) Failure-grounded weights α ℓ ( t ) \alpha_{\ell}(t) shift toward the epistemic layer after compromise (Lemma 2 ).
arXivInterpretation
Introduces a cross-layer synergy operator and a generalized-mean composite trust that propagate prerequisite-layer deficits into dependent layers and recover the weakest-link rule as a limiting case. Existing layered trust models are largely descriptive, naming which trust dimensions matter without specifying how layers combine; this work provides a provably bounded composition and shows composite trust is pinned to the weakest effective layer up to a bounded slack. Well-posedness, monotonicity, and bounds are given as definitions, lemmas, and theorems, with zero bound violations reported across 20,000 random draws.
Uses a no-regret online estimator (Hedge) to set per-layer importance weights from runtime failure attribution, making composite trust most sensitive to the layer currently most responsible for harm. Prior trust systems typically fix weights or track a single scalar trust; this work links the system's failure record directly into trust computation and inherits the Hedge no-regret guarantee. A no-regret layer-tracking lemma is provided, and the numerical study observes weights shifting toward the epistemic layer once it becomes the dominant failure source.
Proposes Trust-Gated Capability Control, which issues short-lived, revocable capability grants only when composite trust and the relevant prerequisite layers clear capability-specific thresholds, and proves it breaks the Trust-Vulnerability Paradox. Zero-trust and protocol work establish that trust must be revocable and capabilities least-privilege, but leave open how a multi-dimensional trust signal should be aggregated into authorization; this work supplies that policy layer and gives a closed-form conservative trust fixed point with a bounded-latency revocation guarantee. Theorems show that under a stealthy compromise high-risk capabilities are revoked and the granted set is non-increasing in compromise depth, with a corollary giving logarithmic-step self-revocation latency; in simulation the analytic fixed point matches simulation with mean absolute error 0.007 and revocation occurs within a few steps on average.
Perspective
The result targets multi-agent LLM collaboration where a controller mediates capability grants, and applies when the deployer can set per-capability risk thresholds and each layer exposes a noisy but informative success signal (e.g., calibration probes, role-conformance checks, audit trails). It makes per-step re-evaluation and prompt revocation practical at constant per-step overhead, so it is directly useful to engineering teams needing to prevent stealthy compromise escalation in long-running collectives and to protocol implementers seeking to ground zero-trust principles in an authorization policy. The framework also sketches extensions to embodied or perceptual agents, using sensor-fusion confidence or perceptual anomaly scores as epistemic-layer evidence, and to an incentive layer for agentic marketplaces.
The simulation instantiates single-agent layer-trust dynamics under a sleeper threat rather than a real multi-agent deployment; the signal quality of each layer's checks, how the dependency structure and coupling matrix are set, and how deployers calibrate per-capability risk thresholds remain open questions for practice. The cross-modal and economic-incentive extensions are sketched rather than independently validated. In addition, several formulas and numerical values (such as specific thresholds and average revocation steps) are not fully rendered in the extracted text, so readers needing exact parameters should consult the original figures and tables.
