Skip to main content
Back to timeline
arXivSource publication:

MIRROR cuts agent-in-the-middle attack success to 0% in LLM multi-agent communication via multipath quorum integrity

Related research and updates

Synopsis

MIRROR is a communication-layer integrity primitive that replicates a single canonicalized payload across k logical routes and accepts a message only when a strict majority of routes report the same digest; under a route-compromise bound of alpha < 0.5, it reduces attack success rate to 0% across MMLU, HumanEval, and MBPP on two frameworks and four topologies and in a MetaGPT deployment against a production API, at 1x LLM token cost, whereas LLM-as-a-Judge costs 35x in the same deployment and blocks up to 44.2% of benign outputs in the topology sweep.

Source-provided article image: MIRROR: Multipath Quorum Integrity for LLM Multi-Agent Communication
Figure 1 ·

Figure 1: Overview of communication threats and the MIRROR defense. (A) Traditional attacks originate from a malicious or compromised sender, which is outside MIRROR’s scope. (B) AiTM attacks manipulate messages in transit via a third-party interceptor. (C) MIRROR replicates across routes and verifies quorum, recovering valid messages or blocking corrupted transmissions.

arXiv

Interpretation

The paper presents MIRROR, a communication-layer integrity primitive that replicates a single canonicalized payload across k logical routes and accepts a message only when a strict majority of routes report the same digest. Unlike existing defenses that rely on semantic validation, which requires additional inference and can block benign outputs, or on transport-layer encryption, which does not help when an intermediary legitimately terminates TLS, MIRROR constrains message integrity at the communication layer through multipath quorum. The abstract gives the primitive's design and its guarantee under a route-compromise bound of alpha < 0.5, evaluated across MMLU, HumanEval, and MBPP on two frameworks and four communication topologies and in a MetaGPT deployment against a production API.

MIRROR uses unkeyed hashing and so authenticates nothing on its own; all integrity derives from the assumption that honest routes form a majority, and the digest serves only to make witness routes constant-size and to bind the recovered payload to the quorum-agreed value under second-preimage resistance. The paper states explicitly that the primitive relies on no key or authentication tag, instead reducing security to the condition that a majority of routes are honest, with a guarantee under alpha < 0.5. This is an explicit statement of the primitive's security assumption, with the route-compromise bound alpha < 0.5 and an extension to correlated routes given in the abstract.

The guarantee extends to correlated routes, where the quantity that matters is the size of the largest shared-failure group rather than the route count; availability and integrity degrade at the same threshold, and below alpha = 0.5 quorum-denial and message-dropping adversaries cannot block honest traffic. The paper generalizes from an independent-route assumption to correlated failures and shows that integrity and availability share the same threshold rather than degrading separately. The abstract states the extension to correlated routes and the non-blockability of honest traffic against quorum-denial and message-dropping adversaries below alpha = 0.5.

Across MMLU, HumanEval, and MBPP on two frameworks and four topologies, and in a MetaGPT deployment against a production API, MIRROR reduces ASR to 0% below the threshold at 1x LLM token cost, while LLM-as-a-Judge costs 35x in the same deployment and blocks up to 44.2% of benign outputs in the topology sweep. The paper directly compares communication-layer integrity against a semantic-validation baseline on attack success rate, token cost, and benign-output blocking rate. The abstract reports evaluation across three benchmarks, two frameworks, four topologies, and one production API deployment, with 1x versus 35x token cost and a 44.2% benign-output blocking rate.

Perspective

The result targets LLM multi-agent communication deployments that can provide multiple logical routes, with an integrity guarantee that holds when the route-compromise fraction is alpha < 0.5 and that extends to correlated routes, where the quantity that matters is the size of the largest shared-failure group rather than the route count. The abstract states that below this threshold quorum-denial and message-dropping adversaries cannot block honest traffic, so availability and integrity degrade at the same threshold. The evaluation covers MMLU, HumanEval, and MBPP on two frameworks and four topologies and a MetaGPT deployment against a production API, reporting 0% ASR below the threshold at 1x LLM token cost.

The abstract does not state the value of k, how routes are constructed, or the concrete form of the canonicalized payload, nor does it give sample sizes or statistical details for each benchmark and topology. How the largest shared-failure group size is measured in the correlated-route extension and how alpha would be estimated in a real deployment are not expanded. The abstract notes that LLM-as-a-Judge blocks up to 44.2% of benign outputs in the topology sweep but does not describe how that rate varies with topology and task. Because the reading scope here is the abstract only, these details require the paper's methods, experiments, and appendix.

Sources