Public articles linked to the same research event.
arXiv MIRROR is a communication-layer integrity primitive that replicates a single canonicalized payload across k logical routes and accepts a message only when a strict majority of routes report the same digest; under a route-compromise bound of alpha < 0.5, it reduces attack success rate to 0% across MMLU, HumanEval, and MBPP on two frameworks and four topologies and in a MetaGPT deployment against a production API, at 1x LLM token cost, whereas LLM-as-a-Judge costs 35x in the same deployment and blocks up to 44.2% of benign outputs in the topology sweep.
MIRROR is a communication-layer integrity primitive that replicates a single canonicalized payload across k logical routes and accepts a message only when a strict majority of routes report the same digest; under a route-compromise bound of alpha < 0.5, it reduces attack success rate to 0% across MMLU, HumanEval, and MBPP on two frameworks and four topologies and in a MetaGPT deployment against a production API, at 1x LLM token cost, whereas LLM-as-a-Judge costs 35x in the same deployment and blocks up to 44.2% of benign outputs in the topology sweep.
MIRROR is a communication-layer integrity primitive that replicates a single canonicalized payload across k logical routes and accepts a message only when a strict majority of routes report the same digest; under a route-compromise bound of alpha < 0.5, it reduces attack success rate to 0% across MMLU, HumanEval, and MBPP on two frameworks and four topologies and in a MetaGPT deployment against a production API, at 1x LLM token cost, whereas LLM-as-a-Judge costs 35x in the same deployment and blocks up to 44.2% of benign outputs in the topology sweep.
MIRROR is a communication-layer integrity primitive that replicates a single canonicalized payload across k logical routes and accepts a message only when a strict majority of routes report the same digest; under a route-compromise bound of alpha < 0.5, it reduces attack success rate to 0% across MMLU, HumanEval, and MBPP on two frameworks and four topologies and in a MetaGPT deployment against a production API, at 1x LLM token cost, whereas LLM-as-a-Judge costs 35x in the same deployment and blocks up to 44.2% of benign outputs in the topology sweep.