Skip to main content
Back to timeline
arXivSource publication:

MeshHeal uses two-timescale peer review to reach 0.839 degraded-phase accuracy at 51k tokens per task on BBH, MATH, and MMLU-Pro, versus Symphony's 0.807 at 115k

Synopsis

The work introduces MeshHeal, a fully decentralized self-healing framework that couples ability-matched peer review across two timescales to address gray failures in decentralized LLM-based multi-agent systems: at the fast timescale, an adaptive hierarchy escalates uncertain or low-scoring outputs from repeated single-reviewer evaluation to committee deliberation and, when needed, correction before use; at the slow timescale, a task- and ability-conditioned peer-relative detector aggregates scores to distinguish persistent degradation from ordinary output variation, trigger mandatory committee review, and eventually exclude degraded agents from ordinary routing, with recovery probes providing fresh evidence for reintegration; the authors also introduce Model-Backed MAS Evaluation, which ti

Source-provided article image: MeshHeal: Two-Timescale Self-Healing for Gray Failures in Decentralized LLM Agent Networks
Figure 1 ·

Figure 1: Overview of MeshHeal . (a) Agents route and complete complementary work. (b) Review by agents with the relevant ability either keeps or replaces the current output, while scores accumulated across tasks update routing eligibility and probes allow recovered agents to return. (c) Model-Backed MAS Evaluation ties execution models to declared abilities and the agent’s healthy or degraded condition, making routing errors affect accuracy.

arXiv

Interpretation

MeshHeal is a fully decentralized self-healing framework aimed at gray failures, where an agent remains responsive while its task-solving quality persistently degrades. Prior work tends to focus on explicit failures or centralized coordination, whereas this framework separates protecting current tasks from altering future routing and allows recovered agents to rejoin. The abstract supports this with a framework description and evaluation across BBH, MATH, and MMLU-Pro, reporting 0.839 degraded-phase accuracy and 51k model tokens per task.

The fast timescale uses an adaptive hierarchy that escalates uncertain or low-scoring outputs from repeated single-reviewer evaluation to committee deliberation and, when needed, correction before use. It ties review intensity to output uncertainty rather than applying a fixed review process to all outputs. The abstract provides the mechanism description and indicates it is tested alongside the overall accuracy results in the degraded-phase evaluation.

The slow timescale uses a task- and ability-conditioned peer-relative detector to aggregate scores, distinguish persistent degradation from ordinary output variation, trigger mandatory committee review, and eventually exclude degraded agents from ordinary routing, with recovery probes providing evidence for reintegration. It bases degradation judgment on peer-relative, task- and ability-conditioned aggregation and pairs it with recovery probes so exclusion is not permanent. The abstract reports that under staggered degradation and recovery, MeshHeal isolates degraded agents, keeps them excluded from ordinary task execution until recovery, and returns them to normal routing.

The authors introduce Model-Backed MAS Evaluation, which ties ability assignments to execution models, because prompt-based ability assignments alone can leave routing errors hidden. It addresses the gap between ability assignment and execution model in multi-agent routing evaluation by providing an evaluation setting closer to execution. The abstract states this evaluation setting is used to faithfully evaluate routing and reports accuracy and token consumption against a baseline across BBH, MATH, and MMLU-Pro.

Perspective

The result is framed for decentralized LLM multi-agent networks, applying to settings where agents coordinate through local interactions and may exhibit normal responsiveness alongside persistently degraded quality. It enables a system to protect current tasks before evidence is sufficient to change future routing, to exclude degraded agents from ordinary routing once degradation is confirmed, and to let recovered agents rejoin through recovery probes. For those building or operating such systems, the directly reusable elements are the division of labor across two timescales, ability-matched peer review, and an evaluation approach that ties ability assignments to execution models.

The abstract reports degraded-phase accuracy and per-task token consumption but does not state per-dataset results, false-positive and false-negative behavior of degradation detection, the trigger conditions for recovery probes, or the scale of the staggered degradation and recovery experiments. Readers concerned about how these mechanisms behave in their own systems would still need the experimental setup and ablation results in the main text. In addition, the abstract does not compare against a wider set of baselines or larger agent networks, so how the framework performs under more complex topologies remains an open question.

Sources