Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1003

MeshHeal: decentralized two-timescale self-healing for gray failures in LLM agent networks (arXiv:2609.29015v1)

MeshHeal is a fully decentralized self-healing framework for LLM-based multi-agent systems that couples ability-matched peer review across two timescales: a fast adaptive escalation path (single reviewer → committee → correction) and a slow peer-relative detector that aggregates scores to identify persistent degradation, trigger mandatory committee review, exclude degraded agents from routing, and support reintegration via recovery probes. Using a new Model-Backed MAS Evaluation to tie ability assignments to execution models, MeshHeal achieves 0.839 degraded-phase accuracy on BBH, MATH, and MMLU-Pro with 51k model tokens per task, outperforming the strongest baseline Symphony (0.807 accuracy at 115k tokens) and isolating/reintegrating degraded agents under staggered degradation and recovery.

KEY POINTS

  1. MeshHeal is a fully decentralized self-healing framework for LLM-based multi-agent systems that couples ability-matched peer review across two timescales: a fast adaptive escalation path (single reviewer → committee → correction) and a slow peer-relative detector that aggregates scores to identify persistent degradation, trigger mandatory committee review, exclude degraded agents from routing, and support reintegration via recovery probes.
  2. Using a new Model-Backed MAS Evaluation to tie ability assignments to execution models, MeshHeal achieves 0.839 degraded-phase accuracy on BBH, MATH, and MMLU-Pro with 51k model tokens per task, outperforming the strongest baseline Symphony (0.807 accuracy at 115k tokens) and isolating/reintegrating degraded agents under staggered degradation and recovery.
  3. It offers a decentralized, resource-efficient method to detect and mitigate persistent 'gray failures' in LLM agent networks, improving task accuracy and routing robustness without central coordination.

WHY IT MATTERS

It offers a decentralized, resource-efficient method to detect and mitigate persistent 'gray failures' in LLM agent networks, improving task accuracy and routing robustness without central coordination.

SOURCES & TIMELINE

1