Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1256

Beyond Symmetric Agents (arXiv:2609.35875v1): MAD gains traced to ensemble-sampling, not cognitive diversity

This paper tests whether cognitive diversity drives improvements from multi-agent debate (MAD) using 23 small open-weight models, five tasks and 5,500+ runs across three diversity axes (personas, temperature, model identity). The authors reject the diversity hypothesis: debate can beat single-agent inference, but at matched compute/token budget it ties or loses to self-consistency sampling; persona prompting typically reduces accuracy (a "persona tax"); mixed-model teams track member ability not heterogeneity; and a context-window overflow bug materially affected prior comparisons — overall recasting reported MAD gains as an ensemble-sampling effect and providing a budget- and contamination-checked baseline.

KEY POINTS

  1. This paper tests whether cognitive diversity drives improvements from multi-agent debate (MAD) using 23 small open-weight models, five tasks and 5,500+ runs across three diversity axes (personas, temperature, model identity).
  2. The authors reject the diversity hypothesis: debate can beat single-agent inference, but at matched compute/token budget it ties or loses to self-consistency sampling; persona prompting typically reduces accuracy (a "persona tax"); mixed-model teams track member ability not heterogeneity; and a context-window overflow bug materially affected prior comparisons — overall recasting reported MAD gains as an ensemble-sampling effect and providing a budget- and contamination-checked baseline.
  3. This matters because it challenges the claimed mechanism for MAD improvements, shows those gains often reduce to ensemble/sampling effects or implementation issues, and sets a stricter baseline that future debate methods must surpass.

WHY IT MATTERS

This matters because it challenges the claimed mechanism for MAD improvements, shows those gains often reduce to ensemble/sampling effects or implementation issues, and sets a stricter baseline that future debate methods must surpass.

SOURCES & TIMELINE

1