Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1122

arXiv:2609.30328v1 — Two label-free tests to detect when multi-agent code judges lack grounding

The paper (arXiv:2609.30328v1) studies when a language-model-based judge of code is actually grounded and introduces two label-free measurements to detect when the judge lacks a basis for its verdict. Evaluating the MARCH multi-agent verification pipeline on two code-judging benchmarks across 80 condition-by-cell measurements, the authors find MARCH declares both solutions equally good on 78–95% of comparisons and attains 4.4% accuracy (versus 43.7% when the same model is asked directly); using one log-derived measurement to gate answers increases accuracy from 20.7% to 36.9% while still answering about half of comparisons.

KEY POINTS

  1. The paper (arXiv:2609.30328v1) studies when a language-model-based judge of code is actually grounded and introduces two label-free measurements to detect when the judge lacks a basis for its verdict.
  2. Evaluating the MARCH multi-agent verification pipeline on two code-judging benchmarks across 80 condition-by-cell measurements, the authors find MARCH declares both solutions equally good on 78–95% of comparisons and attains 4.4% accuracy (versus 43.7% when the same model is asked directly); using one log-derived measurement to gate answers increases accuracy from 20.7% to 36.9% while still answering about half of comparisons.
  3. Provides practical, label-free signals that detect when LLM-based judges lack evidential grounding and shows abstention based on those signals can materially improve judgment reliability.

WHY IT MATTERS

Provides practical, label-free signals that detect when LLM-based judges lack evidential grounding and shows abstention based on those signals can materially improve judgment reliability.

SOURCES & TIMELINE

1