RESEARCH · RESEARCH · #1098
BAER: backbone-adaptive evidence routing improves pairwise LLM judging
The paper on arXiv (2609.30751v1) introduces Backbone-Adaptive Evidence Routing (BAER), a method that adapts how pairwise language-model judges gather evidence while preserving candidate-symmetry. BAER decomposes each expert's signed preference and candidate-invariant reliability into three symmetric heads (evidence stacking, reliability-based expert routing, and candidate-blind reference verification), selects one head per benchmark–backbone condition on development data (frozen before testing), and achieves the highest test accuracy across eight conditions (two 8B judge backbones × four benchmarks) with full prediction coverage and gains of 0.87–7.32 points versus the strongest external baseline.
KEY POINTS
- The paper on arXiv (2609.30751v1) introduces Backbone-Adaptive Evidence Routing (BAER), a method that adapts how pairwise language-model judges gather evidence while preserving candidate-symmetry.
- BAER decomposes each expert's signed preference and candidate-invariant reliability into three symmetric heads (evidence stacking, reliability-based expert routing, and candidate-blind reference verification), selects one head per benchmark–backbone condition on development data (frozen before testing), and achieves the highest test accuracy across eight conditions (two 8B judge backbones × four benchmarks) with full prediction coverage and gains of 0.87–7.32 points versus the strongest external baseline.
- BAER shows that adapting the evidence-gathering protocol per backbone/benchmark yields consistent accuracy gains and full coverage, which matters for building more reliable automated LLM judges.
WHY IT MATTERS
BAER shows that adapting the evidence-gathering protocol per backbone/benchmark yields consistent accuracy gains and full coverage, which matters for building more reliable automated LLM judges.