RESEARCH · RESEARCH · #839
MAWILE: workbench to audit LLM evaluators (arXiv v1)
Megagon Labs released MAWILE, a developer-facing workbench (arXiv:2609.22599v1) that audits sensitivity of LLM judges across four surfaces—judge prompt, judge rubric, target-system input, and target-system output—by constructing controlled perturbations and re-executing the judge. MAWILE supports binary, ordinal, and pairwise judges without requiring gold labels, and the code is available on GitHub.
KEY POINTS
- Megagon Labs released MAWILE, a developer-facing workbench (arXiv:2609.22599v1) that audits sensitivity of LLM judges across four surfaces—judge prompt, judge rubric, target-system input, and target-system output—by constructing controlled perturbations and re-executing the judge.
- MAWILE supports binary, ordinal, and pairwise judges without requiring gold labels, and the code is available on GitHub.
- MAWILE helps developers detect and localize fragility and bias in automated LLM judges, improving reliability of model evaluation without needing gold labels.
WHY IT MATTERS
MAWILE helps developers detect and localize fragility and bias in automated LLM judges, improving reliability of model evaluation without needing gold labels.