Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

TOPIC · ENTITY #6852

judge prompt

Related event timeline, sources and context from the news index.

EVENT TIMELINE

1

RESEARCH · 1 SOURCE · arXiv cs.AI

MAWILE: workbench to audit LLM evaluators (arXiv v1)

Megagon Labs released MAWILE, a developer-facing workbench (arXiv:2609.22599v1) that audits sensitivity of LLM judges across four surfaces—judge prompt, judge rubric, target-system input, and target-system output—by constructing controlled perturbations and re-executing the judge. MAWILE supports binary, ordinal, and pairwise judges without requiring gold labels, and the code is available on GitHub.

6.0