Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

COMPANY · ENTITY #6851

LLM judges

Related event timeline, sources and context from the news index.

EVENT TIMELINE

4

RESEARCH · 1 SOURCE · arXiv cs.AI

Paper (arXiv:2610.02492v1) introduces O*NET-BENCH to audit LLM judges for occupational measurement

The authors introduce O*NET-BENCH, an audit suite derived from a survey of 45,796 worker ratings, and evaluate 33 existing judge configurations across six model families on 4,501 test ratings. They find many judges achieve tie-aware pair accuracy ≥0.60 (and a TF-IDF response-only baseline nearly matches the best), yet judges’ estimates of acceptable responses range from 3.0%–97.9% versus 61.1% for occupation-matched workers; calibration reduces mean bias but explains at most 8.5% of individual worker-rating variance.

6.0

MODELS · 1 SOURCE · arXiv cs.AI

Hapi: a U-Net Swin Transformer for continental-scale 24–72h hydrological forecasts

Hapi is a multivariable U-Net Swin Transformer that forecasts discharge, surface runoff, snow water equivalent, and soil wetness across the contiguous United States at 0.05° resolution for 24–72 hour lead times. On 2024 test data using reconstructed ERA5‑Land inputs it outperformed an operational physics-based model and a state-of-the-art AI model in flood detection, was validated against 3,881 USGS gauges and a Hurricane Helene case study, and runs a four-variable 72-hour forecast in an average 0.11 s on a single A100 GPU.

8.0

RESEARCH · 1 SOURCE · arXiv cs.AI

AutoGym: blueprint-first framework to synthesize verifiable agent gyms

The paper (arXiv:2609.22592v1) introduces AutoGym, a framework that automatically generates complete reinforcement-learning gyms—tasks, executable environments, and verifiers—starting from a minimal domain seed or prior trajectories. AutoGym combines (1) blueprint-first generation that specifies valid solution spaces and verification criteria before materializing environments, (2) explicit generation parameters for fine-grained difficulty and topology control, and (3) active curriculum synthesis that adapts parameter distributions based on model performance; authors report generation of gyms across the capability spectrum, including instances that challenge frontier models in productivity and temporal-reasoning settings.

7.0

RESEARCH · 1 SOURCE · arXiv cs.AI

MAWILE: workbench to audit LLM evaluators (arXiv v1)

Megagon Labs released MAWILE, a developer-facing workbench (arXiv:2609.22599v1) that audits sensitivity of LLM judges across four surfaces—judge prompt, judge rubric, target-system input, and target-system output—by constructing controlled perturbations and re-executing the judge. MAWILE supports binary, ordinal, and pairwise judges without requiring gold labels, and the code is available on GitHub.

6.0