Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1506

Paper (arXiv:2610.02492v1) introduces O*NET-BENCH to audit LLM judges for occupational measurement

The authors introduce O*NET-BENCH, an audit suite derived from a survey of 45,796 worker ratings, and evaluate 33 existing judge configurations across six model families on 4,501 test ratings. They find many judges achieve tie-aware pair accuracy ≥0.60 (and a TF-IDF response-only baseline nearly matches the best), yet judges’ estimates of acceptable responses range from 3.0%–97.9% versus 61.1% for occupation-matched workers; calibration reduces mean bias but explains at most 8.5% of individual worker-rating variance.

KEY POINTS

  1. The authors introduce O*NET-BENCH, an audit suite derived from a survey of 45,796 worker ratings, and evaluate 33 existing judge configurations across six model families on 4,501 test ratings.
  2. They find many judges achieve tie-aware pair accuracy ≥0.60 (and a TF-IDF response-only baseline nearly matches the best), yet judges’ estimates of acceptable responses range from 3.0%–97.9% versus 61.1% for occupation-matched workers; calibration reduces mean bias but explains at most 8.5% of individual worker-rating variance.
  3. This matters because relying on ranking agreement alone can produce highly biased estimates of acceptance rates and occupational aggregates when LLM judges are used for workplace measurement.

WHY IT MATTERS

This matters because relying on ranking agreement alone can produce highly biased estimates of acceptance rates and occupational aggregates when LLM judges are used for workplace measurement.

SOURCES & TIMELINE

1