RESEARCH · RESEARCH · #1506
Paper (arXiv:2610.02492v1) introduces O*NET-BENCH to audit LLM judges for occupational measurement
The authors introduce O*NET-BENCH, an audit suite derived from a survey of 45,796 worker ratings, and evaluate 33 existing judge configurations across six model families on 4,501 test ratings. They find many judges achieve tie-aware pair accuracy ≥0.60 (and a TF-IDF response-only baseline nearly matches the best), yet judges’ estimates of acceptable responses range from 3.0%–97.9% versus 61.1% for occupation-matched workers; calibration reduces mean bias but explains at most 8.5% of individual worker-rating variance.
KEY POINTS
- The authors introduce O*NET-BENCH, an audit suite derived from a survey of 45,796 worker ratings, and evaluate 33 existing judge configurations across six model families on 4,501 test ratings.
- They find many judges achieve tie-aware pair accuracy ≥0.60 (and a TF-IDF response-only baseline nearly matches the best), yet judges’ estimates of acceptable responses range from 3.0%–97.9% versus 61.1% for occupation-matched workers; calibration reduces mean bias but explains at most 8.5% of individual worker-rating variance.
- This matters because relying on ranking agreement alone can produce highly biased estimates of acceptance rates and occupational aggregates when LLM judges are used for workplace measurement.
WHY IT MATTERS
This matters because relying on ranking agreement alone can produce highly biased estimates of acceptance rates and occupational aggregates when LLM judges are used for workplace measurement.