Paper (arXiv:2610.02492v1) introduces O*NET-BENCH to audit LLM judges for occupational measurement
The authors introduce O*NET-BENCH, an audit suite derived from a survey of 45,796 worker ratings, and evaluate 33 existing judge configurations across six model families on 4,501 test ratings. They find many judges achieve tie-aware pair accuracy ≥0.60 (and a TF-IDF response-only baseline nearly matches the best), yet judges’ estimates of acceptable responses range from 3.0%–97.9% versus 61.1% for occupation-matched workers; calibration reduces mean bias but explains at most 8.5% of individual worker-rating variance.