RESEARCH · RESEARCH · #1010
Human–AI hypothesis testing: SCALE policy for cost-aware selective AI scoring and human escalation
This arXiv preprint (arXiv:2609.28859v1) studies using AI judgments together with selective human verification to perform valid hypothesis tests with controlled type‑I/II errors at minimum cost. The authors derive an information‑theoretic lower bound on cost, characterize a report‑dependent information frontier, and introduce SCALE, a sequential cost‑aware policy that adaptively combines selective AI scoring with human escalation; SCALE is finite‑sample valid, matches the bound to first order as target error probabilities vanish, and can be extended using paired AI–human pilot data.
KEY POINTS
- This arXiv preprint (arXiv:2609.28859v1) studies using AI judgments together with selective human verification to perform valid hypothesis tests with controlled type‑I/II errors at minimum cost.
- The authors derive an information‑theoretic lower bound on cost, characterize a report‑dependent information frontier, and introduce SCALE, a sequential cost‑aware policy that adaptively combines selective AI scoring with human escalation; SCALE is finite‑sample valid, matches the bound to first order as target error probabilities vanish, and can be extended using paired AI–human pilot data.
- Provides a principled, cost‑aware framework and a near‑optimal sequential policy for combining inexpensive AI judgments with selective human verification to guarantee valid statistical inference.
WHY IT MATTERS
Provides a principled, cost‑aware framework and a near‑optimal sequential policy for combining inexpensive AI judgments with selective human verification to guarantee valid statistical inference.