RESEARCH · RESEARCH · #1109
HARDEN: constrained evolutionary search to create harder, answer-preserving evaluation cases
HARDEN is a constrained evolutionary search method that adapts inputs of existing evaluation cases into more challenging variants while preserving expected outputs and enforcing feasibility constraints (semantics, realism, execution validity). Applied to FinQA, PubMedQA, and ContractNLI on three Qwen3.5 scales, HARDEN reduced model accuracy by 22.7% on average and up to 49.9% versus single-pass baselines using the same feasibility checks.
KEY POINTS
- HARDEN is a constrained evolutionary search method that adapts inputs of existing evaluation cases into more challenging variants while preserving expected outputs and enforcing feasibility constraints (semantics, realism, execution validity).
- Applied to FinQA, PubMedQA, and ContractNLI on three Qwen3.5 scales, HARDEN reduced model accuracy by 22.7% on average and up to 49.9% versus single-pass baselines using the same feasibility checks.
- Demonstrates that constrained evolutionary search can generate substantially harder yet valid evaluation cases, indicating many benchmarks may underrepresent real-world problem difficulty and suggesting stronger robustness testing methods.
WHY IT MATTERS
Demonstrates that constrained evolutionary search can generate substantially harder yet valid evaluation cases, indicating many benchmarks may underrepresent real-world problem difficulty and suggesting stronger robustness testing methods.