RESEARCH · RESEARCH · #717
CoVer: information-gain rewards and diversity-pruned tests for more reliable code generation
arXiv:2609.21208v1 introduces CoVer, a single-policy GRPO framework that co-trains a coder and verifier using an information-gain reward (mutual information with a graded ground-truth correctness signal) combined with a three-stage diversity-aware pruning of self-generated tests. The authors report improvements in one-shot pass@1 over a Qwen2.5-Instruct backbone (+5.8 points at 7B, +7.1 points at 14B) across five benchmarks and a +3.5-point lift when integrated into the CodeT ranking pipeline with a 7B backbone.
KEY POINTS
- arXiv:2609.21208v1 introduces CoVer, a single-policy GRPO framework that co-trains a coder and verifier using an information-gain reward (mutual information with a graded ground-truth correctness signal) combined with a three-stage diversity-aware pruning of self-generated tests.
- The authors report improvements in one-shot pass@1 over a Qwen2.5-Instruct backbone (+5.8 points at 7B, +7.1 points at 14B) across five benchmarks and a +3.5-point lift when integrated into the CodeT ranking pipeline with a 7B backbone.
- By rewarding tests that provide real information about correctness and pruning redundant cases, CoVer targets two key failure modes in self-play RL for code generation, improving estimator reliability and pass@1 performance.
WHY IT MATTERS
By rewarding tests that provide real information about correctness and pruning redundant cases, CoVer targets two key failure modes in self-play RL for code generation, improving estimator reliability and pass@1 performance.