Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #717

CoVer: information-gain rewards and diversity-pruned tests for more reliable code generation

arXiv:2609.21208v1 introduces CoVer, a single-policy GRPO framework that co-trains a coder and verifier using an information-gain reward (mutual information with a graded ground-truth correctness signal) combined with a three-stage diversity-aware pruning of self-generated tests. The authors report improvements in one-shot pass@1 over a Qwen2.5-Instruct backbone (+5.8 points at 7B, +7.1 points at 14B) across five benchmarks and a +3.5-point lift when integrated into the CodeT ranking pipeline with a 7B backbone.

KEY POINTS

  1. arXiv:2609.21208v1 introduces CoVer, a single-policy GRPO framework that co-trains a coder and verifier using an information-gain reward (mutual information with a graded ground-truth correctness signal) combined with a three-stage diversity-aware pruning of self-generated tests.
  2. The authors report improvements in one-shot pass@1 over a Qwen2.5-Instruct backbone (+5.8 points at 7B, +7.1 points at 14B) across five benchmarks and a +3.5-point lift when integrated into the CodeT ranking pipeline with a 7B backbone.
  3. By rewarding tests that provide real information about correctness and pruning redundant cases, CoVer targets two key failure modes in self-play RL for code generation, improving estimator reliability and pass@1 performance.

WHY IT MATTERS

By rewarding tests that provide real information about correctness and pruning redundant cases, CoVer targets two key failure modes in self-play RL for code generation, improving estimator reliability and pass@1 performance.

SOURCES & TIMELINE

1