RESEARCH · RESEARCH · #1013
Paper shows RLVR (with GRPO) works on Qwen3.5-0.8B without distillation
The arXiv preprint applies Reinforcement Learning with Verifiable Rewards (RLVR) using Group Relative Policy Optimization (GRPO) and an interleaved Wikipedia search tool to train Qwen3.5-0.8B on MuSiQue and a seven-benchmark QA suite. The best run reaches 0.352 average exact-match (vs 0.092 untrained), a 3.8× improvement with no distillation, and the authors find reward shape matters—sparse exact-match rewards perform worst for this model scale.
KEY POINTS
- The arXiv preprint applies Reinforcement Learning with Verifiable Rewards (RLVR) using Group Relative Policy Optimization (GRPO) and an interleaved Wikipedia search tool to train Qwen3.5-0.8B on MuSiQue and a seven-benchmark QA suite.
- The best run reaches 0.352 average exact-match (vs 0.092 untrained), a 3.8× improvement with no distillation, and the authors find reward shape matters—sparse exact-match rewards perform worst for this model scale.
- Shows RLVR can improve small (≈0.8B) models without teacher distillation but that reward design, not just scaling down large-model recipes, is critical.
WHY IT MATTERS
Shows RLVR can improve small (≈0.8B) models without teacher distillation but that reward design, not just scaling down large-model recipes, is critical.