Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1013

Paper shows RLVR (with GRPO) works on Qwen3.5-0.8B without distillation

The arXiv preprint applies Reinforcement Learning with Verifiable Rewards (RLVR) using Group Relative Policy Optimization (GRPO) and an interleaved Wikipedia search tool to train Qwen3.5-0.8B on MuSiQue and a seven-benchmark QA suite. The best run reaches 0.352 average exact-match (vs 0.092 untrained), a 3.8× improvement with no distillation, and the authors find reward shape matters—sparse exact-match rewards perform worst for this model scale.

KEY POINTS

  1. The arXiv preprint applies Reinforcement Learning with Verifiable Rewards (RLVR) using Group Relative Policy Optimization (GRPO) and an interleaved Wikipedia search tool to train Qwen3.5-0.8B on MuSiQue and a seven-benchmark QA suite.
  2. The best run reaches 0.352 average exact-match (vs 0.092 untrained), a 3.8× improvement with no distillation, and the authors find reward shape matters—sparse exact-match rewards perform worst for this model scale.
  3. Shows RLVR can improve small (≈0.8B) models without teacher distillation but that reward design, not just scaling down large-model recipes, is critical.

WHY IT MATTERS

Shows RLVR can improve small (≈0.8B) models without teacher distillation but that reward design, not just scaling down large-model recipes, is critical.

SOURCES & TIMELINE

1