Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1344

GAD-RL: adaptive gated distillation to improve OCR faithfulness

The paper (arXiv:2609.38282v1) introduces GAD-RL, a method that adaptively gates and attenuates teacher distillation during joint post-training to improve OCR faithfulness in vision–language models. On Qwen3.5-2B, GAD-RL disables distillation for high-reward response groups, weights forward KL by the student's probability of the teacher's top token, and achieves 59.92% Micro Recall on CHAOS-Bench (beating GRPO and GRPO+OPD by 8.45 and 4.43 points) and an Overall score of 91.18 on OmniDocBench v1.6.

KEY POINTS

  1. The paper (arXiv:2609.38282v1) introduces GAD-RL, a method that adaptively gates and attenuates teacher distillation during joint post-training to improve OCR faithfulness in vision–language models.
  2. On Qwen3.5-2B, GAD-RL disables distillation for high-reward response groups, weights forward KL by the student's probability of the teacher's top token, and achieves 59.92% Micro Recall on CHAOS-Bench (beating GRPO and GRPO+OPD by 8.45 and 4.43 points) and an Overall score of 91.18 on OmniDocBench v1.6.
  3. Adaptive regulation of teacher supervision preserves transcription faithfulness while letting the student improve, yielding measurable OCR accuracy gains on standard benchmarks.

WHY IT MATTERS

Adaptive regulation of teacher supervision preserves transcription faithfulness while letting the student improve, yielding measurable OCR accuracy gains on standard benchmarks.

SOURCES & TIMELINE

1