NEWS · RESEARCH · #346
ReDraft: reference-driven revision for continual VLLM post-training
This arXiv paper (2609.16639v1) introduces ReDraft, a continual post-training method that has a model revise its own incorrect rollouts using an expert response only as a reference, accepts revised targets via a verifier, and fine-tunes on retained revisions. On Qwen2.5-VL-3B/7B across Counting, Clock Reading, and Jigsaw, ReDraft yields larger target-task gains than SFT while dramatically reducing forgetting (prior-task loss much lower) and improving OPSD metrics.
KEY POINTS
- This arXiv paper (2609.16639v1) introduces ReDraft, a continual post-training method that has a model revise its own incorrect rollouts using an expert response only as a reference, accepts revised targets via a verifier, and fine-tunes on retained revisions.
- On Qwen2.5-VL-3B/7B across Counting, Clock Reading, and Jigsaw, ReDraft yields larger target-task gains than SFT while dramatically reducing forgetting (prior-task loss much lower) and improving OPSD metrics.
- ReDraft creates explicit but policy-proximal training targets from the model's own failures, enabling better new-task learning while greatly reducing catastrophic forgetting in continual multimodal post-training.
WHY IT MATTERS
ReDraft creates explicit but policy-proximal training targets from the model's own failures, enabling better new-task learning while greatly reducing catastrophic forgetting in continual multimodal post-training.