Run SkyRL GRPO post-training of Qwen3‑VL‑8B on Amazon SageMaker HyperPod
Amazon demonstrates running the open-source SkyRL framework on SageMaker HyperPod (EKS) to post-train a Qwen3‑VL‑8B vision‑language model with Group Relative Policy Optimization (GRPO). Using a HyperPod Ray cluster with shared FSx storage, the workflow improved maze solve rate from 43.75% to over 95% on a fixed 64‑maze evaluation set and uses colocated vLLM inference and FSDP-sharded policy training with LoRA adapters.