NEWS · RESEARCH · #244
REFACTOR-VLA: Unsupervised library learning of typed motor programs
The paper proposes REFACTOR-VLA, an unsupervised approach for learning a library of typed motor programs to modularize vision-language-action (VLA) models. It frames current VLA models (e.g., OpenVLA, π0, RT-2, RDT-1B) as 'monolithic'—producing raw commands or short action sequences—arguing this limits long-horizon performance and interpretability, and it criticizes prior skill-discovery work for not resolving when two action sequences are 'behaviorally equivalent.'
KEY POINTS
- The paper proposes REFACTOR-VLA, an unsupervised approach for learning a library of typed motor programs to modularize vision-language-action (VLA) models.
- It frames current VLA models (e.g., OpenVLA, π0, RT-2, RDT-1B) as 'monolithic'—producing raw commands or short action sequences—arguing this limits long-horizon performance and interpretability, and it criticizes prior skill-discovery work for not resolving when two action sequences are 'behaviorally equivalent.'
- If successful, learning reusable, well-typed motor primitives could improve long-horizon task performance, interpretability, and reuse in VLA systems.
WHY IT MATTERS
If successful, learning reusable, well-typed motor primitives could improve long-horizon task performance, interpretability, and reuse in VLA systems.