RESEARCH · RESEARCH · #1242
InfoNCE identifies efficient layer choices for VLA policies (arXiv:2609.36118v1)
This arXiv preprint studies which layers of frozen vision-language backbones should be exposed to an action head for vision-language-action (VLA) policies. Across three pretrained models, two manipulation benchmarks (LIBERO, CALVIN) and 54 layer/fusion configurations, the authors find that most fusion strategies underperform the best single-layer policy, and that an InfoNCE-based proxy score correlates best with downstream policy success—allowing layer selection with 9–33× less GPU compute and reducing mean selection regret from 17.89 to 3.71 percentage points in their evaluations.
KEY POINTS
- This arXiv preprint studies which layers of frozen vision-language backbones should be exposed to an action head for vision-language-action (VLA) policies.
- Across three pretrained models, two manipulation benchmarks (LIBERO, CALVIN) and 54 layer/fusion configurations, the authors find that most fusion strategies underperform the best single-layer policy, and that an InfoNCE-based proxy score correlates best with downstream policy success—allowing layer selection with 9–33× less GPU compute and reducing mean selection regret from 17.89 to 3.71 percentage points in their evaluations.
- This matters because InfoNCE provides a cheap, practical proxy for selecting backbone layers for VLA policies, cutting compute needs and making exhaustive policy sweeps unnecessary in many settings.
WHY IT MATTERS
This matters because InfoNCE provides a cheap, practical proxy for selecting backbone layers for VLA policies, cutting compute needs and making exhaustive policy sweeps unnecessary in many settings.