InfoNCE identifies efficient layer choices for VLA policies (arXiv:2609.36118v1)
This arXiv preprint studies which layers of frozen vision-language backbones should be exposed to an action head for vision-language-action (VLA) policies. Across three pretrained models, two manipulation benchmarks (LIBERO, CALVIN) and 54 layer/fusion configurations, the authors find that most fusion strategies underperform the best single-layer policy, and that an InfoNCE-based proxy score correlates best with downstream policy success—allowing layer selection with 9–33× less GPU compute and reducing mean selection regret from 17.89 to 3.71 percentage points in their evaluations.