Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1242

InfoNCE identifies efficient layer choices for VLA policies (arXiv:2609.36118v1)

This arXiv preprint studies which layers of frozen vision-language backbones should be exposed to an action head for vision-language-action (VLA) policies. Across three pretrained models, two manipulation benchmarks (LIBERO, CALVIN) and 54 layer/fusion configurations, the authors find that most fusion strategies underperform the best single-layer policy, and that an InfoNCE-based proxy score correlates best with downstream policy success—allowing layer selection with 9–33× less GPU compute and reducing mean selection regret from 17.89 to 3.71 percentage points in their evaluations.

KEY POINTS

  1. This arXiv preprint studies which layers of frozen vision-language backbones should be exposed to an action head for vision-language-action (VLA) policies.
  2. Across three pretrained models, two manipulation benchmarks (LIBERO, CALVIN) and 54 layer/fusion configurations, the authors find that most fusion strategies underperform the best single-layer policy, and that an InfoNCE-based proxy score correlates best with downstream policy success—allowing layer selection with 9–33× less GPU compute and reducing mean selection regret from 17.89 to 3.71 percentage points in their evaluations.
  3. This matters because InfoNCE provides a cheap, practical proxy for selecting backbone layers for VLA policies, cutting compute needs and making exhaustive policy sweeps unnecessary in many settings.

WHY IT MATTERS

This matters because InfoNCE provides a cheap, practical proxy for selecting backbone layers for VLA policies, cutting compute needs and making exhaustive policy sweeps unnecessary in many settings.

SOURCES & TIMELINE

1