Tech Meridian ← LIVE FEED
RU

NEWS · MODELS · #308

VideoPrism — Google Research's large foundational visual encoder for diverse video understanding

Google Research introduces VideoPrism, a video foundation model (ViFM) that produces frozen video representations for a wide range of tasks (classification, localization, retrieval, captioning, QA). It is pre-trained on a hybrid corpus of 36 million high-quality video-text pairs plus 582 million additional clips with noisy or machine-generated text and uses a two-stage training scheme (video-text contrastive learning followed by masked video modeling).

KEY POINTS

  1. Google Research introduces VideoPrism, a video foundation model (ViFM) that produces frozen video representations for a wide range of tasks (classification, localization, retrieval, captioning, QA).
  2. It is pre-trained on a hybrid corpus of 36 million high-quality video-text pairs plus 582 million additional clips with noisy or machine-generated text and uses a two-stage training scheme (video-text contrastive learning followed by masked video modeling).
  3. A single, widely pre-trained frozen video encoder that reportedly achieves state-of-the-art results could simplify and standardize many downstream video understanding tasks and research efforts.

WHY IT MATTERS

A single, widely pre-trained frozen video encoder that reportedly achieves state-of-the-art results could simplify and standardize many downstream video understanding tasks and research efforts.

SOURCES & TIMELINE

1