NEWS · MODELS · #308
VideoPrism — Google Research's large foundational visual encoder for diverse video understanding
Google Research introduces VideoPrism, a video foundation model (ViFM) that produces frozen video representations for a wide range of tasks (classification, localization, retrieval, captioning, QA). It is pre-trained on a hybrid corpus of 36 million high-quality video-text pairs plus 582 million additional clips with noisy or machine-generated text and uses a two-stage training scheme (video-text contrastive learning followed by masked video modeling).
KEY POINTS
- Google Research introduces VideoPrism, a video foundation model (ViFM) that produces frozen video representations for a wide range of tasks (classification, localization, retrieval, captioning, QA).
- It is pre-trained on a hybrid corpus of 36 million high-quality video-text pairs plus 582 million additional clips with noisy or machine-generated text and uses a two-stage training scheme (video-text contrastive learning followed by masked video modeling).
- A single, widely pre-trained frozen video encoder that reportedly achieves state-of-the-art results could simplify and standardize many downstream video understanding tasks and research efforts.
WHY IT MATTERS
A single, widely pre-trained frozen video encoder that reportedly achieves state-of-the-art results could simplify and standardize many downstream video understanding tasks and research efforts.