VideoPrism — Google Research's large foundational visual encoder for diverse video understanding
Google Research introduces VideoPrism, a video foundation model (ViFM) that produces frozen video representations for a wide range of tasks (classification, localization, retrieval, captioning, QA). It is pre-trained on a hybrid corpus of 36 million high-quality video-text pairs plus 582 million additional clips with noisy or machine-generated text and uses a two-stage training scheme (video-text contrastive learning followed by masked video modeling).