Putting Captions to the Test: Evaluating Video Caption Quality via Multiple-Choice QA
This research paper proposes redefining video-caption quality in terms of information fidelity and evaluating captions using multiple-choice question answering (MCQA) rather than relying solely on text-overlap with ground-truth references. The approach is intended to address the one-to-many nature of video description and provide a more fine-grained, content-focused assessment of captions for Visual Large Language Models (VLLMs).