MOTIVE: multi-view self-verification framework for vision-language models
The arXiv preprint (arXiv:2610.07018v1) introduces MOTIVE, a Multi-View Self-Verification framework that evaluates candidate answers from complementary verification perspectives and learns a correctness-aligned reliability score to decide whether to accept an answer or trigger history-guided rethinking. Experiments across diverse multimodal benchmarks and VLM backbones show MOTIVE improves over strong self-verification and self-correction baselines, reducing unnecessary reasoning turns while increasing answer reliability.