RESEARCH · RESEARCH · #940
Evidence order shifts multimodal LLM judgments (cross-modal evidence noncommutativity)
A new arXiv paper (arXiv:2609.26986v1) studies how the order of conflicting perceptual (image or speech) and textual evidence affects multimodal large language models. Using a paired-comparison design that fixes instructions and content but swaps source positions, the authors find that placing an image or recording after conflicting text consistently shifts model answers toward the perceptual content, a phenomenon they term cross-modal evidence noncommutativity.
KEY POINTS
- A new arXiv paper (arXiv:2609.26986v1) studies how the order of conflicting perceptual (image or speech) and textual evidence affects multimodal large language models.
- Using a paired-comparison design that fixes instructions and content but swaps source positions, the authors find that placing an image or recording after conflicting text consistently shifts model answers toward the perceptual content, a phenomenon they term cross-modal evidence noncommutativity.
- This matters because evidence order can confound measures of modality bias and change evaluation or prompting outcomes for multimodal models, implying experiment and interface designs must control for ordering effects.
WHY IT MATTERS
This matters because evidence order can confound measures of modality bias and change evaluation or prompting outcomes for multimodal models, implying experiment and interface designs must control for ordering effects.