Evidence order shifts multimodal LLM judgments (cross-modal evidence noncommutativity)
A new arXiv paper (arXiv:2609.26986v1) studies how the order of conflicting perceptual (image or speech) and textual evidence affects multimodal large language models. Using a paired-comparison design that fixes instructions and content but swaps source positions, the authors find that placing an image or recording after conflicting text consistently shifts model answers toward the perceptual content, a phenomenon they term cross-modal evidence noncommutativity.