Labels override definitions in Jev-style typed decision models (arXiv:2610.02586v1)
The paper analyzes Jev-style typed decision interfaces across four open-weight models (Qwen2.5 backbones), eleven classification tasks and a synthetic suite called PolicyBench, and shows that models often follow short option labels rather than their written definitions (an "option-label bias"). Deleting or renaming labels can change accuracy substantially without altering model weights; one codebase ('von') avoided the effect by formatting options as definitions only. The authors provide a two-call test to detect which behavior applies and evaluate four mitigation strategies.