Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1495

Labels override definitions in Jev-style typed decision models (arXiv:2610.02586v1)

The paper analyzes Jev-style typed decision interfaces across four open-weight models (Qwen2.5 backbones), eleven classification tasks and a synthetic suite called PolicyBench, and shows that models often follow short option labels rather than their written definitions (an "option-label bias"). Deleting or renaming labels can change accuracy substantially without altering model weights; one codebase ('von') avoided the effect by formatting options as definitions only. The authors provide a two-call test to detect which behavior applies and evaluate four mitigation strategies.

KEY POINTS

  1. The paper analyzes Jev-style typed decision interfaces across four open-weight models (Qwen2.5 backbones), eleven classification tasks and a synthetic suite called PolicyBench, and shows that models often follow short option labels rather than their written definitions (an "option-label bias").
  2. Deleting or renaming labels can change accuracy substantially without altering model weights; one codebase ('von') avoided the effect by formatting options as definitions only.
  3. The authors provide a two-call test to detect which behavior applies and evaluate four mitigation strategies.

WHY IT MATTERS

This matters because prompt rendering — not model weights — can make classifiers follow labels rather than their intended definitions, affecting reliability and safety of triage, moderation, and routing systems.

SOURCES & TIMELINE

1