RESEARCH · RESEARCH · #1418
Study shows rationales mainly affect verifier judgments, not answer accuracy
The paper introduces a message-intervention diagnostic that holds evidence and candidate answers constant while varying only the rationale passed from a reasoner to a verifier. On 400 examples across MuSiQue, HotpotQA and 2WikiMultiHopQA using DeepSeek as generator and verifier, faithful rationales add almost no answer accuracy versus no rationale, but corrupted rationales substantially change verifier support judgments (10–22% under a blind verifier prompt, 34–55% with explicit rationale-checking), while final answers change less (2–30%); human audits reveal instances of model overtrust.
KEY POINTS
- The paper introduces a message-intervention diagnostic that holds evidence and candidate answers constant while varying only the rationale passed from a reasoner to a verifier.
- On 400 examples across MuSiQue, HotpotQA and 2WikiMultiHopQA using DeepSeek as generator and verifier, faithful rationales add almost no answer accuracy versus no rationale, but corrupted rationales substantially change verifier support judgments (10–22% under a blind verifier prompt, 34–55% with explicit rationale-checking), while final answers change less (2–30%); human audits reveal instances of model overtrust.
- This matters because it reframes rationale sharing as a verification-message channel that can create new failure modes (overtrust or corrupted support) even when answer accuracy is unchanged, affecting evaluation and pipeline design for QA systems.
WHY IT MATTERS
This matters because it reframes rationale sharing as a verification-message channel that can create new failure modes (overtrust or corrupted support) even when answer accuracy is unchanged, affecting evaluation and pipeline design for QA systems.