Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

COMPANY · ENTITY #8553

MuSiQue

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

RESEARCH · 1 SOURCE · arXiv cs.AI

Study shows rationales mainly affect verifier judgments, not answer accuracy

The paper introduces a message-intervention diagnostic that holds evidence and candidate answers constant while varying only the rationale passed from a reasoner to a verifier. On 400 examples across MuSiQue, HotpotQA and 2WikiMultiHopQA using DeepSeek as generator and verifier, faithful rationales add almost no answer accuracy versus no rationale, but corrupted rationales substantially change verifier support judgments (10–22% under a blind verifier prompt, 34–55% with explicit rationale-checking), while final answers change less (2–30%); human audits reveal instances of model overtrust.

7.0

RESEARCH · 1 SOURCE · arXiv cs.AI

Paper shows RLVR (with GRPO) works on Qwen3.5-0.8B without distillation

The arXiv preprint applies Reinforcement Learning with Verifiable Rewards (RLVR) using Group Relative Policy Optimization (GRPO) and an interleaved Wikipedia search tool to train Qwen3.5-0.8B on MuSiQue and a seven-benchmark QA suite. The best run reaches 0.352 average exact-match (vs 0.092 untrained), a 3.8× improvement with no distillation, and the authors find reward shape matters—sparse exact-match rewards perform worst for this model scale.

6.0