Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

COMPANY · ENTITY #7776

HotpotQA

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

RESEARCH · 1 SOURCE · arXiv cs.AI

Study shows rationales mainly affect verifier judgments, not answer accuracy

The paper introduces a message-intervention diagnostic that holds evidence and candidate answers constant while varying only the rationale passed from a reasoner to a verifier. On 400 examples across MuSiQue, HotpotQA and 2WikiMultiHopQA using DeepSeek as generator and verifier, faithful rationales add almost no answer accuracy versus no rationale, but corrupted rationales substantially change verifier support judgments (10–22% under a blind verifier prompt, 34–55% with explicit rationale-checking), while final answers change less (2–30%); human audits reveal instances of model overtrust.

7.0

RESEARCH · 1 SOURCE · arXiv cs.AI

RLDS: decomposing trajectory rewards by subtask improves RL for language-model agents (arXiv:2609.27035v1)

The paper introduces Reinforcement Learning with Decomposed Subtasks (RLDS) and Subtask-Decomposed Advantage Estimation (SDAE), which split trajectory reward into per-subtask shares before policy updates instead of collapsing outcomes to a single scalar as in Group Relative Policy Optimization (GRPO). Evaluated on four benchmarks, RLDS produced substantial gains where subtask heterogeneity is high (ScienceWorld +11.5 points, paired-bootstrap 95% CI [+9.8, +13.3]; FrozenLake +9.8 points, CI [+7.0, +12.8]), showed little effect where diagnostics predicted little recovery (HotpotQA, DeepResearch), and was more compute-efficient on ScienceWorld (-10.9% wall-clock per step).

7.0