Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1405

arXiv preprint 'Predictive Credit' v1 finds mixed/inconclusive effects of AI explanations on experimental forecasts

The arXiv preprint "Predictive Credit" (v1) introduces a protocol to measure whether explanations from research agents add predictive value to planned experiments by comparing paired forecasts that vary only in explanation and donor context. Across controlled learning states, 12 Tox21 endpoints and 24 OpenML tasks the authors report largely inconclusive results for natural-explanation credit (e.g., the v5 frozen credit decision was inconclusive and a preregistered Tox21 ROC AUC interval-score harm test was unmet with D-M = -0.0026, 95% interval [-0.0174, 0.0104]); some model-specific interventions did show effects (DeepSeek V4 Pro matched and donor cards reduced secondary Tox21 drift by 64.5% and 59.1%, while DeepSeek V4 Flash raised matched point MAE from 0.01823 to 0.02020). A researcher-authored mechanism control lowered point MAE by 2.60 percentage points versus description, but matched point-accuracy gains over description were not broadly confirmed and many intervals spanned zero.

KEY POINTS

  1. The arXiv preprint "Predictive Credit" (v1) introduces a protocol to measure whether explanations from research agents add predictive value to planned experiments by comparing paired forecasts that vary only in explanation and donor context.
  2. Across controlled learning states, 12 Tox21 endpoints and 24 OpenML tasks the authors report largely inconclusive results for natural-explanation credit (e.g., the v5 frozen credit decision was inconclusive and a preregistered Tox21 ROC AUC interval-score harm test was unmet with D-M = -0.0026, 95% interval [-0.0174, 0.0104]); some model-specific interventions did show effects (DeepSeek V4 Pro matched and donor cards reduced secondary Tox21 drift by 64.5% and 59.1%, while DeepSeek V4 Flash raised matched point MAE from 0.01823 to 0.02020).
  3. A researcher-authored mechanism control lowered point MAE by 2.60 percentage points versus description, but matched point-accuracy gains over description were not broadly confirmed and many intervals spanned zero.

WHY IT MATTERS

Provides a formal, preregistered protocol and empirical results on whether AI-generated explanations improve experimental forecasting, impacting how research-agent benchmarks and scientific forecasting tools are evaluated and deployed.

SOURCES & TIMELINE

1