Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

COMPANY · ENTITY #8364

ForecastBench

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

RESEARCH · 1 SOURCE · arXiv cs.AI

ReliabilityRoute paper: behavioral stress tests show when forecasting agents should reason

This arXiv paper treats retrieval, reasoning, market priors and historical analogs as observable agent behaviors in ForecastBench-style binary forecasting tasks and finds that which mechanism works best is source-dependent. The authors introduce ReliabilityRoute, a routing intervention that uses reliability features (historical coverage, market-prior availability and sharpness, evidence strength/disagreement, horizon) and show a fixed 2024-fitted rule closely matches a hand taxonomy while a walk-forward self-adjusting rule achieves the best mean Brier score among their deterministic systems across 16 later LLM vintages; gains are modest and historical/search baselines remain competitive; reproducibility artifacts are on GitHub.

6.0

RESEARCH · 1 SOURCE · The Decoder

FRI interim report finds experts underestimated recent AI progress

An interim report from the Forecasting Research Institute (FRI) finds that senior AI experts and superforecasters significantly underestimated recent AI advances across benchmarks and some adoption/economic metrics. Notable gaps include AI reaching IMO gold-medal level in July 2025 earlier than forecasted, a possible (but unverified) solution to a Millennium Prize Problem, faster-than-expected virology and cybersecurity capabilities, and higher-than-expected company revenue figures (FRI cites roughly $100 billion for Anthropic in September 2026).

7.0