Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

MODEL · ENTITY #6541

Claude Opus 4.5

Related event timeline, sources and context from the news index.

EVENT TIMELINE

3

RESEARCH · 1 SOURCE · arXiv cs.AI

DeReAct paper: modular Critic and Context Manager for more reliable ReAct-style agents

The arXiv preprint introduces DeReAct, a modular agent architecture that externalizes two gating policies—a Critic to validate proposed actions and a Context Manager to reconstruct environment-supported state and certify task completion. On GAIA and SWE-bench Verified, DeReAct raises Pass@1 most for weaker "Brain" models (6.5–7.0 points for Qwen3-Coder-480B, 4.2–5.2 points for Claude Sonnet 4.5), with smaller gains as model capability increases; with Claude Opus 4.5 overall Pass@1 is comparable to ReAct but trajectories are more evidence-complete and constraint-satisfying.

7.0

MODELS · 1 SOURCE · The Decoder

OpenAI's GPT-6 Astra scores 80% on Epoch AI's furniture assembly benchmark

Epoch AI's Furniture Assembly Benchmark (FAB) tests models on identifying deliberate assembly errors in photos of IKEA furniture. Previously the best model (Claude Opus 4.5) scored 28% in November 2025; ten months later OpenAI's GPT-6 Astra reaches 80% with a processing time of about three minutes per photo, followed by Claude Fable 5.1 at 70% and Claude Opus 5 at 61%, while some Chinese open-weight models like Kimi K3 lag substantially.

8.0

RESEARCH · 1 SOURCE · Hugging Face

UK AISI publishes verified benchmark results via EvalEval's Evaluation Cards

The UK AI Security Institute (AISI) is using the EvalEval Coalition's Every Eval Ever schema and Evaluation Cards platform to publicly share verified evaluation runs tied to its Terminal-Bench 2.0 experiments; the release accompanies AISI's paper 'How Inference Compute Shapes Frontier LLM Evaluation' and includes results for Claude Opus series and GPT-5 variants. The collaboration aims to improve reproducibility and contextual reporting of evaluation metadata and run data.

7.0