Tech Meridian ← ENTITY INDEX
RU

MODEL · ENTITY #490

GPT-5.5

Related event timeline, sources and context from the news index.

EVENT TIMELINE

5

RESEARCH · 1 SOURCE · arXiv cs.AI

arXiv paper evaluates two-agent vs one-call résumé screening with GPT-5.5 and Claude Opus 4.7

The arXiv preprint (arXiv:2609.19530v1) compares traditional one-call résumé screening to a two-agent protocol where employer- and candidate-side agents exchange evidence, using GPT-5.5 and Claude Opus 4.7 on 600 constructed résumé-job pairs. Two-agent screening advanced a larger share of applications overall (GPT-5.5: 33.3%→39.3%; Opus 4.7: 34.0%→35.5%), substantially increased pass rates on a 191-pair borderline pool (GPT-5.5: 4.5%→26.2%; Opus 4.7: 6.5%→16.1%), changed some one-call decisions in both directions, and produced selections that recurred less often on re-runs (notably under GPT-5.5).

7.0

MODELS · 1 SOURCE · The Decoder

Vals AI’s GPT-6 Astra completes multiple long-horizon game milestones and posts big ARC-AGI-3 gains

According to Vals AI and related community runs, GPT-6 Astra reached far-end Minecraft goals (built a Nether portal and gathered end resources before a Creeper destroyed its chest), won Pokemon FireRed in about 18 hours (vs. ~96 hours for GPT‑5.6 Sol), launched a Factorio rocket in ~10 hours, and more, while ARC Prize reported Astra scored ~62.7% on the ARC‑AGI‑3 benchmark (vs. ~7.78% for GPT‑5.6 Sol). ARC Prize and Vals AI attribute the jump to Astra’s ability to form compact symbolic descriptions from observations and turn them into reusable plans while operating through general screen/mouse/keyboard interfaces.

8.0

RESEARCH · 1 SOURCE · arXiv cs.AI

arXiv paper introduces SAFE benchmark to test whether frontier models seek safety evidence before acting

The paper (arXiv:2609.17865v1) introduces SAFE, a controlled benchmark where models decide whether to retrieve optional safety-relevant evidence that varies in cost, probability, severity, and presentation. Evaluating GPT-5.5, o3, Claude Opus 4.8, and Claude Sonnet 4.6, the authors find distinct evidence-acquisition policies (Opus inspects by default, o3 skips most, GPT-5.5 and Sonnet intermediate), strong sensitivity to severity and retrieval cost, weaker sensitivity to probability, and a mismatch between stated rationales and actual causal influences.

7.0

MODELS · 1 SOURCE · Cohere

Cohere debuts Parse — a high-throughput enterprise document parsing model

Cohere has released Parse, a vision-language document parser optimized for enterprise workloads and available via the Cohere API, Compass search stack, and private single-tenant Model Vault. Cohere positions Parse as cost‑effective ($1.50 per 1,000 pages) with high throughput and stronger price–performance on ParseBench versus several specialized parsers and hyperscaler document services.

7.0

RESEARCH · 1 SOURCE · arXiv cs.AI

Continual Search: iterative approach improves root-cause attribution for long-horizon agent failures

New arXiv paper (arXiv:2609.13463v1) frames root-cause attribution (RCA) for long-horizon agent failures as a large search problem and introduces Continual Search, an iterative framework that prompts LLM-based judges to repeatedly search for unresolved diagnostic evidence. The authors also release MegaRCA-Mix, a 50-trial benchmark of long-horizon, execution-heavy failures, and report that Continual Search boosts attribution performance across benchmarks—for example improving GPT-5.5's F1 from 0.349 to 0.498 (over 40%).

7.0