Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

NEWS · MODELS · #1070

OpenAI's GPT-6 Astra scores 80% on Epoch AI's furniture assembly benchmark

Epoch AI's Furniture Assembly Benchmark (FAB) tests models on identifying deliberate assembly errors in photos of IKEA furniture. Previously the best model (Claude Opus 4.5) scored 28% in November 2025; ten months later OpenAI's GPT-6 Astra reaches 80% with a processing time of about three minutes per photo, followed by Claude Fable 5.1 at 70% and Claude Opus 5 at 61%, while some Chinese open-weight models like Kimi K3 lag substantially.

KEY POINTS

  1. Epoch AI's Furniture Assembly Benchmark (FAB) tests models on identifying deliberate assembly errors in photos of IKEA furniture.
  2. Previously the best model (Claude Opus 4.5) scored 28% in November 2025; ten months later OpenAI's GPT-6 Astra reaches 80% with a processing time of about three minutes per photo, followed by Claude Fable 5.1 at 70% and Claude Opus 5 at 61%, while some Chinese open-weight models like Kimi K3 lag substantially.
  3. The result marks a large jump in vision–language capabilities with potential applications in repair, robotics and real-world visual troubleshooting, though current latency limits real-time use.

WHY IT MATTERS

The result marks a large jump in vision–language capabilities with potential applications in repair, robotics and real-world visual troubleshooting, though current latency limits real-time use.

SOURCES & TIMELINE

1