NEWS · MODELS · #1070
OpenAI's GPT-6 Astra scores 80% on Epoch AI's furniture assembly benchmark
Epoch AI's Furniture Assembly Benchmark (FAB) tests models on identifying deliberate assembly errors in photos of IKEA furniture. Previously the best model (Claude Opus 4.5) scored 28% in November 2025; ten months later OpenAI's GPT-6 Astra reaches 80% with a processing time of about three minutes per photo, followed by Claude Fable 5.1 at 70% and Claude Opus 5 at 61%, while some Chinese open-weight models like Kimi K3 lag substantially.
KEY POINTS
- Epoch AI's Furniture Assembly Benchmark (FAB) tests models on identifying deliberate assembly errors in photos of IKEA furniture.
- Previously the best model (Claude Opus 4.5) scored 28% in November 2025; ten months later OpenAI's GPT-6 Astra reaches 80% with a processing time of about three minutes per photo, followed by Claude Fable 5.1 at 70% and Claude Opus 5 at 61%, while some Chinese open-weight models like Kimi K3 lag substantially.
- The result marks a large jump in vision–language capabilities with potential applications in repair, robotics and real-world visual troubleshooting, though current latency limits real-time use.
WHY IT MATTERS
The result marks a large jump in vision–language capabilities with potential applications in repair, robotics and real-world visual troubleshooting, though current latency limits real-time use.