Vibeleaderboard
← All Intel
Intel / post

Can AI tell if you've built your IKEA furniture wrong?

Source
Epoch AI
Date
Epoch AI@EpochAIResearch
Thread · 6 parts

Can AI tell if you've built your IKEA furniture wrong? Our new benchmark, the Furniture Assembly Benchmark (FAB), gives models the manual and a photo of a half-completed piece of furniture and asks them to spot the mistake. The top score has gone from 28% to 80% in just 10 months.

On FAB, AI must assess builds based on photos like the one shown below. The benchmark consists of 60 photos from 3 different furniture builds.

There are two ways to get a question wrong: miss a real mistake, or flag a build that's fine. Many recent frontier models favor one failure mode over another. Gemini 3.1 Pro, for example, passes almost no correct builds, whereas GPT-5.4 catches almost no mistakes.

For practical utility, speed matters as well as accuracy. Here GPT-6 Astra is a clear outlier. Not only is it the most accurate, but, of all the models we tested, it is also the fastest.

Read the full thread on X
Key takeaways · AI-distilled
  • Epoch AI's Furniture Assembly gives a model the assembly manual plus a photo of a half-completed piece and asks it to spot the mistake; it is small, 60 photos drawn from 3 furniture builds.
  • Epoch counts two ways to fail, missing a real mistake or flagging a build that is fine, and says Gemini 3.1 Pro passes almost no correct builds while GPT-5.4 catches almost no mistakes.
  • Epoch argues speed matters alongside accuracy for practical use, and reports GPT-6 Astra was the fastest of all models it tested as well as the most accurate.
  • Epoch pitches FAB as a proxy for economically important visual-reasoning work such as fixing a car or repairing household appliances.
Terms in this piece · Glossary
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters

Epoch AI's Furniture Assembly Benchmark tracks how well vision models catch real build errors from photos: top accuracy went from 28% to 80% in 10 months, but models split into missing real mistakes or wrongly flagging correct builds.

More from Epoch AI
Recommended reads
Comments

Checking sign-in…

Loading comments…