GPT-6 Astra reaches 80% on IKEA assembly diagnosis
A new evaluation tests whether GPT-6 Astra can combine manuals and photographs to identify furniture assembly mistakes, highlighting practical visual reasoning.
From answering questions to finding the missing dowel
GPT-6 Astra can identify furniture assembly mistakes from an instruction manual and photographs with an accuracy rate reported at about 80%, according to a newly published evaluation discussed by The Decoder and Chinese technology media. The test asks the model to inspect a partly assembled IKEA-style product, compare the visible state with written instructions and point out where the construction went wrong.
The task is deliberately mundane. It requires more than recognizing objects: the system must align diagrams with a three-dimensional scene, track assembly order and distinguish a plausible-looking mistake from a harmless variation. In practice, that means reasoning over images, text and procedural state at the same time.
The reported result is notable because the comparison is framed against an earlier generation that achieved roughly 40% on the same type of task. That gap suggests progress in multimodal reasoning may be visible in household workflows before it appears in conventional academic benchmarks. A model that can diagnose a misplaced panel or connector could eventually support repair, maintenance, inspection and workplace training.
The evidence remains narrow. The evaluation concerns a constrained class of objects and instructions, and the public reports do not establish how well Astra performs with unfamiliar furniture, poor lighting, missing pages or genuinely ambiguous photographs. An 80% score also leaves a substantial failure rate when the model is asked to give physical guidance.
The broader importance is that visual intelligence is being measured through action-relevant diagnosis rather than image description alone. If these results generalize, multimodal models are moving closer to useful assistants for physical tasks. If they do not, the number may remain an impressive but highly specialized benchmark result.