Research

LEGO-Anything AI builds 3D scenes from single photos

University of Maryland and AWS researchers have built LEGO-Anything, an AI system that writes Blender code to reconstruct 3D scenes from photos but struggles to judge its own accuracy.

The Decoder3 days agoResearch
Image: The Decoder

The LEGO-Anything project uses an iterative "Image-to-Code" method where a coding agent writes and refines Blender code to generate editable 3D scenes from a single image. Because the output is an executable program, developers can inspect, query, and modify the resulting geometry directly. To evaluate this approach, the research team created LEGO-Bench, a benchmark containing 208 images from 104 indoor and outdoor simulator scenes using 443 registered assets.

Testing across six GPT configurations revealed that while models successfully output usable code, their geometric accuracy varies. GPT-6 Astra led the pack, scoring 53.4 percent on indoor scenes and 39.6 percent on outdoor scenes, while weaker setups scored around 15 percent. When given a larger reasoning budget, Astra's score on an office subset rose from 32.3 to 61.8 percent. However, the agents frequently ruined their own progress during revisions, sometimes dropping Astra's score from 33.9 to 4.4 percent, because their geometric self-assessment performed near or below chance level.

To fix this blind spot, the team developed LEGO-Plugin, which replaces the agent's subjective judgment with concrete measurements. This tool boosted weaker models by up to 62.7 percent, though Astra only gained about two percentage points. When the generated scenes were tested on standard vision tasks, they achieved usable but mediocre results. Object detection reached about half the performance of the specialized DINO model, while segmentation and depth estimation lagged further behind specialized tools like SAM 3 and Depth Anything 3.

For 3D developers and vision practitioners, this research highlights both the promise and the current limits of code-generating agents. While platforms like Unity are already integrating agents via plugins for Claude Code and Codex, and competitors like World Labs' Atlas and Google Deepmind's GenCeption pursue alternative reconstruction methods, LEGO-Anything proves that reliable 3D coding agents will require external verification tools rather than relying on their own flawed self-evaluation.

This is our own summary of reporting by The Decoder

More in Research