Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1469

LEGO-Anything agents generate editable 3D scenes from single photos but misjudge correctness

Researchers at the University of Maryland and AWS introduced LEGO-Anything, an Image-to-Code approach where coding agents iteratively write Blender code to reconstruct an editable 3D scene from a single photo. They also released LEGO-Bench (208 images from 104 simulated scenes, 443 assets) to evaluate validity, reconstruction, and appearance, finding GPT-6 Astra achieved the best scores (≈53.4% indoor, 39.6% outdoor) but models struggle with self-assessment and geometric accuracy; a no-training extension, LEGO-Plugin, fixes some failure modes and substantially improves weaker agents.

KEY POINTS

  1. Researchers at the University of Maryland and AWS introduced LEGO-Anything, an Image-to-Code approach where coding agents iteratively write Blender code to reconstruct an editable 3D scene from a single photo.
  2. They also released LEGO-Bench (208 images from 104 simulated scenes, 443 assets) to evaluate validity, reconstruction, and appearance, finding GPT-6 Astra achieved the best scores (≈53.4% indoor, 39.6% outdoor) but models struggle with self-assessment and geometric accuracy; a no-training extension, LEGO-Plugin, fixes some failure modes and substantially improves weaker agents.
  3. This shows coding agents can produce inspectable, editable 3D scene programs from single images but that unreliable self-assessment and geometric accuracy remain major barriers, pointing to the need for measurement-based refinement and tool integrations.

WHY IT MATTERS

This shows coding agents can produce inspectable, editable 3D scene programs from single images but that unreliable self-assessment and geometric accuracy remain major barriers, pointing to the need for measurement-based refinement and tool integrations.

SOURCES & TIMELINE

1