RESEARCH · RESEARCH · #705
PolyBridgeBench: benchmarking MLLMs for physics-grounded bridge design
PolyBridgeBench is an executable benchmark that asks multimodal LLMs to produce complete node–member–material bridge topologies from a visual scene and engineering constraints, then runs deterministic legality checks and a native dynamic physics simulation. The benchmark reports deterministic validity, dynamic functional success, and post-failure recovery under a fixed interaction budget; experiments on six representative MLLMs across 189 levels reveal a large gap between deterministic validity and dynamic success, strong sensitivity to material budgets, and limited ability to repair after failures.
KEY POINTS
- PolyBridgeBench is an executable benchmark that asks multimodal LLMs to produce complete node–member–material bridge topologies from a visual scene and engineering constraints, then runs deterministic legality checks and a native dynamic physics simulation.
- The benchmark reports deterministic validity, dynamic functional success, and post-failure recovery under a fixed interaction budget; experiments on six representative MLLMs across 189 levels reveal a large gap between deterministic validity and dynamic success, strong sensitivity to material budgets, and limited ability to repair after failures.
- The benchmark tests whether MLLMs can produce executable, load-bearing engineering designs and adapt after simulator-exposed failures, exposing a practical gap between symbolic validity and real-world functional success.
WHY IT MATTERS
The benchmark tests whether MLLMs can produce executable, load-bearing engineering designs and adapt after simulator-exposed failures, exposing a practical gap between symbolic validity and real-world functional success.