acceptodds
Under review as a conference paper at ICLR 2027

LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures

Abstract

Large Language Model (LLM)-based agents are increasingly capable of generating complex 3D structures, with the potential to reshape how objects are designed and realized in the physical world. Yet, producing elegant geometry is fundamentally different from producing objects that can actually be built and perform their intended functions. Existing evaluations largely focus on geometric or visual quality while overlooking structural soundness, functional affordances, and physical realizability. We introduce LMBuild, a benchmark for evaluating LLM agents on generating buildable and functional structures. LMBuild represents generated objects as assembled structures comprising meaningful part decompositions, joints, materials, and assembly sequences. To support reproducible evaluation, we provide a unified framework consisting of: (1) an interactive environment in which agents can use tools to retrieve, create, and place components to construct an object; (2) a curated benchmark that repurposes established CAD datasets and augments them with real-world object knowledge from Wikipedia; and (3) a four-level evaluation framework covering structural soundness, functional affordance, design quality, and physical realization. Evaluations across 30 domain-specific models, open-source LLMs, and frontier APIs reveal several intriguing findings: (a) Structural soundness and visual alignment are no longer the primary bottlenecks for frontier closed-source models, while functional affordance and physical operability remain substantially more challenging; (b) stronger models more effectively exploit part creation to construct structure-specific components, whereas weaker models tend to rely on retrieval and can struggle with the additional flexibility introduced by creation; and (c) providing explicit functional specifications substantially improves functional-part completeness, kinematic correctness, and overall physical operability. These results show that generating real-world structures requires capabilities beyond geometric synthesis, including deeper reasoning about functional affordances, mechanics, and physical constraints, as well as the ability to design and create novel components when existing ones are insufficient. We expect LMBuild to provide a foundation for measuring progress and incentivizing research toward agents that generate buildable and functional structures, while helping narrow the gap between open-source communities and frontier labs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.