Tuning Multimodal Large Language Model for Microstructure Generation
Abstract
Although 3D material microstructure generation has achieved notable progress, existing works typically rely on domain-specific constraints, leaving research oriented to natural languages largely unexplored. Inspired by the recent success of Large Language Models (LLMs) across diverse domains, this work introduces a multi-conditional generation framework that supports text-, sketch-, and property-driven microstructure design. The proposed method employs a Vector Quantized Variational AutoEncoder (VQ-VAE) to encode microstructural representations and a visual encoder to process sketches, allowing multi-modal alignment between text, sketch and structure within an LLM-based architecture. Moreover, to ensure that the generated microstructures satisfy desired physical properties, property constraints are incorporated into the optimization process. Experimental results demonstrate that the proposed framework can produce high-quality and semantically consistent microstructures, suggesting its potential to support more intuitive and accessible microstructure design workflows.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.