MatMicroBench: Towards Realistic Agentic Materials Microstructure Characterization
Abstract
Real-world microstructural characterization requires a continuous analysis process that observes relevant scientific evidence, models analysis rules as executable procedures, measures quantitative characteristics, and reasons from the resulting evidence and measurements to scientific conclusions. Existing benchmarks, however, largely focus on knowledge-centric or single-step tasks and provide limited coverage of such end-to-end analysis in realistic experimental settings. We introduce MatMicroBench, a benchmark toward realistic agentic materials microstructural characterization, where agents operate over experimentally acquired data and complete long-horizon scientific analysis with multiple interdependent outputs. Built from controlled laboratory experiments, MatMicroBench contains 64 instances based on SEM and EBSD data and spans four complementary tasks: static defect-state analysis, temporal defect-evolution analysis, inter-grain structure analysis, and intra-grain orientation-variation analysis. Each task requires agents to jointly produce spatial evidence, executable code, numerical results, and scientific conclusions, enabling capability-resolved evaluation of Observe, Model, Measure, and Reason. We evaluate 15 frontier models under a fixed execution harness and further compare six representative models across multiple agent harnesses. Under the fixed harness, even the strongest models achieve only around 53 points overall, and different models lead on different tasks, indicating substantial room for improving end-to-end microstructural analysis. Capability analysis further reveals task-specific bottlenecks, while strong code execution does not necessarily translate into accurate scientific outputs. These findings expose distinct failure modes across the characterization process and establish MatMicroBench as a testbed for developing more reliable agents for real-world materials analysis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.