Agentic Simulator Construction with Falsification-Guided Structural Diagnosis
Abstract
Large language models can turn natural-language briefs into runnable simulation code, yet automated simulator refinement remains vulnerable to unreliable struc- tural diagnosis: the same macro-level mismatch may admit multiple plausible explanations, while an LLM-generated diagnosis can be acted upon before it is tested. We present FASim, an agentic framework for automated simulator con- struction with falsification-guided structural diagnosis. FASim constructs and calibrates an executable simulator, routes directly verifiable implementation is- sues to repair, and treats structurally ambiguous failures as competing falsifiable hypotheses. A Hypothesis Agent proposes alternative explanations, while a sepa- rate Falsification Agent derives their observable counterfactual signatures, freezes the test plan before execution, and challenges them through controlled simulator probes. Directly observable signatures are evaluated from execution traces, while stochastic or borderline responses use repeated matched-seed executions. Only sufficiently discriminated diagnoses are allowed to guide structural repair; falsified or inconclusive explanations are blocked from triggering persistent edits. Across user modeling, intervention-driven mask adoption, and mobility under ID and OOD shifts, FASim achieves strong multi-metric simulator fidelity. Ablation and process analyses further show that pre-repair falsification improves refinement quality, pro- duces actionable structural diagnoses, and prevents unsupported explanations from being converted into edits. These results support falsify before repair as a practical principle for more reliable and auditable agentic simulator construction. Code and data: https://anonymous.4open.science/status/FASim-8D87.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.