AFSIM-Agent: Hierarchical Self-Reflective Coding Agent for Cross-Language Adaptation to Unseen Scripting Languages
Abstract
Code agents solve many tasks, but their competence rests on pretraining that supplies the language's syntax and abundant examples. Domain-specific languages (DSLs) break both: models lack the domain's grammar, cannot interpret compiler diagnostics, and have few examples to learn from. We study AFSIM, an open-source multi-domain modeling and simulation framework from the US Air Force Research Laboratory (AFRL), whose scripting language is rare in training data. General coding agents fail often when writing it; we aim for an agent that masters an unfamiliar DSL without retraining. We propose AFSIM-Agent, with five modules. (1) A curated domain knowledge surface turns the DSL into retrievable priors: a read-only document tree of grammar, documentation, and worked examples, plus skill cards linking engineering tasks to source files. (2) An executable feedback ladder supplies four levels of signal—static constraints, minimal-harness compilation, full simulation, and independent fidelity audit—constructing the trace-level signals general agents never observe. (3) A verification-gated repair loop admits only edits verified at the current file revision, so each round builds on verified state. (4) A locally in-distribution (LID) context discipline keeps every LLM call in-distribution and the main context limited to decomposition strategy, control flow, and normalized state, so long-horizon tasks do not dilute critical signals. (5) A deterministic plan-audit gate and anti-reward-hacking guardrails constrain the plan before code is written and the result against the specification at submission. We also build AFSIM-Bench: 114 tasks with compilable reference code and a reproducible measurement instrument for this domain. On the 34 executable tasks, the bare model passes 17.6%, the three general harnesses reach 41.2%–47.1%, and AFSIM-Agent reaches 58.8%; the harness buys the ability to make a project run, and domain grounding buys the ability to make it satisfy intent.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.