Recursive Transformer with State Injection and Per-Slice Supervision for Thermo-Mechanical Surrogate Modeling
Abstract
Transformer-based surrogate models are increasingly replacing expensive first-principles simulation in engineering design, but conventional architectures are often over-parameterized for the small, low-dimensional datasets typical of this setting, where large simulation data is expensive to generate. Under these conditions, excess parameter capacity leads to overfitting and unnecessary memory/compute overhead — this motivates architectures that trade additional compute for additional parameters, an idea explored by weight-shared recursive architectures such as the Universal Transformer (UT). We adapt this recursive paradigm to surrogate thermo-mechanical analysis of semiconductor advanced packages, where each sample consists of a fixed set of Finite Element Analysis (FEA) material and design parameters paired with a sequence of stress contours across the package's layers, and simulation data is far too scarce to support a large-parameter model. However, UT's recursion is "unguided": every recursion step performs the same undifferentiated refinement, with no inherent meaning attached to being at loop 3 versus loop 7— and although each frame in the sequence shares the same FEA parameters, UT treats each frame as an independent sample rather than part of a continuous progression through the sequence. We introduce DEPTH, an RNN-like recursive transformer, that instead gives each recursion step genuine meaning: its single transformer block is applied recursively as a sequence, consuming the static design parameters once and, through recursion produces a matching stress contours at each z-axis (which is the sequence), using exactly one forward pass per layer. An ablation study shows DEPTH's two key ingredients — state injection and per-step supervision are synergistic, not additive— removing either collapses retrieval to near-chance and only their combination recovers strong performance. Since UT must loop per-frame to be competitive, we sweep its internal recursion from 1x to 9x (till it collapse) to quantify the compute overhead needed to match DEPTH: it requires several times DEPTH's per-frame cost to reach parity, then collapses sharply beyond that, while DEPTH matches or exceeds this peak at a fraction of the compute. Gating experiments show carry-gating modestly helps while injection-gating hurts. We benchmark DEPTH against a Tiny Recursive Model and UT (which instead loops internally within each frame) on predictive performance, parameter count, and FLOPs on actual Ansys simulated Stress and Warpage datasets. We validate generalization on a third task, a Laplace PDE capacitance-field solver, confirming recursive weight-sharing transformers as an effective, parameter-efficient choice for small-data engineering surrogate modeling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.