Derivative Cheating in Fusion-Based PINNs: Understanding and Preventing
Abstract
Physics-Informed Neural Networks (PINNs) increasingly incorporate inter-point fusion mechanisms such as batch normalization, PointNet, and attention. In this paper, we identify a concealed but dangerous failure mode that arises when PDE residuals are derived via Automatic Differentiation (AD) in models that adopt such fusion mechanisms, in which AD may include fusion-induced paths unrelated to spatial derivatives, making the residual loss deceptively small while the absolute error remains large. We formalize this discrepancy, analyze the semantic distinction between AD-computed residual and true spatial derivatives, and introduce a diagnostic metric, Derivative Cheating Intensity (DCI), to quantify this discrepancy. We further propose a gradient-decoupled residual correction strategy that treats the fusion input as auxiliary context when computing spatial derivatives, rather than including it in the computational graph of AD. Experiments on Poisson, Helmholtz, and Heat benchmarks across BatchNorm, neural operator, PointNet, and Transformer backbones show that the proposed correction eliminates the failure and recovers accurate solutions. Our code is available at [THIS LINK](https://anonymous.4open.science/r/Derivative_Cheating-8CFE/readme.md).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.