Structural Gaming in Financial Safety Alignment: A Dynamic Strategic Evasion Framework via Tabular Representations
Abstract
As large language models (LLMs) are increasingly integrated into financial systems to enforce regulatory guardrails, they fundamentally alter the safety alignment landscape into a co-evolutionary game between platforms and strategic agents. While conventional safety mechanisms are optimized to detect overt semantic risks, they remain largely oblivious to the structural topology of information. This paper formalizes this vulnerability by introducing Structural Gaming into financial safety alignment. We propose Structural Risk Obfuscation Attack (SO-STAG), a red-teaming framework demonstrating how strategic agents can achieve structural compliance arbitrage. SO-STAG systematically camouflages regulatory violations through deep hierarchical layouts and adversarial tabular representations without falsifying the underlying factual content. To operationalize this, we develop a judge-guided two-phase paradigm incorporating a Hidden Chain-of-Table-Thought (HCoTT) mechanism. By exploiting the inherent information asymmetry, SO-STAG forces a pooling equilibrium that traps the defender in an incomplete information state—effectively creating a structural "cognitive fog" where the alignment guardrails can only observe surface tabular representations while remaining blind to the latent malicious reasoning paths. We model this interaction as a dynamic strategic evasion process where the agent iteratively refines its structural representation based on the defender's feedback. Empirical evaluations across six leading LLMs show that SO-STAG achieves a 97.78% average evasion success rate, exposing a pervasive systemic blindness to structurally concealed risks. Crucially, from an algorithmic efficiency perspective, SO-STAG minimizes the agent's computational cost, utilizing over 80% fewer input tokens compared to state-of-the-art semantic multi-turn baselines. Our work highlights a foundational flaw in current alignment paradigms: by failing to account for structural incentives, existing safeguards invite low-cost strategic manipulation, necessitating a shift toward structure-aware mechanism design in financial AI safety.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.