acceptodds
Under review as a conference paper at ICLR 2027

VERA: Verifiable Feasibility Representations with Counterfactual Credit for Constrained Multi-Agent Control

Abstract

Constrained multi-agent control requires more than predicting rewarding actions: an action can cease to be executable as contact windows, shared capacity, and deadlines change. We introduce VERA, a centralized-training, decentralized-execution framework that separates feasibility estimation from credit assignment. Each actor predicts a five-dimensional verifiable feasibility representation (VFR). After an action is proposed, exact action-conditioned margins available only during training supervise that representation, while a counterfactual group-relative advantage (CGRA) ranks candidate representation–action pairs. Execution uses one actor pass and no privileged state. In a dynamic space–air–ground integrated network (SAGIN), VERA obtains success with coverage violation, within points of a privileged-mask reference. With rewards matched over ten paired seeds, VERA improves success over the strongest baseline by points () and reduces violation by points (). A ten-seed factorial attributes a – point gain to CGRA across handcrafted, learned, random, and latent representations; evaluation on seven unseen topologies preserves a – point advantage over multi-agent proximal policy optimization. From 10 to 40 users, success remains –, and VFR adds only ms to a central processing unit (CPU) actor step. Cross-domain tests further identify the governing condition: counterfactual credit succeeds when candidate scores respect shared constraints and fails under incompatible reward geometries. These results establish action-conditioned feasibility as an auditable training interface and counterfactual credit as a geometry-dependent optimization mechanism.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.