Steady, as she goes: Rethinking Observations and Rewards for Stability Regulation in Cooperative Multi-Agent Systems
Abstract
Continuing multi-agent systems must maintain stable operation and recover safely and promptly from disturbances. Previous works mainly formulate it based on instantaneous observations and tracking-error-type rewards, which may not be appropriate. This mismatch arises from the information available to the learner and the recovery behaviors favored by the reward. Without explicit disturbance context, value learning can conflate external effects with those of agents' actions. Missing temporal information can hinder action selection when the same instantaneous deviation occurs during recovery or renewed departure. Tracking-error style deviation penalties can also favor slower recovery over a safe, faster return that involves temporary overshoot. We study observation and reward grounding for continuing multi-agent regulation to align learning with stable-operation and recovery requirements. We propose an observation design that combines disturbance context available before action selection with windowed deviation and its change, providing information about external conditions and recent evolution. We also propose a recovery-oriented reward that uses time spent outside the normal operating region as the base cost and a potential difference as auxiliary recovery feedback, with safety and continued task performance as requirements. We organize the evaluation around MAPPO in three environments: cooperative payload transport, power-grid voltage control, and traffic signal control, with comparisons that isolate the roles of observation and reward design. The evaluation examines recovery success and time alongside normal operation, failures, and task performance to assess whether improved recovery preserves safe and effective operation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.