acceptodds
Under review as a conference paper at ICLR 2027

Polyhedral Reinforcement Learning: Deep RL with Geometric State Representations

Abstract

In many reinforcement learning (RL) problems, the state already has geometric structure: feasible regions, guards, reachable sets, or collections of objects are represented explicitly by linear constraints. Standard deep RL often discards this structure by flattening the state into a fixed-length vector, forcing the network to recover relationships already explicit in the problem description. We introduce polyhedral reinforcement learning, in which states are finite unions of polytopes and actions induce transitions between them. Representing such states poses a distinctive learning problem: a polytope is a variable-sized, unordered object with several mathematically equivalent but computationally different descriptions. It may be specified by its vertices, by the half-spaces that bound it, by the incidence relations between vertices and facets, or, in structured domains, by a canonical matrix or spatial rendering. We therefore develop representation-aware actor–critic architectures matched to these descriptions: DeepSets over vertices and half-spaces, graph neural networks over relational structure, and convolutional encoders over canonical matrices and rasterised projections. Timed automata provide a controlled comparison in which the same system can be presented through real-valued clock valuations or through its symbolic polyhedral zone. Under the same training budget, point-based agents perform poorly when decisions depend on relations between clocks, while several polyhedral encoders achieve near-perfect success; this advantage largely disappears when such relations are irrelevant. Across five additional domains with intrinsically polyhedral states, no representation dominates universally: performance is strongest when the encoding directly exposes the geometric relationships relevant to the decision. These results show that geometric state representation can provide a powerful inductive bias for RL, and that its benefit depends on the task's relational structure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.