acceptodds
Under review as a conference paper at ICLR 2027

Learning-Guided Formal Abstractions for Policy Synthesis in Stochastic Systems

Abstract

We study nonlinear stochastic control problems with reach-avoid objectives, modeled as Markov decision processes (MDPs) over continuous state-action spaces. Our goal is to compute a policy with a certified lower bound on its probability of satisfying a given reach-avoid objective. Formal abstraction methods provide such guarantees but often scale poorly, as discretizing the state-action space leads to prohibitively large abstract models. By contrast, reinforcement learning (RL) scales better but provides no such guarantees. We introduce learning-guided abstractions, which combine the learning abilities of RL with the guarantees of formal abstractions. Specifically, we use an RL policy to identify a small region of promising state-action pairs and construct a sound abstraction only over this region. The resulting abstraction is represented as a set-valued MDP, on which we synthesize an optimal policy that can be refined to the original continuous MDP with a certified lower bound on its reach-avoid probability. Crucially, RL guides where abstraction effort is spent, while the formal abstraction guarantees the correctness of the synthesized policy. Across our benchmarks, learning-guided abstractions reduce the number of abstract state-action pairs by up to three orders of magnitude while retaining tight probabilistic guarantees, enabling policy synthesis at resolutions and problem sizes for which standard exhaustive abstractions are prohibitive.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.