Under review as a conference paper at ICLR 2027
Model-Free Interpretable Trust Region Policy Optimization Algorithm
Abstract
To make machine learning widely applicable in today's society, we need to create models that are understood by laypeople. Creation of techniques that provide decision support would be highly beneficial in areas such as healthcare, finance, and robotics. Towards that goal, we design a model-free, interpretable version of the Trust Region Policy Optimization method. The method works by creating policies that have the structure of a differentiable decision tree. We lay out the construction of the method and list the subproblems that have to be solved at each leaf and internal node. We also provide convergence bounds to get this algorithm to work effectively.
open until 14 Dec 2026
est. 32% chance this paper gets accepted at ICLR 2027.
Reject 68%Accept 32%
What do you think this paper will get?
All positions stay anonymous.
Related papers
Loading the map…
Discussion (0)
Sign in to comment.