Cost Of Interpretability - Blocked Value Iteration
Abstract
In fields such as healthcare, lawmaking, or industrial design, there is a need to develop machine learning algorithms that are not black-box functions and can be understood by experts in those fields. In such a scenario, interpretability becomes a key requirement for machine learning algorithms. In this paper, we define interpretable policies as being regionwise constant maps. Under this notion, we define the cost of interpretability as the average difference in the expected value of the optimal and the interpretable value functions. With this notion of interpretability, we compute upper bounds on the cost of interpretability. The bound is seen to be a linear function of the function summarization error for the final optimal policy. We run experiments to compute bounds on the cost of interpretability for different reinforcement learning environments under different hyperparameter configurations. Further, we analyze the behaviour of the function approximation method on the dimension and size of the state space for a given family of functions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.