Intent-aware Off-policy Evaluation For Shared-Realizer Dialogue Controllers
Abstract
Asking for clarification can rescue an ambiguous request, but each extra question costs the user a turn. We study how to compare dialogue controllers that make this choice using logged conversations, without treating a user simulator as ground truth. The controllers share an utterance realizer, and their decisions can use a belief over user intent. Our estimator, IM-DR, weights a compact state built from that belief and conversation progress, then corrects a user model’s continuation value with logged rewards. We show how errors in the user model, density ratio, intent posterior and state abstraction enter the resulting bias. For the unnormalised estimator, we also establish a population double robustness identity. An oracle variance comparison holds with fixed independently fitted nuisance functions and bounded or clipped weights. On a synthetic dialogue testbed with Monte Carlo value references from 40,000 rollouts and 50 independent seeds, IM-DR has 0.23× the MSE of per decision importance sampling and 0.50× that of sequential DR using the same user model. Paired comparisons clustered by seed favour it over all four baselines. Its MSE advantage persists under a separate synthetic policy shift, but coverage of nominal 95% intervals falls to 0.760. Whether these results carry over to human dialogue requires future research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.