Dynamic Incentive Design for Managing Strategic Manipulation under Continual Learning
Abstract
When algorithmic predictions affect users' outcomes, people may benefit from manipulating their inputs. How should an organization design incentives when these inputs are also used to train future models? We study a setting where a firm chooses costly incentives around a continually updated predictor, taking the training procedure as given. Incentives affect both current operating payoff and future learning. We characterize this trade-off and derive bounds on the dynamic value of departing from the myopic incentive. In a continual learning system, retained information also affects future training: organizations with the same predictor can share the same myopic optimum, yet different histories can make their optimal dynamic incentives lie on opposite sides of the myopic point. We also examine how incentive design interacts with strategy-aware learning. For small training updates, the two interventions are complementary when incentives increase the gain from the update, and substitutes when they reduce it. To examine the economic implications, we use a stylized model to connect our analysis to managerial agency problems and worker welfare. We then provide simulation evidence on the gains from dynamic incentive design. In our benchmark, an approximate history-aware oracle improves net discounted payoff by 6.53% over no intervention, compared with 1.38% for myopic control and 3.19% for current-information control. We further test whether policy learning can recover oracle gains using only observable information, and find that including retained-history information in fitted Q-iteration can achieve outcomes similar to the oracle benchmark.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.