acceptodds
Under review as a conference paper at ICLR 2027

TAXIGPT: DECOUPLING MOBILITY PRIORS AND ORDER-DERIVED FEEDBACK FOR PASSENGER SEEKING

Abstract

Passenger seeking is central to taxi and autonomous ride-service operations as a vacant vehicle need to decide where to move before its next request is assigned. An effective policy must generate realistic routes and identify movements that lead to passengers. Learning both from a single reward signal is difficult be- cause explicitly modeling both movement and pickup/service dynamics is costly, and coarse approximations can admit unrealistic reward-seeking behavior. We propose TAXIGPT, a knowledge-decoupled framework that separates mobility learning, order-environment reconstruction, and seeking-policy adaptation. A decoder-only Transformer learns a mobility prior from serving trajectories through next-location prediction. An order-derived environment reconstructs passenger availability and joint service outcomes from historical orders, returning rewards and state transitions for variable-length route proposals. Group-relative policy optimization adapts the prior across successive seeking and serving cycles, while locality constraints and reference regularization help preserve movement structure. In Shenzhen simulations, TAXIGPT outperforms five passenger-seeking baselines, increasing cumulative episode profit by 160.7% and reducing average seeking distance by 63.3% relative to the respective best baselines. Applying locality during adaptation further improves profit by 5.6% over inference-only filtering under the same deployment constraint. These results support learning movement priors from real trajectories and seeking preferences from order-derived interaction, enabling effective passenger seeking with less reliance on explicitly modeled urban movement dynamics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.