TAIRL: Optimizing Multi-Turn Open-Ended Dialogue via Temporal Patterns in Human Feedback
Abstract
Multi-turn open-ended dialogue (MOD) is a core scenario in human-machine interaction, where subjective factors such as emotion, interest, knowledge, and topic shift play a decisive role in dialogue quality. However, many interactive reinforcement learning (RL) methods use human feedback as turn-level scalar supervision, leaving its temporal structure underexplored. Motivated by an empirical analysis of human feedback in financial-domain MOD, we propose TAIRL, a temporal-aware interactive RL framework that models three recurring feedback patterns: turn-point concentration, asymmetric timing, and adaptive frequency. TAIRL operationalizes these patterns through TW, ALR, and AER. In an offline-feedback-driven simulator trained from 4,287 real human feedback samples, TAIRL reduces simulator convergence turns by 36–45% and improves sample efficiency by approximately 50% over several RL baselines. On held-out human-evaluated financial dialogues with Qwen2.5-7B, TAIRL-enhanced methods also improve final task success and user-experience metrics, including a 22.8% relative improvement in user satisfaction and a 56% relative improvement in emotional alignment. The proposed mechanisms are lightweight and can be integrated into PPO, GRPO, REINFORCE, and TAMER-style interactive learning. These results suggest that temporal feedback structure is a useful design signal for improving interactive RL in MOD, while broader cross-domain validation remains an important direction for future work. The code and data are available at https://anonymous.4open.science/r/TAIRL-2957.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.