STEER: Steerable Recommendation System with User Memory and Think-Then-Recommend
Abstract
Users are no longer resigned to passive consumption of algorithmic recommendations; they increasingly demand active agency over their digital feeds. Recent product launches—Instagram's "Tune Your Algorithm", Spotify's "Taste Profile", Threads' "Dear Algorithm" —all point to a fundamental gap: mainstream recommendation systems lack a principled mechanism for users to steer recommendations through open, natural-language prompts. We identify a steerable recommendation trilemma: existing paradigms are optimized for at most two of steerability, stateful personalization, and scalability—while hybrid approaches can partially address all three, no prior framework provides a principled, unified solution —collaborative filtering is personalized and scalable but not steerable; text retrieval is steerable and scalable but stateless with limited personalization; LLM-as-recommender is steerable and personalized but not scalable. We present STEER, a framework that resolves this trilemma through two core innovations: (1) STEER User Memory, an LLM-powered user profile that makes the recommendation system's understanding of each user transparent, interpretable, and editable—serving as a shared representation between user and system; and (2) Think-Then-Recommend (TTR), a two-stage architecture that explicitly reasons over the user's natural-language steering prompt and memory, then executes a structured recommendation plan via scalable embedding-based retrieval. We show empirically that personalization (from user memory) and steering (from prompts) provide strong complementary signals: their combination exceeds their individual contributions with large margins. We evaluate STEER at multiple scales: a new human-annotated benchmark (Open-Steer, 2,000+ participants across four domains), large-scale public benchmarks, and online deployment to hundreds of millions of users. Against methods consuming the same steering signal, STEER demonstrates significant gains (e.g., +37.3 points absolute HR@10 from 18.9% to 56.1%, over prompt-only retrieval on our largest dataset), confirming that explicit reasoning over transparent user memory provides value beyond simple query-item matching. Online A/B tests show statistically significant improvements in time spent (+0.40%) and app visits (+0.21%), validating that the framework scales to latency-sensitive production environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.