Winning Without Composing, Losing While Generalizing: Separating Better Scores from Better Choices in Recommendation
Abstract
Recommendation models may need to suggest items from unseen combinations of familiar attributes, such as brand and product category. We examine whether withholding interactions with selected attribute combinations, while retaining their constituent attributes elsewhere in training, provides evidence of compositional generalization. A score that credits only the withheld group can reward recommending more items from that group rather than choosing better items within it. We introduce a model-agnostic audit that keeps the number and positions of group recommendations fixed for each user while changing which items fill those positions. When both models rank every item, our audit splits the difference in their recall exactly into a part due to how many group items each recommends and a part due to which group items it chooses. In experiments with known user preferences, a preference-blind rule that recommends more items from the withheld combination can win the group score against models that rank its unseen items better. Real-data comparisons reveal both cases where higher performance is driven mainly by recommending more items from the withheld combination and cases where the advantage persists after matching the number and positions of recommendations from that combination. Consequently, our audit distinguishes whether a model recommends more items from the withheld combination or chooses better items within it providing a more informative basis for evaluating compositional generalization in recommendation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.