MosaicRec: Diagnosing Specification Grounding and Set Construction in Recommendation Agents
Abstract
Modern recommendation agents can draw on user histories, item catalogs, and external tools to satisfy complex user requests. Their task is not merely to rank plausible items, but to construct recommendation sets that jointly satisfy constraints grounded in distributed evidence. End-to-end success alone cannot reveal whether an agent correctly grounded these constraints, successfully constructed a compliant set, or recovered from an earlier error. We introduce MosaicRec, a benchmark that separates the post-retrieval recommendation pipeline into specification grounding and set construction, enabling separate evaluation of each stage and diagnosis of their interaction. MosaicRec contains held-out tasks from Goodreads and an independently constructed Amazon split, each paired with a solver-certified feasible specification. Its constraints include personalized predicates that require joining user histories with catalog metadata. The evaluation crosses predicted and gold specifications with language-model and exact constructors, checks specification safety through counterexample search, and compares model-generated recommendations with exact execution of the same cached specifications. This design reveals stage-specific weaknesses and cross-stage recovery. Under identical common gold specifications, Gemma 4 31B substantially outperforms a similar-sized Qwen3-32B model, revealing a gap in constrained set construction. Separately evaluating intermediate specifications and final recommendations shows that GPT-5.6-sol can produce valid recommendations despite errors in its own specification by retaining access to the original request and tools; replacing this stage with exact execution of its predicted specifications can therefore reduce end-to-end success. These findings demonstrate how MosaicRec makes individual stages diagnosable while exposing how their interaction shapes final outcomes. Code and data are available at an anonymous repository: https://anonymous.4open.science/r/MosaicRec_Anonymous.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.