PFR-Bench: Benchmarking Personalization in Federated Recommender Systems
Abstract
Recent studies on personalized federated recommendation have explored diverse personalization designs, but inconsistent implementations and evaluation protocols make existing results difficult to compare. More importantly, accuracy-centric evaluation provides limited insight into how these designs behave across users and under federated conditions. We therefore introduce PFR-Bench, a unified benchmark that standardizes six representative methods across ten datasets under a common experimental protocol. Beyond aggregate recommendation accuracy, PFR-Bench examines how personalization varies across heterogeneous users, depends on local interaction evidence, and behaves when client information becomes unreliable. Our evaluation reveals substantial differences that are obscured by overall performance alone. In particular, methods that use limited local evidence effectively are not necessarily those that remain stable when that evidence becomes unreliable, suggesting that local evidence should be modeled jointly through its quantity and reliability. We further find that stronger use of cross-client knowledge can compensate for insufficient local evidence while also increasing exposure to manipulated information from other clients, highlighting the need to balance collaborative benefit against propagation risk. PFR-Bench provides a reproducible basis for further evaluation and development of personalization mechanisms in federated recommendation. Code and benchmark resources are available at https://anonymous.4open.science/r/PFR-BENCH-AD63/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.