Residual Fusion for Cold-Start Matrix Completion and Reward-Model-Guided Generation: A Case Study in Drug Response Prediction with Specification Gaming
Abstract
Content-conditioned sparse matrix completion — predicting an interaction between a query entity and a target entity from content alone, without an identity-indexed lookup table — is a recurring cold-start problem across recommendation, matching, and scientific prediction, since a lookup table is by construction undefined for an entity absent from training. We study this problem in drug response prediction, where both the compound and cell-line axes are extremely sparse (median one observed interaction per compound) and only the cell-line axis must generalize zero-shot to unseen entities. We present a Two-Tower Residual Late Fusion architecture in which a hard-memorization branch resolves fine-grained interaction structure while a gated residual branch, conditioned on cold-start content alone, supplies interaction corrections. Because the residual branch only adds to the memorization branch rather than replacing it, the architecture admits a structural non-degradation property that does not depend on the residual branch being useful at test time. Evaluated on 51 breast cancer cell lines (136,342 records), the architecture outperforms a hard-memorization-only baseline in 48/51 cell lines (94.1%) under leave-cell-line-out evaluation (mean ΔR² = +0.016). It also generalizes zero-shot to 601 externally held-out cell lines across 27 tissue types with zero cell-line overlap with training (median R² = 0.627). Prediction error correlates significantly with embedding-space distance to the nearest training entity (r = 0.159, p = 9.2 × 10⁻⁵), a significant signal that the content encoder has learned a distance-aware, not merely memorized, representation. We then ask a distinct question: is a validated predictor of this kind safe to repurpose, frozen, as a reward model for a downstream generative policy? A six-run reinforcement learning (RL) study (v1–v6) found no instance, under reward-model-blind checks, of the generator gaming the frozen predictor itself. It did, however, uncover a specification-gaming failure in an auxiliary diversity objective, which we diagnosed mechanistically and resolved with a calibrated, similarity-clustered penalty. A final sample-efficiency variant then achieved a reproducible reward improvement, replicated across two independent seeds, from an identical predictor-call budget without eroding this fix. Together, these results indicate that gated residual late fusion offers a practical route to cold-start generalization without sacrificing memorized accuracy, and that reward models built on such architectures can be repurposed for RL fine-tuning with failure modes that are mechanistically diagnosable and correctable rather than silently exploited. Code, trained checkpoints, and the RL evaluation pipeline will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.