acceptodds
Under review as a conference paper at ICLR 2027

Residual Fusion for Cold-Start Matrix Completion and Reward-Model-Guided Generation: A Case Study in Drug Response Prediction with Specification Gaming

Abstract

Content-conditioned sparse matrix completion — predicting an interaction between a query entity and a target entity from content alone, without an identity-indexed lookup table — is a recurring cold-start problem across recommendation, matching, and scientific prediction, since a lookup table is by construction undefined for an entity absent from training. We study this problem in drug response prediction, where both the compound and cell-line axes are extremely sparse (median one observed interaction per compound) and only the cell-line axis must generalize zero-shot to unseen entities. We present a Two-Tower Residual Late Fusion architecture in which a hard-memorization branch resolves fine-grained interaction structure while a gated residual branch, conditioned on cold-start content alone, supplies interaction corrections. Because the residual branch only adds to the memorization branch rather than replacing it, the architecture admits a structural non-degradation property that does not depend on the residual branch being useful at test time. Evaluated on 51 breast cancer cell lines (136,342 records), the architecture outperforms a hard-memorization-only baseline in 48/51 cell lines (94.1%) under leave-cell-line-out evaluation (mean ΔR² = +0.016). It also generalizes zero-shot to 601 externally held-out cell lines across 27 tissue types with zero cell-line overlap with training (median R² = 0.627). Prediction error correlates significantly with embedding-space distance to the nearest training entity (r = 0.159, p = 9.2 × 10⁻⁵), a significant signal that the content encoder has learned a distance-aware, not merely memorized, representation. We then ask a distinct question: is a validated predictor of this kind safe to repurpose, frozen, as a reward model for a downstream generative policy? A six-run reinforcement learning (RL) study (v1–v6) found no instance, under reward-model-blind checks, of the generator gaming the frozen predictor itself. It did, however, uncover a specification-gaming failure in an auxiliary diversity objective, which we diagnosed mechanistically and resolved with a calibrated, similarity-clustered penalty. A final sample-efficiency variant then achieved a reproducible reward improvement, replicated across two independent seeds, from an identical predictor-call budget without eroding this fix. Together, these results indicate that gated residual late fusion offers a practical route to cold-start generalization without sacrificing memorized accuracy, and that reward models built on such architectures can be repurposed for RL fine-tuning with failure modes that are mechanistically diagnosable and correctable rather than silently exploited. Code, trained checkpoints, and the RL evaluation pipeline will be released upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.