acceptodds
Under review as a conference paper at ICLR 2027

Neither Sycophantic nor Stubborn: Preserving Feedback Discernment in Merged Reasoning Models

Abstract

Model merging offers a practical approach to reducing reasoning length while preserving answer accuracy. However, accuracy and output length alone do not capture how merged models respond to external feedback. Across diverse mathematical reasoning benchmarks, we find that some evaluated merging configurations improve accuracy and reduce output length yet degrade feedback discernment, the ability to retain correct answers under misleading advice and correct erroneous answers under valid feedback. This degradation is not inevitable, as other configurations preserve this capability. Motivated by these findings, we formulate discernment-aware merging as a constrained optimization problem that maximizes feedback discernment subject to accuracy and token-cost constraints. We propose **Dis**cernment-guided **C**onditional Fl**o**w Merge (DISCO-Merge), an active optimization framework that couples conditional flow matching with Bayesian optimization. Conditioned on a parent configuration and its behavioral profile, a residual flow learns a conditional distribution of updates to module-wise merging coefficients from evaluated neighboring model pairs. Gaussian-process surrogates estimate candidate performance and uncertainty, and an acquisition function selects candidates for expensive model evaluations. We select the model with the highest discernment among confirmed candidates that meet the token-cost and accuracy constraints. Experiments on eight mathematical benchmarks with Qwen2.5, Qwen3, and Phi-4-mini show that DISCO-Merge surpasses all evaluated merging baselines in macro-averaged discernment and accuracy–efficiency–discernment hypervolume (HV).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.