When Divergence Emerges, Look Back: Exploring Solution-Domain Topology in Collaborative Decoding
Abstract
Token-level small-large model collaborative decoding has a small language model (SLM) generate most reasoning steps and invokes a large language model (LLM) at local disagreements. Existing token-level correction methods mainly alter the current output, but cannot revoke a committed prefix and its internal state; state repair alone also does not establish whether the repaired state can produce more reliable subsequent reasoning. We therefore formulate small–large model collaboration as a post-injection recovery problem under a finite budget and propose the Solution-Domain Topology Tree. When a disagreement is detected, the system rolls back to an earlier reasoning anchor and uses learnable Cache-to-Cache (C2C) projection to construct a repaired state while retaining the unrepaired state as a paired reference. A frozen SLM then evaluates the future solution domains induced by both states under a matched budget and determines whether to adopt the intervened reasoning result. On five mathematical reasoning benchmarks, SDTT improves the Qwen3-1.7B SLM's Math Avg. by 34.64% relative to greedy SLM decoding. This result indicates that effective small-large model collaboration depends not on optimizing the current token in isolation, but on whether the intervention changes the state on which the SLM relies for continued reasoning in a way that improves access to reliable solution domains within a finite budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.