acceptodds
Under review as a conference paper at ICLR 2027

From Teacher Imitation to Failure Diagnosis: Causality-Inspired Test-Time Evolution of Small Language Models with LLM Feedback

Abstract

Test-time evolution of small language models (SLMs) with large language model (LLM) guidance has emerged as an important paradigm for LLM-SLM collaboration. Existing methods typically distill complete LLM-generated solutions into SLMs, while recent approaches selectively transfer knowledge at difficult tokens or intermediate reasoning steps identified from the SLM's reasoning state. Despite differing in what to transfer, these methods largely follow the same paradigm: when an SLM fails, it imitates the correct solution or reasoning provided by a stronger LLM. However, substantial differences in model capacity, training data, and acquired knowledge may make such direct imitation difficult for SLMs to effectively absorb. We take a different perspective: rather than only showing an SLM what the correct solution is, the LLM should first diagnose why the SLM fails. Such diagnosis serves two purposes. First, it guides the LLM to minimally revise the SLM's erroneous reasoning, preserving what the SLM already does correctly while correcting only failure-relevant parts. This provides a learning target closer to the SLM's current reasoning capability than an independently generated LLM solution, facilitating more effective knowledge transfer. Second, the diagnosis itself provides an additional learning signal, enabling the SLM to learn not only how its reasoning should be corrected but also why it failed. Based on this insight, we model LLM-guided SLM evolution from a causal perspective and propose CausalBridge. Reasoning State Probing samples and clusters multiple reasoning trajectories to characterize the SLM's reasoning state. Causal Diagnostic Mediation diagnoses the underlying failures and generates minimal, diagnosis-guided revisions. Finally, Mediator-Guided Evolution uses the diagnostic information as a mediator to improve failure-relevant reasoning. Extensive experiments across diverse reasoning tasks show that CausalBridge consistently outperforms both independent SLM evolution and existing LLM-SLM collaborative evolution methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.