SUPERVISION MECHANISMS IN IMPLICIT CHAIN-OF-THOUGHT: TARGETS, ALLOCATION, AND SPACE
Abstract
We study supervision mechanisms in implicit chain-of-thought models along three dimensions: the quantities to align (Targets), the updates to supervise directly (Allocation), and the dimension and directions of the alignment space (Space).Mathematical analysis reveals how different supervision targets organize reasoning errors: local update targets constrain discrepancies in individual state changes,whereas cumulative update targets aggregate these discrepancies across steps.Controlled comparisons in a shared recurrent student show that, with local update targets, extending direct supervision to intermediate updates in the answer subspace further improves reasoning accuracy. The choice of supervision space also matters: the answer subspace outperforms full-space alignment and rank-matched random subspaces, highlighting the role of direction selection beyond dimensionality reduction. Based on these findings, we introduce Answer-guided Local Update Supervision (ALUS), which aligns teacher and student changes between consecutive hidden states at every valid update. Alignment uses a fixed subspace derived from output-weight rows for training-answer tokens and adds no inference-time modules.ALUS reaches 41.17 ± 0.20% on GPT-2/GSM8K,8.31 percentage points above no alignment and 4.88 above the strongest adapted supervision objective within the shared student. The benefits of all-valid supervision and the answer subspace persist with six updates; experiments with Pythia-1B and nonnumeric ProsQA provide further evidence for the effectiveness of ALUS.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.