Project Before You Distill: Carrying Privileged Chain-of-Thought Across Sensor Gap
Abstract
Large vision–language models provide rich reasoning priors, but physical sensors often lack large paired corpora and intermediate reasoning supervision. Directly distilling a rationale generated from a privileged sensor can introduce claims that are not supported by the deployment sensor. To address this gap, this paper introduces Cross-sensor Rationale Projection (CRP) which treats the privileged rationale as a reasoning proposal and re-derives it against the deployment observation. CRP keeps supported information, abstracts fine details into sensor-supported structure, hedges partially supported claims, and removes unsupported content, followed by a consistency screen. We further analyze information loss across the sensor gap and derive an Estimated Projection Value (EPV), computable even before training the deployment student model, to indicate when projection helps with consistent empirical support. Two purpose-built sensor reasoning benchmarks, RadarQA-300 and AudioQA-300, are constructed for RGB-to-radar and audio-visual-to-audio reasoning. CRP consistently outperforms answer-only supervision, direct privileged-rationale distillation, and deployment-sensor rationale supervision, with gains of and percentage points over answer-only training on RadarQA-300 and AudioQA-300, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.