Beyond Labels: PAC Learning with Intermediate Supervision
Abstract
The traditional learning paradigm relies on input-label pairs but as concept classes become increasingly complex, learning becomes statistically and computationally expensive. Modern machine learning bypasses this problem using intermediate supervision like chain-of-thought and distillation to speed up learning. Although widely explored in practice, its theoretical understanding remains limited. We study the PAC learnability of deep models composed of intermediate gates with partial information about their outputs, revealed only during training. Our flexible concept class captures a wide range of architectures, including deep ReLU networks and transformers and uses gates in layers to map a dimensional feature to a label. While learning is computationally intractable even for one-hidden-layer networks, we show that mild intermediate supervision makes deep models efficiently learnable. We study different degrees of intermediate supervision ranging from exact intermediate traces to partial and noisy information. We provide tight characterizations of the computational complexity; the exact case is fully polynomial in , while the partial and noisy cases are polynomial for constant which we show that is unavoidable by proving a corresponding cryptographic hardness result. Finally, experiments demonstrate the benefits of intermediate supervision in cases where classical label-only approaches fail.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.