acceptodds
Under review as a conference paper at ICLR 2027

Same Answers, Different Targets: The Hidden Joint Law of Structured Supervision

Abstract

Teaching the same answers together can produce different training targets. Two Python workers must return complementary subsets of an input, but the task does not require them to share an implementation style. We construct datasets with identical local programs, marginals, relations, and legal support, yet different frequencies of complete responses: their joint laws. Across three paired Qwen3-8B runs, this change moves held-out same-style generation from 24.3–25.1% to 72.0–72.6%; a common-update-denominator rerun reproduces the contrast at one existing seed. Changing only the weights on the same 1,024 scored validation responses also reverses likelihood-based checkpoint selection on all 32 held-out specifications. A classical minimum-information projection separates task-required from constructor-added dependence. On released programs, upweighting same-source pairings raises a prespecified voter's wrong acceptance by 4.31 points over 75 tasks with unseen inputs. A separate experiment holds the actual first program fixed: a new second attempt shares more errors and supplies less rescue, reducing at-least-one success by about one point. The joint law is thus a learner-visible part of supervision with consequences for both model selection and what one answer adds to another.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.