acceptodds
Under review as a conference paper at ICLR 2027

Gains Stay Where the Loss Went: A Channel-by-Condition Map of Physical-Plausibility Training in Vision-Language Models

Abstract

A vision-language model shown a shadow that points toward the light calls the picture physically possible. Fine-tuning on plausibility repairs this, and prior work reports the repair on one readout, usually an accuracy. On ImPhysible, a rendered testbed of matched normal and violated images, we read one model's judgment on four channels, a 0 to 100 score, the first-token logit, the generated verdict, and a probe on the hidden state. We vary which channel and which condition the training loss passed through. Every (channel, condition) cell the loss passed through rose, in every variant we trained. The cells it did not pass through rose, stayed at chance, or moved the other way, and no variant made them rise reliably. A score-order reward lifts score pair-rank accuracy from 0.469 to 0.576 and pushes the logit of the same answers from 0.504 to 0.456. An alignment loss that holds out one domain raises that domain's representation cosine to between 0.556 and 0.729 while its free answer stays between 0.463 and 0.500. Decide which channel you will read, put the loss through it, and test untrained conditions separately. One model, rendered stimuli, and an earlier data version behind part of the channel-axis evidence bound the claim.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.