acceptodds
Under review as a conference paper at ICLR 2027

When Do Cheap Probes Predict Expensive Training? Probing 3D-CT Encoders for Text Generation

Abstract

Building a 3D CT vision language model begins with a choice of which image encoder to build on. Today that choice is made by fine-tuning every candidate through the full language model and comparing downstream scores, an enormously expensive search. A cheap probe on the encoder's representation promises a way out, but whether it forecasts the expensive outcome has never been tested. We test this with CheapCT on report generation and on MeasureVQA, a new VQA dataset we build. MeasureVQA scores the outcome one capability at a time, its answers measured from segmentation masks and Hounsfield units. Report generation scores the whole report at once and reflects mostly disease. CheapCT forecasts expensive training across every capability. The forecast survives changing the probe readout and the language-model backbone. The rank agreement between CheapCT and fine-tuning stays high throughout, from ρ = 0.90 to 0.97. Used to choose an encoder, CheapCT picks one nearly as good as the best while fine-tuning a single candidate, at orders of magnitude less compute. We release the code and MeasureVQA at https://anonymous.4open.science/r/CheapCT-anon-33B5j.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.