AskLight: Grounding Active-Light Reasoning for Multi-Illuminant White Balance
Abstract
Multi-illuminant white balance requires a spatially varying illumination estimate, but visible light sources poorly indicate the actual illumination: they may be switched off, miss sources outside the view, or share a region with other sources. We present AskLight, which connects semantic source reasoning to continuous illumination decomposition through an explicit source-to-region representation. A shared vision–language model (VLM) classifies detected proposals, then revises the active inventory by removing false actives, adding missed-visible or hidden hypotheses, and associating every retained hypothesis with a dominant illumination region. These relations initialize a variable-cardinality set of illuminant colors and spatial priors, which a recurrent decomposition backend refines into mixture weights and candidate chromaticities through differentiable physical composition. A central obstacle is that no real dataset provides per-source supervision of which sources are active, which are visible, and where each one casts light. We therefore introduce Infinigen-Light, a procedurally rendered multi-illuminant dataset with per-source annotations of activity, visibility, and region influence, which makes active-light reasoning both learnable and separately evaluable from final white balance accuracy. By grounding each illuminant candidate in an explicit source hypothesis rather than an unstructured slot or a global estimate, our method achieves the lowest mean illumination error on Infinigen-Light, LSMI, and MIIW, reducing the LSMI illumination error by 16.8% over the strongest baseline. We will release our source code.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.