TextBokeh: Language-Guided Depth-of-Field Editing with Target and Strength Control
Abstract
Depth of field controls which subject appears sharp and how strongly the rest of a scene is defocused. Existing post-capture interfaces commonly require a click, mask, focal distance, or numerical blur setting, and many are designed only for all-in-focus inputs. We present TextBokeh, a language-guided depth-of-field editor conditioned on a focal-object description and a relative defocus level. We train one model to synthesize bokeh from all-in-focus images and refocus already-defocused images at a requested defocus level. We construct its supervision by choosing the Chebyshev center of the target's estimated disparity interval as the focal plane, then holding this plane fixed while solving for the range of f-numbers that keeps the target sharp, produces visible off-target defocus, limits excessive blur, and respects the source lens range. Sampling this interval yields an ordered, scene-adaptive stack of training targets. Across the retained aperture levels, TextBokeh achieves substantially higher full-reference and perceptual fidelity than level-agnostic general-purpose editors on both AIF-to-Bokeh synthesis and cross-target refocusing. We will publicly release the trained LoRA adapter weights, the complete data synthesis pipeline, and the evaluation resources upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.