Action-Stable HDR Interfaces for Frozen Vision-Language-Action Models
Abstract
Vision-language-action (VLA) policies fail when an exposure change or a bright light patch clips the camera image, even though the task is unchanged: an EV+3 shift takes OpenVLA-OFT, OpenVLA and from 82/90 to 0/90 successful episodes on LIBERO. We ask which camera interface preserves the actions of a frozen VLA and derive a minimum-intervention answer from the observation model. A common task weight cancels from the per-pixel choice of source, and clipping makes single-image recovery non-identifiable. With a second, shorter exposure under an ideal noise-free, unquantized two-level lighting model, four image operations follow: threshold the near-saturated pixels, close them to the smallest compatible lit region, dilate by one pixel to cover a one-pixel localization error, and write the recorded protected values inside. The resulting training-free interface, HDR-Guard, restores 82/90 under the EV+3 shift, where five single-image corrections reach at most 1/45. Under a sensor-aware simulation of a hard-edged light patch brighter than the room, it reaches 76/90 successes, against 70/90 for a soft saturation blend, 57/90 for full-frame Mertens fusion and 21/90 for the raw camera; with 500 episodes per rule, OpenVLA reaches 335/500 with HDR-Guard against 156/500 with Mertens fusion. Against the nominal frame, the lit-region error of HDR-Guard is 6.1 gray levels, against 15.7 for the soft blend; the 1.1% of the frame that the closing adds accounts for 75% of the reconstruction gain, and over 1,100 paired states, writing protected values there has an estimated +3.4-point equal-suite success effect (95% task-cluster interval [-0.4, +7.7]). In a simulation of trigger-free acquisition with the protected frame 50 ms old, OpenVLA reaches 362/500, against 336/500 without delay. On tripod-bracketed DSLR RAW frames, valid protected samples normalized by one calibrated global gain have a 4.8% median relative error against an independent short-exposure reference, and in 178,200 frames of real firefighting footage, 49% of frames with a person show a person with over 5% of their pixels near saturation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.