acceptodds
Under review as a conference paper at ICLR 2027

SEDiT: Mask-Free Video Subtitle Erasure via One-step Diffusion Transformer

Abstract

Existing video subtitle removal methods inpaint masked frames, so they must localize the subtitles first, and the accuracy of that segmentation limits the final quality. We present SEDiT, a one-stage video Subtitle Erasure approach built on a one-step Diffusion Transformer. SEDiT is mask-free at inference: it requires no mask, OCR, or segmentation at test time, and a coarse subtitle box is used only for subtitle-region loss weighting during training. Our key observation is that subtitle removal is often conditionally near-deterministic: once the subtitled video is given, the subtitle-free video is strongly constrained. Under an ideal conditional velocity field, a single Euler step corresponds to a conditional-mean estimator whose expected squared error equals the conditional variance of the target, which explains why one-step inference can work well for subtitle erasure while a multi-modal task such as stylization may still require multiple steps. Together with chunk-wise streaming inference, this enables one-step erasure on native 1080p videos of minute-scale duration. On our paired VSR-Bench-400 benchmark, SEDiT attains dB PSNR and FVD, achieving stronger raw-output fidelity than the evaluated mask-based baselines, which are given ground-truth subtitle boxes rather than detected ones, while taking s per 65-frame 1080p clip—about faster than Minimax-Remover and faster than EffectErase, and without step distillation. On real clips of roughly two minutes, primarily from television dramas and short-form dramas, subtitles are completely removed in under our strict full-clip criterion; supplementary videos provide further qualitative evidence on real video.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.