acceptodds
Under review as a conference paper at ICLR 2027

When Transcript Supervision Cannot Locate Characters: Set-Valued Cuts and Sparse Anchors for CTC OCR

Abstract

Connectionist temporal classification (CTC) learns transcripts by summing monotone alignments, but its character peaks need not identify physical separators. We formalize adjacent-character segmentation as an ordered, one-dimensional, set-valued cut task: a cut is accepted between the projected right edge of one character and the left edge of the next, including their overlap. We establish a two-world lower bound when identical observable word images and transcripts admit separated valid-cut sets, and derive an anchor-recovery relation under fixed linear features and full-rank design. A recognition-conditioned SVTRv2 system combines CTC decoding, a bounded residual cut head, interval supervision from sparse character boxes, and an optional soft interval-gated CTC objective. On 61,842 source-disjoint SynthText words of length 2–5, 1% cut supervision raises midpoint ±2-frame all-cut word accuracy from 69.52% to 90.32%; dense supervision reaches 95.17%. The corresponding strict transcription-and-all-cuts rates are 64.78%, 84.15%, and 88.76%. Controlled synthetic experiments reproduce the lower-bound and rank-transition predictions, while OCR results quantify the gain from sparse geometric anchors.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.