acceptodds
Under review as a conference paper at ICLR 2027

TeXForm: Auditing and Improving LaTeX Formula Normalization in Data and Evaluation Pipelines

Abstract

LaTeX normalization maps the many equivalent forms of a formula to one canonical string. Formula-recognition data pipelines and the formula tracks that rank today's vision-language models all silently rely on normalizers, yet not one of these normalizers states what it preserves. Our experiments reveal failures accumulated invisibly for years: normalization in dataset construction changes what its inputs render to, with case studies exposing symbol loss and structural truncation; inside three widely used benchmarks (UniMER-Test, OCRBench v2, and OmniDocBench), the official scoring pipelines falsely reject 19–26% of render-equivalent prediction–label pairs, with subset rates spanning 6–55%, while accepting some pairs that render differently as false-accept candidates. Beyond traditional text metrics, CDM, the recently proposed render-aware metric, also embeds a normalizer that alters what 14% of inputs render to. To address this systematically, we introduce LaTeX normalization as a formal task with four measurable properties (safety, canonicity, idempotence, and coverage), and characterize when normalize-then-match is a valid equivalence test. Guided by this formalization, we build TeXForm, a LaTeX normalizer driven by an explicit knowledge base of command argument structure and rewrite rules. Among all normalizers we evaluate, TeXForm is the only one to combine substantial canonicity (37–47%) with render-safety above 97%, the property that normalized training labels require. On UniMER-Test and OmniDocBench, replacing built-in normalization with TeXForm recovers 38.05% and 89.83% of previously rejected render-equivalent pairs, respectively. The audited normalization configurations change model rankings, including a leader reversal on UniMER-Test. We release TeXForm, including its knowledge base, as open source.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.