acceptodds
Under review as a conference paper at ICLR 2027

AnimeText: A Large-Scale Dataset and Benchmark for Text Detection in Anime-Style Images

Abstract

Text in anime-style images includes dialogue, titles, sound effects, signatures, and lettering drawn into the artwork. We introduce AnimeText, a dataset of 735,060 images and 4,360,393 text boxes spanning comics, illustrations, posters, and edited images. AnimeText includes four kinds of annotation: fine boxes for compact text regions, coarse boxes for larger groups, polygons for irregular boundaries, and rejected candidates for text-like patterns and poorly localized proposals. We build these annotations with detector proposals and a fine-tuned candidate classifier, and annotators correct the boxes and add text the detector missed. Trained on 5,965 AnimeText images, YOLO11m reaches 0.850 F1 on the full test split of 73,725 images, compared with 0.545 when trained on the same number of Manga109s images. In a labeled sample of test images, non-dialogue text such as sound effects, titles, and signatures accounts for most of the newly detected regions, and its recall rises from 18.3% to 69.8%; dialogue recall also rises from 63.4% to 91.3%. Training on 513,159 AnimeText images raises F1 from 0.850 to 0.896 and AP50:95 from 0.781 to 0.882. Very small text and crowded pages remain the hardest cases.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.