Universal Braille-to-Text Translation
Abstract
For visually impaired readers, text is transcribed into braille under a transcription code, a set of rules specific to a language and a standard, and every code writes into the same 256 Unicode braille codepoints. Given the code, transcription is deterministic, but its inverse is not: several texts can produce the same cells, and the cells do not mark which of hundreds of codes produced them. This inverse, braille-to-text translation, is how transcriptions are proofread and how documents that survive only in braille are recovered. Existing systems require the code, most are built one code at a time, and even the widely used rule-based engine, LibLouis, fails to invert its own output when given the code. In response, we present a system that infers the code and reads braille across 156 codes in 85 languages. We fine-tune an LLM on pairs generated by the deterministic text-to-braille transcriber, and its candidate readings are verified by re-transcription. Our system achieves 0.39% per-cell error, against 11.4% for LibLouis given the correct code. We also analyze the remaining error by whether re-transcription can detect it, and find that it concentrates in a few code families, each for a different cause. On an English Bible, with the code given, 80% of readings that reproduce the input cells read at 0.29% character error, against 2.86% for the rest. We open-source our implementation, corpus, and checkpoints: https://github.com/anonytoucan/ubt.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.