DEBATE, NOT VOTING: AGENTIC ARBITRATION OVER WEAK OPTICAL CHEMICAL STRUCTURE RECOGNITION MODELS
Abstract
Optical chemical structure recognition (OCSR) turns a picture of a molecule into a machine-readable string, and it is the bottleneck between decades of printed chemistry and any computational use of it. Individual OCSR models remain unreliable on real documents: on a 1000-molecule sample of MolParser, three recent readers score 46.0%, 43.4% and 42.4% exact match, yet at least one of them is right on 60.8% of molecules. That gap is an ensembling problem, but it is not one that voting can solve, because the readers rarely agree and are rarely all wrong in the same way. We instead place the readers inside a multi-agent scaffold in which four language-model agents with cheminformatics tools argue about a candidate structure, write their findings to an append-only knowledge base, and are constrained by render-and-compare gates that admit a correction only when it is demonstrably better. The debate reaches 61.9%, beating the best single reader by 15.9 points and plurality voting by 9.2 points ( for both, robust to correction for all 27 tests we report), and is statistically indistinguishable from the any-reader-right ceiling. That aggregate hides two opposing effects: on the 661 molecules carrying no stereochemistry the debate is 3.9 points above the ceiling (), constructing structures no reader proposed, while on the 339 stereochemical ones it falls below. We then swap the arbitrating model. A stronger arbitrator adds 2.2 points (nominal ; neither this nor the +3.9 survives Bonferroni correction over all 27 tests, so we treat both as suggestive), and its mechanism is visible: it opens a correction on 423 of 1000 molecules against 120, and an interaction test indicates that it depends less on the individual readers (nominal , not correction-robust). We argue that the useful unit of progress here is the arbitration policy, not the reader, and that reported OCSR accuracy should be read against the union oracle rather than against the best single model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.