ChemHarness: A Self-Evolving System for Chemical Reasoning via Asymmetric Dual-Stream Verification
Abstract
Self-evolving harnesses improve a language model by consolidating its verified experience into reusable memory, but real-world chemical reasoning problems lack ground-truth answers for verification, and repeated samples of one model tend to repeat the same errors. We present ChemHarness, a closed-loop self-evolving harness for chemical reasoning driven by dual-stream self-verification. ChemHarness solves each problem through two streams of one language model: a neural stream derives the answer in natural language, and a symbolic stream writes a program that an interpreter executes. Over repeated rounds, cross-stream verification compares the two answers and turns them into label-free feedback in the form of cross-verified answers, intermediate lemmas, and discrepancies. ChemHarness organizes this feedback through a procedural tool library across problems, an episodic short-term memory within a problem, and an asymmetric information flow that withholds both memories from the neural stream, which therefore remains an independent verifier. On the chemistry datasets of SciBench, ChemHarness improves direct reasoning by 6.4 to 12.7 points on Qwen2.5-7B, 14B, 32B, and DeepSeek-V4.1-Flash, and exceeds ChemAgent by 23.3 and 22.9 points on Qwen2.5-7B and 14B. We further find that cross-stream verification is markedly more precise than single-stream consistency, with a precision of 63.0% to 75.4% against 46.2% to 55.7%, and that memory reaching the neural stream lowers the precision. Dual-stream self-verification thus allows a harness to evolve in chemical reasoning without ground-truth labels, parameter updates, or external supervision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.