acceptodds
Under review as a conference paper at ICLR 2027

What Feedback Does a Prompt-Optimized Logic Translator Need Under a Proof-Based Judge?

Abstract

Translating natural language into formal logic lets a theorem prover do the inference: reasoning becomes a checkable procedure, and an entailment verdict comes with a proof. What such a translator needs is rarely measured, because published systems are evaluated as packages, each under its own data, models, and judge. We hold the judge fixed, one proof-based judge for every condition, and vary the feedback a prompt-optimized LLM translator receives, its weights frozen. Several models are evaluated on lexical, structural, and monotonicity entailment datasets. We compare feedback sources (target formulas from a symbolic pipeline or entailment labels with the prover's diagnostics), feedback channels, delegation to notation or code, and sentence-level versus pair-level translation. With a strong reflection LM, label feedback fixes the conventions the training pool penalizes, but leaves the two calls disagreeing on argument structure and on predicate names: under a matched training pool, the label-trained translator trails the formula-trained one by eight points on lexical entailment. The condition that shows the formulas as text in the feedback closes that gap to the one that scores against them. A code layer or a one-page notation in the seed prompt each make up a large part of the shortfall without the formulas. Applied to the pair instead, the same prover feedback scores highest on lexical entailment and lowest on the non-entailment list; but on lexical entailment two of every five pairs it proves are merged, the hypothesis formula mapping into the premise formula although the hypothesis sentence has a word the premise lacks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.