Code‑Executable Judge: Mitigating Arithmetic Hallucinations for Wallet‑Agent Evaluation
Abstract
Large‑language‑model‑driven wallet agents often generate answers derived from thousands of billing records, requiring heavy numerical verification. Conventional LLM‑as‑Judge suffers severe arithmetic hallucinations, leading to inconsistent and unreliable evaluation results. To tackle this problem, we present Code‑Executable Judge: instead of letting LLMs directly compute numeric results for judgement, our approach dynamically generates Python verification code and validates answers via program execution. We conduct extensive experiments on wallet‑agent benchmarks. Compared with vanilla LLM‑as‑Judge baselines, our method achieves substantially higher agreement with human annotations. We further systematically analyze failure modes of code‑driven judging and discuss its practical limitations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.