acceptodds
Under review as a conference paper at ICLR 2027

Real Code, Synthetic Mistakes: Teaching Small Code Models Not to Hallucinate

Abstract

Code autocomplete runs on small open models: every keystroke needs an answer within milliseconds, and much source code is not allowed to leave the company that owns it. These models often hallucinate. They call methods that do not exist, pass arguments an API does not accept, or import packages that were never published. The usual remedies do not fit autocomplete: execution-based filtering needs tests that a half-written file does not have, and preference learning needs labelled pairs that nobody collects for individual cursor positions. We start from a simple observation. Every fill-in-the-middle (FIM) example mined from a real repository already contains a correct answer, the code the developer wrote, so a frontier model only has to write a convincing mistake. We turn this into an execution-free pipeline that mines 400K realistic FIM holes across eight languages and produces 2.47M hard negatives in four hallucination categories. With this data we study two questions. First, fine-tuning a 7B model on only 5K curated holes raises exact match on the Delulu hallucination benchmark from 49.6% to 61.2% for a FIM-pretrained model (42.9% to 59.7% for its Instruct variant) and more than halves the rate of completions that fail Delulu's compiler check. Randomly placed holes and cutting outputs to the right length fall well short of this, and the gains transfer to three model families and to languages never seen in training. Second, the synthetic negatives add value when chosen well. Contrastive training on them improves control-flow completion by 1.5–3.5 points over fine-tuning alone, although it does not help long multi-line completions. A cheap, label-free difficulty signal, the fraction of blind LLM judges a negative fools, finds the pairs that help: ORPO trained on the pairs that fool the judges beats fine-tuning on Delulu by 1.5–1.9 points at the same budget, while easy or rule-based negatives add almost nothing. We release the code for every stage of the pipeline after hole mining.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.