Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
Abstract
Coding agents must integrate external tool returns into ongoing reasoning. We draw a structural analogy between the action → observation → continuation loop of a coding agent and a function call site, where a caller binds arguments, a callee returns a value computed elsewhere, and downstream code consumes that value. This analogy motivates learning from function-level dependencies in ordinary code through function-aware fill-in-the-middle (FIM) mid-training: an infilling objective that masks functions selected via program dependency graph analysis and a complexity–inferability double criterion, with teacher-generated rationale-plus-implementation targets. We mid-train Qwen2.5-Coder-Instruct (7B/14B) and Qwen3-8B on a 2.6B-token decontaminated corpus drawn from 968 GitHub repositories, then apply existing agentic post-training pipelines. Mid-training improves SWE-Bench-Verified by +2.8/+3.0 at 7B/14B and by +3.2 on Qwen3-8B; SWE-Bench-Lite gains are +3.7/+4.0/+5.3 on the same models. The improvement holds across two post-training pipelines (R2E-Gym, SWE-Smith) and extends to Qwen3-8B with SWE-Lego. Beyond in-domain gains, mid-training improves final-agent performance on non-agent coding (e.g., LiveCodeBench) and non-coding tool-use benchmarks (τ-bench, BFCL). A single Python-derived corpus thus improves the resulting agents across coding and tool-use tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.