acceptodds
Under review as a conference paper at ICLR 2027

Post-Training Language Models for Gold-Medal Performance in Coding Competitions

Abstract

Competitive programming has become a key test of language model reasoning, with IOI and ICPC among its hardest settings. We present an end-to-end specialization pipeline combining problem curation, synthetic reasoning traces, supervised fine-tuning (SFT), reinforcement learning (RL), and test-time compute. Using 22,000 problems, we train Nano-CC (30B-A3B) with SFT and RL and Ultra-CC (550B-A55B) with SFT alone. Across Nano-CC, SFT drives most single-sample gains and RL adds modest improvements, while feedback-driven GenCorrect iteratively generates and refines diverse solutions and provides the largest system-level gain. On IOI 2025, Nano-CC improves from 130 points to 291 after post-training and to 468 with GenCorrect, exceeding the gold threshold of 438.3 while Ultra-CC reaches 502. Guided by these results, we develop a competition-specific Ultra-CC system and evaluate it prospectively during IOI 2026. Under the same time, internet-access, and submission constraints as human contestants, it scores 535.4 out of 600, exceeding both the gold threshold of 361.12 and the top human score of 498.27. To our knowledge, this is the first AI system to outscore the highest-scoring human contestant on an IOI problem set.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.