Entropy-Guided Adaptive Diagnostics: Enhancing Code Reasoning via Strategic Inference-Time Refinement
Abstract
Large Language Models (LLMs) show great potential in code generation, yet achieving functional correctness remains fundamentally challenging. Two critical issues persist: (1) verifying code typically requires comprehensive but expensive external execution feedback; and (2) while generative entropy correlates with logical errors, it is a noisy indicator that often conflates real flaws with legitimate algorithmic variations. To address these, we present Entropy-Guided Adaptive Diagnostics (EGAD), a training-free framework for efficient inference-time refinement. Instead of resource-heavy parallel sampling, EGAD identifies latent inconsistencies by monitoring structural entropy fluctuations. It utilizes these signals as diagnostic guides rather than rigid mandates: a Suggestion LLM performs targeted analysis on flagged segments to differentiate logical anomalies from semantic noise, providing surgical revisions only where necessary. EGAD requires no additional training and integrates seamlessly into existing inference pipelines. Evaluated on HumanEval, LiveCodeBench, and MLE-bench, EGAD delivers substantial gains across multiple base models. Notably, on LiveCodeBench, EGAD boosts Qwen3-80B-A3B from 65.89% to 76.56% (surpassing Claude 4.1 Opus) while maintaining a minimal inference budget. Our results confirm that spatially-targeted intervention is significantly more compute-efficient than path-level sampling method, demonstrating a superior trade-off between inference cost and functional correctness in test-time code reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.