The Function-Style Trade-off: Understanding LLM Code Generation from a Novel Perspective
Abstract
LLMs have made significant progress on code generation, yet their outputs often suffer from poor readability, inconsistent naming, missing documentation, and jumbled control flow. Current benchmarks overemphasize execution correctness while lacking scalable code-style metrics, hindering the development of style-aware code generation. To fill this gap, we construct a fine-grained code-style evaluation benchmark and propose a graph-based framework using Augmented AST that captures structural style features beyond the token level and enables node-level attribution of style decisions. Experiments reveal a Function-Style Trade-off: improving style naturalness degrades functional performance. We provide an information-theoretic interpretation of this phenomenon, framing fine-tuning as reallocation of representational capacity between natural language and code distributions. Extensive experiments demonstrate that the proposed structural style signal transforms the Function-Style Trade-off from an uncontrollable statistical fluctuation into an analyzable and optimizable engineering problem. This work underscores the necessity of incorporating software quality into LLM code evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.