: A systematic evaluation of algorithmic diversity for code generation models
Abstract
Code generation models are actively being used in professional and educational settings. However, the competence of large language models (LLMs) to generate algorithmically diverse code remains largely unknown. We propose the DivAlgo benchmark, a large-scale dataset constituting algorithmically diverse solutions to programming problems. We also introduce a lightweight similarity metric to conduct an evaluation at scale. Our experiments spanning 24 modern code generation models found that uncontrolled model-generated code is significantly less diverse than the human solutions (captures 31% of human diversity), and even the best diverse code generation strategy could only achieve 70% of the human diversity. Moreover, while it was possible to steer the models to generate solutions based on human strategies, they were unable to reproduce the diversity of the human solutions. These findings indicate the limitations of code generation models to capture the algorithmic diversity covered by humans, and overexposure to model-generated code can narrow the algorithmic vocabulary of programmers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.