HDPO: Aligning Large Language Models to Software Architecture via Hierarchical Preference Optimization
Abstract
Large Language Models struggle with software evolution in complex repositories, often taking a “Semantic Shortcut” that ignores architectural hierarchy. We introduce Hierarchical Direct Preference Optimization (HDPO), a novel alignment framework that enforces structural consistency by decomposing feature generation into a rigorous three-tier mapping of Modules, Tasks, and Files. Across major open-source ecosystems (e.g., Django, Kafka, Kubernetes), HDPO achieves an average structural accuracy of 40.3%, drastically outperforming Step-SFT (14.7%), Step-Weighted-SFT (16.5%), and Step-DPO (19.0%). Similarly, regarding Average Weighted Impact across these ecosystems, HDPO (51.0%) decisively outperforms Step-DPO (25.0%), Step-Weighted-SFT (22.3%), and Step-SFT (20.7%).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.