OpenArch: A Unified PyTorch Framework for Implementing Modern Open-Source LLM Architectures
Abstract
Large Language Models (LLMs) have emerged as foundational systems for natural language understanding and generation tasks. While production-scale implementations of modern LLM architectures are often optimized for speed and scalability, they present significant barriers to learning and reproducibility. This paper introduces OpenArch, an open-source PyTorch framework providing clean, educational implementations of 20+ modern LLM architectures. Our implementations span diverse architectural innovations including Grouped Query Attention (GQA), Mixture-of-Experts (MoE) routing, Rotary Position Embeddings (RoPE) and advanced normalization techniques (QK-Norm, RMSNorm). Each model is implemented as a single readable file with explicit design choices and comprehensive documentation, facilitating learning, comparison and rapid prototyping. We provide detailed architectural analyses, reference comparisons against official implementations and empirical validation across 15+ distinct model families ranging from GPT-2 to Kimi K2 (1T tokens). OpenArch serves as both an educational resource for practitioners entering LLM systems research and a reproducible reference implementation for architecture ablation studies. The codebase achieves functional parity with reference implementations while prioritizing readability over performance optimization. We discuss implementation patterns, common pitfalls in architecture replication and recommendations for practitioners implementing novel LLM architectures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.