acceptodds
Under review as a conference paper at ICLR 2027

ArchRSBench: Enhance LLM Reasoning for Repository Resolution

Abstract

While Large Language Models (LLMs) excel at localized code generation, scaling to repository-level resolution remains a formidable challenge. In practice, monolithic implementations are intractable; developers must decompose features into a structured “Chain of Patches” managed as a Directed Acyclic Graph (DAG) for safe parallelization and review. However, existing benchmarks evaluate end-to-end execution, entirely overlooking the a priori architectural reasoning required to plan these structures. We introduce ArchRSBench, a rigorous benchmark designed to isolate and evaluate system-level architectural decomposition. It formalizes feature planning as a DAG generation task, challenging models to define semantic subtasks, localize files, and map dependencies. We propose a Dual-Evaluation Framework combining deterministic graph-theoretic metrics with LLM-as-a-Judge assessments, supported by a strict anti-contamination pipeline. Our curation yields 1,525 high-fidelity instances across 13 diverse repositories. Empirical evaluations expose critical vulnerabilities: the leading frontier model achieves a zero-shot deterministic score of only 67%. Crucially, architectural planning requires rigorous constraint scaffolding; our specialized Tri-Agent pipeline unlocks latent reasoning capacity, yielding a 12% absolute improvement. Furthermore, augmenting prompts with file contents provides negligible gains over operating strictly on the directory tree topology. Thus, current LLMs over-rely on localized lexical pattern matching, fundamentally lacking the higher-order topological reasoning required for true architectural intelligence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.