Identity-Paired Progressive Depth Training: When Trainability Persists Beyond Expressibility
Abstract
Variational Quantum Algorithms (VQAs) are a leading paradigm for near-term quantum computing, yet their training suffers from sensitivity to circuit depth, initialization, and landscape pathologies such as barren plateaus. We study progressive depth training (PDT)—a layerwise curriculum that trains a shallow circuit before appending new layers—and identify a fundamental obstacle: fixed entangling gates (CNOTs) in hardware-efficient ansatze cause initialization shock, an energy spike when new layers are added. We propose identity-paired progressive depth training (IP-PDT), which appends forward/inverse block pairs—each consisting of a standard rotationCNOT block followed by its reverse—that compose to the identity at initialization. Because the adjacent CNOT rings cancel, the effective circuit retains only a single entangling layer surrounded by overparameterized local rotations. We prove a simple Reachable Set Saturation Theorem: under this construction the variational manifold expands exactly once (when post-entangler rotations are first introduced) and then saturates; all subsequent depth increases provide pure overparameterization of single-qubit unitaries. Despite this saturation, progressive addition of rotation parameters can continue to improve optimization outcomes—a phenomenon we term trainability beyond expressibility. We formalize IP-PDT as a continuation method on nested manifolds, prove monotone energy guarantees under an acceptance rule, and connect energy error to ground-state fidelity through spectral-gap inequalities. A detailed resource analysis shows that IP-PDT achieves lower total gate cost than both baselines by eliminating most CNOT gates. Experiments span six-qubit Transverse-Field Ising, Tilted Ising, and Random Ising Hamiltonians, a broader nine-Hamiltonian benchmark, and system sizes up to qubits. IP-PDT matches or outperforms full-depth baselines on Hamiltonians whose ground states are well-approximated by the single-entangler reachable set, with particularly strong gains in limited-budget regimes. At the hardware-efficient baseline degrades sharply while the single-entangler methods remain trainable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.