Understanding Single- and Multi-Agent Learning: Complexity, Alignment, and Error Propagation
Abstract
When does decomposing a task across specialized agents improve learning efficiency? Empirical studies report mixed outcomes across tasks, training procedures, and resource budgets. We develop a comparative learning-theoretic framework for monolithic and modular sequential policies, organizing established tools around three mechanisms: statistical class complexity, objective and policy-class mismatch, and stagewise learning errors. Under conditional-value access, independent policy and reward factorization yields a local-complexity guarantee; dependent modular policies can match the single-agent sufficient budget under matched dimension, horizon, and radius allocations. Reward discrepancy and class-optimum gaps qualify these comparisons, while a separate reference-demonstration analysis relates local learning rates to downstream error propagation. Finite-class and exact Gaussian regression examples show why neither agent count nor dependency graph alone determines learning costs. The framework makes explicit the observation models and comparators needed to attribute efficiency differences to decomposition. Its PAC results compare sufficient budgets, while the regression example gives exact expected excess risks under stated assumptions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.