PRODUCT-SET DEFICIENCY: A TRAINING -FREE MEASURE OF JOINT-ACTION COUPLING IN MULTI -ROBOT TASKS
Abstract
Cooperative multi-agent reinforcement learning often relies on value factorisation to enable decentralised execution, but this can restrict the combinations of actions that multiple robots can select jointly. We apply product-set deficiency, δ, the complement of a known bounding-box fill ratio, as a training-free measure of this restriction. Given the feasible joint actions for an instruction, obtained by enu- merating joint actions in a contact simulator, δ measures how far this set deviates from the Cartesian product of the individually feasible actions; δ = 0 when ev- ery combination of individually feasible actions is jointly feasible. We evaluate δ on a two-robot rod-pushing task. Instructions requiring coordination through the object have substantially higher δ than an instruction that can be solved in- dependently. Although the absolute value of δ depends on action discretisation, the relative ordering of the instructions remains stable across K ∈ [6, 24], and the same pattern holds for three and four robots pushing a box. We then test whether higher δ translates into greater learning difficulty. Across our experiments, it does not: δ does not predict the cost of factorised learning, and coordination graphs show no detectable learning advantage over independent learners despite substan- tially higher computational cost. An edge-gating variant guided by δ gains where δ ≈ 0 and beats both on one set of fresh seeds, but not on a second, and controls show the same gain without δ. These results suggest that δ is useful for charac- terising non-factorisable joint-action structure, but should not be interpreted as a direct measure of learning difficulty.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.