Knowing More, Completing Less: Reliability Learning in Lifelong Robot Fleets
Abstract
A robot that stays still may have failed, or it may be waiting for another robot that failed. Learning from such motion is a statistical problem; using what is learned is a coordination problem. We study their connection in lifelong multi-agent path finding. Our central observation is a reversal: under the same planning rule, replacing fleet-average reliability with exact individual reliabilities increases completed tasks by 7.96% in a room but decreases them by 13.09% in a warehouse, with the corresponding sign in all six initial-state averages per setting. To understand this gap, we derive the position-feedback likelihood and characterize exactly which immediate task-value comparisons are identifiable. A task-value difference can be identifiable even when individual reliabilities, or both competing values, are not. We then develop a motion-censored release preference that learns from unambiguous actuator outcomes and refines equally preferred next moves, yielding a 9.76% gain in a two-room evaluation with a frozen method and previously unused seeds. A trajectory decomposition associates the room gain with fewer propagated stalls and the warehouse loss with a growing deficit in proposed moves. A four-robot construction then shows how better immediate task completion, more movement, and fewer robots in cells with at most two neighbors can all lead to fewer tasks after two steps. Together, these results connect what a fleet can learn from motion to the future its planner creates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.