cuBCR: Batched GPU Evaluation for Counterfactual Responsibility
Abstract
Multi-agent systems coordinate multiple AI agents to complete tasks, but attributing responsibility becomes challenging when collaborative execution fails. In time-sensitive services, slow root-cause analysis delays mitigation and degrades service quality. Explaining team failure requires analyzing how individual actions constrain the overall system's ability to recover. In this paper, we formalize responsibility attribution in finite stochastic models: holding a subset of agent policies fixed, where a policy specifies how an agent chooses its actions, we evaluate whether the remaining agents can still coordinate an intervention to prevent task failure before a deadline. If the failure is avoided through intervention, we use Shapley values to allocate responsibility among agents who contributed to the failure. Nevertheless, exhaustively evaluating every coalition of agent policies requires an exponentially large number of counterfactual simulations. We tackle this challenge with CUDA-accelerated Batched Counterfactual Responsibility (cuBCR), which accelerates the process through Monte Carlo sampling and GPU parallelism. Sampling curbs the combinatorial explosion in the number of required counterfactual checks, while batched GPU kernels evaluate the recovery simulations concurrently. On an event-driven ticketing benchmark, cuBCR achieves a median 233× speedup over the CPU-based exhaustive failure attribution across 36 workflow tasks. By significantly reducing attribution latency, cuBCR makes iterative counterfactual analysis practical for production debugging, enabling system designers to evaluate how individual policy revisions alter multi-agent failure dynamics quickly.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.