acceptodds
Under review as a conference paper at ICLR 2027

Byte-Matched Evaluation of Rank-Split Federated Adaptation for Tool-Using Vision-Language Agents

Abstract

Tool-using language agents interleave task reasoning with tool-specific execution. In federated deployments, clients own heterogeneous tool environments and cannot share raw trajectories, which makes it natural to give every client **private adaptation capacity** alongside a shared, uploaded partition. We test that premise under a **fixed upload budget** rather than a fixed parameter count. Rank-split federated adaptation (DREL) spends a client's per-event upload on a shared rank-32 reasoning partition while keeping a rank-64 execution partition private, so at the identical **87,293,952-parameter (333 MiB)** upload it deploys exactly **three times** the adapter parameters of a plain rank-32 FedAvg control. On a four-client GTA protocol with two seeds, the threefold capacity does not convert into accuracy: DREL does not show an accuracy advantage over the byte-matched rank-32 FedAvg control (**58.4% versus 60.1%** on the pre-specified eight-cell macro endpoint, **-1.74 points**, 95% CI [-3.56, +0.22]). We then ask where that capacity goes by intervention rather than argument: zeroing the private partition at inference, with the uploaded partition byte-identical, **raised mean accuracy from 58.53% to 60.85%** across the 3 cells evaluated (3 of 3 improving), where the byte-matched control scores 59.30%. We also expose a confounded comparison in our own earlier work that had motivated dropping the method's routing component, and report lower observed GAIA accuracy than local adaptation on the one client evaluated. The reusable lesson is a protocol one: **federated rank comparisons are confounded unless upload is matched exactly.**

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.