acceptodds
Under review as a conference paper at ICLR 2027

FedHVLN: Heterogeneous Federated Vision-and-Language Navigation

Abstract

Federated vision-language navigation (VLN) trains agents from private environments, but averaging native models requires matching parameter shapes. We introduce FedHVLN, which shares compatible language-side modules while retaining each client’s navigator. Four stages train a human-instruction encoder and transformation, a federated Speaker that turns local route context into pseudo-instructions, a feature generator aligned with human features, and pseudo-instruction navigation. Clients upload only the shared modules designated at each stage; inference uses the human-instruction path. On the Room-to-Room (R2R) validation split (1,021 held-out instructions from 56 previously seen environments), the mixed federation reaches 64.45% success rate (SR) and 61.69% success weighted by path length (SPL). At fixed round 991, aggregating the shared interface across both model sizes improves pooled SR and SPL by 3.14 and 3.23 percentage points over matched single-size controls. Both evaluated subsets have positive point differences on all five navigation metrics, although exploratory scene-resampling intervals for the compact-model subset include zero. In homogeneous federations, FedHVLN improves all five metrics over FedVLN, with SR/SPL gains of 1.37/1.49 and 2.15/2.74 points for the wide and compact models, respectively. These results support sharing a compatible language interface across the two tested navigator structures on R2R validation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.