Capability Cards for Compiling Feature Interfaces between Heterogeneous, Independently Trained Collaborative Perception Systems
Abstract
Collaborative perception uses complementary observations from multiple agents to mitigate occlusion and limited sensing coverage. In practice, however, systems are often developed independently: they may use different encoders, learn different representations even under the same architecture, or rely on different fusion mechanisms. As a result, features produced by one system may not be directly usable by another. We therefore investigate whether such systems can establish a compatible feature interface from their independently generated capability descriptions. We cast this problem as feature interface compilation and propose Rosetta, a framework for it. Each system builds a Capability Card from its own model and data alone. The card records how feature restrictions affect the sender’s detection performance and how sensitive the receiver’s detection is to perturbed features; together these give complementary cues for choosing an interface. When two systems meet, a learned compiler reads both cards to select and configure an interface program, whose pair specific parameters are then fitted without labels on a small set of public frames. The sender’s features are delivered only if the fitted mapping passes a feature level check; otherwise the receiver detects alone. Both systems stay frozen: each keeps its own representations and fusion module, and the interface translates features between them rather than imposing a common representation space. On OPV2V, direct collaboration degrades the receiver’s detection in our heterogeneous settings, whereas Rosetta improves detection on average for LiDAR systems that differ in training seed, fusion module or encoder, including systems unseen in compiler training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.