acceptodds
Under review as a conference paper at ICLR 2027

Triage Before Fine-Tuning: Feasibility-Gated Selection and Monitored Adaptation of Embodied Foundation Models for a Fixed Physical Platform

Abstract

Embodied foundation models are increasingly deployed by selection: a practitioner picks a public checkpoint by benchmark rank and fine-tunes it on local data. The platform, however, is fixed before the model is chosen: its sensors, onboard compute and operating site are set by the task, and at the start there is little or no local data. We ask what such a checkpoint loses on a fixed physical platform and which part of the loss local data can buy back. Our answer is to split the gap by what can repair it. A sensing gap, information the sensors do not capture, cannot be closed with their data; a latency gap, output arriving after the control loop needs it, belongs to the loop, not the weights; only a distribution gap is fine-tuning's to repair. We turn this into TRIAGE, a procedure for feasibility-gated selection and monitored adaptation in five stages: gates on the sensor contract and on latency that no score can override, ten-minute local probes that place candidates in tiers, a target-relevant pick within the top tier, and fine-tuning monitored for tracking error, perturbation sensitivity and reliance on the image. In Bench2Drive, a 500 ms deadline below its latency makes the closed-loop champion collide on 80 to 90% of routes; on NAVSIM real logs, losing field of view moves plans 2.2 to 5.7 times more than losing resolution at equal pixels (1.6 to 2.4 on Bridge V2 manipulation), 20k local frames close 75% of a stability gap but raise perturbation sensitivity up to 3.4 times, and picking by score or by champion violates a gate in two and four of four deployment profiles. In a field test on a full-scale vehicle, the factory camera failed the sensing gate, the prescribed added camera repaired the view, laboratory latencies did not carry over on board, and in closed loop plans rested on images over two seconds old, the one stale-command takeover coming under command hold. More broadly, it lets organizations that cannot train foundation models from scratch find out, before buying data or hardware, whether a public checkpoint can run on the robot or vehicle they own, and what to fix if not.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.