From Fixed Keys to Readable Schemas: Small Language Models for Vehicle Agent Function Calls
Abstract
In-vehicle assistants are beginning to act on what drivers say, turning a request such as “turn on cruise control” into a call to one of the vehicle's functions. This Vehicle Function Calling must run on the vehicle's own hardware, which calls for small language models (SLMs), and must abstain when no available function fits, because a wrong call acts on hardware. Since the functions differ across vehicles and change with software updates, the key design question is how to represent them to a small model. A Functional Token (FT) model learns one token per function, whereas a Schema-in-Prompt (SIP) model reads the offered functions' schemas in its prompt. Vehicle Function Calling is a new domain: prior vehicle work uses FT on private data, and no public benchmark compares the two. We answer the question in three connected steps. A formal analysis proves that FT can neither call an Unseen Function, one never used as a training target, nor base its abstention on which functions are offered, whereas SIP can do both. Because abstention needs a guarantee, we calibrate it with split conformal risk control. We release a benchmark from the Android Automotive specification and train 10 SLMs under two billion parameters with matched training for each representation. FT scores zero on Unseen Functions, while SIP reaches 15.9 to 76.6%. When a function trained as a target is not offered, only SIP, which reads the offered set, abstains reliably, though it is much slower to its first call. Across model families, size alone does not order Unseen accuracy. The conformal bound holds on random splits, but not when the requested functions shift. How functions are represented, more than model size, decides what a small in-vehicle model can call.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.