acceptodds
Under review as a conference paper at ICLR 2027

LatentNutri: Anchor-Driven Latent Computation for Efficient Food Nutrition Estimation in Vision-Language Models

Abstract

Image-based food nutrition estimation requires open-domain visual understanding and reliable continuous prediction. Existing systems typically rely on CNN/ViT regressors, which often generalize poorly beyond their training distribution, or on vision-language models (VLMs), which predict nutrition values through slow autoregressive text decoding. We introduce LatentNutri, an Anchor-Driven Latent Computation framework for VLM-based food nutrition estimation. LatentNutri uses Single-Pass Latent Forward to compute target-specific anchor states in one VLM forward pass, and Head-Aware Latent Regression to map these states to continuous nutrition values with lightweight regression heads. This design predicts a structured nutrition vector without autoregressive decoding. Training combines an autoregressive SFT warm-up with Scale-aware Physics-guided Regression, which jointly optimizes scale-normalized numerical accuracy and physical plausibility. Experiments on Nutrition5K and DiningBench show that LatentNutri improves prediction quality over CNN/ViT and autoregressive VLM baselines while achieving a 14.9× inference speedup over autoregressive decoding. Code is provided in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.