LAM++: Enhanced LocateAnything via High-Resolution Post-Training and Geometry-Aware Speculative Decoding
Abstract
Recent vision-language models such as the LocateAnything Model (LAM) have enabled universal localization through generative parallel box decoding, yet they still have gaps with strong dedicated vision models such as SAM3 on dense scenes and small targets. We present **LAM++**, an enhanced LAM with higher accuracy and faster speed via post-training and speculative decoding. We first identify dense scenes and small targets as major bottlenecks, and construct a high-quality corpus of over 20M localization examples. A failure-driven SFT recipe that integrates high-resolution visual processing with single-query training is used to substantially improve the localization accuracy on all domains, and a sequence-level GRPO strategy is used to further strengthen performance. We then introduce a geometry-aware speculative decoding (GeoSpec) framework that reuses LAM's native parallel decoder (MTP) for drafting and employs the autoregressive (AR) policy for verification. Unlike exact token-level verification of speculative decoding, GeoSpec exploits the geometric structure of localization outputs, accepting spatially equivalent coordinate proposals under a relaxed criterion and invoking AR correction only when necessary. This yields a favorable accuracy–efficiency Pareto frontier: GeoSpec achieves a 68% speedup with ΔmF1 ≤ 0.1. Across a diverse suite of localization benchmarks that span different task formulations and visual domains, LAM++ achieves state-of-the-art performance, outperforming SAM3 as well as larger or closed-source VLMs. We will release our models, the full training corpus, and complete post-training recipes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.