Bolo: A Verified Model Hub with Ready-to-Run Inference Programs for Agents and Humans
Abstract
Both AI agents and human engineers can greatly benefit from the ability to use any public machine learning model. Unfortunately, most public models, including those on Hugging Face, are not usable out of the box, and manually writing inference programs does not scale because failures follow diverse, model-specific patterns. Even worse, existing coding agents often produce faulty or misleading inference programs that run successfully yet silently behave differently than intended. Addressing this requires more sophisticated, model-specific verification strategy that can identify the ways those generated programs deviate from user intentions. To this end, we introduce \bolo, a large-scale model hub for constructing and verifying model inference programs. \bolo verifies each generated program at two levels: behavioral checks that test whether outputs are valid for the model's task, and content-level verification that checks the implementation. Across five model-dependent tasks, agents using \bolo achieve higher-quality outputs than agents with direct Hugging Face access while reducing LLM cost by up to 26%. On 8,570 most widely downloaded Hugging Face models, \bolo covers 65.2%, compared with 41.6% for Hugging Face Transformers \verb|pipeline| API, including 1,195 models outside its support. Finally, among 2,019 programs passing behavioral checks, \bolo's content-level verification judges 88.6% clean, compared with 57.2% for a binary LLM judge. Qualitative case studies further demonstrate that \bolo avoids unsupported rejections while identifying genuine implementation issues that output checks alone cannot detect.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.