acceptodds
Under review as a conference paper at ICLR 2027

ModelLakeFishing: Efficient Retrieval over Million-Scale Model Lakes

Abstract

Open model lakes may contain millions of reusable models, making it costly to identify suitable models for a new dataset. This setting requires learning indexable representations of models from heterogeneous evidence, including models, training or evaluation datasets, model cards, and historical evaluations. An important challenge is that this heterogeneous evidence is sparse and concentrated around popular models and benchmarks. We present ModelLakeFishing, a model-retrieval framework for queries that specify a target dataset, a prediction task, and an evaluation metric. It consolidates model lake metadata and historical evaluations into a model–dataset evidence graph, learns model and query embeddings with a structure-aware graph encoder, and indexes model embeddings with Hierarchical Navigable Small World (HNSW) search. At query time, HNSW retrieves 1,000 candidates without scoring every model in the lake. A task-level prior computed only from training-side evaluations then reranks the candidates for the requested metric and returns the top 10 models. Our evaluated lake We evaluate the framework on a lake containing 3,016,439 models and 247,803 model–dataset–task performance edges. In our evaluation, the models and datasets remain in the graph, while selected historical performance records are held out to test whether their best-performing models can be retrieved. Across three such evaluation splits, ModelLakeFishing achieves a mean eligible-query of 0.2968: the best model observed in the held-out records appears among the ten returned models for 29.68% of eligible queries. This retaisn 93.47% of the achieved by an exact baseline that applies the same combined scoring procedure to every model in the lake. Given precomputed query embeddings, candidate retrieval and reranking take 0.747ms at the median and 1.102ms at the 95th percentile. These results demonstrate the feasibility of efficient model retrieval across a multi-million-model lake under sparse supervision.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.