FARE: Fingerprint-Aware Representative Evaluation for LLM-Based Heuristic Discovery
Abstract
LLM-based heuristic discovery iteratively proposes, scores, and selects programs using feedback from an evaluation harness. The harness does not merely measure quality; it defines the objective that discovery optimizes. Evaluating every deployment scenario is infeasible because operating conditions vary continuously and interact combinatorially, so discovery must optimize against a finite scenario bank. Over successive iterations, the discovery loop can exploit gaps in this bank, so its high-scoring programs may fail under unseen conditions. Enlarging the bank raises the cost of evaluating every candidate program and may waste budget on behaviorally redundant scenarios. We introduce FARE (Fingerprint-Aware Representative Evaluation). FARE runs diverse probe programs offline on a scenario pool and fingerprints each scenario by its normalized probe scores. It then selects a compact training bank of scenarios by -medoids under program disagreement distance, so that every pool scenario has a nearby selected scenario as its representative. FARE reserves the unselected scenarios as a held-out test bank that discovery never sees, so the train–held-out gap measures generalization beyond the training bank. FARE also reports the representation error: the mean distance from each pool scenario to its representative. When a program's score is a smooth function of the fingerprint, the representation error bounds the average error of predicting a pool scenario's score from its representative's score. We evaluate FARE against four alternative bank-construction methods on two case studies, scheduling and congestion control, using 5G-Eval, our automated full-stack 5G platform built on real network software. Across three LLMs, FARE's top-ranked candidate program has the highest pool score in almost all model–pool combinations compared with the four alternatives.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.