acceptodds
Under review as a conference paper at ICLR 2027

Predicting Performance of Symbolic and Prompt Programs with Examples

Abstract

A prompt program may succeed on a few test cases yet fail on new inputs at deployment time. We study performance prediction: given a program—either symbolic (e.g., Python) or a prompt executed on an LLM—and a few labeled examples, predict its success rate on unseen instances from the same task distribution. We model each pass/fail execution as a Bernoulli random variable whose success probability is the program's unknown performance. Under this model, the posterior over performance is determined by the observed execution outcomes and a performance prior. We estimate empirical performance priors from a corpus of programs and tasks. In this corpus, symbolic-program success rates concentrate near zero and one, whereas prompt programs have a more diffuse prior with many programs that pass most, but not all, evaluated tests. These priors explain why a few passing tests support greater confidence in near-perfect performance for symbolic programs than for prompt programs in the studied setting. We introduce RAP (Retrieved Approximate Prior), which retrieves related labeled examples and prompt programs from a corpus to construct a proxy prior. With few target observations, RAP achieves higher mean posterior density at held-out reference success rates and narrower 95% credible intervals than no-corpus and all-corpus baselines in the evaluated in-domain and out-of-domain settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.