Prompt Search Is Not Prompt Expressivity: Continuous Prompts Beat GEPA, MIPROv2 and SIMBA, and an Exact Vocabulary Frontier
Abstract
When a prompt optimizer fails to find a good text prompt, it is tempting to conclude that none exists. We separate the two questions and give each its own evidence. Empirically, under a protocol frozen before any test row was scored, a 32-vector continuous prompt on three frozen 3-4B models achieves higher test accuracy than GEPA, MIPROv2 and SIMBA in all 18 model-dataset-regime cells (margins 7.8-34.3 points; all 18 H1 hypotheses rejected by the pre-registered procedure, with Holm adjustment over the 36 H1/H2 hypotheses at nominal level 0.05), and its lowest-accuracy run exceeds the highest-accuracy individual optimizer run in 17 of 18 cells. The ordering holds with every arm scored by one fixed decoder, and when GEPA and MIPROv2 receive the selected continuous configuration's example budget (still 5.9-18.4 points behind in the three LOW cells tested). Audits on frozen records document the scoring differences and show that removing normalised exact train-test duplicates changes the gaps by at most 0.22 points; the comparison is one of resource-unmatched pipelines with different search objectives, so it measures what the pipelines achieve, not what text can express. Theoretically, a representation limit can nonetheless be exact. In a fully specified coupled Gaussian attention model, the risk of the best legal text prompt for a tanh target converges to q - theta^2 + (theta - s)_+^2 in the dictionary's reach s, uniformly over all query-independent masses, so for s < theta no sequence of prompt lengths removes the reach penalty; the model and proof are given in full. A controlled block checks a finite paired-response certificate and shows which premise carries the gap. In a pretrained model the vocabulary constraint is exact at the first layer, where the median ratio of token to soft directional reach over sampled readouts is 11.7-12.5%; over continuation classes satisfying our stated label-consistency and readout assumptions, such task-agnostic first-layer quantities alone cannot yield a positive task-loss lower bound.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.