acceptodds
Under review as a conference paper at ICLR 2027

Should I Optimize Prompts for This Task? Measuring Task Suitability with CARAT

Abstract

Reflective prompt optimization, in which a reflector model revises instructions from task feedback, improves a language model on specific tasks efficiently and without updating its weights, but its outcomes are unstable. Repeated runs of the same setup with different random seeds can shift held-out accuracy considerably, and prompts with nearly the same accuracy can still solve different cases, so a single run's returned prompt says little about the task itself. We introduce CARAT, two task signals for a fixed answerer, reflector and dataset-represented task: CARAT-A measures whether prompt preferences persist across disjoint case samples, and CARAT-C measures whether the observed successes are compatible with a shared prompt. We also propose Multi-Root Full-Retention Sampling (MRFS), which explores the prompt space from independent roots under a fixed proposal budget and retains every valid proposal. Across ten tasks with Qwen3-8B as answerer and reflector, we validate the signals against collections from five prompt optimization methods over multiple seeds: both signals persist across the entire task distribution and are consistent across methods, and MRFS yields the task ordering the optimizers' collections share most. CARAT thus quantifies the task-suitability question that run-level evaluations leave open, putting prompt optimization results on a firmer footing. Code is available at https://anonymous.4open.science/r/carat-1B00.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.