MODEPrompt: Taxonomy-Guided Pareto Evolution for Prompt Optimization
Abstract
Large language models (LLMs) are highly sensitive to prompts specifying task instructions, making automatic prompt optimization (APO) important for reliable performance. Existing APO methods typically guide search using scalar validation scores, which can obscure performance imbalances across heterogeneous reasoning patterns and produce prompts with competitive overall accuracy but low worst-group accuracy. We propose MODEPrompt, a reflective multi-objective evolutionary framework that induces a task-specific taxonomy of mutually exclusive reasoning subgroups from development questions without using answer labels, and treats subgroup accuracies as distinct Pareto objectives. MODEPrompt introduces semantic differential evolution, which contrasts Pareto-front and non-front prompts to extract directional semantic patches that guide mutation and crossover. It further induces outcome-grounded meta-rules from trial acceptance outcomes and subgroup-level score changes, accumulating useful search experience across iterations. Extensive experiments across diverse benchmarks show that MODEPrompt achieves strong overall performance, improves worst-group accuracy, and reduces performance disparities across reasoning subgroups. Further analyses support the stability, semantic interpretability, and performance relevance of the induced taxonomy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.