Which Post Will Be More Engaging? Discovering Evaluation Criteria via LLM-based Reflection
Abstract
Predicting the relative popularity of social media posts plays an important role in understanding the relationship between content and user interactions, while informing content creation, candidate selection, and evaluative feedback for generative models. LLM-based evaluation and feedback-driven prompt optimization support text assessment and iterative instruction refinement, yet historical popularity outcomes provide only indirect guidance for revising evaluation criteria. The combined influence of content, creators, and platform distribution makes individual prediction errors an ambiguous basis for revision, increasing the risk of learning example-specific associations rather than reusable criteria. We introduce Grounded Reflection for Automated Criteria Exploration (GRACE), a framework that leverages LLM-based reflection to discover evaluation criteria from historical engagement feedback. GRACE examines discrepancies between model judgments and observed relative popularity across examples to refine and select topic-specific criteria for predicting the relative popularity of new posts. Across six topic-specific datasets, GRACE achieves 60.43% selective directional accuracy at 84.49% directional coverage, improving on the initial task prompt by 8.46 percentage points and outperforming the general prompt optimization methods evaluated in this study. These results support using historical engagement feedback to improve text-based relative popularity prediction, providing a methodological foundation for evaluation tools designed to aid content creation and selection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.