acceptodds
Under review as a conference paper at ICLR 2027

ProCrit: Self-Elicited Multi-Perspective Reasoning with Critic-Guided Revision for Multimodal Sarcasm Detection

Abstract

Multimodal sarcasm detection requires reasoning over cross-modal incongruities between literal expression and intended meaning, yet the analytical perspectives needed to uncover these incongruities vary across samples. While recent methods make this analytical process explicit, they still rely on fixed, predefined perspectives that operate independently under hand-crafted routing rules. We argue that multimodal sarcasm detection instead calls for self-elicited multi-perspective reasoning, where a model autonomously generates the perspectives needed for each sample and progressively integrates them into a coherent analysis. To realize this goal, we propose ProCrit, a Proposal–Critic two-agent framework with a proposal agent for multi-perspective reasoning and a critic agent for external evaluation and targeted revision guidance. First, to overcome the lack of process-level supervision in existing sarcasm datasets, we synthesize reasoning annotations through a dynamic-role agentic rollout, in which a strong vision-language model sequentially spawns analytical roles within a shared context; the resulting trajectories are flattened into single sequences that preserve cross-perspective dependencies while enabling efficient autoregressive generation. Second, ProCrit adopts a draft–critique–revise paradigm in which the proposal agent dynamically elicits and progressively integrates sample-specific perspectives into a coherent analysis, while the critic agent evaluates the resulting reasoning and provides targeted natural-language feedback for directed revision. Finally, we develop a reciprocal training framework that jointly optimizes the proposal agent's drafting and feedback-guided revision via dual-stage reinforcement learning, while refining the critic agent according to the actual effectiveness of its feedback. Extensive experiments on three benchmarks demonstrate that ProCrit consistently outperforms reasoning-based methods, and human evaluation confirms substantially better-grounded rationales.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.