Beyond Correctness: Benchmarking Response Behaviors and Improving Non-Thinking Outputs in Hybrid-Thinking MLLMs
Abstract
Hybrid-thinking multimodal large language models (MLLMs) support thinking and non-thinking modes, but answer correctness alone does not capture the quality of their final responses. We introduce PatternEval, a failure-enriched diagnostic benchmark of 2,415 multimodal prompts across nine task categories. It evaluates correctness separately from four response-pattern failures: chain-of-thought leakage, response repetition, logical contradiction, and performative reasoning. Across 17 evaluated model configurations, non-thinking response-pattern failure rates exceed those in thinking mode by 2.86–48.64 percentage points, motivating direct evaluation of the response delivered through each interface. We further develop PatternRM, a reward model that supplies pattern-specific penalties to augment verifier-based group relative policy optimization (GRPO). Incorporating PatternRM into verifier-based reinforcement learning reduces non-thinking response-pattern failures while largely preserving accuracy on PatternEval. Comparison with an explicit length penalty reveals a trade-off: length regularization achieves lower aggregate failure rates, whereas GRPO with PatternRM achieves higher downstream scores while reducing failures relative to the verifier-based baseline. Together, our findings highlight the value of response-pattern metrics for evaluating reinforcement learning beyond task accuracy and demonstrate that integrating pattern-aware reward modeling with verifier-based training can improve both response quality and downstream task performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.