From Language Reasoning to Cross-Modal Sentiment Evidence: Recursive Harness Search for Multimodal Sentiment Recognition
Abstract
The performance of multimodal sentiment analysis depends not only on the model itself, but also on how information is organized and on how the model is coordinated with external reasoning. Recursive self-improvement (RSI) provides an automated route to this by iteratively searching, executing, and modifying executable programs around a base model. However, existing RSI work has focused almost exclusively on text-based reasoning tasks. Unlike purely textual settings, sentiment cues in multimodal sentiment analysis are distributed across text, audio, and visual modalities, where they may complement or conflict with one another and vary from sample to sample. Consequently, the acquisition and utilization of sentiment cues must be searched jointly with the reasoning strategy and carried by the same executable program. To operationalize this view, we present SentiHarness, a recursive Harness search framework that jointly edits sentiment-cue acquisition, the organization of modality-specific evidence across model calls, cross-modal evidence use, and reasoning within one executable program. With the inner-loop model fixed, a matched CH-SIMS run improves the validation frontier from 81.14% under Fixed-Exposure Reasoning Search to 89.0% when multimodal evidence handling is made searchable. We further find that Harness effectiveness depends strongly on the state of the inner-loop model: after Plain SFT, the marginal benefit of the same Harness can shrink substantially or even reverse, revealing a concrete model–Harness interaction. Motivated by this coupling, we further explore Harness-aware SFT, which improves retained-Harness Acc-2 from 88.18% with Plain SFT to 90.37% under a matched-budget CH-SIMS training setting. These results identify executable multimodal evidence handling and model–Harness interaction as two central dimensions of recursive Harness optimization for multimodal sentiment analysis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.