acceptodds
Under review as a conference paper at ICLR 2027

Relation-Aware Visual–Physiological Affect Regression with an Evidence-Prefixed Large Language Model

Abstract

Multimodal affect regression aims to predict continuous emotional states by combining complementary yet imperfect evidence from observable behavior and physiological bodily responses. Facial video captures visible expressions which can be subtle or intentionally regulated, while blood‑volume pulse (BVP) measurements are susceptible to motion artifacts, sensor contact noise, and varying individual physiological baselines. We introduce a visual–physiological affect dataset of concurrently‑recorded, temporally paired facial‑video and BVP streams, with participant‑specific, continuous‑valued valence–arousal ratings collected for each stimulus presentation. We further propose an evidence‑prefixed large‑language‑model architecture for direct numerical affect regression. Visual and BVP encoders produce modality‑specific token sequences, and a continuous learned relation token summarizes their learned‑space interactions through modality summaries, absolute difference, element‑wise product, and cosine discrepancy. These evidence tokens precede a learnable valence‑arousal query in a QLoRA‑adapted Qwen model, whose final query hidden state directly predicts valence and arousal. Optimized solely via affect supervision, the relation token requires no manually annotated modality‑relation or signal‑quality labels. We perform controlled comparisons under a subject-dependent, source-trial-disjoint split and report window-level results for unimodal, non-LLM visual–BVP fusion, and no-relation variants. The relation token represents geometric discrepancy between learned representations rather than verified psychological conflict.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.