acceptodds
Under review as a conference paper at ICLR 2027

Structured Auxiliary Supervision for Continuous Imbalanced Financial Sentiment Analysis

Abstract

Financial sentiment regression with large language models (LLMs) trains on neutral-dominated data in which the extreme scores that drive financial decisions are rare, so fine-tuned models pull their predictions toward the dense center, a failure mode we term neutral collapse. Existing remedies reweight samples by label density or classify over a partition fixed before training, leaving the category structure of the continuous label space unlearned, and an auxiliary classification task added to supply it interferes with regression and converges at a different rate, so no fixed task weight serves both. We propose Structured Auxiliary Supervision (SAS), a plug-and-play framework that makes the label partition a trained parameter and schedules its supervision by the convergence trajectory of each task. Differentiable Structured Partitioning (DSP) learns thresholds that split the sentiment space into balanced auxiliary categories through differentiable soft assignments, giving the sparse extremes their own supervision targets. Trajectory-based Convergence Balancing (TCB) sets the task weights from windowed convergence velocities, shifting gradient toward the slower-converging task, and weights the auxiliary categories by a softmax over their losses. The two components are coupled in one loss and one backward pass, and the auxiliary head is discarded after training, so SAS adds no inference cost to the backbone. On three financial datasets and three backbone scales, fixed-weight multi-task fine-tuning (Fixed-MTL) is worse than Single-task regression in all nine settings, whereas SAS attains the lowest MSE in all nine, 10.5% below Fixed-MTL, a reduction that stays between 9.9% and 11.6% across a 24-fold increase in backbone size. On RoBERTa-base it has the lowest MSE among twelve baselines under identical protocols, and on public financial and cross-domain benchmarks it lowers MSE by 9.4% to 10.7%; DSP alone yields 52% of the joint MSE gain and TCB alone 30%. On NEU/RoBERTa-base, SAS compresses the extreme-to-neutral MSE ratio from 3.85 to 1.95, and its advantage widens as imbalance grows: under 20 label skew, its MSE rises by 25% versus 80% for Single-task regression and 92% for Fixed-MTL.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.