Align Before Aggregation: Component-Conditioned Evidence Alignment for Multimodal Financial Forecasting
Abstract
Different temporal dynamics within a market window may require different news, yet text-enhanced forecasting typically organizes evidence at the stock or window level. While frequency-aware methods exploit multiscale structure to improve representations and fusion, we examine whether evidence should be aligned separately to temporal components before aggregation. Under dot-product matching and linear query aggregation over a shared candidate set, we show that global alignment produces a normalized weighted geometric mean of component-level alignment distributions, allowing opposing evidence preferences to cancel. Motivated by this observation, CANER follows an align-before-aggregation principle. It decomposes the observed price–volume window into and high-frequency components and calibrates their queries using shared semantic anchors constructed by deterministically verbalizing observable market states. Each component aligns with and selectively absorbs evidence from the same historical news pool before the enhanced states are aggregated for prediction. On CMIN-US, CMIN-CN, and FNSPID, CANER achieves the highest ACC in all nine dataset–horizon settings, exceeding the strongest baseline by 0.34 percentage points on average, and the highest MCC in seven settings. Comparisons with matched parameter counts and news-slot budgets, together with cross-component mismatch analysis, support the predictive value of preserving component–evidence correspondence before final aggregation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.