Beyond Cartesian Scanning: Unifying Conflict and Geometry in State Modeling for Fine-Grained Medical Image Classification
Abstract
Fine-grained medical image classification depends on subtle local variations whose diagnostic significance is closely related to surrounding anatomical context. Effective recognition therefore requires refining weak discriminative features during contextual integration and organizing surrounding evidence through meaningful structural relationships. Although Vision Mamba efficiently captures long-range dependencies, global connectivity alone does not explicitly address these requirements. We propose Conflict Recalibration and Geometric Aggregation (CRGA), a framework that enhances pretrained VMamba through directional-response-guided feature refinement and circular–radial context aggregation. Specifically, Conflict-Conditioned Recalibration (CCR) compares spatially aligned horizontal and vertical scan responses before fusion and uses their discrepancies as conditioning information for subsequent representation refinement. Through a one-hop connection, this signal modulates selective state writing and readout in the next stage, enabling contextual integration to account explicitly for relationships between directional responses. To organize surrounding structural context, Geometric Circular–Radial Structured Aggregation (GCRSA) introduces a geometric prior inspired by angular–radial formulations in medical imaging. It organizes existing tokens into closed circumferential and open radial sequences around the feature-grid center, distinguishing contextual interactions within radial bands from those across radial distances. Both components retain the original feature grid and Cartesian scan order, while zero-initialized modulation projections and residual gates preserve the pretrained backbone mapping at initialization. Experiments on a clinically acquired EUS dataset show that CRGA improves VMamba accuracy by 2.2 percentage points, while evaluations on two public medical image benchmarks demonstrate competitive performance. Same-backbone ablations on EUS show accuracy gains from each component, with the combined model achieving the highest accuracy and F1 among the evaluated variants.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.