acceptodds
Under review as a conference paper at ICLR 2027

Tracing Object Evidence Across Depth: Text-Modulated Memory Framework for Weakly Supervised Semantic Segmentation

Abstract

Weakly supervised semantic segmentation (WSSS) learns pixel-level predictions from image-level category labels. While CLIP provides rich semantic priors for localization, recovering complete objects requires integrating complementary evidence across visual layers. Conventional multi-layer aggregation lacks an explicit mechanism for retaining and revising concept-specific evidence as semantic context deepens, leaving local structures weakly tied to category semantics and producing incomplete object coverage and spurious background responses. We propose a text-modulated cross-layer xConvLSTM framework that formulates visual-layer integration as concept-conditioned memory evolution across network depth. Treating selected CLIP layers as an ordered sequence, the framework uses text representations to control how spatial memory incorporates, retains, and suppresses visual evidence, progressively connecting local structures and object parts to category semantics. To adapt this process to each image while preserving CLIP's semantic priors, we introduce anchored adaptive prompting, which augments frozen text embeddings with bounded learnable and image-conditioned residuals. The resulting concept-specific memories calibrate visual tokens, guide decoding, and refine pixel relations, coupling evidence accumulation across depth with response propagation across space. Extensive experiments on WSSS benchmarks demonstrate more complete object localization, higher-quality pseudo-labels, and consistent segmentation gains.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.