HOLOGRAM: A Graph-Conditioned Context Steering Framework for Multilingual, Multimodal, and Multi-Granular Toxic Content Detection
Abstract
In this work, we introduce a novel graph-conditioned steering framework HOLOGRAM, which is among the first content moderating unified architecture that adapts a frozen VLM to these three axes of moderation heterogeneity through the same mechanism. For each post, HOLOGRAM constructs a semantic moderation neighborhood from training. Subsequently, a graph-based-encoder models how nearby examples support, duplicate, or contrast with one another. Its representations select representative artifacts to generate a query-conditioned low-rank context vector inside every frozen decoder layer of the VLMs. The resulting framework supports multiple modalities, languages, and supports labels at both, binary and fine-grained levels. We evaluate our framework across an array of established vector (and memory)-steering methods and prove that our graph-conditioned steering framework not only outperforms prior methods, but also yields high consistency. The results demonstrate the potential of local relational residual context addition as a general interface for heterogeneous social-media moderation and proves its strong usability, even towards production deployment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.