Information-Guided Representation Learning from GUI Pixels for Misinformation Detection
Abstract
Most misinformation detectors operate on text and metadata, whereas users encounter posts through rendered interfaces that also expose account cues and selected replies. Screenshots make these cues available but entangle decision-relevant evidence with visual appearance. We study GUI-derived representations using source-aware structured fields, a frozen screenshot encoder, and trainable content, identity, appearance, and exposure branches. The model keeps selected cross-type interactions explicit and applies a consistency objective to low-level transformations that preserve the post and its visible evidence. To separate event generalization from repeated event content, we construct a post-grouped split of 1,878 PHEME and Twitter15/16 posts into 183 adjudicated global event groups. On 360 posts from 37 held-out groups, a structured model using GUI-extracted and legacy-derived fields reaches 71.67% Macro-F1, versus 64.81% for extracted-text features and 54.01% for a frozen screenshot embedding alone. Yet ordinary fusion, which leads under a random-post split, falls behind structured fields on unseen events; neither typed interactions nor appearance-pair consistency closes this gap. This reversal shifts the focus from fusing more GUI information to identifying which interface-derived evidence generalizes across events.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.