acceptodds
Under review as a conference paper at ICLR 2027

Understanding Feature Learning in Graph Contrastive Learning: Neighborhood Correlation and Theory-Guided Augmentation

Abstract

Graph contrastive learning (GCL) learns representations by contrasting augmented graph views, where each node representation aggregates information from its local neighborhood. Despite the empirical success of GCL, how neighborhood structure shapes the feature-learning process remains theoretically underexplored. This paper provides a theoretical characterization of feature learning in GCL under a sparse coding model that separates task-relevant features with stronger neighborhood correlation from task-irrelevant features with weaker neighborhood correlation. Our analysis reveals a selective feature-learning mechanism: neighborhood aggregation preferentially amplifies features with stronger neighborhood correlation while suppressing less correlated features. We quantitatively characterize how feature learning and downstream performance depend on feature overlap, the fraction of correlated neighbors, and neighborhood size. Guided by these theoretical insights, we propose a plug-in augmentation strategy, Theory-Guided Neighborhood Augmentation (TGNA), that constructs and selects graph views using an empirical measure of neighborhood correlation. Experiments on synthetic and real-world graphs support our theoretical predictions and show that TGNA generally improves downstream performance across various GCL frameworks while reducing contrastive training cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.