acceptodds
Under review as a conference paper at ICLR 2027

Where Does Protein Structure Emerge? Layer-Wise Native 3D Organization Across Protein and General Language Models

Abstract

Protein language models (PLMs) are trained primarily from amino-acid sequences, yet their representations support tasks closely tied to protein structure. We ask where, at what scale, and how consistently native three-dimensional organization becomes visible across network depth. We study dedicated PLMs spanning Transformer, convolutional, and state-space architectures, together with general-purpose language models and protein-oriented checkpoints. We analyze structure from two complementary perspectives: how the geometry of whole-protein representations changes across layers, and whether residues that are nearby in native 3D space remain nearby in representation space. We find strongly depth-dependent and model-specific organization. Whole-protein shape complexity often compresses across depth, but its trajectory varies across model families. In dedicated PLMs, Local Topological Fidelity (LTF) consistently peaks before the final layer, whereas general-purpose LLMs show more heterogeneous behavior. Secondary-structure conditioning further reveals that intermediate-scale neighborhoods (–) centered on helical residues exhibit the strongest fidelity. Finally, LTF enables downstream-label-free, model-adaptive layer selection, with particularly strong utility on structure-sensitive evaluations. Across ten PLMs and four evaluations, it is the only downstream-label-free selector in the statistically top-ranked group alongside supervised layer-selection methods. The selected region remains stable across structural reference sets, supporting reuse without task-specific layer search.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.