acceptodds
Under review as a conference paper at ICLR 2027

Multi-scale intrinsic dimension estimators for text embeddings

Abstract

Persistent homology dimension (PHDim) characterizes the geometry of text embeddings, but its single value combines structures at different scales. We introduce scale-resolved PHDim based on quantile-trimmed minimum spanning tree energies and prove consistent dimension recovery for every fixed quantile interval under regular sampling assumptions. In contextual token embeddings, we identify an approximate cell–skeleton geometry: repeated token types form local cells connected by a type-level skeleton. Building on this structure, we propose a three-band estimator with a fixed sampling budget for stable, informative comparisons between texts. Controlled interventions and synthetic experiments link its responses to contextual variability, lexical diversity, and the organization of longer-range geometry. Scale separation reveals a more consistent signal for AI-text detection: classical PHDim can be either higher or lower for generated texts, whereas removing the lexically dominated lower band yields a consistently nonnegative AI–human difference across the tested generators. Beyond AI-text detection, scale separation provides a more informative characterization of text complexity and defects. Different defects produce distinct three-band profiles: recognition errors predominantly raise fine and middle dimensions, repetitive boilerplate lowers middle and coarse dimensions, and disrupted semantic coherence yields a different pattern from phrase repetition. Together with scale-dependent differences in writing proficiency, these profiles enable a more granular analysis of text collections than a single dimension estimate.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.