acceptodds
Under review as a conference paper at ICLR 2027

ARIAD: Adaptive Trajectory Matching for Prompt Reconstruction

Abstract

Distributed inference exposes intermediate states even when prompt text remains on the client. We introduce Adaptive Prompt Reconstruction from Internal Anchor Dynamics (ARIAD), which reconstructs a request by fitting and verifying a compact internal trajectory, without training an inverse model. An endpoint fit supplies provisional tokens; self-generated substitutions select internal anchors, and fixed target statistics define a shared metric for embedding optimization and causal token verification. We analyze the standardized cosine used by the attack: different depths can resolve different local ambiguities, while candidate omissions limit exact decoding. Controlled experiments separate conditioning, trajectory coverage, verification, and adaptive placement. Across five white-box initializations, ARIAD reaches .9806 accuracy and .6067 exact match, exceeding dual endpoints by .0069 and .0387. On 150 additional requests after independent validation selection, its accuracy is .9869 versus .9762 for a conditioned boundary; exact-match counts are 101 and 104. Shared-fit verification improves partial reconstruction without changing exact match, while larger proposal sets improve candidate coverage. Permitted-target black-box experiments further test adaptive placement under model mismatch. Experiments across domains, depths, model families, generated identifiers, and longer requests reveal both recovery gains and limits.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.