acceptodds
Under review as a conference paper at ICLR 2027

PTP: Previous-Token Prediction based LLM Inversion for Near-Exact Prompt Reconstruction

Abstract

Large language models (LLMs) generate text by auto-regressively sampling the next token. This inherently leads to a many-to-many mapping between prompts and responses, complicating the task of inferring prompts from observed outputs. Prior work on LLM inversion frames prompt recovery as a semantic reconstruction task. They rely on fine-tuning pretrained sequence-to-sequence models on large external datasets with access to model weights or logits to generate semantically plausible prompts. In contrast, we present a functional approach to inverting a given LLM in a black-box setting by training an explicit inverse language model entirely from scratch on data synthetically generated from the target LLM itself. Analogous to forward next-token prediction, our inverse model is trained using previous-token prediction (PTP), establishing a generative link between the forward and inverse processes that enables faithful prompt reconstruction. Moreover, it naturally supports diverse prompt reconstructions through sampling, whereby all such prompts induce similar responses under the forward, target LLM. Our approach generalizes across datasets and exhibits semantic transferability in reconstructing prompts from responses generated by different LLMs. PTP substantially outperforms prior work on exact prompt recovery with tokenizer access while remaining competitive on semantic reconstruction metrics in fully black-box setup.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.