Beyond Prompt Recovery: A Paired Diagnostic Framework for Behavioral Reuse
Abstract
Recovering recognizable prompt text leaves open which tested responses that text contributes on redeployment. We propose a paired diagnostic framework linking recovered content to gains over blank and random controls, gaps to source-prompt scores, and responses to source-rule edits. We study a fixed four-query recovery procedure across seven checkpoints using 100 controlled prompts and 80 public tests, of which 79 support behavioral evaluation. Recovered prompts add 14.04 percentage points over controls while remaining 23.43 points below the source configuration. A selector comparison finds substantial improvements in controlled content recovery alongside an uncertain change in incremental probe behavior. Rule interventions on two checkpoints yield positive multiplicity-adjusted intervals for two of four primary contrasts, while replacement effects adjusted for unchanged probes remain uncertain. Response traces and scorer sensitivities identify how these measured gains arise. The resulting protocol supports auditable diagnosis of same-checkpoint probe-score reuse.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.