GibberCraft: ADMM-Based System-Prompt Obfuscation for LLM Applications
Abstract
System prompts define the tasks, response contracts, and workflow rules of large-language-model applications. Their readable text is consequently a valuable implementation artifact, yet prompt extraction and reverse engineering can recover secret prompts from deployed services. We study whether we can replace a readable system prompt with a gibberish-like sequence of ordinary tokens that preserves the original behavior. This task is challenging because the replacement prompt must be optimized in a discrete token space. We formulate this task as constrained behavior-preserving optimization and leverage the alternating direction method of multipliers (ADMM) to separate continuous behavioral optimization from the hard vocabulary constraint. We further introduce Gumbel-perturbed hard-token sampling to account for hard-projection effects during optimization, improving optimization robustness. We evaluate behavioral fidelity, success rates, and protection against prompt injection across multiple tasks and model families. We achieve a 100% success rate on all tasks and match or outperform prior work, and achieve up to 9.3 speedup. These results demonstrate the feasibility of replacing readable system prompts with gibberish-like hard-token prompts for prompt protection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.