Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous Prompts
Daniel KhashabiXinxi LyuSewon MinLianhui QinKyle RichardsonSean WelleckHannaneh HajishirziTushar KhotAshish SabharwalSameer Singh
Reveals that continuous prompts optimized for language models can be projected onto completely arbitrary or contradictory discrete text without losing task performance, exposing critical flaws in standard prompt interpretability methods.
Prompt tuning optimizes continuous numerical vectors rather than updating an entire language model's parameters, offering a lightweight approach to controlling AI systems. However, these continuous prompts are not inherently human-readable. Practitioners often attempt to interpret them by mapping continuous vectors to their nearest discrete words. As organizations increasingly rely on prompt-based AI for critical tasks, understanding whether these discrete textual interpretations faithfully reflect how models actually behave is essential for safety, governance, and model reliability.
The main objective of the article is to evaluate whether continuous prompts can be faithfully interpreted into human-readable text by mapping them to their nearest discrete words. The authors formulate and test the Prompt Waywardness hypothesis, which posits a fundamental disconnect between what a continuous prompt instructs a model to do and what its nearest-neighbor discrete text projection appears to say.
To evaluate this behavior, the authors conducted empirical experiments using the GPT-2 language model family across five distinct classification tasks. They modified standard continuous prompt tuning to jointly optimize downstream task accuracy while pulling the continuous prompt toward an arbitrarily chosen target text. The evaluation tested 62 distinct target texts—including instructions from unrelated tasks and random sentences—measuring both task accuracy retention and textual overlap. They also evaluated scaling across model sizes from 124 million to 1.5 billion parameters and tested various prompt lengths.
The findings reveal that continuous prompts can be optimized to solve a given task while simultaneously projecting to virtually any arbitrary target text, retaining performance within 2% of unconstrained prompts while achieving over 94% text overlap with the arbitrary target. Furthermore, continuous prompts forced to project to the true definitions of tasks showed no meaningful performance advantage over those projecting to completely irrelevant text, with both exhibiting a similar 1.5% average accuracy drop compared to unconstrained baselines. The disconnect between continuous prompts and their textual interpretations becomes more pronounced as model size increases and prompt length grows, due to the high expressive power in early network layers and the mathematical reality that infinitely many continuous vectors map to the same discrete token.
These results demonstrate that projecting continuous prompts to nearest-neighbor words does not yield a faithful explanation of model behavior. This creates severe security and governance risks, as malicious or biased model behaviors can be concealed beneath benign, harmless-looking textual descriptions, creating a false sense of security for auditors. Additionally, the findings indicate that attempting to discover human-readable discrete prompts purely through unconstrained continuous optimization will often yield degenerate, unfaithful solutions.
Organizations and developers should avoid relying on nearest-neighbor word projections to interpret or audit continuous prompts. AI teams developing interpretable prompting methods must incorporate domain-specific constraints rather than depending solely on continuous gradients. Further research should focus on developing novel architectural mechanisms and robust explanation frameworks that bridge the continuous-discrete representation gap.
The conclusions are supported by consistent results across multiple benchmark tasks, prompt lengths, and model sizes within the auto-regressive GPT-2 architecture. Readers should note that these findings specifically reflect standard nearest-neighbor projection techniques within existing transformer frameworks; evaluating alternative projection methods or newer model architectures remains an important area for future confirmation.
- Paper: The Power of Scale for Parameter-Efficient Prompt Tuning, Brian Lester et al. (2021). This foundational study introduces continuous soft prompts and their scaling behavior, the prompt-tuning setup that the source tests for interpretability.
- Paper: P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks, Xiao Liu et al. (2022). Its account of P-Tuning v2’s continuous prompts across layers supplies key context for the source’s investigation of how learned prompt vectors relate to text.
- Paper: GPT Understands, Too, Xiao Liu et al. (2021). This work develops P-Tuning’s trainable continuous prompt embeddings, providing the method-level background for the source’s analysis of prompt interpretation.
- Paper: Noisy Channel Language Model Prompting for Few-Shot Text Classification, Sewon Min et al. (2022). Its GPT-2 experiments on parameter-efficient prompt tuning provide useful context for the source’s evaluation of continuous prompts on classification tasks.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). By showing that chain-of-thought rationales can conceal the cues driving a model’s answer, this study extends the source’s warning that readable text need not faithfully explain model behavior.
- Paper: Prompting is not a substitute for probability measurements in large language models, Jennifer Hu et al. (2023). This later study extends the critique of prompt-based interpretation by testing whether models’ verbal judgments faithfully reflect their internal probability distributions.
