Text Embeddings Reveal (Almost) As Much As Text
John X. MorrisVolodymyr KuleshovVitaly ShmatikovAlexander M. Rush
Demonstrates that dense text embeddings fail to preserve privacy by introducing an iterative decoding method that reconstructs exact text inputs and extracts sensitive personal information from clinical records.
Modern artificial intelligence systems frequently convert sensitive text documents into numerical vectors, known as text embeddings, and store them in hosted vector databases for fast retrieval and search. Organizations often operate under the assumption that sending only embeddings—rather than raw source documents—to third-party database providers safeguards data privacy. The article investigates whether an adversary with black-box query access to an embedding model can invert these numerical representations and fully recover the original text, thereby exposing significant data security vulnerabilities.
The main objective of the article is to demonstrate and evaluate a controlled generation method that systematically reconstructs original, full-text sequences from dense text embeddings without requiring direct access to the underlying model's internal weights or gradients.
To achieve this, the authors developed Vec2Text, a multi-step framework based on an encoder-decoder transformer architecture. Instead of relying on a single guessing step, Vec2Text iteratively generates hypothesis text, computes the difference between the hypothesis vector and the target vector, and applies discrete corrections across multiple rounds. The researchers evaluated the approach using 5 million training passages across popular commercial and open-source models, including Google's GTR-base and OpenAI's text-embeddings-ada-002, and tested generalization across 15 standard information retrieval datasets and a pseudo-reidentified clinical database.
The findings reveal that dense text embeddings leak substantial amounts of original text. For short passages of 32 words or word pieces, Vec2Text successfully recovered 92% of inputs with an exact match and reached an average sequence similarity score of 97.3 out of 100 on standard Wikipedia data, outperforming basic non-iterative models which achieved 0% exact recovery. On clinical notes, the method extracted 89% of full patient names and recovered 26% of documents word-for-word. While longer text sequences presented greater reconstruction difficulty, the model consistently captured underlying semantic content across all tested domains. Furthermore, initial defense simulations showed that injecting small amounts of calibrated Gaussian noise into embeddings degraded text reconstruction by nearly 87% while reducing search performance by only about 2%.
These results demonstrate that numerical text embeddings create virtually the same privacy and compliance risks as raw plaintext data. Organizations can no longer assume that transmitting embeddings to external vector database vendors preserves confidentiality. A security breach of an embedding repository could lead to direct leaks of personally identifiable information, intellectual property, or confidential clinical records, posing severe legal and compliance hazards under data privacy frameworks.
Organizations handling sensitive text should immediately treat text embeddings with the same rigorous data governance, access controls, and encryption standards applied to raw text. Engineering teams utilizing third-party vector databases should consider evaluating lightweight noise-injection defenses to mitigate straightforward reconstruction attacks, balancing this against minor decreases in search accuracy. System architects must also monitor query access to embedding APIs, as the reconstruction technique relies on repeated queries to refine text hypotheses.
Confidence in these findings is high for short-to-moderate text inputs up to 32 tokens, supported by consistent performance across multiple datasets and embedding models. Readers should note certain boundaries: the evaluation assumed an adversary possessed query access to the exact embedding model used to create the vectors, tested sequences limited to 128 tokens, and did not examine adaptive attackers trained specifically to counteract noisy defenses. Further analysis is required to determine the feasibility of inverting multi-page documents and developing defenses that withstand adaptive reconstruction techniques.
- Paper: The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks, Nicholas Carlini et al. (2018). Introduces foundational methods and metrics for measuring unintended data exposure and memorization extraction in neural sequence models.
- Paper: Extracting Training Data from Large Language Models, Nicholas Carlini et al. (2020). Demonstrates practical black-box data extraction attacks recovering sensitive personal information from language models, establishing the privacy vulnerabilities built upon in embedding inversion.
- Paper: Understanding deep image representations by inverting them, Aravindh Mahendran et al. (2014). Pioneers the core paradigm of understanding deep representation fidelity and information retention by inverting latent codes back to input space.
- Paper: Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval, Lee Xiong et al. (2021). Provides fundamental background on training dense contrastive text embeddings that map sequences into the continuous latent spaces inverted by the source.
- Paper: Deep Models Under the GAN: Information Leakage from Collaborative Deep Learning, Briland Hitaj et al. (2017). Examines early generative reconstruction and model inversion attacks targeting representation leakage in decentralized machine learning architectures.
- Paper: Language Models are Injective and Hence Invertible, Giorgos Nikolaou et al. (2025). Extends representation inversion by theoretically proving the injectivity of Transformer hidden states and developing exact, non-iterative token recovery algorithms.
- Paper: Stealing part of a production language model, Nicholas Carlini et al. (2024). Broadens model extraction and parameter inversion to commercial black-box APIs by recovering the projection layers and embedding dimensions of production LLMs.
- Paper: Large-scale online deanonymization with LLMs, Simon Lermen et al. (2026). Applies embedding-based user representations and language models to full-scale automated online deanonymization and identity matching.
- Paper: On the Theoretical Limitations of Embedding-Based Retrieval, Orion Weller et al. (2026). Analyzes the fundamental capacity and information-theoretic limits of single-vector dense text representations for retrieval.
- Paper: Extracting alignment data in open models, Federico Barbero et al. (2025). Leverages semantic embeddings and text generation techniques to systematically extract proprietary fine-tuning and alignment data from open language models.
- Paper: NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models, Chankyu Lee et al. (2025). Applies instruction-tuned decoder architectures to train state-of-the-art generalist text embedding models across diverse downstream tasks.
