FastText.zip: Compressing text classification models
Armand JoulinEdouard GravePiotr BojanowskiMatthijs DouzeHervé JégouTomas Mikolov
Proposes a product quantization approach to compress text classification models by two orders of magnitude with minimal accuracy loss, enabling fastText deployment on memory-constrained devices.
Text classification is vital for applications such as search ranking, email sorting, and spam filtering. However, standard high-performing models often require gigabytes of memory to store large vocabularies and embedding matrices, making on-device deployment difficult on resource-constrained platforms like smartphones.
The article demonstrates an effective compression pipeline that shrinks text classification models by multiple orders of magnitude while preserving competitive accuracy and fast processing speeds.
The researchers evaluated their approach on eight standard benchmark datasets and a large-scale hashtag prediction dataset with over 300,000 output labels. They combined several techniques: product quantization (a method that splits vectors into smaller pieces to store them compactly), separate encoding of vector magnitude and angle, vocabulary pruning based on feature importance and training-set coverage, and dictionary hashing.
The key findings show that standard product quantization compresses models by a factor of 10 without meaningful accuracy loss. When paired with structured vocabulary pruning and hashing, the pipeline achieves extreme compression—reducing model footprints by a factor of 1,000 to 4,000 (often shrinking models from over 100 megabytes down to less than 64 kilobytes) with an average accuracy drop of less than one percent. Retraining the output layer after compressing input matrices compensates for quantization distortions, and max-coverage pruning maintains strong predictive performance on massive datasets where naive pruning methods fail.
These results demonstrate that engineering teams can deploy lightweight, accurate language classifiers locally on edge hardware, substantially reducing server infrastructure costs, cloud latency, and memory overhead compared to heavier neural network alternatives.
Decision-makers should consider adopting this compression pipeline for embedded or mobile text processing tasks. If deployment demands extreme size reduction, teams should test model-specific cut-offs to ensure required feature coverage. Future work could explore scaling vector dimensions based on feature frequencies and splitting low-value words into character sub-units to further optimize performance on short inputs.
- Paper: Bag of Tricks for Efficient Text Classification, Armand Joulin et al. (2017). It introduces the baseline fastText linear text classification architecture and n-gram embedding framework that FastText.zip directly compresses using product quantization.
- Paper: Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding, Song Han et al. (2015). It introduces fundamental model compression techniques combining pruning, quantization, and coding that form the conceptual foundation for compressing embedding models.
- Paper: Feature hashing for large scale multitask learning, Kilian Q. Weinberger et al. (2009). It provides foundational principles for feature hashing and dimensionality reduction in large-scale text classification that motivate FastText.zip's hashing-based compression approaches.
- Paper: Distributed Representations of Words and Phrases and their Compositionality, Tomas Mikolov et al. (2013). It details foundational techniques for learning dense distributed representations and word embeddings that serve as the primary memory bottleneck addressed in FastText.zip.
- Paper: Accelerating Large-Scale Inference with Anisotropic Vector Quantization, Ruiqi Guo et al. (2020). It builds upon product quantization for large-scale embeddings by introducing an anisotropic, score-aware quantization loss to accelerate inference and retrieval.
- Paper: A Survey of Quantization Methods for Efficient Neural Network Inference, Amir Gholami et al. (2021). It provides an extensive survey synthesizing modern quantization paradigms, theoretical trade-offs, and inference acceleration strategies across deep learning architectures.
- Paper: To prune, or not to prune: exploring the efficacy of pruning for model compression, Michael Zhu et al. (2017). It systematically investigates the trade-offs between pruning and architectural downsizing under strict memory budgets across language and vision tasks.
- Paper: Stop Indexing at Full Precision: Revisiting Clustering for Vector Embeddings, Leonardo Kuffo et al. (2026). It explores aggressive quantization and clustering techniques directly on vector embeddings to optimize memory footprint and search efficiency.
