Quantization Beyond Uniform Bit Allocation

K. S. SreeramjiSabyasachi BasuRavishankar KrishnaswamyKirankumar ShiragurYujia Wang

article2026arXiv0 citations

Introduces a variable bit allocation framework that exploits the structural properties of Matryoshka embeddings to improve retrieval recall by up to 18% over standard uniform quantization under fixed memory budgets.

Listen

Modern machine learning systems rely heavily on vector representations, known as embeddings, to power large-scale search and recommendation platforms. However, modern embeddings often span thousands of dimensions, resulting in massive memory footprints that reach terabytes of storage for billion-scale datasets. While data compression techniques such as quantization reduce storage requirements by converting high-precision numbers into compact bit codes, traditional approaches allocate bits uniformly across all dimensions. This uniform treatment fails to account for modern embedding models that exhibit Matryoshka representation learning properties, where information is disproportionately concentrated in the leading coordinates.

The main objective of the article is to demonstrate that a variable bit allocation framework improves information retrieval accuracy over uniform quantization baselines under identical memory budgets. It evaluates how non-uniform bit distribution across embedding dimensions impacts nearest neighbor search across multiple standard retrieval benchmarks.

To evaluate this framework, the authors conducted computational experiments across six standard information retrieval datasets—including MS MARCO, DBpedia, and Quora—using embeddings generated by commercial models from OpenAI and Cohere. The authors partitioned embeddings into contiguous buckets and employed a data-driven greedy local search algorithm to assign bit budgets iteratively based on search accuracy improvements. The evaluations tested both Product Quantization and Scalar Quantization across constrained memory budgets ranging from approximately 0.17 to 1.0 bits per dimension, measuring accuracy using exact brute-force search over dequantized vectors to isolate quantization error.

The analysis revealed three key findings. First, variable bit allocation consistently outperformed standard uniform allocation at identical storage constraints across all evaluated datasets and models. Second, the performance advantages were most substantial in highly compressed, sub-1-bit-per-dimension regimes, achieving relative accuracy gains of up to 18% for Scalar Quantization and up to 8% for Product Quantization. Third, the data-driven greedy allocation automatically assigned a higher concentration of storage to the initial embedding dimensions, empirically validating that the framework successfully captures and exploits the underlying variance structure of Matryoshka-style representations.

These findings indicate that retrieval platforms can significantly reduce their hardware footprint and infrastructure costs without sacrificing search accuracy. By aligning compression schemes with the natural structure of modern embeddings, engineering teams can achieve better quality per byte than standard industry quantizers offer. However, because variable bit allocation departs from uniform memory layouts, hardware-level optimizations such as vector processing instructions face integration challenges that must be addressed during system deployment.

Organizations operating large-scale vector search systems should consider structure-aware compression as a viable pathway to reduce indexing overhead, particularly when operating under tight memory budgets. Before broad implementation, engineering teams should evaluate closed-form or heuristic allocation strategies, as the proposed greedy search process is computationally expensive to run over massive datasets. Further work is also needed to integrate variable quantization into production approximate nearest-neighbor graph indexes and to optimize custom data layouts for low-latency query processing.

Cover for Quantization Beyond Uniform Bit Allocation

Abstract

Quantization is a fundamental technique to handle the growing sizes of embeddings generated by modern models. Existing quantization schemes are largely embedding agnostic and allocate bits uniformly across dimensions. However, recent models produce embeddings with significant geometric structure. In this work, we investigate whether a variable bit allocation scheme can improve quantization quality under a fixed memory budget. We propose a simple variable bit allocation framework that partitions an embedding into contiguous buckets and allocates storage non-uniformly across them. Using a greedy allocation strategy, we instantiate this framework for both Product Quantization (PQ) and Scalar Quantization (SQ). We perform a series of experiments on embeddings known to have the Matryoshka property (MRL), and consistently observe that non-uniform allocations outperform uniform baselines at identical storage budgets. The largest improvements occur in the low-bit regime, where uniform allocation is particularly inefficient for MRL embeddings. At the same compression rates, variable allocation improves recall by up to 8% for PQ and up to 18% for SQ. Our results suggest a new direction for structure-aware compression and indexing techniques for large-scale retrieval systems.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 A Greedy Bit Allocation Scheme
  • 3.1 Matryoshka embeddings
  • 3.2 Greedy Allocation Framework
  • 3.2.1 Greedy Product Quantization
  • 3.2.2 Greedy Scalar Quantization
  • 4 Experiments
  • 5 Limitations and Future Work
  • References

Knowls

  1. Knowl 1 — Greedy Bit Allocation Framework for Variable Quantization

    algorithm

    The greedy bit allocation framework optimizes the distribution of storage across contiguous dimension segments of an embedding vector under a fixed global memory budget.

    Given an embedding of total dimension DD partitioned into KK contiguous buckets B1,B2,…,BKB_1, B_2, \dots, B_K with dimensions dk=∣Bk∣d_k = |B_k| such that ∑k=1Kdk=D\sum_{k=1}^K d_k = D, the algorithm searches for an allocation vector b⃗=(b1,…,bK)\vec{b} = (b_1, \dots, b_K), where bk≥0b_k \ge 0 denotes the budget in bytes allocated to bucket BkB_k.

    Input: Initial uniform byte allocation b⃗(0)=(binit,…,binit)∈RK\vec{b}^{(0)} = (b_{\text{init}}, \dots, b_{\text{init}}) \in \mathbb{R}^K, step-size increment δ∈N\delta \in \mathbb{N} (bytes), validation set V\mathcal{V}
    Output: Optimized byte allocation vector b⃗=(b1,…,bK)\vec{b} = (b_1, \dots, b_K)
    b⃗←b⃗(0)\vec{b} \leftarrow \vec{b}^{(0)}
    repeat
        k∗←arg⁡max⁡k∈{1,…,K}R(b⃗+δe⃗k,V)k^* \leftarrow \arg\max_{k \in \{1, \dots, K\}} R(\vec{b} + \delta \vec{e}_k, \mathcal{V})
        b⃗←b⃗+δe⃗k∗\vec{b} \leftarrow \vec{b} + \delta \vec{e}_{k^*}
    until target total memory budget is achieved
    return b⃗\vec{b}

    The search initializes from a uniform baseline b⃗(0)\vec{b}^{(0)} across all buckets. At each step, candidate configurations are evaluated by adding a fixed increment δ\delta bytes to bucket BkB_k (using standard basis vector e⃗k\vec{e}_k) and measuring search recall R(b⃗+δe⃗k,V)R(\vec{b} + \delta \vec{e}_k, \mathcal{V}) on validation query set V\mathcal{V} using exact Euclidean distance over dequantized database vectors. The winning bucket k∗k^* receives δ\delta bytes permanently. The overall runtime scales as O(K⋅Nsteps)O(K \cdot N_{\text{steps}}), where NstepsN_{\text{steps}} is the number of increments needed to achieve the target memory budget.

  2. Knowl 2 — Variable-Budget Product Quantization with Front-Loaded Remainder Partitioning

    algorithm

    In variable-budget Product Quantization (PQ), a bucket BkB_k of dimension dkd_k with an assigned budget of bkb_k bytes is decomposed into exactly bkb_k contiguous subvectors. Each subvector is quantized using an independent codebook Ck,j\mathcal{C}_{k,j} containing C=256C = 256 centroids, requiring exactly log⁡2(256)=8\log_2(256) = 8 bits (1 byte) per subvector index.

    Input: Sub-vector x(k)∈Rdkx^{(k)} \in \mathbb{R}^{d_k} of bucket BkB_k, byte budget bkb_k, bucket training data Xtrain(k)X_{\text{train}}^{(k)}
    Output: Quantized code vector q⃗(k)∈{0,…,255}bk\vec{q}^{(k)} \in \{0, \dots, 255\}^{b_k}
    s←⌊dk/bk⌋s \leftarrow \lfloor d_k / b_k \rfloor
    r←dk mod bkr \leftarrow d_k \bmod b_k
    Partition x(k)x^{(k)} and Xtrain(k)X_{\text{train}}^{(k)} into bkb_k contiguous subvectors, where the first rr subvectors have dimension s+1s + 1 and the remaining bk−rb_k - r subvectors have dimension ss
    for each subvector index j=1j = 1 to bkb_k do
        Train codebook Ck,j\mathcal{C}_{k,j} with 256 centers using kk-means on the jj-th partition of Xtrain(k)X_{\text{train}}^{(k)}
        q⃗(k)[j]←arg⁡min⁡c∈Ck,j∥xj(k)−c∥2\vec{q}^{(k)}[j] \leftarrow \arg\min_{c \in \mathcal{C}_{k,j}} \| x_j^{(k)} - c \|_2
    end for
    return q⃗(k)\vec{q}^{(k)}

    When dkd_k is not evenly divisible by bkb_k, remainder dimensions r=dk mod bkr = d_k \bmod b_k are front-loaded by assigning s+1=⌊dk/bk⌋+1s + 1 = \lfloor d_k / b_k \rfloor + 1 dimensions to each of the first rr subvectors and s=⌊dk/bk⌋s = \lfloor d_k / b_k \rfloor dimensions to the remaining bk−rb_k - r subvectors. During search, trained codebooks Ck(b)\mathcal{C}_{k}(b) are cached per bucket index kk and byte budget bb so that kk-means clustering is executed at most once per configuration.

  3. Knowl 3 — Variable-Budget Scalar Quantization with Uniform Strided Bit Allocation

    algorithm

    Variable-budget Scalar Quantization (SQ) discretizes each coordinate in bucket BkB_k (of dimension dkd_k) into a bit-width wi∈S={0,2,4,8}w_i \in S = \{0, 2, 4, 8\} subject to a total budget Nbits=8bkN_{\text{bits}} = 8 b_k bits. A 1-bit quantization width is excluded because splitting at the mode destroys the unimodal coordinate distribution of embeddings; a 0-bit width prunes the coordinate entirely.

    Input: Sub-vector x(k)∈Rdkx^{(k)} \in \mathbb{R}^{d_k} of bucket BkB_k, byte budget bkb_k
    Output: Quantized code vector q⃗(k)\vec{q}^{(k)}
    S←{0,2,4,8}S \leftarrow \{0, 2, 4, 8\}
    Nbits←8⋅bkN_{\text{bits}} \leftarrow 8 \cdot b_k
    Wbase←max⁡{w∈S∣w⋅dk≤Nbits}W_{\text{base}} \leftarrow \max \{ w \in S \mid w \cdot d_k \le N_{\text{bits}} \}
    if Wbase=8W_{\text{base}} = 8 then
        Wnext←8W_{\text{next}} \leftarrow 8
        u←0u \leftarrow 0
    else
        Wnext←min⁡{w∈S∣w>Wbase}W_{\text{next}} \leftarrow \min \{ w \in S \mid w > W_{\text{base}} \}
        u←(Nbits−dk⋅Wbase)/(Wnext−Wbase)u \leftarrow (N_{\text{bits}} - d_k \cdot W_{\text{base}}) / (W_{\text{next}} - W_{\text{base}})
    end if
    for i=1i = 1 to dkd_k do
        wi←Wbasew_i \leftarrow W_{\text{base}}
    end for
    if u>0u > 0 then
        Δ←dk/u\Delta \leftarrow d_k / u
        for j=0j = 0 to u−1u - 1 do
            w⌊j⋅Δ⌋+1←Wnextw_{\lfloor j \cdot \Delta \rfloor + 1} \leftarrow W_{\text{next}}
        end for
    end if
    for each dimension i=1i = 1 to dkd_k do
        q⃗(k)[i]←ScalarQuantize(xi(k),wi)\vec{q}^{(k)}[i] \leftarrow \text{ScalarQuantize}(x_i^{(k)}, w_i)
    end for
    return q⃗(k)\vec{q}^{(k)}

    The algorithm sets a baseline bit-width WbaseW_{\text{base}} across all coordinates, and then promotes uu dimensions to WnextW_{\text{next}} using a uniform spatial stride Δ=dk/u\Delta = d_k / u. Each coordinate xi(k)x_i^{(k)} is then independently quantized to wiw_i bits.

  4. Knowl 4 — Dimension-Wise Variance Decay as a Geometric Proxy for MRL Embeddings

    model/method

    Matryoshka Representation Learning (MRL) trains a feature extractor parameter θF\theta_F on a labeled dataset D={(x1,y1),…,(xN,yN)}\mathcal{D} = \{(x_1, y_1), \dots, (x_N, y_N)\} with xi∈Xx_i \in \mathcal{X} and yi∈[L]y_i \in [L] across nested dimensions m∈Mm \in \mathcal{M} using linear classifiers W(m)∈RL×mW^{(m)} \in \mathbb{R}^{L \times m}:

    min⁡{W(m)}m∈M, θF1N∑i=1N∑m∈Mcm⋅L(W(m)⋅F(xi;θF)1:m, yi)\min_{\{W^{(m)}\}_{m \in \mathcal{M}}, \, \theta_F} \frac{1}{N} \sum_{i=1}^N \sum_{m \in \mathcal{M}} c_m \cdot \mathcal{L}\left(W^{(m)} \cdot F(x_i; \theta_F)_{1:m}, \, y_i\right)

    where F(xi;θF)1:mF(x_i; \theta_F)_{1:m} denotes truncation of the extracted feature to its first mm coordinates, L\mathcal{L} is multi-class softmax cross-entropy loss, and cm≥0c_m \ge 0 is the importance weight of dimension mm (default cm=1c_m = 1).

    While the MRL loss does not enforce an explicit closed-form coordinate geometry, the nested optimization concentrates information in the initial coordinates. This produces an empirical step-like decay in coordinate variance as dimension index increases. When evaluated via a rolling window average (window length 128 dimensions) on models such as OpenAI text-embedding-3-large (D=3072D = 3072) and Cohere embed-v4 (D=1536D = 1536), embeddings exhibit discrete iso-variance levels across dimensions. This variance hierarchy provides a geometric proxy that justifies allocating more quantization bits to leading coordinates.

  5. Knowl 5 — Recall Improvements of Variable Bit Allocation over Uniform Allocation

    empirical result

    Across benchmark retrieval datasets (MS Marco, DBpedia-Entity, Quora, FiQA, SciDocs, SciFact) and embedding models (OpenAI text-embedding-3-large and Cohere embed-v4), non-uniform variable bit allocation consistently outperforms uniform bit allocation at identical memory budgets in the sub-1 bit per dimension (bpd) regime (approx0.17\\approx 0.17 bpd to 1.001.00 bpd):

    • Product Quantization (PQ): Variable bit allocation achieves up to an 8% relative improvement in 100-recall@100 over uniform PQ.
    • Scalar Quantization (SQ): Variable bit allocation achieves up to an 18% relative improvement in 100-recall@100 over uniform SQ.

    The largest gains occur in the extreme low-bit regime (≤0.50\le 0.50 bpd), where uniform quantization is highly constrained. OpenAI embeddings yield higher relative improvements than Cohere embeddings, corresponding to the more pronounced iso-variance step transitions present in the OpenAI representations.

  6. Knowl 6 — Emergence of Leading-Dimension Concentration in Greedy Bit Search

    empirical result

    Without explicitly encoding the Matryoshka objective into the optimization criteria, the greedy bit allocation algorithm discovers bit distributions that concentrate memory in the initial dimension buckets across all evaluated datasets and models:

    • In an 8-bucket configuration (B1B_1 to B8B_8 from leading to trailing dimensions), the initial buckets B1B_1 and B2B_2 receive up to 40–50% of the total bit budget at low overall bit allocations.
    • In Scalar Quantization (SQ), the discovered allocation is sharply asymmetric, assigning 4-bit or 8-bit precision to the earliest buckets while assigning 0-bit (pruning) or 2-bit precision to trailing buckets.
    • In Product Quantization (PQ), the allocation exhibits a smoother, monotonic decay in byte budget from the initial buckets toward trailing buckets.

    This behavior confirms that optimizing search recall directly on validation queries recovers the underlying latent MRL hierarchy.

  7. Knowl 7 — Retrieval Recall for Unquantized Dimension Truncation Across Compression Rates

    data/table

    Truncating full-precision (32-bit float) embeddings to their leading dtruncd_{\text{trunc}} dimensions according to equivalent storage budgets yields the following baseline 100-recall@100 (%) performance across datasets:

    Cohere embed-v4 (Full D=1536D=1536)
    Dataset 0.17 bpd 0.33 bpd 0.50 bpd 0.67 bpd 0.83 bpd 1.00 bpd
    (32 B, 8d) (64 B, 16d) (96 B, 24d) (128 B, 32d) (160 B, 40d) (192 B, 48d)
    MS Marco 0.24 1.04 2.30 3.87 5.92 8.64
    DBpedia 0.20 0.68 1.54 2.38 3.76 5.47
    Quora 1.95 7.13 12.59 17.46 22.62 27.62
    FiQA 1.19 2.98 5.52 7.72 9.70 12.61
    SciFact 5.73 8.22 10.66 12.66 15.36 17.79
    SciDocs 2.08 4.15 6.75 8.76 10.61 12.93
    OpenAI text-embedding-3-large (Full D=3072D=3072)
    Dataset 0.17 bpd 0.33 bpd 0.50 bpd 0.67 bpd 0.83 bpd 1.00 bpd
    (64 B, 16d) (128 B, 32d) (192 B, 48d) (256 B, 64d) (320 B, 80d) (384 B, 96d)
    MS Marco 1.80 6.65 13.53 19.98 25.97 31.38
    DBpedia 1.21 4.60 9.63 14.86 20.02 24.73
    Quora 7.18 19.32 31.13 39.85 46.05 50.97
    FiQA 5.67 13.59 21.34 27.93 34.02 38.96
    SciFact 10.68 17.93 24.52 29.58 34.02 38.19
    SciDocs 5.84 12.71 20.77 27.41 32.96 37.63

    In this baseline, vectors are truncated to dtrunc=Budget (bytes)/4d_{\text{trunc}} = \text{Budget (bytes)} / 4 float32 coordinates without quantization. Pure dimension truncation yields substantially lower recall (e.g., 5–51% at 1.00 bpd) compared to quantizing over the full vector dimensions at identical total byte budgets (which reach 40–80% recall).

  8. Knowl 8 — Experimental Evaluation Setup for Variable Bit Allocation Quantization

    experimental setup

    The empirical evaluation of variable bit allocation is configured as follows:

    • Hardware & Software: Linux x86_64 host with 16 hardware threads on an Intel Xeon Platinum 8272CL CPU (2.60 GHz), 32 GiB RAM, with no GPU. Quantization routines are implemented in C++17 branched from Microsoft DiskANN (using OpenMP, Boost, and Intel MKL). DiskANN graph indexing is disabled.
    • Distance Computation & Metric: Brute-force exact L2 (Euclidean) nearest-neighbor distance computation over dequantized (inflated) database vectors to isolate quantization error from index approximations. Performance is measured using average 100-recall@100.
    • Datasets: Six datasets from the BEIR benchmark: MS Marco (500,000 base, 6,980 dev queries), DBpedia-Entity (500,000 base, 400 test queries), Quora (500,000 base, 10,000 test queries), FiQA (57,638 base, 648 test queries), SciDocs (25,657 base, 1,000 test queries), and SciFact (5,183 base, 300 test queries). Large base corpora are truncated to 500,000 vectors.
    • Embedding Models: OpenAI text-embedding-3-large (D=3072D = 3072) and Cohere embed-v4 (D=1536D = 1536).
    • Partitioning & Budget Sweeps: Dimensions are split into K=8K = 8 equal contiguous buckets (384 dims/bucket for OpenAI, 192 dims/bucket for Cohere). Search starts at uniform allocations of 64 bytes (OpenAI) or 32 bytes (Cohere) and runs for 40 iterations with increments of δ=8\delta = 8 bytes (OpenAI) and δ=4\delta = 4 bytes (Cohere), spanning target budgets from ≈0.17\approx 0.17 bpd to 1.001.00 bpd in steps of 1/121/12 bpd. Training codebooks and scalar parameters uses a 10% random sample of base corpus vectors.
  9. Knowl 9 — Limitations of Greedy Variable Bit Allocation Quantization

    limitation

    The variable bit allocation methodology presents several explicit limitations:

    1. High Computational Complexity: Evaluating KK candidate buckets at each greedy step requires repeated distance evaluations and codebook training, making the local search slow and impractical for large-scale production dataset construction.
    2. Coarse Fixed Bucket Discretization: Dimensions are partitioned into a small, fixed number of equal contiguous buckets (K=8K = 8). Evaluating finer-grained or non-uniform bucket granularities exponentially increases the greedy search space.
    3. Hardware and Systems Inefficiency: Non-uniform bit-widths and variable subvector dimensionalities produce irregular memory layouts, breaking alignment required for SIMD vectorization and reducing memory access throughput during query-time scanning.
    4. Absence of Closed-Form Formulation: The bit allocations are obtained purely empirically via greedy validation evaluation; a theoretical closed-form mapping from MRL loss weights to optimal bit allocations is not established.

Coverage note — No substantial contributed material was omitted from the extracted knowls.

References

  1. 1.Cecilia Aguerrebere, Ishwar Singh Bhati, Mark Hildebrand, Mariano Tepper, and Theodore Willke. 2023. Similarity Search in the Blink of an Eye with Compressed Indices. Proc. VLDB Endow. 16, 11 (July 2023), 3433–3446. https://doi.org/10.14778/3611479.3611537
  2. 2.Mohammad Kalim Akram, Saba Sturua, Nastia Havriushenko, Quentin Herreros, Michael Günther, Maximilian Werk, and Han Xiao. 2026. jina-embeddings-v5-text: Task-Targeted Embedding Distillation. arXiv:2602.15547 [cs.CL] https://arxiv.org/abs/2602.15547
  3. 3.Fabien André, Anne-Marie Kermarrec, and Nicolas Le Scouarnec. 2017. Accelerated Nearest Neighbor Search with Quick ADC. In Proceedings of the 2017 ACM on International Conference on Multimedia Retrieval (Bucharest, Romania) (ICMR ’17). Association for Computing Machinery, New York, NY, USA, 159–166. https://doi.org/10.1145/3078971.3078992
  4. 4.Fabien André, Anne-Marie Kermarrec, and Nicolas Le Scouarnec. 2015. Cache locality is not enough: High-Performance Nearest Neighbor Search with Product Quantization Fast Scan. Proc. VLDB Endow. 9 (2015), 288–299. https://api.semanticscholar.org/CorpusID:5966664
  5. 5.Mingkai Chen, Cheng Liu, Shengwen Liang, Lei Zhang, Xiaowei Li, and Huawei Li. 2026. ANNS-AMP: Accelerating Approximate Nearest Neighbor Search via Adaptive Mixed-Precision Computing. arXiv:2606.07156 [cs.PF] https://arxiv.org/abs/2606.07156
  6. 6.Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S. Weld. 2020. SPECTER: Document-level Representation Learning using Citation-informed Transformers. In ACL.
  7. 7.Cohere. 2025. Announcing Embed Multimodal v4. https://docs.cohere.com/changelog/embed-multimodal-v4 Accessed: 2026-06-11.
  8. 8.Paul Covington, Jay Adams, and Emre Sargin. 2016. Deep Neural Networks for YouTube Recommendations. In Proceedings of the 10th ACM Conference on Recommender Systems (Boston, Massachusetts, USA) (RecSys ’16). Association for Computing Machinery, New York, NY, USA, 191–198. https://doi.org/10.1145/2959100.2959190
  9. 9.Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. 2026. The Faiss Library. IEEE Transactions on Big Data 12, 2 (2026), 346–361. https://doi.org/10.1109/TBDATA.2025.3618474
  10. 10.Kenneth Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos, Ashwin Mathur, David Stap, Jay Gala, Wissam Siblini, Dominik Krzemiński, Genta Indra Winata, Saba Sturua, Saiteja Utpala, Mathieu Ciancone, Marion Schaeffer, Gabriel Sequeira, Diganta Misra, Shreeya Dhakal, Jonathan Rystrøm, Roman Solomatin, Ömer Çağatan, Akash Kundu, Martin Bernstorff, Shitao Xiao, Akshita Sukhlecha, Bhavish Pahwa, Rafał Poświata, Kranthi Kiran GV, Shawon Ashraf, Daniel Auras, Björn Plüster, Jan Philipp Harries, Loïc Magne, Isabelle Mohr, Mariya Hendriksen, Dawei Zhu, Hippolyte Gisserot-Boukhlef, Tom Aarsen, Jan Kostkan, Konrad Wojtasik, Taemin Lee, Marek Šuppa, Crystina Zhang, Roberta Rocca, Mohammed Hamdy, Andrianos Michail, John Yang, Manuel Faysse, Aleksei Vatolin, Nandan Thakur, Manan Dey, Dipam Vasani, Pranjal Chitale, Simone Tedeschi, Nguyen Tai, Artem Snegirev, Michael Günther, Mengzhou Xia, Weijia Shi, Xing Han Lù, Jordan Clive, Gayatri Krishnakumar, Anna Maksimova, Silvan Wehrli, Maria Tikhonova, Henil Panchal, Aleksandr Abramov, Malte Ostendorff, Zheng Liu, Simon Clematide, Lester James Miranda, Alena Fenogenova, Guangyu Song, Ruqiya Bin Safi, Wen-Ding Li, Alessia Borghini, Federico Cassano, Hongjin Su, Jimmy Lin, Howard Yen, Lasse Hansen, Sara Hooker, Chenghao Xiao, Vaibhav Adlakha, Orion Weller, Siva Reddy, and Niklas Muennighoff. 2025. MMTEB: Massive Multilingual Text Embedding Benchmark. arXiv preprint arXiv:2502.13595 (2025). https://doi.org/10.48550/arXiv.2502.13595
  11. 11.Jianyang Gao and Cheng Long. 2023. High-Dimensional Approximate Nearest Neighbor Search: with Reliable and Efficient Distance Comparison Operations. Proc. ACM Manag. Data 1, 2, Article 137 (June 2023), 27 pages. https://doi.org/10.1145/3589282
  12. 12.Jianyang Gao and Cheng Long. 2024. RaBitQ: Quantizing High-Dimensional Vectors with a Theoretical Error Bound for Approximate Nearest Neighbor Search. Proc. ACM Manag. Data 2, 3, Article 167 (May 2024), 27 pages. https://doi.org/10.1145/3654970
  13. 13.Allen Gersho and Robert M. Gray. 1992. Vector Quantization and Signal Compression. The Springer International Series in Engineering and Computer Science, Vol. 159. Springer.
  14. 14.Ruiqi Guo, Philip Sun, Erik Lindgren, Quan Geng, David Simcha, Felix Chern, and Sanjiv Kumar. 2020. Accelerating Large-Scale Inference with Anisotropic Vector Quantization. In International Conference on Machine Learning. https://arxiv.org/abs/1908.10396
  15. 15.Faegheh Hasibi, Fedor Nikolaev, Chenyan Xiong, Krisztian Balog, Svein Erik Bratsberg, Alexander Kotov, and Jamie Callan. 2017. DBpedia-Entity V2: A Test Collection for Entity Search. In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’17). ACM, 1265–1268. https://doi.org/10.1145/3077136.3080751
  16. 16.Jui-Ting Huang, Ashish Sharma, Shuying Sun, Li Xia, David Zhang, Philip Pronin, Janani Padmanabhan, Giuseppe Ottaviano, and Linjun Yang. 2020. Embedding-based Retrieval in Facebook Search. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (Virtual Event, CA, USA) (KDD ’20). Association for Computing Machinery, New York, NY, USA, 2553–2561. https://doi.org/10.1145/3394486.3403305
  17. 17.J. Y. Huang and Peter M. Schultheiss. 1963. Block Quantization of Correlated Gaussian Random Variables. IEEE Transactions on Communications 11 (1963), 289–296.
  18. 18.Piotr Indyk and Rajeev Motwani. 1998. Approximate nearest neighbors: towards removing the curse of dimensionality. In Proceedings of the Thirtieth Annual ACM Symposium on Theory of Computing (Dallas, Texas, USA) (STOC ’98). Association for Computing Machinery, New York, NY, USA, 604–613. https://doi.org/10.1145/276698.276876
  19. 19.Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko. 2018. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2704–2713. https://doi.org/10.1109/CVPR.2018.00286
  20. 20.Herve Jégou, Matthijs Douze, and Cordelia Schmid. 2011. Product Quantization for Nearest Neighbor Search. IEEE Transactions on Pattern Analysis and Machine Intelligence 33, 1 (2011), 117–128. https://doi.org/10.1109/TPAMI.2010.57
  21. 21.Aditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford, Aditya Sinha, Vivek Ramanujan, William Howard-Snyder, Kaifeng Chen, Sham Kakade, Prateek Jain, and Ali Farhadi. 2022. Matryoshka Representation Learning. In Advances in Neural Information Processing Systems. 30233–30249.
  22. 22.Ying Liu, Dengsheng Zhang, Guojun Lu, and Wei-Ying Ma. 2007. A survey of content-based image retrieval with high-level semantics. Pattern Recognition 40, 1 (2007), 262–282. https://doi.org/10.1016/j.patcog.2006.04.045
  23. 23.Macedo Maia, Siegfried Handschuh, André Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur. 2018. WWW’18 Open Challenge: Financial Opinion Mining and Question Answering. In Companion Proceedings of the The Web Conference 2018 (Lyon, France) (WWW ’18). International World Wide Web Conferences Steering Committee, Republic and Canton of Geneva, CHE, 1941–1942. https://doi.org/10.1145/3184558.3192301
  24. 24.Yu A. Malkov and D. A. Yashunin. 2020. Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence 42, 4 (2020), 824–836. https://doi.org/10.1109/TPAMI.2018.2889473
  25. 25.Yusuke Matsui, Yusuke Uchida, Hervé Jégou, and Shin’ichi Satoh. 2018. A Survey of Product Quantization. ITE Transactions on Media Technology and Applications 6, 1 (2018), 2–10.
  26. 26.Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers. 2022. MTEB: Massive Text Embedding Benchmark. arXiv preprint arXiv:2210.07316 (2022). https://doi.org/10.48550/ARXIV.2210.07316
  27. 27.Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. In Proceedings of the Workshop on Cognitive Computation: Integrating neural and symbolic approaches 2016 co-located with the 30th Annual Conference on Neural Information Processing Systems (NIPS 2016), Barcelona, Spain, December 9, 2016 (CEUR Workshop Proceedings), Tarek Richard Besold, Antoine Bordes, Artur S. d’Avila Garcez, and Greg Wayne (Eds.). CEUR-WS.org. https://ceur-ws.org/Vol-1773/CoCoNIPS_2016_paper9.pdf
  28. 28.Nomic AI. 2024. nomic-embed-text-v1.5 Model Repository. https://huggingface.co/nomic-ai/nomic-embed-text-v1.5 Accessed: 2026-06-11.
  29. 29.OpenAI. 2024. New embedding models and API updates. https://openai.com/blog/new-embedding-models-and-api-updates Accessed: 2026-06-11.
  30. 30.Miloš Radovanović, Alexandros Nanopoulos, and Mirjana Ivanović. 2009. Nearest neighbors in high-dimensional data: the emergence and influence of hubs. In Proceedings of the 26th Annual International Conference on Machine Learning (Montreal, Quebec, Canada) (ICML ’09). Association for Computing Machinery, New York, NY, USA, 865–872. https://doi.org/10.1145/1553374.1553485
  31. 31.Madhuri Shanbhogue, Zhe Li, Shanfeng Zhang, Gustavo Hernández Ábrego, Shih-Cheng Huang, Aashi Jain, Daniel Salz, Sonam Goenka, Chaitra Hegde, Ji Ma, Feiyang Chen, Jiaxing Wu, Tanmaya Dabral, Babak Samari, Kevin Poulet, Daniel Cer, Kaifeng Chen, Paul Suganthan, Hui Hui, Jovan Andonov, Philippe Schlattner, Jay Han, Iftekhar Naim, Wing Lowe, Vladimir Pchelin, Albert Yang, Yi-Ting Chen, Zhongli Ding, Grace Zhang, Georg Heigold, Yichang Chen, Antoine Reveillon, Brendan Mccloskey, Wenlei Zhou, Dahun Kim, Rui Meng, Emma Wang, Jack Zheng, Halley Fede, Zhen Yang, Keegan Mosley, Brian Potetz, Sahil Dua, Henrique Schechter Vera, Shen Gao, Hesen Zhang, Andreas Hess, Hengxuan Ying, Alberto Montes, Karan Gill, Min Choi, Sebastian Russo, Anja Hauth, Jinhyuk Lee, Michael Boratko, Megan Barnes, Vikram Rao, Claudiu Musat, Cyril Allauzen, Ehsan Variani, Shankar Kumar, Tom Bagby, Junyi Jiao, Yang Gu, Tengxin Li, Ayush Agrawal, Roberto Santana, Dev Nath, Stephen Karukas, Shuoxuan Han, Lucia Loher, Alice Twu, Nidhi Vyas, Siddharth Bhai, Frank Palma Gomez, Wangyuan Zhang, Chaoren Liu, Jizheng Yang, Steve Qiu, Shijie Zhang, Sujay Kulkarni, Sascha Rothe, Sean Nakamoto, Raphael Hoffmann, Zach Gleicher, Yunhsuan Sung, Qin Yin, Tom Duerig, and Mojtaba Seyedhosseini. 2026. Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini. arXiv:2605.27295 [cs.CV] https://arxiv.org/abs/2605.27295
  32. 32.Claude E. Shannon. 1959. Coding Theorems for a Discrete Source With a Fidelity Criterion. Vol. 7. Institute of Radio Engineers, International Convention Record, 325–350. https://doi.org/10.1109/9780470544242.ch21
  33. 33.Harsha Vardhan Simhadri, Ravishankar Krishnaswamy, Gopal Srinivasa, Suhas Jayaram Subramanya, Andrija Antonijevic, Dax Pryce, David Kaczynski, Shane Williams, Siddarth Gollapudi, Varun Sivashankar, Neel Karia, Aditi Singh, Shikhar Jaiswal, Neelam Mahapatro, Philip Adams, Bryan Tower, and Yash Patel. 2023. DiskANN: Graph-structured Indices for Scalable, Fast, Fresh and Filtered Approximate Nearest Neighbor Search. https://github.com/Microsoft/DiskANN
  34. 34.Philip Sun, David Simcha, Dave Dopson, Ruiqi Guo, and Sanjiv Kumar. 2023. SOAR: Improved Indexing for Approximate Nearest Neighbor Search. In Neural Information Processing Systems. https://arxiv.org/abs/2404.00774
  35. 35.Ganap Ashit Tewary, Nrusinga Charan Gantayat, and Jeff Zhang. 2026. AQR-HNSW: Accelerating Approximate Nearest Neighbor Search via Density-aware Quantization and Multi-stage Re-ranking. arXiv:2602.21600 [cs.IR] https://arxiv.org/abs/2602.21600
  36. 36.Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021. BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2). https://openreview.net/forum?id=wCu6T5xFjeJ
  37. 37.Henrique Schechter Vera, Sahil Dua, Biao Zhang, Daniel Salz, Ryan Mullins, Sindhu Raghuram Panyam, Sara Smoot, Iftekhar Naim, Joe Zou, Feiyang Chen, Daniel Cer, Alice Lisak, Min Choi, Lucas Gonzalez, Omar Sanseviero, Glenn Cameron, Ian Ballantyne, Kat Black, Kaifeng Chen, Weiyi Wang, Zhe Li, Gus Martins, Jinhyuk Lee, Mark Sherwood, Juyeong Ji, Renjie Wu, Jingxiao Zheng, Jyotinder Singh, Abheesht Sharma, Divyashree Sreepathihalli, Aashi Jain, Adham Elarabawy, AJ Co, Andreas Doumanoglou, Babak Samari, Ben Hora, Brian Potetz, Dahun Kim, Enrique Alfonseca, Fedor Moiseev, Feng Han, Frank Palma Gomez, Gustavo Hernández Ábrego, Hesen Zhang, Hui Hui, Jay Han, Karan Gill, Ke Chen, Koert Chen, Madhuri Shanbhogue, Michael Boratko, Paul Suganthan, Sai Meher Karthik Duddu, Sandeep Mariserla, Setareh Ariafar, Shanfeng Zhang, Shijie Zhang, Simon Baumgartner, Sonam Goenka, Steve Qiu, Tanmaya Dabral, Trevor Walker, Vikram Rao, Waleed Khawaja, Wenlei Zhou, Xiaoqi Ren, Ye Xia, Yichang Chen, Yi-Ting Chen, Zhe Dong, Zhongli Ding, Francesco Visin, Gaël Liu, Jiageng Zhang, Kathleen Kenealy, Michelle Casbon, Ravin Kumar, Thomas Mesnard, Zach Gleicher, Cormac Brick, Olivier Lacombe, Adam Roberts, Qin Yin, Yunhsuan Sung, Raphael Hoffmann, Tris Warkentin, Armand Joulin, Tom Duerig, and Mojtaba Seyedhosseini. 2025. EmbeddingGemma: Powerful and Lightweight Text Representations. arXiv:2509.20354 [cs.CL] https://arxiv.org/abs/2509.20354
  38. 38.Voyage AI. 2026. Voyage 4. https://blog.voyageai.com/2026/01/15/voyage-4/ Accessed: 2026-06-11.
  39. 39.David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or Fiction: Verifying Scientific Claims. In EMNLP.
  40. 40.Roger Weber, Hans-Jörg Schek, and Stephen Blott. 1998. A Quantitative Analysis and Performance Study for Similarity-Search Methods in High-Dimensional Spaces. In Proceedings of the 24rd International Conference on Very Large Data Bases (VLDB ’98). Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 194–205.
  41. 41.Jiaqi Zhang. 2025. Survey of Quantization-Aware Training (QAT) Applications in Deep Learning Quantization. Proceedings of the 2025 International Symposium on Artificial Intelligence and Computational Social Sciences (2025). https://api.semanticscholar.org/CorpusID:284019251
  42. 42.Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang, Pengjun Xie, An Yang, Dayiheng Liu, Junyang Lin, Fei Huang, and Jingren Zhou. 2025. Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models. arXiv:2506.05176 [cs.CL] https://arxiv.org/abs/2506.05176

Citation

MLA
Sreeramji, K. S., et al. “Quantization Beyond Uniform Bit Allocation”. arXiv, 2026, http://arxiv.org/abs/2608.19388v1.
APA
Sreeramji, K. S., Basu, S., Krishnaswamy, R., Shiragur, K., & Wang, Y. (2026). Quantization Beyond Uniform Bit Allocation. arXiv. http://arxiv.org/abs/2608.19388v1
Chicago
Sreeramji, K. S., S. Basu, R. Krishnaswamy, K. Shiragur, and Y. Wang. 2026. “Quantization Beyond Uniform Bit Allocation”. arXiv. http://arxiv.org/abs/2608.19388v1.
Harvard
Sreeramji, K.S. et al. (2026) “Quantization Beyond Uniform Bit Allocation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2608.19388v1.
Vancouver
1. Sreeramji KS, Basu S, Krishnaswamy R, Shiragur K, Wang Y (2026) Quantization Beyond Uniform Bit Allocation. arXiv

BibTeX

@article{sreeramji2026quantization,
  title = {Quantization Beyond Uniform Bit Allocation},
  author = {Sreeramji, K. S. and Basu, Sabyasachi and Krishnaswamy, Ravishankar and Shiragur, Kirankumar and Wang, Yujia},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2608.19388v1},
  eprint = {2608.19388}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission