Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs

Yeonhong ParkJake HyunSangLyul ChoBonggeun SimJae W. Lee

article2024ICML65 citations

Proposes a post-training quantization framework and a specialized GPU serving engine that pack multiple Large Language Models of varying bit-widths into the memory footprint of a single high-precision model without sacrificing inference speed or output quality.

Listen

Deploying large language models often requires serving multiple model sizes concurrently to handle varying response-time requirements, execute multi-task workloads, and support acceleration techniques like speculative decoding. However, storing multiple distinct models requires immense hardware memory and substantial training compute, posing severe bottlenecks for resource-constrained environments like desktop and mobile edge devices.

The article demonstrates an "any-precision" framework that allows multiple models of varying precision levels—ranging from 3-bit to 8-bit—to be served from a single parent model footprint. By combining a lightweight post-training quantization method with a specialized graphics processing unit execution engine, the approach enables dynamic model scaling without storing redundant weight parameters or retraining models from scratch.

To achieve this, the authors developed an incremental upscaling technique using non-uniform, clustering-based quantization. Starting with a compact 3-bit seed model, the method iteratively appends single bits to split weight clusters into higher-precision representations. Complementing this algorithm, the authors designed a software execution engine based on a bitplane memory layout, which allows hardware to load only the exact bit-width required for a specific query instead of reading full bit-vectors.

Key findings show substantial resource and performance gains. Storing a complete suite of precision levels from 3-bit to 8-bit in a single any-precision model reduces total memory consumption by up to 3.56 times compared to deploying each model separately. Models generated through incremental upscaling match the state-of-the-art accuracy and language quality of independently trained models at each respective bit-width, showing negligible quality degradation across standard evaluation benchmarks. Furthermore, the complete quantization process takes under one minute on standard consumer hardware, and the custom execution engine delivers inference speeds that match or exceed existing fixed-precision engines across desktop, laptop, and mobile processors.

These findings indicate that organizations can significantly lower hardware infrastructure costs, improve query throughput, and deploy adaptable on-device language models without costly re-training pipelines. Unlike uniform quantization techniques, which suffer severe quality degradation under incremental upscaling due to cumulative rounding errors, clustering-based non-uniform quantization provides a stable and mathematically robust foundation for flexible-precision serving.

Organizations serving varied artificial intelligence workloads should consider adopting bitplane-based and codebook-quantized architectures when designing low-latency, on-device serving systems. For deployments requiring high batch sizes or long input prompt processing, teams should implement hybrid execution strategies that transition to standard compute kernels to maintain optimal hardware utilization.

Confidence in these findings is high across tested open-source models (including LLaMA, Mistral, and OPT architectures) and consumer hardware tiers. However, decision-makers should note that the current kernel implementation is optimized primarily for smaller batch sizes and memory-bound generation phases, and extended validation on larger model architectures and diverse production workloads is recommended prior to broad organizational rollout.

No sufficiently relevant recommendations were found.

Cover for Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs

Abstract

Recently, considerable efforts have been directed towards compressing Large Language Models (LLMs), which showcase groundbreaking capabilities across diverse applications but entail significant deployment costs due to their large sizes. Meanwhile, much less attention has been given to mitigating the costs associated with deploying multiple LLMs of varying sizes despite its practical significance. Thus, this paper introduces any-precision LLM, extending the concept of any-precision DNN to LLMs. Addressing challenges in any-precision LLM, we propose a lightweight method for any-precision quantization of LLMs, leveraging a post-training quantization framework, and develop a specialized software engine for its efficient serving. As a result, our solution significantly reduces the high costs of deploying multiple, different-sized LLMs by overlaying LLMs quantized to varying bit-widths, such as 3, 4, ..., n bits, into a memory footprint comparable to a single n-bit LLM. All the supported LLMs with varying bit-widths demonstrate state-of-the-art model quality and inference throughput, proving itself to be a compelling option for deployment of multiple, different-sized LLMs. The code is available at https://github.com/SNU-ARC/any-precision-llm.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 2.1. GPU Basics
  • 2.2. LLM Quantization
  • 3. Motivation
  • 3.1. Need for Deploying Multiple, Different-sized LLMs
  • 3.2. Challenges of Deploying Multiple, Different-sized LLMs
  • 3.3. Our Solution: Any-Precision LLM
  • 4. Any-Precision Quantization for LLM
  • 4.1. Incremental Upscaling
  • 4.2. Non-uniform Quantization-based Incremental Upscaling
  • 5. Specialized Software Engine
  • 5.1. Need for New Software Engine
  • 5.2. System Overview
  • 5.3. GPU Kernel Optimization
  • 6. Evaluation
  • 6.1. Any-Precision Quantization Results
  • 6.3. End-to-end Throughput
  • 7. Conclusion
  • Impact Statement
  • Acknowledgement
  • References
  • A. Evaluation Details
  • A.1. Datasets
  • A.2. Perplexity Calculations
  • A.3. Zero-shot Tasks
  • B. Full Results on Zero-shot Tasks
  • C. Additional Kernel Microbenchmark Results
  • C.1. Kernel Latency on Various Matrix Sizes
  • C.2. Comparison with Kernels for Uniform Quantization
  • C.3. Matrix-Matrix Multiplication Performance
  • D. Additional End-to-End Throughput Evaluation Results
  • E. Incremental Upscaling with Uniform Quantization Methods
  • E.1. Incremental Upscaling with GPTQ
  • E.2. Incremental Upscaling with AWQ

Knowls

  1. Knowl 1 — Any-precision LLMs share one stored parent model across bit-widths

    model/method

    An any-precision LLM stores one quantized parent model and supports lower-bit versions by using prefixes of the parent weights’ bit representations, together with the centroid tables for each supported bit-width. This lets a deployment select among models with different quality–latency trade-offs without storing a separate full weight set for each model or training each version independently. In the paper’s Llama-2-7B memory comparison, supporting bit-widths {3,6}\{3,6\} uses 5.6 GB versus 8.3 GB for separate deployment (1.49× savings); {4,8}\{4,8\} uses 7.7 GB versus 10.8 GB (1.40×); {3,4,6}\{3,4,6\} uses 5.6 GB versus 12.1 GB (2.15×); {3,4,8}\{3,4,8\} uses 7.7 GB versus 13.7 GB (1.76×); {3,4,6,8}\{3,4,6,8\} uses 7.9 GB versus 19.1 GB (2.41×); and supporting every width from 3 to 8 bits uses 8.4 GB versus 29.9 GB (3.56×).

  2. Knowl 2 — Incremental upscaling constructs the supported bit-widths from a seed model

    algorithm

    Given candidate bit-widths n1<n2<⋯<nKn_1<n_2<\cdots<n_K, incremental upscaling first quantizes the full-precision model to the minimum supported width n1n_1, producing a seed model. It then raises the model by one bit at a time until it reaches nKn_K: at each step, each quantized parameter retains its existing bits and receives one additional bit. The resulting higher-bit model therefore contains the lower-bit model as a bit-prefix, allowing all intermediate widths to be supported by one parent model. The paper’s evaluated configuration uses a 3-bit seed and upscales through 4–8 bits. The method is post-training quantization and does not require training separate models.

  3. Knowl 3 — Weighted cluster splitting makes incremental upscaling compatible with non-uniform quantization

    model/method

    The paper uses SqueezeLLM’s clustering-based, non-uniform quantization as the backbone for both seed generation and upscaling. In this scheme, each weight is represented by a cluster assignment and reconstructed using that cluster’s centroid. To increase precision by one bit, each existing cluster is divided into two sub-clusters by weighted K-means over the weights in that cluster; a sensitivity metric based on an approximated second-order derivative supplies the weighting. Each resulting sub-cluster has its own centroid. Repeating this operation adds one code bit per step while preserving the assignments needed to recover the lower-bit model from the higher-bit representation.

  4. Knowl 4 — The any-precision GPU engine uses bitplanes and width-specific centroid tables

    model/method

    The engine stores quantized weights as separate bitplanes rather than packing each weight’s bits contiguously into a single array. For a requested bit-width, it loads only the required bitplanes and the corresponding centroid table; for example, the paper’s 4-bit execution example loads the 4-bit centroid table and bitplanes 4–7. With channel-wise quantization, each centroid-table row corresponds to an output channel and contains 2k2^k centroids for a supported kk-bit width. Within a thread block, centroid data is placed in shared memory for repeated lookup, while threads load distinct weight-bitplane regions. Each thread then loads activations and bit-vectors, rearranges the weight bits so each weight’s code is contiguous, forms centroid indices, fetches the centroids, and performs multiply-accumulate operations. Unlike conventional bitpacking, this layout permits memory traffic to fall with the selected precision.

  5. Knowl 5 — Three GPU-kernel optimizations reduce bitplane access and decoding overhead

    model/method

    The engine uses three optimizations for the costs introduced by bitplane weights. First, weight bitplane layout optimization permutes bytes during preprocessing so that threads in a warp access corresponding FP16 activations coalescently; the resulting weight indices are no longer sequential within each thread’s work. Second, an efficient bit-transpose treats a 32-bit vector as packed sub-vectors of the selected weight width and uses bitwise operations to rearrange codes. For 4-bit weights, the paper reports 40 bitwise operations for the operation, versus 76 for applying an 8-by-8 transpose twice; the approach also supports 2- and 8-bit widths and uses the next larger power-of-two width for non-power-of-two cases. Third, for 3-bit weights, table-lookup merging pairs adjacent 3-bit indices into a 6-bit index and uses an expanded 64-entry table of centroid pairs. This reduces the reported bitwise operations per 32-bit vector from 24 to 16, at the cost of additional shared-memory use.

  6. Knowl 6 — Upscaled SqueezeLLM models retain quality close to independently quantized models

    empirical result

    The paper compares 4–8-bit SqueezeLLM models produced by incremental upscaling from a 3-bit seed (SqLLM+IU) with independently quantized SqueezeLLM models (SqLLM). Evaluation covers Llama-2-7B, Mistral-7B, and OPT-6.7B, OPT-2.7B, and OPT-1.3B, using perplexity on WikiText2, PTB, and C4 and zero-shot accuracy averaged over ARC-easy, ARC-challenge, HellaSwag, PIQA, and WinoGrande. Perplexity differences are generally small; the reported exceptions exceeding 0.1 include 4-bit PTB for OPT-6.7B (12.82 versus 12.59, an increase of 0.23) and OPT-1.3B (16.59 versus 16.45, an increase of 0.14). The average zero-shot results are also close: for example, Llama-2-7B at 4 bits scores 68.2% versus 68.3%, and Mistral-7B at 4 bits scores 73.3% versus 73.2%; OPT-6.7B at 4 bits is a larger exception at 60.3% versus 59.9%. Overall, the measurements support using the upscaled set as a near-quality match to independently quantized widths across the evaluated models and tasks.

  7. Knowl 7 — Generating a 3–8-bit model set takes under a minute on a 24-core CPU

    empirical result

    The authors measured the time to generate a 3-bit seed and incrementally upscale it through 8 bits on an Intel i9-13900K CPU with 24 cores. Seed-generation, upscaling, and total times were, respectively: Llama-2-7B, 36.2 s, 15.6 s, and 51.8 s; Mistral-7B, 37.2 s, 18.2 s, and 55.4 s; OPT-6.7B, 37.0 s, 12.8 s, and 49.8 s; OPT-2.7B, 14.0 s, 6.2 s, and 20.2 s; OPT-1.3B, 6.4 s, 3.0 s, and 9.4 s. Thus, the reported full conversion completes in less than a minute even for the evaluated 7B models.

  8. Knowl 8 — The bitplane kernel delivers quantized matrix-vector speedups across three GPU classes

    empirical result

    For Llama-2-7B matrix-vector multiplication, the authors compare the custom kernel with a cuBLAS FP16 baseline on an RTX 4090, RTX 4070 Laptop, and Jetson AGX Orin 64 GB. Across the three tested Llama-2-7B weight-matrix shapes, 3-bit speedups are 3.99–4.67× on the RTX 4090, 4.97–5.29× on the RTX 4070 Laptop, and 3.84–4.35× on the Jetson; at 8 bits, the corresponding ranges are 1.56–1.81×, 1.72–1.87×, and 1.78–1.86×. Against the SqueezeLLM kernel, which supports only 3 and 4 bits, the new kernel is broadly comparable on the two RTX GPUs and improves performance on the Jetson. The results show that bitplane support can deliver increasing speedups as precision falls without sacrificing competitive matrix-vector throughput.

  9. Knowl 9 — End-to-end generation benefits from lower precision, while large prefill batches need a different path

    empirical result

    Integrated with TensorRT-LLM, the engine improves Llama-2-7B generation throughput as weight precision decreases. On an RTX 4090, the reported throughput for generating 128 tokens rises from 68 tokens/s at FP16 to 262 tokens/s at 3 bits; for 1024 tokens it rises from 66 to 245 tokens/s. The longer-generation speedup is somewhat smaller because attention is not accelerated by weight quantization. For matrix-matrix multiplication, the quantized kernel performs well at small batch sizes, but becomes slower than cuBLAS FP16 for sufficiently large row counts, as it does not use tensor cores. The implementation therefore dequantizes weights separately and switches to cuBLAS when the row count exceeds a threshold such as 16; the paper reports that the added latency is about 100 μs and does not scale with the row count.

  10. Knowl 10 — Uniform GPTQ and AWQ do not preserve quality reliably under the tested upscaling procedure

    limitation

    The paper’s incremental-upscaling approach is effective with non-uniform clustering, but attempts to apply it to uniform GPTQ and AWQ produce substantial quality degradation in the tested settings. For Llama-2-7B on WikiText2, GPTQ upscaled from a 3-bit seed has 4-bit perplexity 36.82, compared with 5.71 for independently quantized GPTQ; GPTQ with activation reordering gives 12.70 versus 5.68. AWQ upscaling gives infinite 4-bit perplexity and 22.50 at 5 bits, compared with 5.59 and 5.50 for independent AWQ quantization. The paper attributes the difficulty to incompatibility between the methods’ quantization-specific reconstruction or preprocessing and the requirement that each lower-bit code remain a prefix of the higher-bit code. These results limit the demonstrated quality-preserving method to the compatible non-uniform quantization backbone used in the paper.

Coverage note — Detailed per-task zero-shot scores, additional matrix-size and small-batch tables, and GPU implementation listings are omitted because they extend the reported comparisons or provide implementation detail beyond the main contributed methods and findings.

References

  1. 1.Agarwal, R., Vieillard, N., Zhou, Y., Stanczyk, P., Ramos, S., Geist, M., and Bachem, O. Generalized knowledge distillation for auto-regressive language models. In The Twelfth International Conference on Learning Representations, 2024.
  2. 2.Chee, J., Cai, Y., Kuleshov, V., and Sa, C. D. QuIP: 2-bit quantization of large language models with guarantees. Advances in Neural Information Processing Systems, 36, 2023.
  3. 3.Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., and Jumper, J. Accelerating large language model decoding with speculative sampling, 2023.
  4. 4.Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. Think you have solved question answering? try arc, the AI2 reasoning challenge, 2018.
  5. 5.Cowan, M., Moreau, T., Chen, T., Bornholt, J., and Ceze, L. Automatic generation of high-performance quantized machine learning kernels. In Proceedings of the 18th ACM/IEEE International Symposium on Code Generation and Optimization, CGO 2020, pp. 305–316, New York, NY, USA, 2020. Association for Computing Machinery. ISBN 9781450370479. doi: 10.1145/3368826.3377912.
  6. 6.Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. QLoRA: Efficient finetuning of quantized LLMs. In Thirty-seventh Conference on Neural Information Processing Systems, 2023.
  7. 7.Dettmers, T., Svirschevski, R., Egiazarian, V., Kuznedelev, D., Frantar, E., Ashkboos, S., Borzunov, A., Hoefler, T., and Alistarh, D. SpQR: A sparse-quantized representation for near-lossless llm weight compression. 2024.
  8. 8.Frantar, E. and Alistarh, D. SparseGPT: Massive language models can be accurately pruned in one-shot. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J. (eds.), International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pp. 10323–10337. PMLR, 2023.
  9. 9.Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D. OPTQ: Accurate quantization for generative pre-trained transformers. In The Eleventh International Conference on Learning Representations, 2023.
  10. 10.Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. A framework for few-shot language model evaluation, 12 2023.
  11. 11.Gholami, A., Kim, S., Dong, Z., Yao, Z., Mahoney, M. W., and Keutzer, K. A survey of quantization methods for efficient neural network inference, 2021.
  12. 12.Gu, Y., Dong, L., Wei, F., and Huang, M. MiniLLM: Knowledge distillation of large language models. In The Twelfth International Conference on Learning Representations, 2024.
  13. 13.Han, S., Shen, H., Philipose, M., Agarwal, S., Wolman, A., and Krishnamurthy, A. Mcdnn: An approximation-based execution framework for deep stream processing under resource constraints. In Proceedings of the 14th Annual International Conference on Mobile Systems, Applications, and Services, MobiSys ’16, pp. 123–136, New York, NY, USA, 2016. Association for Computing Machinery. ISBN 9781450342698. doi: 10.1145/2906388.2906396.
  14. 14.Hinton, G., Vinyals, O., and Dean, J. Distilling the knowledge in a neural network, 2015.
  15. 15.Hsieh, C., Li, C., Yeh, C., Nakhost, H., Fujii, Y., Ratner, A., Krishna, R., Lee, C., and Pfister, T. Distilling Step-by-Step! outperforming larger language models with less training data and smaller model sizes. In Rogers, A., Boyd-Graber, J. L., and Okazaki, N. (eds.), Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9-14, 2023, pp. 8003–8017. Association for Computational Linguistics, 2023. doi: 10.18653/V1/2023.FINDINGS-ACL.507.
  16. 16.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. Mistral 7B, 2023.
  17. 17.Kim, J., Lee, J. H., Kim, S., Park, J., Yoo, K. M., Kwon, S. J., and Lee, D. Memory-efficient fine-tuning of compressed large language models via sub-4-bit integer quantization. Advances in Neural Information Processing Systems, 36, 2023a.
  18. 18.Kim, S., Hooper, C., Gholami, A., Dong, Z., Li, X., Shen, S., Mahoney, M. W., and Keutzer, K. SqueezeLLM: Dense-and-sparse quantization. arXiv preprint arXiv:2306.07629, 2023b.
  19. 19.Kim, S., Hooper, C., Wattanawong, T., Kang, M., Yan, R., Genc, H., Dinh, G., Huang, Q., Keutzer, K., Mahoney, M. W., Shao, Y. S., and Gholami, A. Full stack optimization of transformer inference: a survey, 2023c.
  20. 20.Kim, S., Mangalam, K., Moon, S., Malik, J., Mahoney, M. W., Gholami, A., and Keutzer, K. Speculative decoding with big little decoder. In Thirty-seventh Conference on Neural Information Processing Systems, 2023d.
  21. 21.Lee, C., Jin, J., Kim, T., Kim, H., and Park, E. Owq: Outlier-aware weight quantization for efficient fine-tuning and inference of large language models. Proceedings of the AAAI Conference on Artificial Intelligence, 38(12):13355–13364, Mar. 2024. doi: 10.1609/aaai.v38i12.29237.
  22. 22.Leviathan, Y., Kalman, M., and Matias, Y. Fast inference from transformers via speculative decoding. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org, 2023.
  23. 23.Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S. AWQ: Activation-aware weight quantization for llm compression and acceleration. arXiv preprint arXiv:2306.00978, 2023.
  24. 24.Liu, Z., Oguz, B., Zhao, C., Chang, E., Stock, P., Mehdad, Y., Shi, Y., Krishnamoorthi, R., and Chandra, V. LLM-QAT: Data-free quantization aware training for large language models, 2023.
  25. 25.Ma, X., Fang, G., and Wang, X. LLM-Pruner: On the structural pruning of large language models. In Thirty-seventh Conference on Neural Information Processing Systems, 2023.
  26. 26.Marcus, M., Kim, G., Marcinkiewicz, M. A., MacIntyre, R., Bies, A., Ferguson, M., Katz, K., and Schasberger, B. The Penn Treebank: annotating predicate argument structure. In Proceedings of the Workshop on Human Language Technology, HLT ’94, pp. 114–119, USA, 1994. Association for Computational Linguistics. ISBN 1558603573. doi: 10.3115/1075812.1075835.
  27. 27.Merity, S., Xiong, C., Bradbury, J., and Socher, R. Pointer sentinel mixture models, 2016.
  28. 28.Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Wang, Z., Wong, R. Y. Y., Zhu, A., Yang, L., Shi, X., Shi, C., Chen, Z., Arfeen, D., Abhyankar, R., and Jia, Z. SpecInfer: Accelerating generative large language model serving with speculative inference and token tree verification, 2023.
  29. 29.NVIDIA. TensorRT-LLM. URL https://github.com/NVIDIA/TensorRT-LLM.
  30. 30.Park, G., Park, B., Kim, M., Lee, S., Kim, J., Kwon, B., Kwon, S. J., Kim, B., Lee, Y., and Lee, D. LUT-GEMM: Quantized matrix multiplication based on LUTs for efficient inference in large-scale generative language models. In The Twelfth International Conference on Learning Representations, 2024.
  31. 31.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer, 2023.
  32. 32.Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. WinoGrande: an adversarial winograd schema challenge at scale. Commun. ACM, 64(9):99–106, aug 2021. ISSN 0001-0782. doi: 10.1145/3474381.
  33. 33.Santacroce, M., Wen, Z., Shen, Y., and Li, Y. What matters in the structured pruning of generative language models?, 2023.
  34. 34.Sutskever, I., Vinyals, O., and Le, Q. V. Sequence to sequence learning with neural networks. In Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 2, NIPS’14, pp. 3104–3112, Cambridge, MA, USA, 2014. MIT Press.
  35. 35.Tata, S. and Patel, J. M. PiQA: An algebra for querying protein data sets. In Proceedings of the 15th International Conference on Scientific and Statistical Database Management (SSDBM 2003), 9-11 July 2003, Cambridge, MA, USA, pp. 141–150. IEEE Computer Society, 2003. doi: 10.1109/SSDM.2003.1214975.
  36. 36.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, E., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T. Llama 2: Open foundation and fine-tuned chat models, 2023.
  37. 37.turboderp. ExLlamaV2. URL https://github.com/turboderp/exllamav2.
  38. 38.Umuroglu, Y. and Jahre, M. Work-in-progress: towards efficient quantized neural network inference on mobile devices. In 2017 International Conference on Compilers, Architectures and Synthesis For Embedded Systems (CASES), pp. 1–2, 2017. doi: 10.1145/3125501.3125528.
  39. 39.Wan, C., Santriaji, M., Rogers, E., Hoffmann, H., Maire, M., and Lu, S. ALERT: Accurate learning for energy and timeliness. In 2020 USENIX Annual Technical Conference (USENIX ATC 20), pp. 353–369. USENIX Association, July 2020. ISBN 978-1-939133-14-4.
  40. 40.Warren, H. S. Hacker’s Delight. Addison-Wesley Professional, 2nd edition, 2012. ISBN 0321842685.
  41. 41.Yu, H., Li, H., Shi, H., Huang, T. S., and Hua, G. Any-precision deep neural networks. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2-9, 2021, pp. 10763–10771. AAAI Press, 2021. doi: 10.1609/AAAI.V35I12.17286.
  42. 42.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. HellaSwag: Can a machine really finish your sentence? In Korhonen, A., Traum, D., and Màrquez, L. (eds.), Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 4791–4800, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1472.
  43. 43.Zhang, J., Elnikety, S., Zarar, S., Gupta, A., and Garg, S. Model-Switching: Dealing with fluctuating workloads in Machine-Learning-as-a-Service systems. In 12th USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 20). USENIX Association, July 2020.
  44. 44.Zhang, M., Chen, H., Shen, C., Yang, Z., Ou, L., Yu, X., and Zhuang, B. LoRAPrune: Pruning meets low-rank parameter-efficient fine-tuning, 2023.
  45. 45.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L. OPT: Open pre-trained transformer language models, 2022.
  46. 46.Zhao, B., Cui, Q., Song, R., Qiu, Y., and Liang, J. Decoupled knowledge distillation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pp. 11943–11952. IEEE, 2022. doi: 10.1109/CVPR52688.2022.01165.

Citation

MLA
Park, Y., et al. “Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs”. arXiv, 2024, http://arxiv.org/abs/2402.10517v4.
APA
Park, Y., Hyun, J., Cho, S., Sim, B., & Lee, J. W. (2024). Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs. arXiv. http://arxiv.org/abs/2402.10517v4
Chicago
Park, Y., J. Hyun, S. Cho, B. Sim, and J. W. Lee. 2024. “Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs”. arXiv. http://arxiv.org/abs/2402.10517v4.
Harvard
Park, Y. et al. (2024) “Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.10517v4.
Vancouver
1. Park Y, Hyun J, Cho S, Sim B, Lee JW (2024) Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs. arXiv

BibTeX

@article{park2024any,
  title = {Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs},
  author = {Park, Yeonhong and Hyun, Jake and Cho, SangLyul and Sim, Bonggeun and Lee, Jae W.},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.10517v4},
  eprint = {2402.10517}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/