Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs
Yeonhong ParkJake HyunSangLyul ChoBonggeun SimJae W. Lee
Proposes a post-training quantization framework and a specialized GPU serving engine that pack multiple Large Language Models of varying bit-widths into the memory footprint of a single high-precision model without sacrificing inference speed or output quality.
Deploying large language models often requires serving multiple model sizes concurrently to handle varying response-time requirements, execute multi-task workloads, and support acceleration techniques like speculative decoding. However, storing multiple distinct models requires immense hardware memory and substantial training compute, posing severe bottlenecks for resource-constrained environments like desktop and mobile edge devices.
The article demonstrates an "any-precision" framework that allows multiple models of varying precision levels—ranging from 3-bit to 8-bit—to be served from a single parent model footprint. By combining a lightweight post-training quantization method with a specialized graphics processing unit execution engine, the approach enables dynamic model scaling without storing redundant weight parameters or retraining models from scratch.
To achieve this, the authors developed an incremental upscaling technique using non-uniform, clustering-based quantization. Starting with a compact 3-bit seed model, the method iteratively appends single bits to split weight clusters into higher-precision representations. Complementing this algorithm, the authors designed a software execution engine based on a bitplane memory layout, which allows hardware to load only the exact bit-width required for a specific query instead of reading full bit-vectors.
Key findings show substantial resource and performance gains. Storing a complete suite of precision levels from 3-bit to 8-bit in a single any-precision model reduces total memory consumption by up to 3.56 times compared to deploying each model separately. Models generated through incremental upscaling match the state-of-the-art accuracy and language quality of independently trained models at each respective bit-width, showing negligible quality degradation across standard evaluation benchmarks. Furthermore, the complete quantization process takes under one minute on standard consumer hardware, and the custom execution engine delivers inference speeds that match or exceed existing fixed-precision engines across desktop, laptop, and mobile processors.
These findings indicate that organizations can significantly lower hardware infrastructure costs, improve query throughput, and deploy adaptable on-device language models without costly re-training pipelines. Unlike uniform quantization techniques, which suffer severe quality degradation under incremental upscaling due to cumulative rounding errors, clustering-based non-uniform quantization provides a stable and mathematically robust foundation for flexible-precision serving.
Organizations serving varied artificial intelligence workloads should consider adopting bitplane-based and codebook-quantized architectures when designing low-latency, on-device serving systems. For deployments requiring high batch sizes or long input prompt processing, teams should implement hybrid execution strategies that transition to standard compute kernels to maintain optimal hardware utilization.
Confidence in these findings is high across tested open-source models (including LLaMA, Mistral, and OPT architectures) and consumer hardware tiers. However, decision-makers should note that the current kernel implementation is optimized primarily for smaller batch sizes and memory-bound generation phases, and extended validation on larger model architectures and diverse production workloads is recommended prior to broad organizational rollout.
- Paper: SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models, Guangxuan Xiao et al. (2023). SmoothQuant establishes training-free LLM quantization and the accuracy–efficiency trade-offs that frame this paper’s more flexible precision strategy.
- Paper: GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, Elias Frantar et al. (2023). GPTQ provides a key post-training low-bit LLM quantization baseline, clarifying the accuracy and deployment challenges this paper’s incremental method addresses.
- Paper: SqueezeLLM: Dense-and-Sparse Quantization, Sehoon Kim et al. (2024). SqueezeLLM shows how non-uniform low-bit quantization can preserve LLM quality, preparing readers for this paper’s clustering-based precision upscaling.
No sufficiently relevant recommendations were found.
