SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks

Jiwon SongKyungseok OhTaesu KimHyungjun KimYulhwa KimJae-Joon Kim

article2024ICML81 citations

Proposes a block-level pruning method called SLEB that removes redundant transformer blocks based on layer output similarity, achieving real hardware acceleration and low perplexity degradation without requiring specialized sparse matrix support.

Listen

Large language models provide powerful natural language processing capabilities, but their immense size creates substantial memory and computational challenges during real-world deployment. Traditional compression techniques like weight pruning often fail to deliver practical acceleration because modern hardware struggles with sparse operations, while dynamic methods like early exiting cannot reduce memory footprints and perform poorly when processing multiple user requests simultaneously.

The article demonstrates and evaluates SLEB, a streamlined pruning method designed to accelerate large language models by identifying and removing entire redundant transformer blocks without requiring expensive model retraining.

The authors developed an iterative, training-free approach that evaluates the importance of each architectural block based on total network output impact, pruning the most redundant blocks one at a time using a small set of 128 calibration text samples. They tested this method across several model architectures, including the OPT and LLaMA-2 model families ranging from 6.7 billion to 70 billion parameters, and benchmarked processing latency, multi-user throughput, language quality, and performance on standard reasoning tasks against leading pruning alternatives.

The evaluation yielded several key findings:

  1. Removing entire blocks translates directly into runtime speedups. On a 70-billion-parameter model, pruning 20% of the blocks yielded a 1.27 times improvement in throughput during generation and a 1.26 times reduction in prompt latency, whereas existing weight pruning methods produced negligible speedup or even slowed down generation.
  2. Language quality and reasoning accuracy remained well-preserved up to a 20% pruning ratio, outperforming existing fine-grained pruning methods across zero-shot benchmarks.
  3. The method exhibited high stability across datasets and executed quickly, compressing a 70-billion-parameter model in approximately 1.5 hours on two enterprise graphics processors.
  4. Pruning entire blocks demonstrated complete compatibility with 4-bit post-training quantization, enabling compound compression without additional loss in language fluency.

These findings indicate that architectural redundancy at the block level can be exploited to achieve tangible operational cost reductions and lower response latency. Because the approach permanently deletes entire blocks, organizations can simultaneously shrink hardware memory footprints and improve batched request efficiency without needing specialized sparsity hardware or costly retraining cycles.

Organizations seeking to optimize large language model serving should consider adopting block-level pruning as an alternative to fine-grained weight pruning, particularly when targeting compression rates up to 20%. For maximum memory savings, teams should combine this approach with standard 4-bit weight quantization.

Confidence in these findings is high for models within the 6.7-billion to 70-billion parameter range up to moderate pruning targets. However, stakeholders should note that performance degrades significantly if more than 20% to 30% of blocks are removed without fine-tuning, and results may vary when applied to fundamentally different model architectures.

Song et al (2024).pdf

No sufficiently relevant recommendations were found.

Cover for SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks

Table of Contents

  • Abstract
  • 1. Introduction
  • 2. Motivation
  • 2.1. Pruning
  • 2.2. Early Exit
  • 3. Proposed SLEB
  • 3.1. Output Similarity across Transformer Blocks
  • 3.2. Redundancy Verification of Transformer Blocks
  • 3.3. Proposed SLEB Algorithm
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Elimination of Transformer Blocks using SLEB
  • 4.3. Language Modeling
  • 4.4. Dependency on Calibration Dataset
  • 4.5. Zero-shot Tasks
  • 4.6. Speedup
  • 4.7. Compatibility with Post-Training Quantization
  • 5. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. SLEB details
  • A.1. Runtime
  • A.2. Selected Transformer Blocks
  • B. More Evaluation Results
  • B.1. Limitations of Early Exit
  • B.2. Language Modeling
  • B.3. Dependency on Calibration Dataset
  • B.4. Zero-shot tasks
  • B.5. Speedup
  • B.6. Compatibility with Post-Training Quantization
  • B.7. Perplexity and Accuracy Results of SLEB across Various Sparsity Ratios
  • B.8. Fine-tuning

Knowls

  1. Knowl 1 — SLEB removes whole transformer blocks using static, model-specific pruning

    model/method

    SLEB compresses a pretrained large language model by permanently deleting complete transformer blocks, rather than pruning weights or channels within blocks. It uses calibration data to choose blocks whose removal least affects the model’s token predictions, and recalculates this choice after each deletion because block importance changes as the model changes. The resulting pruned model uses a fixed block structure at inference time: it does not require per-token skip decisions or additional training. Removing whole blocks also removes their parameters and computation, so the intended inference acceleration tracks the number of blocks deleted.

  2. Knowl 2 — Iterative negative log-likelihood is SLEB’s block-removal criterion

    equation

    For a target language model MM, SLEB scores each candidate block by the average token negative log-likelihood after removing that block. Given calibration tokens w1,…,wKw_1,\ldots,w_K, let M−jM_{-j} denote the model obtained by deleting block jj from the current model. The score is

    Sj(M)=−1K∑k=1Klog⁡pM−j(wk∣w<k),S_j(M)=-\frac{1}{K}\sum_{k=1}^{K}\log p_{M_{-j}}(w_k\mid w_{<k}),

    where pM−j(wk∣w<k)p_{M_{-j}}(w_k\mid w_{<k}) is the probability assigned to token wkw_k given its preceding tokens by the candidate model. SLEB removes the block with the smallest score, corresponding to the lowest calibration-set negative log-likelihood among the block-deletion candidates. Importantly, the candidate models are formed from the model as it exists at that iteration, not always from the original model. This iterative reassessment addresses the changing importance of blocks after earlier deletions.

    The paper also tests two alternatives. A local score, 1−cos⁡(Aj,Bj)1-\cos(A_j,B_j), compares a block’s input AjA_j and output BjB_j; it can miss small early changes that later affect predictions. A one-shot score evaluates each block deletion against the original model, rather than the already-pruned model; selecting blocks independently this way can produce harmful runs of consecutive deletions. Both alternatives yield worse perplexity than the iterative criterion in the reported comparisons.

  3. Knowl 3 — SLEB’s iterative block-selection procedure

    algorithm

    Given a pretrained model, a calibration set, and a target number nn of blocks to remove, SLEB tests each remaining block as a deletion candidate, evaluates the candidate model’s average calibration-token negative log-likelihood, deletes the candidate with the lowest score, and repeats on the updated model. The output is the model with nn blocks removed. At each iteration, every remaining block is considered; thus, for an original model with NN blocks, there are at most nNnN candidate-model evaluations, each requiring evaluation on the calibration set. The reported experiments use 128 randomly sampled WikiText-2 training sequences for calibration and do not fine-tune the pruned model.

    Input: pretrained model M with N transformer blocks, calibration tokens C, removal count n
    Output: pruned model M
    for i = 0 to n - 1
        bestScore = infinity
        bestBlock = none
        for each remaining block j in M
            candidate = M with block j removed
            score = average token negative log-likelihood of candidate on C
            if score < bestScore
                bestScore = score
                bestBlock = j
            end if
        end for
        M = M with bestBlock removed
    end for
    return M
  4. Knowl 4 — Adjacent transformer blocks have highly similar outputs

    empirical result

    The paper measures block-output similarity using cosine similarity. In a residual transformer stack, the output after block i+1i+1 is represented as xi+1=Ti+1(xi)+xix_{i+1}=T_{i+1}(x_i)+x_i, where xix_i is the representation after block ii and Ti+1T_{i+1} is the computation contributed by block i+1i+1. For a single random input token, the authors compare cosine similarities between representations at different block positions. In OPT-13B and LLaMA-2-13B, the measured similarity is consistently high between neighboring blocks, while similarity between blocks farther apart varies across the model. This observation motivates checking for block redundancy, but does not by itself establish that any particular block can be removed without affecting model quality.

  5. Knowl 5 — Evaluation uses pretrained OPT and LLaMA-2 models without pruning-time fine-tuning

    experimental setup

    SLEB is evaluated on OPT-6.7B, OPT-13B, OPT-30B, OPT-66B, LLaMA-2-7B, LLaMA-2-13B, and LLaMA-2-70B. The main pruning targets are 10% and 20% of transformer blocks; when the target fraction does not give an integer block count, the count is rounded up. Calibration uses 128 sequences randomly sampled from the WikiText-2 training set, and the reported pruning process is inference-only, without fine-tuning. Perplexity is evaluated on C4 and WikiText-2, and zero-shot accuracy is averaged over PIQA, WinoGrande, HellaSwag, ARC-easy, and ARC-challenge. Pruning experiments use NVIDIA A100 GPUs with 80 GB memory; OPT-66B and LLaMA-2-70B pruning require two GPUs.

  6. Knowl 6 — SLEB preserves C4 perplexity at 10% and 20% block removal

    empirical result

    The following C4 validation perplexities are reported in model order OPT-6.7B, OPT-13B, OPT-30B, OPT-66B, LLaMA-2-7B, LLaMA-2-13B, and LLaMA-2-70B. Dense-model perplexities are 12.71,12.06,11.44,10.99,7.26,6.73,5.7112.71, 12.06, 11.44, 10.99, 7.26, 6.73, 5.71. With SLEB removing 10% of blocks, they are 13.84,12.43,11.65,11.25,9.34,7.80,6.3213.84, 12.43, 11.65, 11.25, 9.34, 7.80, 6.32; with 20% removed, they are 15.99,13.81,12.74,12.54,12.32,9.42,7.3115.99, 13.81, 12.74, 12.54, 12.32, 9.42, 7.31. Thus the reported degradation is moderate at the tested 10% and 20% block-removal levels. The authors report that these results compare favorably with weight- and channel-pruning baselines in most tested cases, despite SLEB removing coarser-grained units.

  7. Knowl 7 — Whole-block removal improves end-to-end inference speed on LLaMA-2

    empirical result

    On LLaMA-2-70B using two NVIDIA A100 GPUs, SLEB improves both prompt-processing latency and token-generation throughput. For prompt processing of a 2,048-token input, dense inference takes 1,718.4 ms; SLEB at 10% and 20% block removal takes 1,529.1 ms (1.12×\times speedup) and 1,364.1 ms (1.26×\times), respectively. For token generation of 128 tokens at batch size 64, throughput is 299 tokens/s for the dense model, 336 tokens/s at 10% removal (1.12×\times), and 381 tokens/s at 20% removal (1.27×\times). In the same setup, 2:4 pruning gives 293 tokens/s (0.98×\times) and 25% channel pruning gives 331 tokens/s (1.11×\times). The reported results illustrate that removing whole blocks yields end-to-end gains in both inference stages, whereas weight-pruning speedups vary with the serving scenario.

  8. Knowl 8 — SLEB retains higher zero-shot accuracy than most pruning baselines

    empirical result

    Mean accuracy (%) across PIQA, WinoGrande, HellaSwag, ARC-easy, and ARC-challenge is reported below in model order OPT-6.7B, OPT-13B, OPT-30B, OPT-66B, LLaMA-2-7B, LLaMA-2-13B, and LLaMA-2-70B. Dense-model means are 60.70,61.79,64.40,66.16,69.00,71.76,76.5760.70, 61.79, 64.40, 66.16, 69.00, 71.76, 76.57. SLEB at 10% block removal scores 60.00,62.07,64.48,65.38,62.24,66.77,73.1460.00, 62.07, 64.48, 65.38, 62.24, 66.77, 73.14; at 20%, it scores 57.61,60.08,62.86,62.53,56.80,62.96,70.8157.61, 60.08, 62.86, 62.53, 56.80, 62.96, 70.81. The best reported non-SLEB baseline means for those models are 56.31,60.20,63.65,65.74,58.23,63.45,72.3456.31, 60.20, 63.65, 65.74, 58.23, 63.45, 72.34, respectively; the best baseline varies among the compared methods. SLEB at 10% exceeds that baseline on six of the seven models, with OPT-66B the exception (65.38 versus 65.74).

  9. Knowl 9 — SLEB’s perplexity is less sensitive to the calibration dataset than SliceGPT’s

    empirical result

    For LLaMA-2-70B, the paper compares perplexity when calibration uses 128 sequences from WikiText-2 or C4 training data and evaluation uses WikiText-2 or C4. SLEB removes 20% of blocks; SparseGPT and Wanda use 2:4 pruning, and SliceGPT removes 25% of channels. On C4 evaluation, changing calibration from WikiText-2 to C4 changes perplexity from 7.31 to 7.26 for SLEB, from 8.42 to 8.16 for SparseGPT, from 8.23 to 8.10 for Wanda, and from 20.03 to 8.11 for SliceGPT. On WikiText-2 evaluation, the corresponding WikiText-2-to-C4 calibration results are 4.88 to 5.16 for SLEB, 4.98 to 5.70 for SparseGPT, 5.26 to 5.48 for Wanda, and 4.89 to 7.76 for SliceGPT. These measurements support the paper’s finding that SLEB is relatively insensitive to which of the two calibration datasets is used, particularly compared with SliceGPT.

  10. Knowl 10 — Quality degrades sharply when SLEB removes substantially more than 20% of blocks

    limitation

    The paper evaluates SLEB at removal levels from 10% to 50% and finds progressively worse language-modeling and zero-shot results at higher levels. For OPT-6.7B, C4 perplexity and mean zero-shot accuracy are 15.99 and 57.61% at 20% removal, 21.11 and 52.93% at 30%, 37.90 and 47.81% at 40%, and 135.49 and 40.95% at 50%. For LLaMA-2-70B, the corresponding pairs are 7.31 and 70.81%, 8.64 and 67.37%, 12.16 and 62.42%, and 19.68 and 56.49%. These results qualify the paper’s favorable findings at 10%–20%: substantially more aggressive block removal carries a large quality cost.

Coverage note — The paper’s 4-bit AWQ compatibility experiment and its LoRA fine-tuning comparison are omitted as secondary extensions; neither changes the central block-selection method or its main inference and quality evaluations.

References

  1. 1.Ashkboos, S., Croci, M. L., do Nascimento, M. G., Hoefler, T., and Hensman, J. SliceGPT: Compress large language models by deleting rows and columns. In The Twelfth International Conference on Learning Representations, 2024.
  2. 2.Bisk, Y., Zellers, R., Le bras, R., Gao, J., and Choi, Y. Piqa: Reasoning about physical commonsense in natural language. Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):7432–7439, Apr. 2020. doi: 10.1609/aaai.v34i05.6239. URL https://ojs.aaai.org/index.php/AAAI/article/view/6239.
  3. 3.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  4. 4.Chen, Y., Pan, X., Li, Y., Ding, B., and Zhou, J. Ee-llm: Large-scale training and inference of early-exit large language models with 3d parallelism. arXiv preprint arXiv:2312.04916, 2023.
  5. 5.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N. Palm: Scaling language modeling with pathways, 2022.
  6. 6.Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018.
  7. 7.Del Corro, L., Del Giorno, A., Agarwal, S., Yu, B., Awadallah, A., and Mukherjee, S. Skipdecode: Autoregressive skip decoding with batching and caching for efficient llm inference. arXiv preprint arXiv:2307.02628, 2023.
  8. 8.Din, A. Y., Karidi, T., Choshen, L., and Geva, M. Jump to conclusions: Short-cutting transformers with linear transformations. arXiv preprint arXiv:2303.09435, 2023.
  9. 9.Frantar, E. and Alistarh, D. Optimal brain compression: A framework for accurate post-training quantization and pruning. Advances in Neural Information Processing Systems, 35:4475–4488, 2022.
  10. 10.Frantar, E. and Alistarh, D. Sparsegpt: Massive language models can be accurately pruned in one-shot. In International Conference on Machine Learning, pp. 10323–10337. PMLR, 2023.
  11. 11.Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D. Gptq: Accurate post-training quantization for generative pretrained transformers. International Conference on Learning Representations, 2023.
  12. 12.Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. A framework for few-shot language model evaluation, 12 2023. URL https://zenodo.org/records/10256836.
  13. 13.Han, S., Pool, J., Tran, J., and Dally, W. Learning both weights and connections for efficient neural network. Advances in neural information processing systems, 28, 2015.
  14. 14.Hassibi, B., Stork, D. G., and Wolff, G. J. Optimal brain surgeon and general network pruning. In IEEE international conference on neural networks, pp. 293–299. IEEE, 1993.
  15. 15.Hu, E. J., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al. Lora: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2021.
  16. 16.LeCun, Y., Denker, J., and Solla, S. Optimal brain damage. Advances in neural information processing systems, 2, 1989.
  17. 17.Lee, C., Jin, J., Kim, T., Kim, H., and Park, E. Owq: Lessons learned from activation outliers for weight quantization in large language models. In Proceedings of the AAAI Conference on Artificial Intelligence, 2024.
  18. 18.Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S. Awq: Activation-aware weight quantization for llm compression and acceleration. arXiv preprint arXiv:2306.00978, 2023.
  19. 19.Liu, Z., Wang, J., Dao, T., Zhou, T., Yuan, B., Song, Z., Shrivastava, A., Zhang, C., Tian, Y., Re, C., et al. Deja vu: Contextual sparsity for efficient llms at inference time. In International Conference on Machine Learning, pp. 22137–22176. PMLR, 2023.
  20. 20.Ma, X., Fang, G., and Wang, X. Llm-pruner: On the structural pruning of large language models. In Thirty-seventh Conference on Neural Information Processing Systems, 2023.
  21. 21.Merity, S., Xiong, C., Bradbury, J., and Socher, R. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843, 2016.
  22. 22.Mishra, A., Latorre, J. A., Pool, J., Stosic, D., Stosic, D., Venkatesh, G., Yu, C., and Micikevicius, P. Accelerating sparse deep neural networks. arXiv preprint arXiv:2104.08378, 2021.
  23. 23.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32, 2019.
  24. 24.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020. URL http://jmlr.org/papers/v21/20-074.html.
  25. 25.Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. Winogrande: An adversarial winograd schema challenge at scale. arXiv preprint arXiv:1907.10641, 2019.
  26. 26.Schuster, T., Fisch, A., Gupta, J., Dehghani, M., Bahri, D., Tran, V., Tay, Y., and Metzler, D. Confident adaptive language modeling. Advances in Neural Information Processing Systems, 35:17456–17472, 2022.
  27. 27.Shi, S., Wang, Q., and Chu, X. Efficient sparse-dense matrix-matrix multiplication on gpus using the customized sparse storage format. In 2020 IEEE 26th International Conference on Parallel and Distributed Systems (ICPADS), pp. 19–26. IEEE, 2020.
  28. 28.Sun, M., Liu, Z., Bair, A., and Kolter, J. Z. A simple and effective pruning approach for large language models. arXiv preprint arXiv:2306.11695, 2023.
  29. 29.Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca, 2023.
  30. 30.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G. Llama: Open and efficient foundation language models, 2023a.
  31. 31.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  32. 32.Varshney, N., Chatterjee, A., Parmar, M., and Baral, C. Accelerating llama inference by enabling intermediate layer decoding via instruction tuning with lite. arXiv e-prints, pp. arXiv–2310, 2023.
  33. 33.Wang, Z. Sparsert: Accelerating unstructured sparsity on gpus for deep learning inference. In Proceedings of the ACM international conference on parallel architectures and compilation techniques, pp. 31–42, 2020.
  34. 34.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 38–45, Online, October 2020. Association for Computational Linguistics. URL https://www.aclweb.org/anthology/2020.emnlp-demos.6.
  35. 35.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.
  36. 36.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L. Opt: Open pre-trained transformer language models, 2022.
  37. 37.Zhang, Y., Zhao, L., Lin, M., Sun, Y., Yao, Y., Han, X., Tanner, J., Liu, S., and Ji, R. Dynamic sparse no training: Training-free fine-tuning for sparse llms. In The Twelfth International Conference on Learning Representations, 2024.

Citation

MLA
Song, J., et al. “SLEB: Streamlining LLMs Through Redundancy Verification and Elimination of Transformer Blocks”. arXiv, 2024, https://doi.org/10.48550/arxiv.2402.09025.
APA
Song, J., Oh, K., Kim, T., Kim, H., Kim, Y., & Kim, J.-J. (2024). SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks. arXiv. https://doi.org/10.48550/arxiv.2402.09025
Chicago
Song, J., K. Oh, T. Kim, H. Kim, Y. Kim, and J.-J. Kim. 2024. “SLEB: Streamlining LLMs Through Redundancy Verification and Elimination of Transformer Blocks”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2402.09025.
Harvard
Song, J. et al. (2024) “SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks”. arXiv. Available at: https://doi.org/10.48550/arxiv.2402.09025.
Vancouver
1. Song J, Oh K, Kim T, Kim H, Kim Y, Kim J-J (2024) SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks. https://doi.org/10.48550/arxiv.2402.09025

BibTeX

@misc{https://doi.org/10.48550/arxiv.2402.09025,
  doi = {10.48550/ARXIV.2402.09025},
  url = {https://arxiv.org/abs/2402.09025},
  author = {Song, Jiwon and Oh, Kyungseok and Kim, Taesu and Kim, Hyungjun and Kim, Yulhwa and Kim, Jae-Joon},
  keywords = {Computation and Language (cs.CL), Machine Learning (cs.LG), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks},
  publisher = {arXiv},
  year = {2024},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/