| 研究生: |
王偉力 Wang, Wei-Li |
|---|---|
| 論文名稱: |
基於位移式步長調整之變換器溢位感知靜態量化 Overflow-Aware Static Quantization for Transformers via Shift-Based Scale Adjustment |
| 指導教授: |
郭致宏
Kuo, Chih-Hung |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 電機工程學系 Department of Electrical Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 102 |
| 中文關鍵詞: | 大語言模型 、訓練後量化 、靜態量化 |
| 外文關鍵詞: | Large language models, Post-training quantization, Static quantization |
| 相關次數: | 點閱:87 下載:1 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
本論文提出一種適用於大型語言模型 (Large Language Model, LLM) 低位元推論之溢位感知靜態量化方法 (Overflow-Aware Static Quantization, OVASQ),在維持靜態量化低推論成本的同時,提高其對激活值範圍變化的適應能力。傳統靜態量化使用固定量化步長,容易因離群值或輸入差異造成數值裁剪與表示解析度下降;動態量化雖可調整尺度,卻需於推論期間重複掃描激活值並計算量化參數。為改善上述問題,OVASQ 於離線階段決定初始量化步長,並於推論期間以區塊為單位偵測溢位,透過位移運算擴大量化範圍,以避免高幅值激活值飽和。此外,本論文採用定點量化步長細化機制,在不增加線上浮點尺度計算的情況下降低量化誤差。實驗結果顯示,在 Llama 2 7B與 WikiText-2 資料集上,OVASQ 於 W4A4 設定下可維持接近浮點模型的困惑度,並於 W4A3 與 W3A4 設定下顯著改善傳統靜態量化的模型效能。相較於需搬運 FP16 激活值進行線上尺度計算的動態量化流程,OVASQ 僅需傳遞少量位移資訊,即可減少約 75% 的激活值資料搬運量。
This paper proposes Overflow-Aware Static Quantization (OVASQ), a low-bit inference method for large language models (LLMs) that improves adaptability to runtime activation-range variations while preserving the low online overhead of static quantization. Conventional static quantization employs fixed quantization step sizes and is therefore susceptible to value clipping and reduced representation resolution caused by activation outliers and input-dependent distribution shifts. Although dynamic quantization can adapt its scales to the current input, it requires repeated activation scans and quantization-parameter computation during inference. To address these limitations, OVASQ determines the initial quantization step sizes offline and performs tile-level overflow detection during inference. When an output exceeds the representable range of the low-bit integer format, bit-shift operations are applied to expand the quantization range and prevent the saturation of large-magnitude activations. In addition, a fixed- point quantization step-size refinement mechanism is introduced to reduce quantization error without incurring online floating-point scale computation. Experimental results on Llama 2 7B with the WikiText-2 dataset show that OVASQ achieves perplexity close to that of the floating-point model under the W4A4 setting and significantly improves the performance of conventional static quantization under the lower-bit W4A3 and W3A4 settings. Compared with dynamic quantization workflows that transfer FP16 activations for online scale computation, OVASQ requires only a small amount of additional shift metadata and reduces activation data movement by approximately 75%.
[1] A. Gholami, Z. Yao, S. Kim, C. Hooper, M. W. Mahoney, and K. Keutzer, “Ai and Memory Wall,” IEEE Micro, vol. 44, no. 3, pp. 33–39, 2024.
[2] T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer, “LLM. int8 (): 8-bit Matrix Multiplication for Transformers at Scale,” arXiv preprint arXiv:2208.07339, 2022.
[3] G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han, “SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models,” in International conference on machine learning, pp. 38087–38099, PMLR, 2023.
[4] Z. Yao, R. Yazdani Aminabadi, M. Zhang, X. Wu, C. Li, and Y. He, “ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers,” Advances in neural information processing systems, vol. 35, pp. 27168–27183, 2022.
[5] Y. Zhao, C.-Y. Lin, K. Zhu, Z. Ye, L. Chen, S. Zheng, L. Ceze, A. Krishnamurthy, T. Chen, and B. Kasikci, “Atom: Low-Bit Quantization for Efficient and Accurate Llm Serving,” Proceedings of Machine Learning and Systems, vol. 6, pp. 196–209, 2024.
[6] Y. Sun, R. Liu, H. Bai, H. Bao, K. Zhao, Y. Li, J. Hu, X. Yu, L. Hou, C. Yuan, et al., “FlatQuant: Flatness Matters for Llm Quantization,” arXiv preprint arXiv:2410.09426, 2024.
[7] D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al., “Mastering the game of go without human knowl- edge,” nature, vol. 550, no. 7676, pp. 354–359, 2017.
[8] J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al., “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774, 2023.
[9] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention Is All You Need,” Advances in neural information processing systems, vol. 30, 2017.
[10] H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al., “Llama 2: Open Foundation and Fine-Tuned Chat Models,” arXiv preprint arXiv:2307.09288, 2023.
[11] S. Han, H. Mao, and W. J. Dally, “Deep Compression: Compressing Deep Neural Networks With Pruning, Trained Quantization and Huffman Coding,” ICLR, 2016.
[12] J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, “Roformer: Enhanced transformer with Rotary Position Embedding,” Neurocomputing, vol. 568, p. 127063, 2024.
[13] B. Zhang and R. Sennrich, “Root Mean Square Layer Normalization,” Advances in neural information processing systems, vol. 32, 2019.
[14] N. Shazeer, “Glu Variants Improve Transformer,” arXiv preprint arXiv:2002.05202, 2020.
[15] Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, and H. Zhang, “{DistServe}: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving,” in 18th USENIX Symposium on Operating Systems Design and Implementa- tion (OSDI 24), pp. 193–210, 2024.
[16] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient Memory Management for Large Language Model Serving with PagedAttention,” in Proceedings of the 29th symposium on operating systems princi- ples, pp. 611–626, 2023.
[17] T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré, “FlashAttention: Fast and Memory- Efficient Exact Attention with IO-Awareness,” Advances in neural information pro- cessing systems, vol. 35, pp. 16344–16359, 2022.
[18] X. Ma, G. Fang, and X. Wang, “LLM-Pruner: On the Structural Pruning of Large Language Models,” Advances in neural information processing systems, vol. 36, pp. 21702–21720, 2023.
[19] X. Wang, Y. Zheng, Z. Wan, and M. Zhang, “SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model compression,” in International Con- ference on Learning Representations, vol. 2025, pp. 19299–19319, 2025.
[20] G. Hinton, O. Vinyals, and J. Dean, “Distilling the Knowledge in a Neural Network,” arXiv preprint arXiv:1503.02531, 2015.
[21] B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko, “Quantization and Training of Neural Networks for Efficient Integer- Arithmetic-Only Inference,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2704–2713, 2018.
[22] M. Nagel, M. Fournarakis, R. A. Amjad, Y. Bondarenko, M. Van Baalen, and T. Blankevoort, “A white paper on Neural Network Quantization,” arXiv preprint arXiv:2106.08295, 2021.
[23] E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh, “GPTQ: Accurate Post- Training Quantization for Generative Pre-trained Transformers,” arXiv preprint arXiv:2210.17323, 2022.
[24] S. Kim, A. Gholami, Z. Yao, M. W. Mahoney, and K. Keutzer, “I-BERT: Integer-only BERT Quantization,” in International conference on machine learning, pp. 5506–5518, PMLR, 2021.
[25] B. A. Motetti, M. Risso, A. Burrello, E. Macii, M. Poncino, and D. J. Pagliari, “Joint Pruning and Channel-Wise Mixed-Precision Quantization for Efficient Deep Neural Networks,” IEEE Transactions on Computers, vol. 73, no. 11, pp. 2619–2633, 2024.
[26] S. Merity, C. Xiong, J. Bradbury, and R. Socher, “Pointer Sentinel Mixture Models,” arXiv preprint arXiv:1609.07843, 2016.