簡易檢索 / 詳目顯示

研究生: 王偉力
Wang, Wei-Li
論文名稱: 基於位移式步長調整之變換器溢位感知靜態量化
Overflow-Aware Static Quantization for Transformers via Shift-Based Scale Adjustment
指導教授: 郭致宏
Kuo, Chih-Hung
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 電機工程學系
Department of Electrical Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 102
中文關鍵詞: 大語言模型 、訓練後量化 、靜態量化
外文關鍵詞: Large language models, Post-training quantization, Static quantization
相關次數: 點閱:87  下載:1 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本論文提出一種適用於大型語言模型 (Large Language Model, LLM) 低位元推論之溢位感知靜態量化方法 (Overflow-Aware Static Quantization, OVASQ),在維持靜態量化低推論成本的同時,提高其對激活值範圍變化的適應能力。傳統靜態量化使用固定量化步長,容易因離群值或輸入差異造成數值裁剪與表示解析度下降;動態量化雖可調整尺度,卻需於推論期間重複掃描激活值並計算量化參數。為改善上述問題,OVASQ 於離線階段決定初始量化步長,並於推論期間以區塊為單位偵測溢位,透過位移運算擴大量化範圍,以避免高幅值激活值飽和。此外,本論文採用定點量化步長細化機制,在不增加線上浮點尺度計算的情況下降低量化誤差。實驗結果顯示,在 Llama 2 7B與 WikiText-2 資料集上,OVASQ 於 W4A4 設定下可維持接近浮點模型的困惑度,並於 W4A3 與 W3A4 設定下顯著改善傳統靜態量化的模型效能。相較於需搬運 FP16 激活值進行線上尺度計算的動態量化流程,OVASQ 僅需傳遞少量位移資訊,即可減少約 75% 的激活值資料搬運量。

    This paper proposes Overflow-Aware Static Quantization (OVASQ), a low-bit inference method for large language models (LLMs) that improves adaptability to runtime activation-range variations while preserving the low online overhead of static quantization. Conventional static quantization employs fixed quantization step sizes and is therefore susceptible to value clipping and reduced representation resolution caused by activation outliers and input-dependent distribution shifts. Although dynamic quantization can adapt its scales to the current input, it requires repeated activation scans and quantization-parameter computation during inference. To address these limitations, OVASQ determines the initial quantization step sizes offline and performs tile-level overflow detection during inference. When an output exceeds the representable range of the low-bit integer format, bit-shift operations are applied to expand the quantization range and prevent the saturation of large-magnitude activations. In addition, a fixed- point quantization step-size refinement mechanism is introduced to reduce quantization error without incurring online floating-point scale computation. Experimental results on Llama 2 7B with the WikiText-2 dataset show that OVASQ achieves perplexity close to that of the floating-point model under the W4A4 setting and significantly improves the performance of conventional static quantization under the lower-bit W4A3 and W3A4 settings. Compared with dynamic quantization workflows that transfer FP16 activations for online scale computation, OVASQ requires only a small amount of additional shift metadata and reduces activation data movement by approximately 75%.

    中文摘要 i 英文延伸摘要 ii 誌謝 xv 第一章 緒論 1 1-1 前言 1 1-2 研究動機 2 1-3 研究貢獻 3 1-4 論文架構 4 第二章 相關研究背景介紹 5 2-1 變換器與大型語言模型架構 6 2-1-1 詞元嵌入與旋轉位置編碼 6 2-1-2 多頭自注意力機制 8 2-1-3 門控前饋網路 8 2-1-4 正規化、殘差連接與 Llama 2 架構 9 2-2 大型語言模型推論流程 10 2-2-1 提示詞預填充階段 10 2-2-2 自回歸解碼階段 11 2-2-3 鍵值快取機制 12 2-3 神經網路模型壓縮技術 12 2-3-1 參數剪枝 13 2-3-2 低秩分解 13 2-3-3 知識蒸餾 14 2-3-4 模型量化 14 第三章 量化相關文獻回顧 16 3-1 神經網路量化技術分類 17 3-1-1 量化表示方式 17 3-1-2 量化感知訓練與訓練後量化 18 3-1-3 僅權重量化與權重–激活值量化 19 3-1-4 量化粒度 20 3-2 大型語言模型激活值量化挑戰 20 3-2-1 激活值分布與離群值現象 20 3-2-2 輸入資料與詞元位置造成的數值範圍差異 21 3-2-3 量化範圍與表示解析度 21 3-2-4 超低位元權重–激活值量化 22 3-3 動態量化與靜態量化 23 3-3-1 動態激活值量化 23 3-3-2 靜態激活值量化 23 3-3-3 模型效能與推論成本比較 24 3-4 代表性大型語言模型量化方法 25 3-4-1 ZeroQuant [4] 25 3-4-2 SmoothQuant [3] 26 3-4-3 Atom [5] 26 3-4-4 FlatQuant [6] 27 3-5 現有方法比較 28 第四章 基於位移式步長調整之變換器溢位感知靜態量化 30 4-1 量化基礎與推論流程比較 32 4-1-1 量化表示與離線量化步長 32 4-1-2 量化範圍與量化誤差 35 4-1-3 動態、靜態與 OVASQ 量化流程 37 4-1-3-1 動態量化流程 39 4-1-3-2 傳統靜態量化流程 39 4-1-3-3 OVASQ 流程 40 4-2 OVASQ 整體架構與離線初始化 41 4-2-1 離線校正與線上推論 42 4-2-2 基於峰值中位數的初始量化步長 42 4-3 區塊溢位偵測與位移式範圍調整 44 4-3-1 區塊矩陣乘法與尺度對齊 45 4-3-2 重新量化與溢位偵測 47 4-3-3 局部位移與參考位移更新 48 4-4 定點量化步長細化 50 4-4-1 量化步長細化因子 50 4-4-2 逐詞元細化因子搜尋 51 4-4-3 OVASQ 推論流程 52 4-5 變換器模型整合與記憶體傳輸量分析 55 4-5-1 OVASQ 於變換器線性層之整合 55 4-5-2 DRAM 記憶體傳輸量模型 58 第五章 實驗結果與分析 60 5-1 實驗設定 60 5-1-1 模型、資料集與評估指標 61 5-1-2 比較方法與量化設定 62 5-2 量化模型效能、細化因子與執行時間分析 63 5-2-1 不同位元設定之模型效能 63 5-2-2 線性層執行時間與區塊大小分析 65 5-2-3 量化步長細化因子消融分析 66 5-3 DRAM 記憶體傳輸量分析 67 第六章 結論與未來展望 71 6-1 結論 71 6-2 未來展望 71 參考文獻 73

    [1] A. Gholami, Z. Yao, S. Kim, C. Hooper, M. W. Mahoney, and K. Keutzer, “Ai and Memory Wall,” IEEE Micro, vol. 44, no. 3, pp. 33–39, 2024.
    [2] T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer, “LLM. int8 (): 8-bit Matrix Multiplication for Transformers at Scale,” arXiv preprint arXiv:2208.07339, 2022.
    [3] G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han, “SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models,” in International conference on machine learning, pp. 38087–38099, PMLR, 2023.
    [4] Z. Yao, R. Yazdani Aminabadi, M. Zhang, X. Wu, C. Li, and Y. He, “ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers,” Advances in neural information processing systems, vol. 35, pp. 27168–27183, 2022.
    [5] Y. Zhao, C.-Y. Lin, K. Zhu, Z. Ye, L. Chen, S. Zheng, L. Ceze, A. Krishnamurthy, T. Chen, and B. Kasikci, “Atom: Low-Bit Quantization for Efficient and Accurate Llm Serving,” Proceedings of Machine Learning and Systems, vol. 6, pp. 196–209, 2024.
    [6] Y. Sun, R. Liu, H. Bai, H. Bao, K. Zhao, Y. Li, J. Hu, X. Yu, L. Hou, C. Yuan, et al., “FlatQuant: Flatness Matters for Llm Quantization,” arXiv preprint arXiv:2410.09426, 2024.
    [7] D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al., “Mastering the game of go without human knowl- edge,” nature, vol. 550, no. 7676, pp. 354–359, 2017.
    [8] J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al., “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774, 2023.
    [9] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention Is All You Need,” Advances in neural information processing systems, vol. 30, 2017.
    [10] H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al., “Llama 2: Open Foundation and Fine-Tuned Chat Models,” arXiv preprint arXiv:2307.09288, 2023.
    [11] S. Han, H. Mao, and W. J. Dally, “Deep Compression: Compressing Deep Neural Networks With Pruning, Trained Quantization and Huffman Coding,” ICLR, 2016.
    [12] J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, “Roformer: Enhanced transformer with Rotary Position Embedding,” Neurocomputing, vol. 568, p. 127063, 2024.
    [13] B. Zhang and R. Sennrich, “Root Mean Square Layer Normalization,” Advances in neural information processing systems, vol. 32, 2019.
    [14] N. Shazeer, “Glu Variants Improve Transformer,” arXiv preprint arXiv:2002.05202, 2020.
    [15] Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, and H. Zhang, “{DistServe}: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving,” in 18th USENIX Symposium on Operating Systems Design and Implementa- tion (OSDI 24), pp. 193–210, 2024.
    [16] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient Memory Management for Large Language Model Serving with PagedAttention,” in Proceedings of the 29th symposium on operating systems princi- ples, pp. 611–626, 2023.
    [17] T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré, “FlashAttention: Fast and Memory- Efficient Exact Attention with IO-Awareness,” Advances in neural information pro- cessing systems, vol. 35, pp. 16344–16359, 2022.
    [18] X. Ma, G. Fang, and X. Wang, “LLM-Pruner: On the Structural Pruning of Large Language Models,” Advances in neural information processing systems, vol. 36, pp. 21702–21720, 2023.
    [19] X. Wang, Y. Zheng, Z. Wan, and M. Zhang, “SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model compression,” in International Con- ference on Learning Representations, vol. 2025, pp. 19299–19319, 2025.
    [20] G. Hinton, O. Vinyals, and J. Dean, “Distilling the Knowledge in a Neural Network,” arXiv preprint arXiv:1503.02531, 2015.
    [21] B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko, “Quantization and Training of Neural Networks for Efficient Integer- Arithmetic-Only Inference,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2704–2713, 2018.
    [22] M. Nagel, M. Fournarakis, R. A. Amjad, Y. Bondarenko, M. Van Baalen, and T. Blankevoort, “A white paper on Neural Network Quantization,” arXiv preprint arXiv:2106.08295, 2021.
    [23] E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh, “GPTQ: Accurate Post- Training Quantization for Generative Pre-trained Transformers,” arXiv preprint arXiv:2210.17323, 2022.
    [24] S. Kim, A. Gholami, Z. Yao, M. W. Mahoney, and K. Keutzer, “I-BERT: Integer-only BERT Quantization,” in International conference on machine learning, pp. 5506–5518, PMLR, 2021.
    [25] B. A. Motetti, M. Risso, A. Burrello, E. Macii, M. Poncino, and D. J. Pagliari, “Joint Pruning and Channel-Wise Mixed-Precision Quantization for Efficient Deep Neural Networks,” IEEE Transactions on Computers, vol. 73, no. 11, pp. 2619–2633, 2024.
    [26] S. Merity, C. Xiong, J. Bradbury, and R. Socher, “Pointer Sentinel Mixture Models,” arXiv preprint arXiv:1609.07843, 2016.

    下載圖示
    校外:立即公開
    QR CODE