簡易檢索 / 詳目顯示

研究生: 詹皓鈞
Chan, Hao-Chun
論文名稱: 適用於預訓練量化模型的低位寬裝置端微調方法
Low Bit-Width On-Device Fine-Tuning Framework Tailored for Pretrained Quantized Models
指導教授: 郭致宏
Kuo, Chih-Hung
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 電機工程學系
Department of Electrical Engineering
論文出版年: 2024
畢業學年度: 113
語文別: 中文
論文頁數: 91
中文關鍵詞: 神經網路 、裝置端學習 、低位寬訓練
外文關鍵詞: Neural Network, On-device learning, Low bit-width training
相關次數: 點閱:110  下載:0 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 傳統的神經網路訓練需要將訓練資料傳輸到雲端伺服器,以進行大量且高精度的計算。然而,這樣的資料傳輸會消耗大量能量並增加延遲,同時也會帶來資料隱私的問題。為了解決這些問題,出現了在裝置端學習 (On-device learning) 的方法。裝置端學習不僅能減少能耗和延遲,還能針對用戶訓練客製化的模型。然而,由於終端裝置的運算資源和記憶體有限,傳統的神經網路訓練並不適用於終端裝置上的學習。本文提出了一種適用於預訓練量化模型的低位寬裝置端微調方法。這種方法可以將在雲端伺服器上預訓練好的量化模型,轉移到資源受限的裝置上進行微調,以適應具體應用場景的需求。低位寬微調方法能顯著降低訓練過程中的計算成本和記憶體需求,使得在邊緣裝置和嵌入式系統中應用深度學習模型成為可能。在 VGG-16 模型於 CIFAR-10 資料集的遷移學習 (Transfer learning) 實驗中,我們的方法使用全整數 INT8 訓練達到93.59%的準確率,相比浮點數訓練僅下降 0.4%。這表明我們的方法能在大幅度降低資源消耗的情況下,仍然保持高準確率,展現了卓越的效能與創新性。

    Traditional neural network training requires transferring training data to cloud servers for extensive and high-precision computations. However, these data transfers consume significant energy and increase latency, while also posing data privacy issues. To address these challenges, on-device learning methods have emerged. On-device learning not only reduces energy consumption and latency but also allows for the customization of models for individual users. Nevertheless, the limited computational resources and memory of edge devices make traditional neural network training unsuitable for on-device learning. To solve this problem, this paper proposes a low bit-width on-device fine-tuning framework tailored for pretrained quantized models. This method enables the continuous optimization of the quantized model directly on resource-constrained edge devices. The low bit-width fine-tuning method significantly reduces the computational cost and memory demand during training, making it feasible to train models on edge devices. In transfer learning experiments with VGG-16 on CIFAR-10, our method achieved 93.59% accuracy using full integer INT8 training, only 0.4% lower than floating-point training. This demonstrates that our method maintains high precision while significantly reducing resource consumption, showcasing exceptional performance and innovation.

    中文摘要 II 誌謝 XV 目錄 XVI 表目錄 XIX 圖目錄 XX 第一章 緒論 1 1-1 前言 1 1-2 研究動機 2 1-3 研究貢獻 3 1-4 論文架構 4 第二章 相關研究背景介紹 5 2-1 深度學習 5 2-1-1 深度神經網路 5 2-1-2 反向傳播法 7 2-1-3 線性層的反向傳播 8 2-1-4 卷積神經網路 10 2-2 經典卷積神經網路架構 12 2-2-1 LeNet-5 12 2-2-2 VGG-16 13 2-2-3 MobileNet 13 2-3 神經網路壓縮介紹 14 2-3-1 網路量化(Network quantization) 14 2-3-2 網路剪枝(Network pruning) 15 2-3-3 知識蒸餾(Knowledge distillation) 15 2-3-4 高效架構設計(Efficient architecture design) 16 2-4 裝置端訓練 (On-device training) 介紹 17 第三章 低位寬訓練相關文獻回顧 18 3-1 網路量化壓縮技術 18 3-1-1 對稱式與非對稱式量化 18 3-1-2 訓練後量化 (Post-Training Quantization, PTQ) 19 3-1-3 量化感知訓練 (Quantization-Aware Training, QAT) 20 3-2 量化技術的應用方式 22 3-2-1 靜態 (Static) 量化 22 3-2-2 動態 (Dynamic) 量化 23 3-2-3 後見之明量化 (In-Hindsight Quantization) 24 3-3 裝置端訓練相關技術 24 3-3-1 量化感知縮放 (Quantization-Aware Scaling) 24 3-4 低位寬訓練方法比較 26 第四章 適用於預訓練量化模型的低位寬微調方法 30 4-1 低位寬微調 (Low Bit-Width Fine-tuning) 32 4-1-1 前向傳播 (Forward propagation) 33 4-1-2 反向傳播 (Backward propagation) 33 4-1-3 權重更新 (Weight update) 34 4-1-4 重新量化器 (Requantizers) 35 4-1-5 量化方案 36 4-2 近似 softmax 網路 (Approximate softmax network, ASN) 38 4-2-1 Softmax 操作在分類任務中的應用與計算 38 4-2-2 近似 softmax 網路 (ASN) 基本概念 39 4-2-3 近似 softmax 網路 (ASN) 權重與偏差計算 42 第五章 實驗環境與數據分析 47 5-1 資料集 (Dataset) 47 5-2 實驗操作細節 48 5-3 低位寬微調實驗 50 5-3-1 低位寬微調有效性分析 50 5-3-2 低位寬訓練方法比較 52 5-3-3 儲存權重位寬分析 56 5-4 近似 softmax 網路 (ASN) 效能評估 58 5-4-1 ASN 與標準 Softmax 的準確率與誤差分析 58 5-4-2 ASN 與現有硬體架構的運算資源比較 60 5-5 整合低位寬微調與近似 softmax 網路 (ASN) 實驗 63 5-6 全整數 INT8 微調實驗 64 第六章 結論與未來展望 66 6-1 結論 66 6-2 未來展望 66 參考文獻 67

    [1] D. Silver et al., “Mastering the Game of Go with Deep Neural Networks and Tree Search.” Nature, vol. 529, no. 7587, pp.484-489, 2016.
    [2] OpenAI, “GPT-4 Technical Report”, arXiv e-prints, 2023. doi:10.48550/arXiv.2303.08774.
    [3] A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet Classification with Deep Convolutional Neural Networks,” Advances in Neural Information Processing Systems, vol. 25, pp. 1097-1105, 2012.
    [4] J. Lin, L. Zhu, W. M. Chen, W. C. Wang, C. Gan, and S. Han, "On-Device Training Under 256KB Memory," Advances in Neural Information Processing Systems, vol. 35, pp. 22941-22954, 2022.
    [5] D. E. Rumelhart, G. E. Hintion, and R. J. Williams, “Learning Representations by Back-Propagation Errors,” Cognitive Modeling, vol. 5, no. 3, p. 1, 1998.
    [6] J. Johnson. "Backpropagation for a Linear Layer." cs231n.stanford.edu. https://cs231n.stanford.edu/handouts/linear-backprop.pdf (accessed July 4, 2024).
    [7] Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, "Gradient-Based Learning Applied to Document Recognition," Proceedings of the IEEE, vol. 86, no. 11, pp. 2278-2324, 1998.
    [8] K. Simonyan, and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,” 2014, arXiv:1409.1556.
    [9] A. G. Howard et al., "MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications." arXiv preprint arXiv:1704.04861 (2017).
    [10] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov and L. C. Chen, "MobileNetV2: Inverted Residuals and Linear Bottlenecks," 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 2018, pp. 4510-4520, doi: 10.1109/CVPR.2018.00474.
    [11] A. Howard et al., "Searching for MobileNetV3," in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 1314-1324.
    [12] M. Nagel, M. Fournarakis, R. A. Amead, Y. Bondarenko, M. Van Baalen, and T. Blankevoort, "A White Paper on Neural Network Quantization," arXiv preprint arXiv:2106.08295, 2021.
    [13] H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf, “Pruning Filters for Efficient ConvNets,” Advances in Neural Information Processing Systems, 2016.
    [14] S. Han, J. Pool, J. Tran, and W. J. Dally, “Learning both Weights and Connections for Efficient Neural Network,” Advances in Neural Information Processing Systems, vol. 28, 2015.
    [15] G. Hinton, O. Vinyals, and J. Dean, “Distilling the Knowledge in a Neural Network,” arXiv preprint arXiv:1503.02531, 2015.
    [16] X. Zhang, X. Zhou, M. Lin, and J. Sun, "ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 6848-6856.
    [17] M. Nagel, M. V. Baalen, T. Blankevoort and M. Welling, “Data-Free Quantization Through Weight Equalization and Bias Correction,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 1325-1334.
    [18] B. Jacob et al., “Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2704-2713, 2018.
    [19] M. Fournarakis and M. Nagel, "In-Hindsight Quantization Range Estimation for Quantized Training," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 3063-3070.
    [20] S. Wu, G. Li, F. Chen, and L. Shi, "Training and Inference with Integers in Deep Neural Networks," in International Conference on Learning Representations, 2018.
    [21] M. Wang, S. Rasoulinezhad, P. H. Leong, and H. K.-H. So, "NITI: Training Integer Neural Networks Using Integer-Only Arithmetic," IEEE Transactions on Parallel and Distributed Systems, vol. 33, no. 11, pp. 3249-3261, 2022.
    [22] A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, and A. Lerer, “Automatic Differentiation in PyTorch,” in Advances in Neural Information Processing Systems Workshops, 2017.
    [23] S. B. Ali, S. I. Filip, and O. Sentieys, "A Stochastic Rounding-Enabled LowPrecision Floating-Point MAC for DNN Training," in 2024 Design, Automation & Test in Europe Conference & Exhibition (DATE), 2024: IEEE, pp. 1-6.
    [24] N. A. Koca, A. T. Do, and C. H. Chang, "Hardware-efficient Softmax Approximation for Self-Attention Networks," in 2023 IEEE International Symposium on Circuits and Systems (ISCAS), 2023: IEEE, pp. 1-5.

    下載圖示
    2026-10-09公開
    QR CODE