| 研究生: |
詹皓鈞 Chan, Hao-Chun |
|---|---|
| 論文名稱: |
適用於預訓練量化模型的低位寬裝置端微調方法 Low Bit-Width On-Device Fine-Tuning Framework Tailored for Pretrained Quantized Models |
| 指導教授: |
郭致宏
Kuo, Chih-Hung |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 電機工程學系 Department of Electrical Engineering |
| 論文出版年: | 2024 |
| 畢業學年度: | 113 |
| 語文別: | 中文 |
| 論文頁數: | 91 |
| 中文關鍵詞: | 神經網路 、裝置端學習 、低位寬訓練 |
| 外文關鍵詞: | Neural Network, On-device learning, Low bit-width training |
| 相關次數: | 點閱:110 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
傳統的神經網路訓練需要將訓練資料傳輸到雲端伺服器,以進行大量且高精度的計算。然而,這樣的資料傳輸會消耗大量能量並增加延遲,同時也會帶來資料隱私的問題。為了解決這些問題,出現了在裝置端學習 (On-device learning) 的方法。裝置端學習不僅能減少能耗和延遲,還能針對用戶訓練客製化的模型。然而,由於終端裝置的運算資源和記憶體有限,傳統的神經網路訓練並不適用於終端裝置上的學習。本文提出了一種適用於預訓練量化模型的低位寬裝置端微調方法。這種方法可以將在雲端伺服器上預訓練好的量化模型,轉移到資源受限的裝置上進行微調,以適應具體應用場景的需求。低位寬微調方法能顯著降低訓練過程中的計算成本和記憶體需求,使得在邊緣裝置和嵌入式系統中應用深度學習模型成為可能。在 VGG-16 模型於 CIFAR-10 資料集的遷移學習 (Transfer learning) 實驗中,我們的方法使用全整數 INT8 訓練達到93.59%的準確率,相比浮點數訓練僅下降 0.4%。這表明我們的方法能在大幅度降低資源消耗的情況下,仍然保持高準確率,展現了卓越的效能與創新性。
Traditional neural network training requires transferring training data to cloud servers for extensive and high-precision computations. However, these data transfers consume significant energy and increase latency, while also posing data privacy issues. To address these challenges, on-device learning methods have emerged. On-device learning not only reduces energy consumption and latency but also allows for the customization of models for individual users. Nevertheless, the limited computational resources and memory of edge devices make traditional neural network training unsuitable for on-device learning. To solve this problem, this paper proposes a low bit-width on-device fine-tuning framework tailored for pretrained quantized models. This method enables the continuous optimization of the quantized model directly on resource-constrained edge devices. The low bit-width fine-tuning method significantly reduces the computational cost and memory demand during training, making it feasible to train models on edge devices. In transfer learning experiments with VGG-16 on CIFAR-10, our method achieved 93.59% accuracy using full integer INT8 training, only 0.4% lower than floating-point training. This demonstrates that our method maintains high precision while significantly reducing resource consumption, showcasing exceptional performance and innovation.
[1] D. Silver et al., “Mastering the Game of Go with Deep Neural Networks and Tree Search.” Nature, vol. 529, no. 7587, pp.484-489, 2016.
[2] OpenAI, “GPT-4 Technical Report”, arXiv e-prints, 2023. doi:10.48550/arXiv.2303.08774.
[3] A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet Classification with Deep Convolutional Neural Networks,” Advances in Neural Information Processing Systems, vol. 25, pp. 1097-1105, 2012.
[4] J. Lin, L. Zhu, W. M. Chen, W. C. Wang, C. Gan, and S. Han, "On-Device Training Under 256KB Memory," Advances in Neural Information Processing Systems, vol. 35, pp. 22941-22954, 2022.
[5] D. E. Rumelhart, G. E. Hintion, and R. J. Williams, “Learning Representations by Back-Propagation Errors,” Cognitive Modeling, vol. 5, no. 3, p. 1, 1998.
[6] J. Johnson. "Backpropagation for a Linear Layer." cs231n.stanford.edu. https://cs231n.stanford.edu/handouts/linear-backprop.pdf (accessed July 4, 2024).
[7] Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, "Gradient-Based Learning Applied to Document Recognition," Proceedings of the IEEE, vol. 86, no. 11, pp. 2278-2324, 1998.
[8] K. Simonyan, and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,” 2014, arXiv:1409.1556.
[9] A. G. Howard et al., "MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications." arXiv preprint arXiv:1704.04861 (2017).
[10] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov and L. C. Chen, "MobileNetV2: Inverted Residuals and Linear Bottlenecks," 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 2018, pp. 4510-4520, doi: 10.1109/CVPR.2018.00474.
[11] A. Howard et al., "Searching for MobileNetV3," in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 1314-1324.
[12] M. Nagel, M. Fournarakis, R. A. Amead, Y. Bondarenko, M. Van Baalen, and T. Blankevoort, "A White Paper on Neural Network Quantization," arXiv preprint arXiv:2106.08295, 2021.
[13] H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf, “Pruning Filters for Efficient ConvNets,” Advances in Neural Information Processing Systems, 2016.
[14] S. Han, J. Pool, J. Tran, and W. J. Dally, “Learning both Weights and Connections for Efficient Neural Network,” Advances in Neural Information Processing Systems, vol. 28, 2015.
[15] G. Hinton, O. Vinyals, and J. Dean, “Distilling the Knowledge in a Neural Network,” arXiv preprint arXiv:1503.02531, 2015.
[16] X. Zhang, X. Zhou, M. Lin, and J. Sun, "ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 6848-6856.
[17] M. Nagel, M. V. Baalen, T. Blankevoort and M. Welling, “Data-Free Quantization Through Weight Equalization and Bias Correction,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 1325-1334.
[18] B. Jacob et al., “Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2704-2713, 2018.
[19] M. Fournarakis and M. Nagel, "In-Hindsight Quantization Range Estimation for Quantized Training," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 3063-3070.
[20] S. Wu, G. Li, F. Chen, and L. Shi, "Training and Inference with Integers in Deep Neural Networks," in International Conference on Learning Representations, 2018.
[21] M. Wang, S. Rasoulinezhad, P. H. Leong, and H. K.-H. So, "NITI: Training Integer Neural Networks Using Integer-Only Arithmetic," IEEE Transactions on Parallel and Distributed Systems, vol. 33, no. 11, pp. 3249-3261, 2022.
[22] A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, and A. Lerer, “Automatic Differentiation in PyTorch,” in Advances in Neural Information Processing Systems Workshops, 2017.
[23] S. B. Ali, S. I. Filip, and O. Sentieys, "A Stochastic Rounding-Enabled LowPrecision Floating-Point MAC for DNN Training," in 2024 Design, Automation & Test in Europe Conference & Exhibition (DATE), 2024: IEEE, pp. 1-6.
[24] N. A. Koca, A. T. Do, and C. H. Chang, "Hardware-efficient Softmax Approximation for Self-Attention Networks," in 2023 IEEE International Symposium on Circuits and Systems (ISCAS), 2023: IEEE, pp. 1-5.