簡易檢索 / 詳目顯示

研究生: 陳錫峰
Chen, Hsi-Feng
論文名稱: 針對以記憶體內運算為核心且支援晶片內訓練與推論之神經網路加速器之設計
Design of Compute-In-Memory-based Neural Network Accelerator with On-chip Training and Inference
指導教授: 邱瀝毅
Chiou, Lih-Yih
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 電機工程學系
Department of Electrical Engineering
論文出版年: 2021
畢業學年度: 109
語文別: 中文
論文頁數: 60
中文關鍵詞: 卷積神經網路 、加速器 、晶片內訓練 、電子系統層級 、記憶體內運算
外文關鍵詞: Convolutional Neural Network (CNN), Accelerator, On-Chip Training, Electronic System Level (ESL), Computing-In-Memory (CIM)
相關次數: 點閱:1186  下載:2 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 近年來人工智慧成為最熱門的主題,且其中的機器學習更是現今成長最快速的領域之一,各種不同的演算法誕生也進入了許多領域中,如市場預測、醫療輔助、圖像及語音辨識皆被廣泛探討。而若提到機器學習就必須提到最早開始發展且最主流的神經網路,其中又以圖片辨識的卷積神經網路發展最為成熟且最受歡迎,卷積神經網路藉由輸入的圖片以及神經網路中的權重(Weight)對圖片進行辨識也被稱為推論(Inference)並將誤差回傳至伺服器端進行訓練(Training),隨著卷積神經網路的發展,設計具有高能量效率且同時支援推論以及晶片內訓練的加速器的必要性與日俱增。
    本論文中不僅提出一針對卷積神經網路之加速器架構,並且我們採用了高能量效益(Energy Efficiency)的記憶體內運算作為加速器中的主要運算單元,且同步支援神經網路推論以及晶片內訓練,透過晶片內訓練機制加強使用者資料之隱私安全性。此加速器以電子系統層級建置,能有彈性地提供設計者於加速器內進行架構探勘與效能評估。

    Nowadays, Artificial Intelligence (AI) has become a popular topic and machine learning is one of the fastest growing fields. Various kinds of algorithms are applied in many different applications such as market prediction, medical assistance as well as image or speech recognition. In machine learning, the earliest and mainstream algorithm is neural network. Moreover, convolutional neural network (CNN) widely applied in the image recognition field is more popular than other networks. CNN implement image prediction as known as inference by input images and weights in CNN model then upload prediction error to server for CNN training. As the CNN development evolves, it is becoming essential to design a high energy efficient hardware that can support inference and on-chip training at the same time.

    In this work, we not only proposed an accelerator architecture focusing on CNN applications, but also adopt computing-in-memory (CIM) as the main processing unit of the accelerator. When personal data can be processed by on-chip training mechanism, users’ data privacy can be protected. The proposed accelerator is built at electronic system level and capable of making designers to explore accelerator architecture and evaluate performance.

    摘 要 i 表目錄 ix 圖目錄 x 第1章 緒論 1 1.1 研究概觀 1 1.2 研究動機 3 1.3 研究貢獻 5 1.4 論文架構 5 第2章 相關研究背景 6 2.1 卷積神經網路 6 2.2 卷積神經網路相關硬體元件 10 2.2.1 專用硬體加速器內核心運算單元 12 2.2.2 專用硬體加速器內周邊單元 14 2.3 卷積神經網路訓練概觀 15 2.3.1 卷積神經網路訓練種類簡介 16 2.3.2 晶片內訓練 18 2.3.3 硬體取向量化訓練簡介 19 2.4 電子系統層級設計 20 第3章 相關文獻探討 22 3.1 以卷積神經網路為目標之硬體加速器 22 3.1.1 利用數位運算單元為核心之加速器 22 3.1.2 利用記憶體內運算單元為核心之加速器 25 3.2 相關文獻總結 29 第4章 以記憶體內運算為核心之神經網路加速器設計 30 4.1 問題描述 30 4.2 目標虛擬平台之環境與架構 32 4.2.1 資料集與目標神經網路 32 4.2.2 虛擬平台介紹 33 4.2.3 卷積神經網路加速器內部架構 34 4.2.4 記憶體內運算架構 37 4.3 卷積神經網路推論與晶片內訓練機制 38 4.3.1 加速器系統之推論加速機制 38 4.3.2 加速器晶片內訓練機制 40 第5章 實驗結果與分析 44 5.1 實驗環境設置 44 5.2 實驗一、硬體取向量化訓練機制下神經網路準確率分析與比較 45 5.2.1 實驗環境設定 45 5.2.2 有無採用硬體取向量化訓練機制之準確率分析 46 5.3 實驗二、晶片內訓練機制之準確率分析與比較 47 5.3.1 實驗環境設定 47 5.3.2 不同訓練參數配置下對準確率之比較分析 48 5.4 實驗三、加速器之神經網路推論與晶片內訓練硬體細節 52 5.4.1 實驗環境設定 52 5.4.2 加速器對神經網路推論時之硬體分析 53 第6章 結論與未來研究 56 6.1 結論 56 6.2 未來工作 57 參考文獻 58

    [1] M. I. Jordan and T. M. Mitchell, “Machine learning: Trends, perspectives, and prospects,” Science, vol. 349, no. 6245, pp. 255–260, 2015.
    [2] K. O.‘Shea, and R. Nash. (2015). ‘‘An introduction to convolutional neural networks.’’ [Online]. Available: https://arxiv.org/abs/1511.08458.
    [3] A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” ACM Communication, vol. 60, no. 6, pp. 84–90, 2017.
    [4] Y. Lecun, L. Bottou, Y. Bengio and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE, vol. 86, no. 11, pp. 2278-2324, Nov. 1998.
    [5] S. Albawi, T. A. Mohammed, and S. Al-Zawi, “Understanding of a convolutional neural network,” in Proc. 2017 Int. Conf. Eng. Technol. ICET 2017, vol. 2018-Janua, pp. 1–6, 2018.
    [6] S. Choi, J. Sim, M. Kang, Y. Choi, H. Kim, and L. S. Kim, “An Energy-Efficient Deep Convolutional Neural Network Training Accelerator for in Situ Personalization on Smart Devices,” IEEE J. Solid-State Circuits, vol. 55, no. 10, pp. 2691–2702, 2020.
    [7] C. Lu, Y. Wu and C. Yang, "A 2.25 TOPS/W Fully-Integrated Deep CNN Learning Processor with On-Chip Training," in Proc. 2019 IEEE Asian Solid-State Circuits Conference (A-SSCC), 2019, pp. 65-68. in
    [8] S. K. Gonugondla, M. Kang and N. Shanbhag, "A 42pJ/decision 3.12TOPS/W robust in-memory machine learning classifier with on-chip training," in Proc. 2018 IEEE International Solid - State Circuits Conference - (ISSCC), 2018, pp. 490-492.
    [9] E.Cengil, A.Çinar, and Z.Güler, “A GPU-based convolutional neural network approach for image classification,” in Proc. Int. Artif. Intell. Data Process. Symp. (IDAP), 2017.
    [10] SuperDataScience team, “Convolutional Neural Networks (CNN): Step 4 - Full Connection,” 2018. [Online]. Available: https://www.superdatascience.com/blogs/convolutional-neural-networks-cnn-step-4-full-connection. [Access: 18-August-2021]
    [11] S. H. Han and K. Y. Lee, “Implementation of image classification CNN using multi thread GPU,” Proc. - Int. SoC Des. Conf. 2017, ISOCC 2017, pp. 296–297, 2018.
    [12] Y.H. Chen, T. Krishna, J.S. Emer, and V. Sze, “Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks,” IEEE J. Solid-State Circuits, vol. 52, no. 1, pp. 127–138, 2017.
    [13] M. Kuzin, Y. Shmelev, V. Kuskov, “New Trends in the World of IoT Threat,” 2018. [Online]. Available: https://securelist.com/new-trends-in-the-world-of-iot-threats/87991/. [Access: 18-August-2021]
    [14] S. J. Pan and Q. Yang, "A Survey on Transfer Learning," IEEE Transactions on Knowledge and Data Engineering, vol. 22, no. 10, pp. 1345-1359, Oct. 2010.
    [15] N. A. Syed, H. Liu, and K. K. Sung, “Incremental learning with support Vector Machines,” Proc. Int. Joint Conf. on Artificial Intelligence (IJCAI-99), 1999.
    [16] B. Jacob, S. Kligys, B. Chen, M. Zhu, “Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference,” in Proc. IEEE Computer Vision and Pattern Recognition, pp. 2704–2713, 2018.
    [17] Y.–L. Tsai, “Energy Efficiency Exploration at Electronic System Level for Bus-Based Embedded Systems,” M.S. thesis, EE, National Cheng Kung University, Tainan, Taiwan, 2017
    [18] D. Han, J. Lee, and H. J. Yoo, “A Low-Power Deep Neural Network Online Learning Processor for Real-Time Object Tracking Application,” IEEE Trans. Circuits System I, vol. 66, no. 5, pp. 1794–1804, 2019.
    [19] B. Fleischer, S. Shukla, M. Ziegler, J. Silberman, J. Oh, V. Srinivasan, J. Choi, S. Mueller, A. Agrawal, T. Babinsky, N. Cao, Chia-Yu Chen, P. Chuang, T. Fox, G. Gristede, M. Guillorn, H. Haynie, M. Klaiber, D. Lee, Shih-Hsieh Lo, G. Maier, M. Scheuermann, S. Venkataramani, C. Vezyrtzis, N. Wang, F. Yee, C. Zhou, Pong-Fei Lu, B. Curran, L. Chang, K. Gopalakrishnan, “A Scalable Multi-TeraOPS Core for AI Training and Inference,” IEEE Solid-State Circuits Letters, 2018, vol. 1, no. 12, pp. 217–220.
    [20] H. Jia, H. Valavi, Y. Tang, J. Zhang, and N. Verma, “A programmable heterogeneous microprocessor based on bit-scalable in-memory computing,” IEEE J. Solid-State Circuits, vol. 55, no. 9, pp. 2609–2621, 2020.
    [21] H. Jiang, X. Peng, S. Huang, and S. Yu, “CIMAT: A compute-in-memory architecture for on-chip training based on transpose SRAM arrays,” IEEE Trans. Computer, vol. 69, no. 7, pp. 944–954, 2020.
    [22] W.–S. Ling, “An Energy-Efficient 8T-SRAM-based Compute-In-Memory using Weight Partition for 8-Bit Precision Convolutional Neural Network” M.S. thesis, EE, National Cheng Kung University, Tainan, Taiwan, 2021

    下載圖示
    2026-08-20公開
    QR CODE