| 研究生: |
陳錫峰 Chen, Hsi-Feng |
|---|---|
| 論文名稱: |
針對以記憶體內運算為核心且支援晶片內訓練與推論之神經網路加速器之設計 Design of Compute-In-Memory-based Neural Network Accelerator with On-chip Training and Inference |
| 指導教授: |
邱瀝毅
Chiou, Lih-Yih |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 電機工程學系 Department of Electrical Engineering |
| 論文出版年: | 2021 |
| 畢業學年度: | 109 |
| 語文別: | 中文 |
| 論文頁數: | 60 |
| 中文關鍵詞: | 卷積神經網路 、加速器 、晶片內訓練 、電子系統層級 、記憶體內運算 |
| 外文關鍵詞: | Convolutional Neural Network (CNN), Accelerator, On-Chip Training, Electronic System Level (ESL), Computing-In-Memory (CIM) |
| 相關次數: | 點閱:1186 下載:2 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
近年來人工智慧成為最熱門的主題,且其中的機器學習更是現今成長最快速的領域之一,各種不同的演算法誕生也進入了許多領域中,如市場預測、醫療輔助、圖像及語音辨識皆被廣泛探討。而若提到機器學習就必須提到最早開始發展且最主流的神經網路,其中又以圖片辨識的卷積神經網路發展最為成熟且最受歡迎,卷積神經網路藉由輸入的圖片以及神經網路中的權重(Weight)對圖片進行辨識也被稱為推論(Inference)並將誤差回傳至伺服器端進行訓練(Training),隨著卷積神經網路的發展,設計具有高能量效率且同時支援推論以及晶片內訓練的加速器的必要性與日俱增。
本論文中不僅提出一針對卷積神經網路之加速器架構,並且我們採用了高能量效益(Energy Efficiency)的記憶體內運算作為加速器中的主要運算單元,且同步支援神經網路推論以及晶片內訓練,透過晶片內訓練機制加強使用者資料之隱私安全性。此加速器以電子系統層級建置,能有彈性地提供設計者於加速器內進行架構探勘與效能評估。
Nowadays, Artificial Intelligence (AI) has become a popular topic and machine learning is one of the fastest growing fields. Various kinds of algorithms are applied in many different applications such as market prediction, medical assistance as well as image or speech recognition. In machine learning, the earliest and mainstream algorithm is neural network. Moreover, convolutional neural network (CNN) widely applied in the image recognition field is more popular than other networks. CNN implement image prediction as known as inference by input images and weights in CNN model then upload prediction error to server for CNN training. As the CNN development evolves, it is becoming essential to design a high energy efficient hardware that can support inference and on-chip training at the same time.
In this work, we not only proposed an accelerator architecture focusing on CNN applications, but also adopt computing-in-memory (CIM) as the main processing unit of the accelerator. When personal data can be processed by on-chip training mechanism, users’ data privacy can be protected. The proposed accelerator is built at electronic system level and capable of making designers to explore accelerator architecture and evaluate performance.
[1] M. I. Jordan and T. M. Mitchell, “Machine learning: Trends, perspectives, and prospects,” Science, vol. 349, no. 6245, pp. 255–260, 2015.
[2] K. O.‘Shea, and R. Nash. (2015). ‘‘An introduction to convolutional neural networks.’’ [Online]. Available: https://arxiv.org/abs/1511.08458.
[3] A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” ACM Communication, vol. 60, no. 6, pp. 84–90, 2017.
[4] Y. Lecun, L. Bottou, Y. Bengio and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE, vol. 86, no. 11, pp. 2278-2324, Nov. 1998.
[5] S. Albawi, T. A. Mohammed, and S. Al-Zawi, “Understanding of a convolutional neural network,” in Proc. 2017 Int. Conf. Eng. Technol. ICET 2017, vol. 2018-Janua, pp. 1–6, 2018.
[6] S. Choi, J. Sim, M. Kang, Y. Choi, H. Kim, and L. S. Kim, “An Energy-Efficient Deep Convolutional Neural Network Training Accelerator for in Situ Personalization on Smart Devices,” IEEE J. Solid-State Circuits, vol. 55, no. 10, pp. 2691–2702, 2020.
[7] C. Lu, Y. Wu and C. Yang, "A 2.25 TOPS/W Fully-Integrated Deep CNN Learning Processor with On-Chip Training," in Proc. 2019 IEEE Asian Solid-State Circuits Conference (A-SSCC), 2019, pp. 65-68. in
[8] S. K. Gonugondla, M. Kang and N. Shanbhag, "A 42pJ/decision 3.12TOPS/W robust in-memory machine learning classifier with on-chip training," in Proc. 2018 IEEE International Solid - State Circuits Conference - (ISSCC), 2018, pp. 490-492.
[9] E.Cengil, A.Çinar, and Z.Güler, “A GPU-based convolutional neural network approach for image classification,” in Proc. Int. Artif. Intell. Data Process. Symp. (IDAP), 2017.
[10] SuperDataScience team, “Convolutional Neural Networks (CNN): Step 4 - Full Connection,” 2018. [Online]. Available: https://www.superdatascience.com/blogs/convolutional-neural-networks-cnn-step-4-full-connection. [Access: 18-August-2021]
[11] S. H. Han and K. Y. Lee, “Implementation of image classification CNN using multi thread GPU,” Proc. - Int. SoC Des. Conf. 2017, ISOCC 2017, pp. 296–297, 2018.
[12] Y.H. Chen, T. Krishna, J.S. Emer, and V. Sze, “Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks,” IEEE J. Solid-State Circuits, vol. 52, no. 1, pp. 127–138, 2017.
[13] M. Kuzin, Y. Shmelev, V. Kuskov, “New Trends in the World of IoT Threat,” 2018. [Online]. Available: https://securelist.com/new-trends-in-the-world-of-iot-threats/87991/. [Access: 18-August-2021]
[14] S. J. Pan and Q. Yang, "A Survey on Transfer Learning," IEEE Transactions on Knowledge and Data Engineering, vol. 22, no. 10, pp. 1345-1359, Oct. 2010.
[15] N. A. Syed, H. Liu, and K. K. Sung, “Incremental learning with support Vector Machines,” Proc. Int. Joint Conf. on Artificial Intelligence (IJCAI-99), 1999.
[16] B. Jacob, S. Kligys, B. Chen, M. Zhu, “Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference,” in Proc. IEEE Computer Vision and Pattern Recognition, pp. 2704–2713, 2018.
[17] Y.–L. Tsai, “Energy Efficiency Exploration at Electronic System Level for Bus-Based Embedded Systems,” M.S. thesis, EE, National Cheng Kung University, Tainan, Taiwan, 2017
[18] D. Han, J. Lee, and H. J. Yoo, “A Low-Power Deep Neural Network Online Learning Processor for Real-Time Object Tracking Application,” IEEE Trans. Circuits System I, vol. 66, no. 5, pp. 1794–1804, 2019.
[19] B. Fleischer, S. Shukla, M. Ziegler, J. Silberman, J. Oh, V. Srinivasan, J. Choi, S. Mueller, A. Agrawal, T. Babinsky, N. Cao, Chia-Yu Chen, P. Chuang, T. Fox, G. Gristede, M. Guillorn, H. Haynie, M. Klaiber, D. Lee, Shih-Hsieh Lo, G. Maier, M. Scheuermann, S. Venkataramani, C. Vezyrtzis, N. Wang, F. Yee, C. Zhou, Pong-Fei Lu, B. Curran, L. Chang, K. Gopalakrishnan, “A Scalable Multi-TeraOPS Core for AI Training and Inference,” IEEE Solid-State Circuits Letters, 2018, vol. 1, no. 12, pp. 217–220.
[20] H. Jia, H. Valavi, Y. Tang, J. Zhang, and N. Verma, “A programmable heterogeneous microprocessor based on bit-scalable in-memory computing,” IEEE J. Solid-State Circuits, vol. 55, no. 9, pp. 2609–2621, 2020.
[21] H. Jiang, X. Peng, S. Huang, and S. Yu, “CIMAT: A compute-in-memory architecture for on-chip training based on transpose SRAM arrays,” IEEE Trans. Computer, vol. 69, no. 7, pp. 944–954, 2020.
[22] W.–S. Ling, “An Energy-Efficient 8T-SRAM-based Compute-In-Memory using Weight Partition for 8-Bit Precision Convolutional Neural Network” M.S. thesis, EE, National Cheng Kung University, Tainan, Taiwan, 2021