| 研究生: |
洪胤期 Hung, Yin-Chi |
|---|---|
| 論文名稱: |
強化深度運算單元之二維卷積運算 2D Convolution Enhancements for Neural Computing Unit |
| 指導教授: |
侯廷偉
Hou, Ting-Wei |
| 學位類別: |
碩士 Master |
| 系所名稱: |
工學院 - 工程科學系 Department of Engineering Science |
| 論文出版年: | 2020 |
| 畢業學年度: | 108 |
| 語文別: | 中文 |
| 論文頁數: | 58 |
| 中文關鍵詞: | 卷積神經網路 、輔助運算處理器 、指令集 |
| 外文關鍵詞: | convolution, co-processor, instruction |
| 相關次數: | 點閱:111 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
本研究延續全連接網路輔助運算處理器(Neural Computing Unit, NCU)之基礎,延伸探討全連接與卷積運算之硬體實作,分析卷積與全連接關係之轉換,最後在NCU硬體平行運算的基礎上,提出提升卷積運算效能之方案。同時本研究承襲NCU之指令集架構特色,提出擴充指令集,使其更適用於卷積運算模型。有別於其他研究提出之加速器架構須由主處理器下命令執行運算,基於指令集架構的輔助處理器能夠讓主處理器更專注於其他任務控制,輔助處理器任務中止前不需要主處理器再對輔助處理器下達命令。
本研究提出兩種輔助處理器協同方案以提升卷積運算效能並具體實現。方案一,針對卷積資料型態與全連接網路資料型態分析,找出兩者間之異同,並藉由資料型態之預先轉換,使NCU通用於這兩種計算型態。
方案二,針對NCU硬體實作強化設計,將資料型態轉換任務從主處理器轉移至輔助處理器運行,並同時維持其原有平行運算能力,藉以提升效能以及降低主處理器負擔,此方案利用擴充指令集具體實現。
目前版本受限於開發板硬體資源有限,實驗比較中,設定卷積核2*2、卷積核位移1,針對不同輸入陣列執行卷積運算,其中卷積核權重值取自Keras預先訓練完成模型,參考基準為應用程式執行時間。與Raspberry Pi相比,方案一,提升約1.4~93.8倍運算效能,最為顯著之效能提升為輸入陣列為6*6。方案二則提升1.9~162.2倍,最為顯著之效能提升同樣發生在輸入陣列為6*6時。另設定卷積核4*4、卷積核位移1之實驗中,方案一,提升約1.5~116.3倍運算效能,方案二則提升1.1~215.9倍,兩方案同樣於輸入陣列為6*6時效能提升最為顯著。
To speed up neural networks computations in embedded systems, this research proposes an approach to enhance a co-processor architecture which is to accelerate the fully connected neural network. This research analyzed differences between fully connected and convolution operations, and uses the parallel computing design of our lab’s NCU (neural computing unit) as a base to speed up convolution operations. NCU is an instruction-based coprocessor designed to cooperate with ARM hard-core. The instruction-based co-processor NCU architecture makes NCU more flexible, and NCU runs parallelly with the ARM hard-core.
This research inherits the characteristics of the NCU and expands NCU’s instruction set by proposing an enhanced set of instructions (and hardware) for convolution operation models. The computing logic of proposed approach was implemented using Verilog and deployed on Field Programmable Gate Array (FPGA). The logic circuits for calculating convolutions uses the parallel computing capabilities of NCU and enhances NCU’s finite state machines for controlling the operations.
In the first experiment, the convolution kernel was 2*2, and the stride was set to 1.Different arrays were inputted. The kernel data were pre-trained by Keras. For comparison, Raspberry Pi 3 Model B+ was used. The proposed implementation has a speedup of 1.9 to 162.2 times as compared with Raspberry Pi 3 Model B+. For another experiment, the convolution kernel was changed to 4*4. The proposed implementation has a speedup of 1.1 to 215.9 times as compared with Raspberry Pi 3 Model B+.
[1] J. Wang, J. Lin and Z. Wang, "Efficient hardware architectures for deep convolutional neural network," in IEEE Transactions on Circuits and Systems I: Regular Papers, vol. 65, no. 6, pp. 1941-1953, June 2018, doi: 10.1109/TCSI.2017.2767204.
[2] T. Abtahi, C. Shea, A. Kulkarni and T. Mohsenin, "Accelerating convolutional neural network with FFT on embedded hardware," in IEEE Transactions on Very Large Scale Integration (VLSI) Systems, vol. 26, no. 9, pp. 1737-1749, Sept. 2018, doi: 10.1109/TVLSI.2018.2825145.
[3] A. Ardakani, C. Condo, M. Ahmadi and W. J. Gross, "An architecture to accelerate convolution in deep neural networks," in IEEE Transactions on Circuits and Systems I: Regular Papers, vol. 65, no. 4, pp. 1349-1362, April 2018, doi: 10.1109/TCSI.2017.2757036.
[4] A. Ardakani, C. Condo and W. J. Gross, "A convolutional accelerator for neural networks with binary weights," 2018 IEEE International Symposium on Circuits and Systems (ISCAS), Florence, Italy, 2018, pp. 1-5, doi: 10.1109/ISCAS.2018.8350945.
[5] C. Zhang, P. Li, G. Sun, Y. Guan, B. Xiao and J. Cong, "Optimizing FPGA-based accelerator design for deep convolutional neural networks", Proceedings of the 2015 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays, Monterey California, USA, pp. 161-170, 2015.
[6] L. Bai, Y. Zhao and X. Huang, "A CNN accelerator on FPGA using depthwise separable convolution," in IEEE Transactions on Circuits and Systems II: Express Briefs, vol. 65, no. 10, pp. 1415-1419, Oct. 2018, doi: 10.1109/TCSII.2018.2865896.
[7] L. Zong-ling, W. Lu-yuan, Y. Ji-yang, C. Bo-wen and H. Liang, "The design of lightweight and multi parallel CNN accelerator based on FPGA," 2019 IEEE 8th Joint International Information Technology and Artificial Intelligence Conference (ITAIC), Chongqing, China, 2019, pp. 1521-1528, doi: 10.1109/ITAIC.2019.8785800.
[8] Y. Ma, Y. Cao, S. Vrudhula and J. Seo, "Optimizing the convolution operation to accelerate deep neural networks on FPGA," in IEEE Transactions on Very Large Scale Integration (VLSI) Systems, vol. 26, no. 7, pp. 1354-1367, July 2018, doi: 10.1109/TVLSI.2018.2815603.
[9] Huimin Li, Xitian Fan, Li Jiao, Wei Cao, Xuegong Zhou and Lingli Wang, "A high performance FPGA-based accelerator for large-scale convolutional neural networks," 2016 26th International Conference on Field Programmable Logic and Applications (FPL), Lausanne, 2016, pp. 1-9, doi: 10.1109/FPL.2016.7577308.
[10] 余誌偉(2019). 適用於深度神經網路推論之嵌入式異質多核心架構之設計與實作 國立成功大學碩士學位論文