簡易檢索 / 詳目顯示

研究生: 林志泓
Lin, Chih-Hung
論文名稱: 類自主式神經網路及通用可重構加速器之設計與模擬
An Autonomous-like Neural Network and A Reconfigurable Accelerator Design and Simulation
指導教授: 周哲民
Jou, Jer-Min
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 電機工程學系
Department of Electrical Engineering
論文出版年: 2021
畢業學年度: 109
語文別: 中文
論文頁數: 58
中文關鍵詞: 生成對抗網路 、對偶學習 、自主式神經網路 、可重構加速器 、模擬器
外文關鍵詞: Generative Adversarial Network, Dual Learning, Autonomous Neural Network, Reconfigurable Accelerator, Simulator
相關次數: 點閱:148  下載:0 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 神經網路已被廣泛應用在各種領域,傳統神經網路訓練需要大量有標籤訓練集,然而有標籤訓練集需要人工為數據進行標記,造成極高的成本,因此不需要人類介入,能夠自主學習的自主式神經網路成為未來發展的趨勢。本文結合對偶學習與生成對抗式網路,提出一種新的生成對抗式對偶學習方法,只需要無標籤訓練集即可進行訓練,達到類自主式學習之效果。由於神經網路之構成包含RNN、CNN等多種不同架構,且往往需要大量運算,因此需要一個通用的加速器來加速計算,本文提出一個通用可重構加速器,其支援多種data flow,且藉由可重構特性,能夠將硬體資源分割以平行執行多個任務,提升資源利用率。目前沒有方法能夠自動化的找到最佳的可重構配置,因此必須由人工決定如何分割PE Array,為了快速驗證不同神經網路在不同大小形狀的PE Array上執行的速度用以決定適合的可重構配置,本文提出通用型可重構硬體架構模擬器,此模擬器根據前述的硬體架構設計,模擬特定data flow下,資料在Buffer和PE之間的傳遞,並在PE內做乘加運算。模擬器的輸入參數化,可以任意調整不同的神經網路參數和PE Array,做cycle accurate的模擬。我們模擬了不同大小的PE Array和不同平行加速方法的執行時間。

    Neural networks have been widely used in various fields. Traditional neural network training requires a large number of labeled training sets. However, labeled training sets need to manually label the data, resulting in extremely high costs, so it does not require human intervention and can learn independently the autonomous neural network has become the trend of future development. This paper combines dual learning and generative adversarial networks, and proposes a new generative adversarial dual learning method, which only needs unlabeled training set to be trained to achieve the effect of autonomous-like learning. Since the composition of neural networks includes a variety of different architectures such as RNN and CNN, and often requires a lot of calculations, a general accelerator is needed to accelerate calculations. This paper proposes a general reconfigurable accelerator that supports multiple data flows and can be the reconfigurable feature can divide hardware resources to execute multiple tasks in parallel and improve resource utilization. At present, there is no method to automatically find the best reconfigurable configuration. Therefore, it is necessary to manually decide how to divide the PE Array. In order to quickly verify the execution speed of different neural networks on PE Arrays of different sizes and shapes, determine the appropriate reconfigurable configuration. This paper proposes a general-purpose reconfigurable hardware architecture simulator. This simulator simulates the transfer of data between Buffer and PE under a specific data flow based on the aforementioned hardware architecture design, and performs multiplication and addition operations in PE. The inputs of the simulator are parameterizing, it can adjust different neural network parameters and PE Array arbitrarily to make cycle accurate simulation. We simulated the execution time of different sizes of PE Array and different parallel acceleration methods.

    摘要 II SUMMARY III OUR PROPOSED DESIGN III EXPERIMENTS V CONCLUSION VII 誌謝 VIII 目錄 IX 表目錄 X 圖目錄 X 第一章 緒論 1 1.1研究背景 1 1.2研究動機與目的 1 1.3論文架構 2 第二章 背景知識與相關研究 3 2.1神經網路 3 2.2神經機器翻譯 7 2.3神經網路平行加速方法 11 第三章 新型類自主式神經網路研究 15 3.1對偶學習 15 3.2生成對抗網路 18 3.3 生成對抗式對偶學習 22 第四章 通用型可重構硬體架構設計 25 4.1通用型可重構硬體架構概述與挑戰 25 4.2資料流 26 4.3硬體架構設計 29 4.4緩衝器大小 35 4.5 資料存取 37 第五章 模擬器 39 5.1 模擬器輸出輸入 39 5.2 模擬器架構 40 5.3模擬流程 46 第六章 實驗結果與討論 48 6.1 PE Array大小對執行時間之影響 48 6.2 生成對抗網路之訓練時間分析 50 6.3 PE Array形狀對執行時間之影響 52 6.4 可重構分割之平行加速 54 第七章 結論與未來展望 56 參考文獻 57

    [1] RUCK, Dennis W.; ROGERS, Steven K.; KABRISKY, Matthew. Feature selection using a multilayer perceptron. Journal of Neural Network Computing, 2.2: 40-48, 1990.
    [2] MEDSKER, Larry R.; JAIN, L. C. Recurrent neural networks. Design and Applications, 5: 64-67, 2001.
    [3] MIKOLOV, Tomáš, et al. Recurrent neural network based language model. In: Eleventh annual conference of the international speech communication association. 2010.
    [4] WIERING, Marco A.; VAN OTTERLO, Martijn. Reinforcement learning. Adaptation, learning, and optimization, 2012.
    [5] CHO, Kyunghyun, et al. On the properties of neural machine translation: Encoder-decoder approaches. arXiv preprint arXiv:1409.1259, 2014.
    [6] GOODFELLOW, Ian, et al. Generative adversarial nets. Advances in neural information processing systems, 27, 2014.
    [7] LI, Mu, et al. Scaling distributed machine learning with the parameter server. In: 11th {USENIX} Symposium on Operating Systems Design and Implementation ({OSDI} 14). p. 583-598. 2014.
    [8] SIMONYAN, Karen; ZISSERMAN, Andrew. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
    [9] HUANG, Zhiheng; XU, Wei; YU, Kai. Bidirectional LSTM-CRF models for sequence tagging. arXiv preprint arXiv:1508.01991, 2015.
    [10] HE, Kaiming, et al. Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition. p. 770-778. 2016.
    [11] HE, Di, et al. Dual learning for machine translation. Advances in neural information processing systems, 29: 820-828, 2016.
    [12] WU, Yonghui, et al. Google's neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144, 2016.
    [13] ALBAWI, Saad; MOHAMMED, Tareq Abed; AL-ZAWI, Saad. Understanding of a convolutional neural network. In: 2017 International Conference on Engineering and Technology (ICET). Ieee, p. 1-6. 2017.
    [14] GOYAL, Priya, et al. Accurate, large minibatch sgd: Training imagenet in 1 hour. arXiv preprint arXiv:1706.02677, 2017.
    [15] WEI, Xuechao, et al. Automated systolic array architecture synthesis for high throughput CNN inference on FPGAs. In: Proceedings of the 54th Annual Design Automation Conference 2017. p. 1-6. 2017.

    [16] XIA, Yingce, et al. Model-level dual learning. In: International Conference on Machine Learning. PMLR, p. 5383-5392. 2018.
    [17] HASSAN, Hany, et al. Achieving human parity on automatic chinese to english news translation. arXiv preprint arXiv:1803.05567, 2018.
    [18] HOLCOMB, Sean D., et al. Overview on deepmind and its alphago zero ai. In: Proceedings of the 2018 international conference on big data and education. p. 67-71. 2018.
    [19] SAMAJDAR, Ananda, et al. Scale-sim: Systolic cnn accelerator simulator. arXiv preprint arXiv:1811.02883, 2018.
    [20] HUANG, Yanping, et al. Gpipe: Efficient training of giant neural networks using pipeline parallelism. Advances in neural information processing systems, 32: 103-112, 2019.
    [21] NARAYANAN, Deepak, et al. PipeDream: generalized pipeline parallelism for DNN training. In: Proceedings of the 27th ACM Symposium on Operating Systems Principles. p. 1-15. 2019.

    下載圖示
    2026-09-24公開
    QR CODE