簡易檢索 / 詳目顯示

研究生: 黃俞紘
Huang, Yu-Hong
論文名稱: 基於RRAM的神經網路加速器的熱感知權重映射方法
A Thermal-Aware Mapping Methodology for RRAM-based Neural Network Accelerator
指導教授: 林英超
Lin, Ing-Chao
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 資訊工程學系
Department of Computer Science and Information Engineering
論文出版年: 2021
畢業學年度: 109
語文別: 英文
論文頁數: 47
中文關鍵詞: 神經網路加速器 、權重量化 、權重剪枝 、可變電阻式記憶體 、權重映射
外文關鍵詞: Neural Network Accelerator, Weight Quantization, Weight Pruning, RRAM, Weight mapping
相關次數: 點閱:234  下載:0 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 電阻式隨機存取存儲記憶體 (Resistive Random Access Memory or RRAM)在記憶體內計算 (Computing in memory) 方面已展現出巨大的潛力,其能支援神經網路計算系統中高記憶體頻寬和低功耗的要求。基於 RRAM 的神經網路加速器被認為是可以用來加速密集使用記憶的應用(例如神經網絡 (Neural Network)中的乘法和累加單元)的解決方案。然而,由於 RRAM 單元電阻固有的阻值飄移特性以及受高溫影響而阻值變化的負面影響,會使基於 RRAM的神經網絡加速器在做推論時的準確度明顯降低。在本論文中,我們提出了一種熱感知的映射方法,在 基於RRAM的神經網路加速器中存在一些實際設計限制的情況下,將 Neural Network 模型的權重映射到 RRAM 子陣列(subarray)中。我們提出的熱感知映射方法包括兩種演算法、熱感知權重重映射 (Weight Remapping) 和權重剪枝和分裂 (Weight pruning and splitting)。實驗結果顯示,使用我們的方法時,即使加速器周圍溫度在 360K 左右,VGG 8、VGG11、Alexnet 和 Resnet34 模型在 CIFAR10 數據集上的推理準確度損失僅低於理想結果的 2%

    Resistive random access memory (Resistive Random Access Memory or RRAM) has
    shown great potential for computing in memory (Computing In Memory) to support the requirements of high memory bandwidth and low power in neuromorphic computing systems. RRAM-based accelerators are considered as a power-efficient design solution in accelerating memory-intensive applications, such as the multiply-and-accumulate units in neural networks (Neural Network or NN). However, the accuracy of RRAM-based NN computing can degrade significantly due to the intrinsic statistical variations of the resistance of RRAM cells, as well as the negative effects of high temperatures. In this thesis, we propose a thermal-aware mapping method to map the weights of the NN model into RRAM subarrays with some practical design limitations in the NN accelerators. The methodology includes two algorithms, the thermal-aware weight remapping (Weight Remapping) and the weight pruning and+ splitting (Weight Pruning and Splitting). Experimental results have shown that using our methodology, inference accuracy losses of the VGG 8, VGG11, Alexnet, and Resnet34 models with the CIFAR-10 dataset are less than 2% compared to the ideal results even when the surrounding temperature is around 360K.

    Table of Contents 摘要 i Abstract ii Table of Contents iii List of Figures v Chapter 1. Introduction 1 Chapter 2. Related Work 5 2.1 RRAM-­based Neural Network Accelerators 5 2.2 Neural network computation and weight mapping 7 2.3 Thermal Impact in RRAM­-based Systems 9 2.4 Thermal­-Aware and Thermal­-Resilient Designs 10 Chapter 3. Method 13 3.1 System Architecture Scenario 13 3.2 Proposed Thermal-­Aware Mapping Method 15 3.3 Weight Remapping Algorithm (WR) 16 3.4 Weight Pruning and Splitting (WPS) 19 Chapter 4. Experimental Setting and Simulation Tools 27 4.1 Experimental Setting and Simulation Tools 27 4.2 Thermal-­Aware Weight Remapping 29 4.3 Thermal-­Aware Weight Pruning and Splitting 33 4.4 Comparisons 37 4.5 Thermal­-Aware Weight Pruning and Splitting from four corner in different temperature ranges 40 Chapter 5. Conclusions 44 References 45

    [1] Majed Valad Beigi and Gokhan Memik. Thermal¬aware optimizations of reram-based neuromorphic computing systems. In ACM/ESDA/IEEE Design Automation Conference, pages 1–6, 2018.
    [2] Pai¬Yu Chen, Xiaochen Peng, and Shimeng Yu. Neurosim+: An integrated device-toalgorithm framework for benchmarking synaptic devices and array architectures. In IEEE International Electron Devices Meeting (IEDM), pages 6.1.1–6.1.4, 2017.
    [3] Tianshi Chen, Zidong Du, Ninghui Sun, Jia Wang, Chengyong Wu, Yunji Chen, and Olivier Temam. A high¬throughput neural network accelerator. volume 35, pages 24– 32, 2015.
    [4] Yu¬Hsin Chen, Tushar Krishna, Joel S. Emer, and Vivienne Sze. Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks. volume 52, pages 127–138, 2017.
    [5] Yunji Chen and Tao Luo. Dadiannao: A machine¬learning supercomputer. In 2014 47th Annual IEEE/ACM International Symposium on Microarchitecture, pages 609– 622, 2014.
    [6] Ping Chi, Shuangchen Li, Cong Xu, Tao Zhang, Jishen Zhao, Yongpan Liu, Yu Wang, and Yuan Xie. Prime: A novel processing¬in¬memory architecture for neural network computation in reram¬based main memory. In ACM/IEEE International Symposium on Computer Architecture, pages 27–39, 2016.
    [7] Alberto Garcia¬Garcia, Sergio Orts¬Escolano, Sergiu Oprea, Victor Villena¬Martinez, and Jose Garcia¬Rodriguez. A review on deep learning techniques applied to semantic segmentation. 2017.
    [8] Kaiyuan Guo, Shulin Zeng, Jincheng Yu, Yu Wang, and Huazhong Yang. A survey of fpga¬based neural network accelerator. 2018.
    [9] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016.
    [10] Julia Hirschberg and Christopher D. Manning. Advances in natural language processing. volume 349, pages 261–266. American Association for the Advancement of Science, 2015.
    [11] Norman P. Jouppi, Cliff Young, Nishant Patil, and David Patterson. In¬datacenter performance analysis of a tensor processing unit. In ACM/IEEE International Symposium on Computer Architecture, pages 1–12, 2017.
    [12] Duckhwan Kim, Jaeha Kung, Sek Chai, Sudhakar Yalamanchili, and Saibal Mukhopadhyay. Neurocube: A programmable digital neuromorphic architecture with high¬density 3d memory. In ACM/IEEE International Symposium on Computer Architecture, pages 380–392, 2016.
    [13] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems, volume 25, 2012.
    [14] Shuying Liu and Weihong Deng. Very deep convolutional neural network based image classification using small training sample size. In IAPR Asian Conference on Pattern Recognition, pages 730–734, 2015.
    [15] Xiao Liu, Mingxuan Zhou, Tajana S. Rosing, and Jishen Zhao. Hr3am: A heat resilient design for rram¬based neuromorphic computing. In IEEE/ACM International Symposium on Low Power Electronics and Design, pages 1–6, 2019.
    [16] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, and James Bradbury. Pytorch: An imperative style, high¬performance deep learning library. 2019.
    [17] Xiaochen Peng, Shanshi Huang, Hongwu Jiang, Anni Lu, and Shimeng Yu. Dnn+neurosim v2.0: An end¬to¬end benchmarking framework for compute¬in¬memory accelerators for on¬chip training. pages 1–1, 2020.
    [18] Xiaochen Peng, Rui Liu, and Shimeng Yu. Optimizing weight mapping and data flow for convolutional neural networks on rram based processing¬in¬memory architecture. In IEEE International Symposium on Circuits and Systems, pages 1–5, 2019.
    [19] Ali Shafiee, Anirban Nag, Naveen Muralimanohar, Rajeev Balasubramonian, John Paul Strachan, Miao Hu, R. Stanley Williams, and Vivek Srikumar. Isaac: A convolutional neural network accelerator with in¬situ analog arithmetic in crossbars. In ACM/IEEE Annual International Symposium on Computer Architecture, pages 14–26, 2016.
    [20] Hyein Shin, Myeonggu Kang, and Lee¬Sup Kim. A thermal¬aware optimization framework for reram¬based deep neural network acceleration. In IEEE/ACM International Conference On Computer Aided Design, pages 1–9, 2020.
    [21] M. Stan, Runjie Zhang, and K. Skadron. Hotspot 6.0: Validation, acceleration and extension. 2015.
    [22] Majed Valad Beigi and Gokhan Memik. Thor: Thermal¬aware optimizations for extending reram lifetime. In IEEE International Parallel and Distributed Processing Symposium, pages 670–679, 2018.
    [23] Christian Walczyk, Damian Walczyk, Thomas Schroeder, and Thomas Bertaud. Impact of temperature on the resistive switching behavior of embedded HfO2¬based rram devices. volume 58, pages 3124–3131, 2011.
    [24] Wei Wen, Chi¬Ruo Wu, Xiaofang Hu, Beiye Liu, Tsung¬Yi Ho, Xin Li, and Yiran Chen. An eda framework for large scale hybrid neuromorphic computing systems. In ACM/EDAC/IEEE Design Automation Conference, pages 1–6, 2015.
    [25] Lixue Xia, Boxun Li, Tianqi Tang, Peng Gu, Pai¬Yu Chen, Shimeng Yu, Yu Cao, Yu Wang, Yuan Xie, and Huazhong Yang. Mnsim: Simulation platform for memristor-based neuromorphic computing system. volume 37, pages 1009–1022, 2018.
    [26] Cheng¬Xin Xue, Wei¬Hao Chen, Je¬Syu Liu, Jia¬Fang Li, and Wei¬Yu Lin. 24.1 a 1mb multibit reram computing¬in¬memory macro with 14.6ns parallel mac computing time for cnn based ai edge processors. In IEEE International Solid¬ State Circuits Conference, pages 388–390, 2019.
    [27] Furqan Zahoor. Resistive random access memory (rram): an overview of materials, switching mechanism, performance, multilevel cell (mlc) storage, modeling, and applications. 2020.
    [28] Shuhang Zhang, Grace Li Zhang, Bing Li, Hai Helen Li, and Ulf Schlichtmann. Lifetime enhancement for rram¬based computing¬in¬memory engine considering aging and thermal effects. In IEEE International Conference on Artificial Intelligence Circuits and Systems, pages 11–15, 2020.
    [29] Zhong¬Qiu Zhao, Peng Zheng, Shou¬Tao Xu, and Xindong Wu. Object detection with deep learning: A review. volume 30, pages 3212–3232, 2019.
    [30] Minxuan Zhou, Mohsen Imani, Saransh Gupta, and Tajana Rosing. Thermal¬aware design and management for search¬based in¬memory acceleration. In ACM/IEEE Design Automation Conference, pages 1–6, 2019.

    下載圖示
    2026-08-31公開
    QR CODE