簡易檢索 / 詳目顯示

研究生: 凌偉碩
Ling, Wei-Shuo
論文名稱: 可應用於八位元精準度的卷積神經網路之具有能量效益且使用權重分割技術的8T-SRAM記憶體內運算
An Energy-Efficient 8T-SRAM-based Compute-In-Memory using Weight Partition for 8-Bit Precision Convolutional Neural Network
指導教授: 邱瀝毅
Chiou, Lih-Yih
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 電機工程學系
Department of Electrical Engineering
論文出版年: 2021
畢業學年度: 109
語文別: 中文
論文頁數: 79
中文關鍵詞: 人工智慧 、神經網路 、邊緣運算 、靜態隨機存取記憶體 、記憶體內運算
外文關鍵詞: Artificial intelligence (AI), Convolutional neural networks (CNNs), Edge computing, Static random access memory (SRAM), Compute-in-memory (CIM)
相關次數: 點閱:197  下載:1 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 近年來,人工智慧的相關技術蓬勃發展,其中的神經網路領域更是被廣泛運用在各個應用,像是影像辨識、物件辨識與語言翻譯等等。為了降低邊緣裝置與雲端的資料傳輸量,近年的趨勢是把AI演算法導入IoT設備,將蒐集到的資料透過AI加速器來執行邊緣運算。另外,為了克服傳統數位加速器的記憶體牆效應,發展出基於記憶體內運算的加速器,藉此避免傳統數位加速器中過多資料搬移的現象,大幅減少功率的消耗與計算的延遲時間。
    本論文為採用40奈米製程設計,提出具有能量效益與支援8-bit精準度之乘加運算的8T-SRAM記憶體內運算,透過本論文提出的後級量化機制,減少多餘的電路成本,使面積與功率均能獲得改善。相比於過去的相關文獻,本論文的記憶體內運算架構,在能量效益、運算精準度與操作速度上都具有一定的競爭優勢。

    In recent years, artificial intelligence-related technologies have flourished. Among them, the field of neural networks is widely used in various applications, such as image recognition, object recognition, language translation, and so on. In order to reduce the amount of data transmission between edge devices and the cloud, the trend in recent years is to introduce AI algorithms into IoT devices, and use AI accelerators to perform edge computing on the collected data. In addition, in order to overcome the memory wall effect of traditional digital accelerators, several compute-in-memory-based accelerators have been developed to avoid excessive data movement in the traditional digital accelerator. This change may greatly reduce power consumption and calculation delay time.
    We proposed an energy-efficient 8T-SRAM-based compute-in-memory macro using 40nm process. Through the backend quantization mechanism proposed in this thesis, both area and power can be reduced. When compared with related works in recent years, the proposed compute-in-memory architecture has competitive advantages in terms of energy efficiency, computing accuracy, and operating speed.

    摘要 i 致謝 vii 目錄 viii 圖目錄 x 表目錄 xiii 第1章 緒論 1 1.1 研究概觀 1 1.1.1 背景介紹 1 1.1.2 記憶體內運算 4 1.1.3 記憶體內運算之分類 7 1.1.4 記憶體內運算之相關名詞解釋 13 1.2 研究動機 17 1.3 論文貢獻 19 1.4 論文架構 19 第2章 相關研究文獻 20 2.1 A Twin-8T SRAM-based CIM [39] 20 2.2 A Two-Way Transpose 6T-SRAM CIM [32] 25 2.3 A 6T-SRAM CIM for 8b MAC operation [31] 30 2.4 相關文獻總結 34 第3章 具權重位元分割累加機制與混合位元量化機制之8T SRAM記憶體內運算 36 3.1 設計概述 36 3.2 設計架構 37 3.3 權重計算區塊 40 3.4 數位時間轉換器 44 3.5 混合位元類比數位轉換器陣列 47 3.6 輸出合成器 52 第4章 實驗結果與分析 55 4.1 數位時間轉換器 55 4.1.1 數位時間轉換器之非線性度分析 55 4.1.2 數位時間轉換器之蒙地卡羅分析 56 4.2 權重計算區塊 57 4.2.1 權重計算區塊之基本運作 57 4.2.2 8T SRAM單元之靜態雜訊邊界分析 58 4.2.3 權重計算區塊之非線性度分析 59 4.3 混合位元類比數位轉換器陣列 60 4.3.1 類比數位轉換器之基本操作 60 4.3.2 類比數位轉換器之非線性度分析 60 4.4 輸出合成器 62 4.5 整體效能分析 63 4.5.1 能量效益分析 63 4.5.2 電路功耗占比 65 4.6 晶片佈局 66 第5章 結論與未來研究方向 68 5.1 結論 68 5.2 未來工作 70 參考文獻 71

    [1] Yann Le Cun, Lawrence D. Jackel, Bernhard Boser, John S. Denker,Hans Peter Graf, Isabelle Guyon, D.Henderson, Richard E. Howard, and W.Hubbard “Handwritten Digit Recognition: Applications of Neural Network Chips and Automatic Learning,” IEEE Communications Magazine, vol. 27, no. November, pp. 41–46, 1989.
    [2] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton, “ImageNet Classification with Deep Convolutional Neural Networks,” in Proc. the 25th International Conference on Neural Information Processing Systems, 2012, pp. 1097–1105.
    [3] Karen Simonyan, and Andrew Zisserman, “Very deep convolutional networks for large-scale image recognition,” in Proc. 3rd International Conference on Learning Representations( ICLR), 2015, pp. 1–14.
    [4] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, “Deep residual learning for image recognition,” in Proc. IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2016, pp. 770–778.
    [5] Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam, “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications,” 2017. [Online]. Available: http://arxiv.org/abs/1704.04861.
    [6] Kou-Hung Lawrence Loh, “1.2 Fertilizing AIoT from Roots to Leaves,” in Proc. IEEE International Solid- State Circuits Conference - (ISSCC), 2020, pp. 15–21.
    [7] IDC FutureScape, “IDC FutureScape : Worldwide IT Industry 2020 Predictions,” Idc, 2020.[Online].Available: https://www.statista.com/statistics/ 1134766/nominal-gdp-driven-by-digitally-transformed-enterprises/.
    [8] Hoi-Jun Yoo, “1.2 Intelligence on Silicon: From Deep-Neural-Network Accelerators to Brain Mimicking AI-SoCs,” in Proc. IEEE International Solid-State Circuits Conference - (ISSCC), 2019, pp. 20–26.
    [9] Yu-Hsin Chen, Tushar Krishna, Joel S. Emer, and Vivienne Sze, “Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks,” IEEE Journal of Solid-State Circuits, vol. 52, no. 1, pp. 127–138, 2017.
    [10] Seongwook Park, Kyeongryeol Bong, Dongjoo Shin, Jinmook Lee, Sungpill Choi, and Hoi-Jun Yoo, “A 1.93TOPS/W Scalable Deep Learning /Inference Processor with Tetra-Parallel MIMD Architecture for Big-Data Applications,” in Proc. IEEE International Solid-State Circuits Conference - (ISSCC), 2015, pp. 80–82.
    [11] Zidong Du, Robert Fasthuber, Tianshi Chen, Paolo Ienne, Ling Li, Tao Luo, Xiaobing Feng, Yunji Chen, and Olivier Temam, “ShiDianNao: Shifting vision processing closer to the sensor,” in Proc. International Symposium on Computer Architecture, 2015, pp. 92–104.
    [12] Chuan-Jia Jhang, Cheng-Xin Xue, Je-Min Hung, Fu-Chun Chang, and Meng-Fan Chang, “Challenges and Trends of SRAM-Based Computing-In-Memory for AI Edge Devices,” IEEE Transactions on Circuit and Systems- I: Regular papers, vol. 68, no. 5, pp. 1773–1786, 2021.
    [13] Kea-Tiong Tang, Wei-Chen Wei, Zuo-Wei Yeh, Tzu-Hsiang Hsu, Yen-Cheng Chiu, Cheng-Xin Xue, Yu-Chun Kuo, Tai-Hsing Wen, Mon-Shu Ho, Chung-Chuan Lo, Ren-Shuo Liu, Chih-Cheng Hsieh, and Meng-Fan Chang,“Considerations of Integrating Computing-In-Memory and Processing-In-Sensor into Convolutional Neural Network Accelerators for Low-Power Edge Devices,” in Proc. IEEE Symposium on VLSI Circuits, Digest of Technical Papers, 2019, pp. T166–T167.
    [14] Haitong Li, Tony F. Wu, Subhasish Mitra, and H.-S. Philip Wong, “Resistive RAM-Centric Computing: Design and Modeling Methodology,” IEEE Transactions on Circuits and Systems I: Regular Papers, vol. 64, no. 9, pp. 2263–2273, 2017.
    [15] Reiji Mochida, Kazuyuki Kouno, Yuriko Hayata, Masayoshi Nakayama, Takashi Ono, Hitoshi Suwa, Ryutaro Yasuhara, Koji Katayama, Takumi Mikawa, and Yasushi Gohou, “A 4M synapses integrated analog ReRAM based 66.5 TOPS/W neural-network processor with cell current controlled writing and flexible network architecture,” in Proc. Symposium on VLSI Technology, 2018, pp. 175–176.
    [16] Tony F. Wu, Haitong Li, Ping-Chen Huang, Abbas Rahimi, Jan M. Rabaey, H.-S. Philip Wong, Max M. Shulaker , and Subhasish Mitra, “Brain-inspired computing exploiting carbon nanotube FETs and resistive RAM: Hyperdimensional computing case study,” in Proc. IEEE International Solid-State Circuits Conference - (ISSCC), 2018, vol. 61, pp. 492–494.
    [17] Cheng-Xin Xue, Wei-Hao Chen, Je-Syu Liu, Jia-Fang Li, Wei-Yu Lin, Wei-En Lin, Jing-Hong Wang, Wei-Chen Wei, Ting-Wei Chang, Tung-Cheng Chang, Tsung-Yuan Huang, Hui-Yao Kao, Shih-Ying Wei, Yen-Cheng Chiu, Chun-Ying Lee, Chung-Chuan Lo, Ya-Chin King, Chorng-Jung Lin, Ren-Shuo Liu, Chih-Cheng Hsieh, Kea-Tiong Tang, and Meng-Fan Chang, “24.1 A 1Mb Multibit ReRAM Computing-In-Memory Macro with 14.6ns Parallel MAC Computing Time for CNN Based AI Edge Processors,” in Proc. IEEE International Solid-State Circuits Conference, 2019, pp. 388–390.
    [18] Cheng-Xin Xue, Wei-Hao Chen, Je-Syu Liu, Jia-Fang Li, Wei-Yu Lin, Wei-En Lin, Jing-Hong Wang, Wei-Chen Wei, Tsung-Yuan Huang, Ting-Wei Chang, Tung-Cheng Chang, Hui-Yao Kao, Yen-Cheng Chiu, Chun-Ying Lee, Ya-Chin King , Chrong-Jung Lin, Ren-Shuo Liu, Chih-Cheng Hsieh , Kea-Tiong Tang, and Meng-Fan Chang, “Embedded 1-Mb ReRAM-Based Computing-in- Memory Macro with Multibit Input and Weight for CNN-Based AI Edge Processors,” IEEE Journal of Solid-State Circuits, vol. 55, no. 1, pp. 203–215, 2020.
    [19] Cheng-Xin Xue, Je-Min Hung, Hui-Yao Kao, Yen-Hsiang Huang, Sheng-Po Huang, Fu-Chun Chang, Peng Chen, Ta-Wei Liu,Chuan-Jia Jhang, Chin-I Su, Win-San Khwa , Chung-Chuan Lo , Ren-Shuo Liu , Chih-Cheng Hsieh , Kea-Tiong Tang, Yu-Der Chih , Tsung-Yung Jonathan Chang, and Meng-Fan Chang, “16.1 A 22nm 4Mb 8b-Precision ReRAM Computing-in-Memory Macro with 11.91 to 195.7TOPS/W for Tiny AI Edge Devices,” in Proc. IEEE International Solid-State Circuits Conference - (ISSCC), 2021, vol. 0, pp. 246–248.
    [20] Yu Pan, Peng Ouyang, Yinglin Zhao,Wang Kang, Shouyi Yin Youguang Zhang, Weisheng Zhao, and Shaojun Wei, “A Multilevel Cell STT-MRAM-Based Computing In-Memory Accelerator for Binary Convolutional Neural Network,” IEEE Transactions on Magnetics, vol. 54, no. 11, pp. 1–5, 2018.
    [21] Yuhan Shi, Sangheon Oh, Zhisheng Huang, Xiao Lu, Seung H. Kang, and Duygu Kuzum, “Performance Prospects of Deeply Scaled Spin-Transfer Torque Magnetic Random-Access Memory for In-Memory Computing,” IEEE Electron Device Letters, vol. 41, no. 7, pp. 1126–1129, 2020.
    [22] Iason Giannopoulos, Abu Sebastian, Manuel Le Gallo, Vara Prasad Jonnalagadda, M. Sousa, M.N. Boon, and Evangelos Eleftheriou, “8-bit Precision In-Memory Multiplication with Projected Phase-Change Memory,” in Proc. International Electron Devices Meeting (IEDM), 2019, pp. 27.7.1-27.7.4.
    [23] Je-Min Hung, Xueqing Li, Juejian Wu, and Meng-Fan Chang, “Challenges and Trends inDeveloping Nonvolatile Memory-Enabled Computing Chips for Intelligent Edge Devices,” IEEE Transactions on Electron Devices, vol. 67, no. 4, pp. 1444–1453, 2020.
    [24] Sujan K. Gonugondla, Mingu Kang, and Naresh R. Shanbhag, “A 42pJ/decision 3.12TOPS/W robust in-memory machine learning classifier with on-chip training,” in Proc. IEEE International Solid-State Circuits Conference - (ISSCC), 2018, vol. 61, pp. 490–492.
    [25] Win-San Khwa, Jia-Jing Chen, Jia-Fang Li, Xin Si, En-Yu Yang, Xiaoyu Sun, Rui Liu, Pai-Yu Chen, Qiang Li, Shimeng Yu, and Meng-Fan Chang, “A 65nm 4Kb algorithm-dependent computing-in-memory SRAM unit-macro with 2.3ns and 55.8TOPS/W fully parallel product-sum operation for binary DNN edge processors,” in Proc. IEEE International Solid-State Circuits Conference, 2018, vol. 61, pp. 496–498.
    [26] Mingu Kang, Sujan K. Gonugondla, Sujan K. Gonugondla, and Naresh R. Shanbhag, “A 19.4-nJ/Decision, 364-K Decisions/s, In-Memory Random Forest Multi-Class Inference Accelerator,” IEEE Journal of Solid-State Circuits, vol. 53, no. 7, pp. 2126–2135, 2018.
    [27] Mingu Kang, Sujan K. Gonugondla, Ameya Patil, and Naresh R. Shanbhag, “A Multi-Functional In-Memory Inference Processor Using a Standard 6T SRAM Array,” IEEE Journal of Solid-State Circuits, vol. 53, no. 2, pp. 642–655, 2018.
    [28] Ruiqi Guo, Yonggang Liu, Shixuan Zheng, Ssu-Yen Wu, Peng Ouyang, Win-San Khwa, Xi Chen, Jia-Jing Chen, Xiudong Li, Leibo Liu, Meng-Fan Chang, Shaojun Wei, and Shouyi Yin, “A 5.1pJ/Neuron 127.3us/Inference RNN-based Speech Recognition Processor using 16 Computing-in-Memory SRAM Macros in 65nm CMOS,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing, 2019, pp. 6351–6355.
    [29] Jinseok Kim, Jongeun Koo, Taesu Kim, Yulhwa Kim, Hyungjun Kim, Seunghyun Yoo, and Jae-Joon Kim, “Area-Efficient and Variation-Tolerant In-Memory BNN Computing using 6T SRAM Array,” in Proc. Symposium on VLSI Circuits, 2019, pp. 5–6.
    [30] Xin Si, Win-SanKhwa , Jia-Jing Chen, Jia-Fang Li, Xiaoyu Sun, Rui Liu, Shimeng Yu, Hiroyuki Yamauchi, Qiang Li, and Meng-Fan Chang, “A Dual-Split 6T SRAM-Based Computing-in-Memory Unit-Macro with Fully Parallel Product-Sum Operation for Binarized DNN Edge Processors,”IEEE Transactions on Circuits and Systems I: Regular Papers, vol. 66, no. 11, pp. 4172–4185, 2019.
    [31] Jian-Wei Su, Xin Si, Yen-Chi Chou, Ting-Wei Chang, Wei-Hsing Huang, Yung-Ning Tu, Ruhui Liu, Pei-Jung Lu, Ta-Wei Liu, Jing-Hong Wang, Zhixiao Zhang, Hongwu Jiang, Shanshi Huang,Chung-Chuan Lo, Ren-Shuo Liu, Chih-Cheng Hsieh, Kea-Tiong Tang, Shyh-Shyuan Sheu, Sih-Han Li, Heng-Yuan Lee, Shih-Chieh Chang, Shimeng Yu, and Meng-Fan Chang, “A 28nm 64Kb 6T SRAM Computing-in-Memory Macro with 8b MAC Operation for AI Edge Chips,” in Proc. IEEE International Solid-State Circuits Conference - (ISSCC), 2020, pp. 246–248.
    [32] Jian-Wei Su, Xin Si, Yen-Chi Chou, Ting-Wei Chang, Wei-Hsing Huang, Yung-Ning Tu, Ruhui Liu, Pei-Jung Lu, Ta-Wei Liu, Jing-Hong Wang, Zhixiao Zhang, Hongwu Jiang, Shanshi Huang, Chung-Chuan Lo, Ren-Shuo Liu, Chih-Cheng Hsieh, Kea-Tiong Tang ,Shyh-Shyuan Sheu, Sih-Han Li, Heng-Yuan Lee, Shih-Chieh Chang, Shimeng Yu, and Meng-Fan Chang, “A 28nm 64Kb Inference-Training Two-Way Transpose Multibit 6T SRAM Compute-in-Memory Macro for AI Edge Chips,” in Proc. IEEE International Solid-State Circuits Conference - (ISSCC), 2020, pp. 240–242.
    [33] Jian-Wei Su, Yen-Chi Chou, Ruhui Liu, Ta-Wei Liu, Pei-Jung Lu, Ping-Chun Wu, Yen-Lin Chung, Li-Yang Hung, Jin-Sheng Ren, Tianlong Pan, Sih-Han Li, Shih-Chieh Chang, Shyh-Shyuan Sheu,Wei-Chung Lo, Chih Wu, Xin Si, Chung-Chuan Lo, Ren-Shuo Liu, Chih-Cheng Hsieh, Kea-Tiong Tang, and Meng-Fan Chang, “A 28nm 384kb 6T-SRAM Computation-in-Memory Macro with 8b Precision for AI Edge Chips,” in Proc. IEEE International Solid-State Circuits Conference - (ISSCC), 2021, pp. 250–252.
    [34] Yu-Der Chih, Po-Hao Lee, Hidehiro Fujiwara, Yi-Chun Shih, Chia-Fu Lee, Rawan Naous, Yu-Lin Chen, Chieh-Pu Lo, Cheng-Han Lu, Haruki Mori, Wei-Chang Zhao, Dar Sun, Mahmut E. Sinangil, Yen-Huei Chen, Tan-Li Chou, Kerem Akarvardar, Hung-Jen Liao, Yih Wang, Meng-Fan Chang, and Tsung-Yung Jonathan Chang, “An 89TOPS/W and 16.3TOPS/mm2 All-Digital SRAM-Based Full-Precision Compute-In Memory Macro in 22nm for Machine-Learning Edge Applications,” in Proc. IEEE International Solid-State Circuits Conference - (ISSCC), 2021, vol. 64, pp. 252–254.
    [35] Jun Yang, Yuyao Kong, Zhen Wang, Yan Liu, Bo Wang, Shouyi Yin, and Longxin Shi,“Sandwich-RAM: An Energy-Efficient In-Memory BWN Architecture with Pulse-Width Modulation,” in Proc. 2019 IEEE International Solid- State Circuits Conference - (ISSCC), 2019, pp. 394–396.
    [36] Qing Dong , Mahmut E. Sinangil , Burak Erbagci, Dar Sun , Win-San Khwa, Hung-Jen Liao, Yih Wang, and Jonathan Chang, “A 351TOPS/W and 372.4GOPS Compute-in-Memory SRAM Macro in 7nm FinFET CMOS for Machine-Learning Applications,” in Proc. IEEE International Solid-State Circuits Conference - (ISSCC), 2020, pp. 488–489.
    [37] Avishek Biswas, and Anantha P. Chandrakasan Massachusetts, “Conv-RAM: An energy-efficient SRAM with embedded convolution computation for low-power CNN-based machine learning applications,” in Proc. IEEE International Solid-State Circuits Conference - (ISSCC), 2018, vol. 61, pp. 488–490.
    [38] Z.Jiang, S.Yin, M.Seok, and J. S.Seo, “XNOR-SRAM: In-Memory Computing SRAM Macro for Binary/Ternary Deep Neural Networks,” in Proc. Symposium on VLSI Technology, 2018, pp. 173–174.
    [39] Xin Si, Jia-Jing Chen, Yung-Ning Tu, Wei-Hsing Huang, Jing-Hong Wang, Yen-Cheng Chiu, Wei-Chen Wei, Ssu-Yen Wu, Xiaoyu Sun, Rui Liu, Shimeng Yu, Ren-Shuo Liu, Chih-Cheng Hsieh , Kea-Tiong Tang, Qiang Li, and Meng-Fan Chang, “A Twin-8T SRAM Computation-In-Memory Macro for Multiple-Bit CNN-Based Machine Learning,” in Proc. IEEE International Solid- State Circuits Conference - (ISSCC), 2019, pp. 396–398.
    [40] Hossein Valavi, Peter J. Ramadge, Eric Nestler, and Naveen Verma, “A 64-Tile 2.4-Mb In-Memory-Computing CNN Accelerator Employing Charge-Domain Compute,” IEEE Journal of Solid-State Circuits, vol. 54, no. 6, pp. 1789–1799, 2019.
    [41] Leland Chang, Robert K. Montoye, Yutaka Nakamura, Kevin A. Batson, Richard J. Eickemeyer, Robert H. Dennard, Wilfried Haensch, and Damir Jamsek, “An 8T-SRAM for Variability Tolerance and Low-Voltage Operation in High-Performance Caches,” IEEE Journal of Solid-State Circuits, vol. 43, no. 4, pp. 956–962, 2008.
    [42] Arne Holst , “Number of Internet of Things IoT connected devices worldwide from 2019 to 2030,” Statista, 2021.

    下載圖示
    2026-09-02公開
    QR CODE