| 研究生: |
鄧光宇 Teng, Kuang-Yu |
|---|---|
| 論文名稱: |
應用於鐵電場效電晶體記憶體內運算系統之元件變異感知穩健關鍵字辨識訓練 Device-Variation-Aware Training for Robust Keyword Spotting in FeFET-Based Compute-in-Memory Systems |
| 指導教授: |
盧達生
Lu, Darsen 張順志 Chang, Soon-Jyh |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 電機工程學系 Department of Electrical Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 147 |
| 中文關鍵詞: | 關鍵字辨識 、記憶體內運算 、鐵電場效電晶體 、元件變異 、變異感知訓練 |
| 外文關鍵詞: | keyword spotting, compute-in-memory, FeFET, device-to-device variation, variation-aware training |
| 相關次數: | 點閱:9 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
隨著邊緣人工智慧的快速發展,超低功耗且可長時間持續運作的關鍵字辨識系統需求日益增加。以鐵電場效電晶體(Ferroelectric Field-Effect Transistor, FeFET)為基礎之記憶體內運算(Compute-in-Memory, CIM)技術,由於具備非揮發性權重儲存與記憶體內運算能力,可有效降低待機功耗與資料搬移能耗,因此成為實現超低功耗邊緣人工智慧的重要技術。然而,FeFET 元件間變異所造成的類比運算誤差,將導致推論準確率下降,成為 FeFET-CIM 實際應用所面臨的重要挑戰。
本論文提出一套適用於 FeFET-CIM 關鍵字辨識系統之硬體感知變異感知訓練(Hardware-Aware Variation-Aware Training)架構。所建立之可微分 FeFET-CIM 模型整合位元序列乘加運算、類比數位轉換器量化以及元件間變異等硬體特性,並導入兩階段訓練流程,包括傳統量化感知訓練(Quantization-Aware Training, QAT)與變異感知微調(Variation-Aware Fine-Tuning),使模型於訓練階段即可學習補償硬體所造成的運算誤差。
本研究分別於 Software Accuracy、Hardware Accuracy 及 Hardware Accuracy with D2D Variation 三種執行條件下評估所提出方法的效能,並透過蒙地卡羅模擬分析不同硬體實例下的穩健性。實驗結果顯示,所提出之方法在理想執行條件下可維持與基準模型相近的推論準確率;在存在元件間變異時,則能顯著提升模型的推論準確率,同時有效降低不同硬體實例所造成的效能波動,展現良好的硬體穩健性。
綜合而言,本研究證實硬體感知變異感知訓練能有效提升 FeFET-CIM 關鍵字辨識系統對元件間變異的容忍能力,為未來超低功耗邊緣人工智慧應用中部署高可靠度 FeFET-CIM 系統提供一種具可行性的解決方案。
The rapid growth of edge artificial intelligence has created a strong demand for ultra-low-power, always-on keyword spotting (KWS) systems. Ferroelectric field-effect transistor (FeFET)-based compute-in-memory (CIM) is a promising solution because it reduces standby power and data-movement energy. However, device-to-device (D2D) variation introduces analog computation errors that degrade inference accuracy.
This thesis proposes a hardware-aware variation-aware training framework for FeFET-CIM-based KWS systems. A differentiable FeFET-CIM model that captures bit-serial MAC computation, ADC quantization, and D2D variation is incorporated into a two-stage training procedure consisting of conventional quantization-aware training and variation-aware fine-tuning.
The proposed framework is evaluated using Software Accuracy, Hardware Accuracy, and Hardware Accuracy with D2D Variation, where hardware robustness is assessed through Monte Carlo simulations. Experimental results show that the proposed method preserves comparable inference accuracy under ideal conditions while significantly improving robustness against D2D variation and reducing performance variability across different hardware realizations.
These results demonstrate that hardware-aware variation-aware training is an effective approach for deploying robust FeFET-CIM-based keyword spotting systems for future ultra-low-power edge AI applications.
[1] T. Wang, J. Guo, B. Zhang, et al., "Deploying AI on edge: Advancement and challenges in edge intelligence," Mathematics, vol. 13, no. 11, Art. no. 1878, 2025.
[2] Murata Manufacturing Co., Ltd., "Understanding edge AI: Artificial intelligence meets IoT," Murata Manufacturing Articles, Sep. 2024.
[3] C. Wolters, X. Yang, U. Schlichtmann, et al., "Memory is all you need: An overview of compute-in-memory architectures for accelerating large language model inference," arXiv:2406.08413, 2024.
[4] A. Gebregiorgis, H. A. Du Nguyen, J. Yu, et al., "A survey on memory-centric computer architectures," ACM Journal on Emerging Technologies in Computing Systems, vol. 18, no. 4, Art. no. 79, 2022.
[5] I. López-Espejo, Z.-H. Tan, J. H. L. Hansen, et al., "Deep spoken keyword spotting: An overview," IEEE Access, vol. 10, pp. 4169–4199, 2022.
[6] W. Shan, M. Yang, J. Xu, et al., "A 510 nW 0.41 V low-memory low-computation keyword-spotting chip using serial FFT-based MFCC and binarized depthwise separable convolutional neural network in 28 nm CMOS," in IEEE Int. Solid-State Circuits Conf. (ISSCC) Dig. Tech. Papers, 2020, pp. 146–148.
[7] Z. K. Abdul and A. K. Al-Talabani, "Mel frequency cepstral coefficient and its applications: A review," IEEE Access, vol. 10, pp. 122136–122158, 2022.
[8] K. Suzuki, K. Hiraga, K. Bessho, et al., "A 40 nm 2 kb MTJ-based non-volatile SRAM macro with novel data-aware store architecture for normally off computing," in 2023 IEEE Symposium on VLSI Technology and Circuits (VLSI Technology and Circuits), 2023.
[9] M. Faghani, H. Rezaee-Dehsorkh, N. Ravanshad, et al., "Ultra-low-power voice activity detection system using level-crossing sampling," Electronics, vol. 12, no. 4, Art. no. 795, 2023.
[10]M. Hellenbrand, I. Teck, and J. L. MacManus-Driscoll, "Progress of emerging non-volatile memory technologies in industry," MRS Communications, vol. 14, no. 6, pp. 1099–1112, 2024.
[11]D. Kau, S. Tang, I. V. Karpov, et al., "A stackable cross point phase change memory," in IEEE Int. Electron Devices Meeting (IEDM), 2009, pp. 1–4.
[12]S. Jung, H. Lee, S. Myung, et al., "A crossbar array of magnetoresistive memory devices for in-memory computing," Nature, vol. 601, no. 7892, pp. 211–216, 2022.
[13]Y. Huang, K. Cao, K. Zhang, et al., "Implementation of 16 Boolean logic operations based on one basic cell of spin-transfer-torque magnetic random access memory," Science China Information Sciences, vol. 66, Art. no. 162402, 2023.
[14]W. Cai, M. Wang, K. Cao, et al., "Stateful implication logic based on perpendicular magnetic tunnel junctions," Science China Information Sciences, vol. 65, Art. no. 122406, 2022.
[15]Y. Ling, Z. Wang, L. Wu, et al., "An RRAM-based hierarchical computing-in-memory architecture with synchronous parallelism for 3D point cloud recognition," IEEE Transactions on Circuits and Systems II: Express Briefs, vol. 71, no. 9, pp. 4412–4416, 2024.
[16]S.-M. Cho, J. Lee, H. Jo, et al., "Binary-weighted neural networks using FeRAM array for low-power AI computing," Nanomaterials, vol. 15, no. 15, Art. no. 1166, 2025.
[17]J. Kang, P. Huang, R. Han, et al., "Flash-based computing in-memory scheme for IoT," in IEEE Int. Conf. ASIC (ASICON), 2019, pp. 1–4.
[18]P. Duhan, T. Ali, P. Khedgarkar, et al., "Endurance study of silicon-doped hafnium oxide (HSO) and zirconium-doped hafnium oxide (HZO)-based FeFET memory," IEEE Transactions on Electron Devices, vol. 70, no. 11, pp. 5645–5650, 2023.
[19]X. Yin, Q. Huang, H. Errahmouni Barkam, et al., "A homogeneous FeFET-based time-domain compute-in-memory fabric for matrix-vector multiplication and associative search," IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 44, no. 5, pp. 1856–1868, 2025.
[20]C. Garg, N. Chauhan, S. Deng, et al., "Impact of random spatial fluctuation in non-uniform crystalline phases on the device variation of ferroelectric FET," IEEE Electron Device Letters, vol. 42, no. 8, pp. 1160–1163, 2021.
[21]K. Ni, W. Chakraborty, J. Smith, et al., "Fundamental understanding and control of device-to-device variation in deeply scaled ferroelectric FETs," in Proc. 2019 Symposium on VLSI Technology, 2019, pp. T40–T41.
[22]B. Manna, A. Saha, Z. Jiang, et al., "Variation-resilient FeFET-based in-memory computing leveraging probabilistic deep learning," IEEE Transactions on Electron Devices, vol. 71, no. 5, pp. 2963–2969, 2024.
[23]Z. He, J. Lin, R. Ewetz, et al., "Noise injection adaption: End-to-end ReRAM crossbar non-ideal effect adaption for neural network mapping," in Proc. 56th ACM/IEEE Design Automation Conference (DAC), 2019.
[24]Q. Wang, Y. Park, and W. D. Lu, "Device variation effects on neural network inference accuracy in analog in-memory computing systems," Advanced Intelligent Systems, vol. 4, no. 8, Art. no. 2100199, 2022.
[25]M. J. Rasch, C. Mackin, M. Le Gallo, et al., "Hardware-aware training for large-scale and diverse deep learning inference workloads using in-memory computing-based accelerators," Nature Communications, vol. 14, no. 1, Art. no. 5282, 2023.
[26]V. Joshi, M. Le Gallo, S. Haefeli, et al., "Accurate deep neural network inference using computational phase-change memory," Nature Communications, vol. 11, no. 1, Art. no. 2473, 2020.
[27]S. Kariyappa, H. Tsai, K. Spoon, et al., "Noise-resilient DNN: Tolerating noise in PCM-based AI accelerators via noise-aware training," IEEE Transactions on Electron Devices, vol. 68, no. 9, pp. 4356–4362, 2021.
[28]W. Jiang, Q. Lou, Z. Yan, et al., "Device-circuit-architecture co-exploration for computing-in-memory neural accelerators," IEEE Transactions on Computers, vol. 70, no. 4, pp. 595–605, 2021.
[29]Y. LeCun, Y. Bengio, and G. Hinton, "Deep learning," Nature, vol. 521, no. 7553, pp. 436–444, 2015.
[30]I. D. Mienye and T. G. Swart, "A comprehensive review of deep learning: Architectures, recent advances, and applications," Information, vol. 15, no. 12, Art. no. 755, 2024.
[31]V. Sze, Y.-H. Chen, T.-J. Yang, et al., "Efficient processing of deep neural networks: A tutorial and survey," Proceedings of the IEEE, vol. 105, no. 12, pp. 2295–2329, 2017.
[32]J. S. Larsen and L. Clemmensen, "Weight sharing and deep learning for spectral data," in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, pp. 4227–4231.
[33]G.-B. Huang, Z. Bai, L. L. C. Kasun, et al., "Local receptive fields based extreme learning machine," IEEE Computational Intelligence Magazine, vol. 10, no. 2, pp. 18–29, 2015.
[34]H. Gholamalinezhad and H. Khosravi, "Pooling methods in deep neural networks, a review," arXiv:2009.07485, 2020.
[35]A. D. Rasamoelina, F. Adjailia, and P. Sinčák, "A review of activation function for artificial neural network," in Proc. IEEE 18th World Symp. Appl. Mach. Intell. Informat. (SAMI), 2020, pp. 281–286.
[36]L. Maas, A. Y. Hannun, and A. Y. Ng, "Rectifier nonlinearities improve neural network acoustic models," in Proc. ICML Workshop Deep Learn. Audio, Speech, Lang. Process., 2013.
[37]S. Ioffe and C. Szegedy, "Batch normalization: Accelerating deep network training by reducing internal covariate shift," in Proc. Int. Conf. Mach. Learn. (ICML), 2015, pp. 448–456.
[38]N. Boyko, K. Boksho, and P. Telishevskyi, "Neural networks: Training with backpropagation and the gradient algorithm," in Proc. IEEE 9th Int. Conf. Problems Infocommun. Sci. Technol. (PIC S&T), 2022, pp. 1–6.
[39]O. Elharrouss, Y. Mahmood, Y. Bechqito, et al., "Loss functions in deep learning: A comprehensive review," arXiv:2504.04242, 2025.
[40]Z. Mo, Z. Zhang, and K.-L. Tsui, "Domain generalization study of empirical risk minimization from causal perspectives," IEEE Transactions on Multimedia, vol. 27, pp. 4284–4296, 2025.
[41]S. Ruder, "An overview of gradient descent optimization algorithms," arXiv:1609.04747, 2016.
[42]S. Damadi, G. Moharrer, M. Cham, et al., "The backpropagation algorithm for a math student," in Proc. Int. Joint Conf. Neural Netw. (IJCNN), 2023, pp. 1–9.
[43]D. P. Kingma and J. Ba, "Adam: A method for stochastic optimization," in Proc. Int. Conf. Learn. Represent. (ICLR), 2015.
[44]X. Ying, "An overview of overfitting and its solutions," Journal of Physics: Conference Series, vol. 1168, Art. no. 022022, 2019.
[45]N. Srivastava, G. Hinton, A. Krizhevsky, et al., "Dropout: A simple way to prevent neural networks from overfitting," Journal of Machine Learning Research, vol. 15, no. 1, pp. 1929–1958, 2014.
[46]A. Krogh and J. A. Hertz, "A simple weight decay can improve generalization," in Proc. Adv. Neural Inf. Process. Syst. (NIPS), vol. 4, 1991, pp. 950–957.
[47]L. Prechelt, "Early stopping—but when?," in Neural Networks: Tricks of the Trade, 2nd ed., G. Montavon, G. B. Orr, and K.-R. Müller, Eds. Springer, Berlin, Germany, pp. 53–67, 2012.
[48]Y. Bengio, N. Léonard, and A. Courville, "Estimating or propagating gradients through stochastic neurons for conditional computation," arXiv:1308.3432, 2013.
[49]I. McLoughlin, L. Pham, Y. Song, et al., "Spectrogram features for audio and speech analysis," Applied Sciences, vol. 16, no. 2, Art. no. 572, 2026.
[50]M. Sahidullah and G. Saha, "A novel windowing technique for efficient computation of MFCC for speaker recognition," IEEE Signal Processing Letters, vol. 20, no. 2, pp. 149–152, 2013.
[51]M. A. Yusnita, M. P. Paulraj, S. Yaacob, et al., "Analysis of accent-sensitive words in multi-resolution mel-frequency cepstral coefficients for classification of accents in Malaysian English," International Journal of Automotive and Mechanical Engineering, vol. 7, pp. 1053–1073, 2013.
[52]Y. Ma, N. Suda, Y. Cao, et al., "Scalable and modularized RTL compilation of convolutional neural networks onto FPGA," in Proc. 26th Int. Conf. Field Programmable Logic and Applications (FPL), 2016, pp. 1–8.
[53]R. Krishnamoorthi, "Quantizing deep convolutional networks for efficient inference: A whitepaper," arXiv:1806.08342, 2018.
[54]E. H. Lee, D. Miyashita, E. Chai, et al., "LogNet: Energy-efficient neural networks using logarithmic computation," in Proc. IEEE Int. Conf. Acoustics, Speech and Signal Processing (ICASSP), 2017, pp. 5900–5904.
[55]A. Zhou, A. Yao, Y. Guo, et al., "Incremental network quantization: Towards lossless CNNs with low-precision weights," in Proc. Int. Conf. Learn. Represent. (ICLR), 2017.
[56]M. Courbariaux, Y. Bengio, and J.-P. David, "BinaryConnect: Training deep neural networks with binary weights during propagations," in Proc. Adv. Neural Inf. Process. Syst. (NIPS), vol. 28, 2015, pp. 3123–3131.
[57]M. Horowitz, "Computing's energy problem (and what we can do about it)," in Proc. IEEE Int. Solid-State Circuits Conf. (ISSCC) Dig. Tech. Papers, 2014, pp. 10–14.
[58]J. Fang, A. Shafiee, H. Abdel-Aziz, et al., "Post-training piecewise linear quantization for deep neural networks," in Proc. Eur. Conf. Comput. Vis. (ECCV), 2020, pp. 69–86.
[59]D. Miyashita, E. H. Lee, and B. Murmann, "Convolutional neural networks using logarithmic data representation," arXiv:1603.01025, 2016.
[60]S. Han, H. Mao, and W. J. Dally, "Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding," in Proc. Int. Conf. Learn. Represent. (ICLR), 2016.
[61]S. P. Lloyd, "Least squares quantization in PCM," IEEE Transactions on Information Theory, vol. 28, no. 2, pp. 129–137, 1982.
[62]T. Dettmers, A. Pagnoni, A. Holtzman, et al., "QLoRA: Efficient finetuning of quantized LLMs," in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 36, pp. 10088–10115, 2023.
[63]Z. Sun and R. Huang, "Time complexity of in-memory matrix-vector multiplication," IEEE Transactions on Circuits and Systems II: Express Briefs, vol. 68, no. 8, pp. 2785–2789, 2021.
[64]S. Sun, J. Bai, H. Chen, et al., "Model quantization for computing-in-memory: A survey," Science China Information Sciences, vol. 68, no. 11, Art. no. 211401, 2025.
[65]X. Si, S. Jain, C. H. Kim, et al., "A Twin-8T SRAM computation-in-memory macro for multiple-bit CNN-based machine learning," in IEEE Int. Solid-State Circuits Conf. (ISSCC) Dig. Tech. Papers, 2019, pp. 396–398.
[66]J.-W. Su, C.-X. Xue, C.-Y. Chuang, et al., "A 28 nm 384 kb 6T-SRAM computation-in-memory macro with 8-bit precision for AI edge chips," in IEEE Int. Solid-State Circuits Conf. (ISSCC) Dig. Tech. Papers, 2021, pp. 250–252.
[67]M. Kang, S. K. Gonugondla, A. Patil, et al., "A multi-functional in-memory inference processor using a standard 6T SRAM array," IEEE Journal of Solid-State Circuits, vol. 53, no. 2, pp. 642–655, 2018.
[68]X. Si, Y. Luo, X. Sun, et al., "A 28 nm 64 Kb 6T SRAM computing-in-memory macro with 8-bit MAC operation for AI edge chips," in IEEE Int. Solid-State Circuits Conf. (ISSCC) Dig. Tech. Papers, 2020, pp. 246–248.
[69]M. E. Sinangil, B. Erbagci, R. Naous, et al., "A 7-nm compute-in-memory SRAM macro supporting multi-bit input, weight and output and achieving 351 TOPS/W and 372.4 GOPS," IEEE Journal of Solid-State Circuits, vol. 56, no. 1, pp. 188–198, 2021.
[70]C.-J. Jhang, C.-X. Xue, J.-M. Hung, et al., "Challenges and trends of SRAM-based computing-in-memory for AI edge devices," IEEE Transactions on Circuits and Systems I: Regular Papers, vol. 68, no. 5, pp. 1773–1786, 2021.
[71]J. Zhang, Z. Wang, and N. Verma, "In-memory computation of a machine-learning classifier in a standard 6T SRAM array," IEEE Journal of Solid-State Circuits, vol. 52, no. 4, pp. 915–924, 2017.
[72]A. Biswas and A. P. Chandrakasan, "CONV-SRAM: An energy-efficient SRAM with in-memory dot-product computation for low-power convolutional neural networks," IEEE Journal of Solid-State Circuits, vol. 54, no. 1, pp. 217–230, 2019.
[73]J.-H. Fu and S.-J. Chang, "A 12 TOPS/W computing-in-memory accelerator for convolutional neural networks," in Proc. IEEE Int. Symp. Circuits and Systems (ISCAS), 2022.
[74]K. Bai, H. Yan, H. Li, et al., "Research on voice activity detection methods based on deep learning," in Proc. 14th Asian Control Conference (ASCC), 2024.
[75]J. Liao, B. Zeng, Q. Sun, et al., "Grain size engineering of ferroelectric Zr-doped HfO₂ for the highly scaled devices applications," IEEE Electron Device Letters, vol. 40, no. 11, pp. 1868–1871, 2019.
[76]A. Vardar, N. Laleni, S. Baskaran, et al., "A 28 nm FeFET-based content-addressable memory for energy-efficient similarity search and few-shot learning," IEEE Journal of the Electron Devices Society, early access, doi: 10.1109/JEDS.2025.3648870, 2025.
[77]T. Soliman, S. Chatterjee, N. Laleni, et al., "First demonstration of in-memory computing crossbar using multi-level Cell FeFET," Nature Communications, vol. 14, no. 1, Art. no. 6348, 2023.
[78]Z. Yan, X. S. Hu, and Y. Shi, "SWIM: Selective write-verify for computing-in-memory neural accelerators," in Proc. 59th ACM/IEEE Design Automation Conference (DAC), 2022.
[79]Z. Yan, X. S. Hu, and Y. Shi, "U-SWIM: Universal selective write-verify for computing-in-memory neural accelerators," IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 43, no. 6, pp. 1822–1833, 2024.
[80]V. Garg, J. Jia, O. Phadke, et al., "A 28-nm FeFET compute-in-memory macro with 64×64 array size and on-chip 4-bit flash ADC," IEEE Solid-State Circuits Letters, vol. 9, pp. 13–16, 2026.
[81]Z. Jiang, H. Zhao, J. Tang, et al., "Strategies of high-accuracy memristor-based analogue computing in memory for artificial intelligence," Nature Materials, vol. 25, no. 7, pp. 1110–1124, 2026.
[82]P. Mannocci, G. Larelli, M. Bonomi, and D. Ielmini, "Achieving high precision in analog in-memory computing systems," npj Unconventional Computing, vol. 3, no. 1, Art. no. 1, 2026.
[83]Y. Zhou, Z. Wang, Y. Li, et al., "Characterizing and demystifying the implicit convolution algorithm on commercial matrix-multiplication accelerators," in Proc. IEEE Int. Symp. Workload Characterization (IISWC), 2021, pp. 214–225.