| 研究生: |
曾奕儒 Tseng, Yi-Ju |
|---|---|
| 論文名稱: |
Novella-NPU 編譯器之通用 Tiling 演算法與記憶體分配最佳化及驗證 Generic Tiling and On-Chip Memory-Aware Allocation: Compiler Optimization and Verification for Novella-NPU |
| 指導教授: |
陳中和
Chen, Chung-Ho |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 電機工程學系 Department of Electrical Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 106 |
| 中文關鍵詞: | 通用 Tiling 演算法 、記憶體分配 、指令集模擬器 |
| 外文關鍵詞: | Generic Tiling, NPU Instruction Set Simulator, On-Chip Memory-aware Allocation |
| 相關次數: | 點閱:70 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
隨著人工智慧技術持續發展,對高效能與低功耗 AI 加速器的需求日益增加,進 而促使 Neural Processing Unit (NPU) 成為重要之硬體加速方案。由於受限於實際硬體 之 On-Chip Memory 容量與運算資源,神經網路模型於部署至 NPU 時,往往必須先 將 Feature Maps 與 Tensor 進行 Tiling,以符合硬體可支援之處理規模。
然而,傳統未經融合 (Fusion) 之獨立算子執行模式在結合 Tiling 後,常導致大量 Intermediate Feature Maps 須頻繁於外部 DRAM 與 On-Chip Memory 之間搬移。此現 象不僅造成嚴重的記憶體存取瓶頸 (Memory Access Bottleneck),亦大幅削弱硬體加 速器於效能與功耗上的優勢。
為突破上述瓶頸,本論文針對 Novella-NPU 提出一套具硬體感知能力之 Compiler Backend 最佳化與驗證方法,並發展三項核心技術。首先,Pattern-based Fusion 使編譯器能夠以可擴展方式支援不同型態之融合節點,並將特殊函數對應之運算 圖樣轉為 Fusion 版本;其次,Generic Tiling 使編譯器能夠針對融合後之複雜運算 圖,分析節點間之生產者與消費者關係,並推導合適之切分方式;最後,On-Chip Memory-Aware Allocation 於融合節點執行過程中有效管理中間資料之生命週期與分 配,使短暫存活之 Intermediate Variables 得以儘可能保留於 On-Chip SRAM 中,以降 低不必要之外部記憶體存取。
本研究亦開發 Novella Andersen 之 Instruction Set Simulator (ISS),作為編譯器後 端之驗證工具,用以檢查 Macro Op Generation 階段所產生之指令正確性,並快速驗 證生成指令之執行結果。此外,本研究亦透過 RTL simulation 取得實際硬體模擬結 果,作為效能分析與功能驗證之依據。藉由結合 ISS 與 RTL simulation,本研究得以 同時兼顧編譯器開發過程中之快速除錯需求,以及最終硬體執行結果之真實性驗證。
實驗結果於 MBv2 中,總 I/O bytes 降低 53.9%,推論效能提升 2.04 倍;於 DeiTtiny 中,總 I/O bytes 降低 61.9%,推論效能提升 1.90 倍,驗證本研究方法可同時有 效降低資料搬移成本並提升整體執行效能。
The rapid growth of artificial intelligence has driven increasing demand for high-performance, energy-efficient AI accelerators, making Neural Processing Units (NPUs) an important hardware platform for efficient inference. However, when neural network models are deployed on NPUs, limited on-chip memory capacity often requires tensor tiling, which in turn causes frequent movement of intermediate feature maps between on-chip memory and external DRAM. This results in significant memory access overhead and reduces the performance benefits of hardware acceleration.
To address this problem, this thesis presents a hardware-aware compiler backend optimization and verification framework for the Novella-NPU. The proposed framework consists of three core optimization techniques and a verification flow. First, Pattern-based Fusion provides an extensible mechanism for supporting diverse fused operators, including special functions. Second, Generic Tiling analyzes producer-consumer relationships in fused computational graphs and derives appropriate spatial tiling sizes for the fusion node. Third, On-Chip Memory-Aware Allocation manages the lifetime and placement of internal tensors within fused nodes, keeping intermediate variables in on-chip Memory whenever possible. Compared with the original backend flow, the proposed framework improves both fusion extensibility and tiling generality, allowing more complex fused graphs to be mapped efficiently onto the target hardware. For verification, a custom C++ Instruction Set Simulator (ISS) was developed to validate macro-op generation and the behavior of generated instructions, while RTL simulation was used to obtain hardware-level results.
Experimental results on MobileNetV2 (MBv2) and DeiT-Tiny show that the proposed framework reduces total I/O bytes by 53.9% and 61.9%, respectively, while improving overall performance by 2.04× and 1.90×. Overall, these results show that the proposed compiler backend optimization enhances on-chip data reuse, reduces external memory traffic, and delivers more efficient execution for both CNN and Transformer-based models.
[1] T. Chen, T. Moreau, Z. Jiang, L. Zheng, E. Yan, H. Shen, M. Cowan, L. Wang, Y. Hu, L. Ceze, et al. TVM: An automated End-to-End optimizing compiler for deep learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18), pages 578–594, 2018.
[2] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
[3] A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017.
[4] P. H. Lee. Memory layout unit design of cnn/transformer unified accelerator and memory subsystem analysis. Master’s thesis, National Cheng Kung University, 2024.
[5] W.-L. Lo. Compiler-hardware co-optimization and msfp-supported compute core design for novella-npu. Master’s thesis, National Cheng Kung University, 2025.
[6] E.-Y. Pong. Non-linear function approximations and their npu operator legalization in tvm compiler. Master’s thesis, National Cheng Kung University, 2025.
[7] Y. Tai. Design of a kernel-agnostic compute core for convolution and gemm. Master’s thesis, National Cheng Kung University, 2024.
[8] H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou. Training data-efficient image transformers & distillation through attention. In International conference on machine learning, pages 10347–10357. PMLR, 2021.
[9] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
[10] C.-H. Wang. Custom compiler instruction generation and scheduling optimization for novella-npu with n:m sparse matrix operations support. Master’s thesis, National Cheng Kung University, 2025.
[11] S. Williams, A. Waterman, and D. Patterson. Roofline: an insightful visual performance model for multicore architectures. Communications of the ACM, 52(4):65–76, 2009.