| 研究生: |
凃韋帆 Tu, Wei-Fan |
|---|---|
| 論文名稱: |
基於記憶體佈局單元增強之編譯器與硬體協同優化實現 Swin Transformer 於 Novella NPU 之端對端部署 End-to-End Swin Transformer Deployment on Novella NPU via Compiler-Hardware Co-Optimization with Memory Layout Unit Enhancement |
| 指導教授: |
陳中和
Chen, Chung-Ho |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 電機工程學系 Department of Electrical Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 109 |
| 中文關鍵詞: | 神經網路加速器 、Swin Transformer 、TVM 編譯器 、硬體軟體協同設計 |
| 外文關鍵詞: | Neural Processing Unit, Swin Transformer, TVM Compiler, Hardware-Software Co-design |
| 相關次數: | 點閱:39 下載:2 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
本研究以 Swin Transformer 部署至 Novella Neural Processing Unit(NPU)(Brown 版本)為核心,完成涵蓋編譯器中端、後端與硬體的端對端協同設計。
Swin Transformer 相較於 DeiT 等標準 Vision Transformer,帶來兩類核心挑戰。其一,Window Partition 與 Cyclic Shift 等產生大量複雜的資料重排搬運,涉及四維張量操作,現有以三維 DMA 為主的指令集無法有效表達;其二,Shifted Window Multi-Head Attention 因視窗切割使 Batch 維度大於 1,相較於 DeiT 的 Multi-Head Attention,需要硬體支援 Batch Matrix Multiply(BMM),同時編譯器亦須對應擴充,以正確映射此類多批次矩陣運算。
針對上述問題,本研究對 TVM 編譯器中端進行系統性擴充:Cleanup 階段負責計算圖清理,將 Swin 特有的複雜子圖化簡為硬體可處理或編譯器後端友善的形式;Partition 階段將計算圖切割為 NPU 子圖與 CPU 子圖,目標是使 Swin 的核心運算可以完全由 NPU 完整執行,降低對 CPU 的依賴;Legalization 階段則將子圖中的原生算子合法化為 NPU 硬體真正支援的操作;在編譯器後端,DRAM Allocator 修正處理 Reshape 的策略,解決 Swin 中大量 Reshape 帶來的記憶體搬運問題;Tiling 與 Lower 階段針對新增的 Spatial Transpose 與 BMM 算子設計對應的切分策略與硬體映射,並修改 dependency tag 排程以支援新算子的功能單元相依性。
針對大量四維資料搬運造成的指令膨脹問題,本研究在 MLU 搬運單元的指令格式中新增 ALen、SrcZStride 與 DstZStride 三個欄位,將三維 MLU 擴充為四維,使相鄰的多筆搬運得以合併為單筆 MacroOp,同時大幅減少對應的 SRC_BASE 與 DST_BASE 等差指令寫入次數。此四維擴充同步實作於 TVM 編譯器後端、ISS 指令集模擬器與 RTL 硬體,並驗證三層行為完全一致。
驗證方面,以 AlgoSim / ISS / RTL 三層架構確認模型端對端推論正確性,針對 Swin Transformer 模型,三層級之間的輸出均達到 bit-exact 一致。在 ImageNet-1k 驗證集(50,000 張)上,BF16 + msfp16-block8 量化推論的 top-1 精度為 81.11%,相較 FP32 原始模型基準(81.19%)僅下降 0.08 個百分點,驗證了編譯器轉換的數值正確性與量化格式的有效性。四維搬運優化在 MobileNet V2、DeiT-tiny、Swin Transformer 與 DeiT+Fusion四個模型上,使總指令數減少 25-35%,其中 LOAD/STORE MacroOp 減少約 50%,SRC_BASE/DST_BASE 等差指令減少 62-72%。硬體代價方面,RTL 合成結果顯示此擴充僅使整體 NPU 面積增加 0.66%,以極小的硬體成本換取顯著的指令效率提升;就記憶體佔用而言,指令壓縮亦使 DeiT 與 Swin 等 Transformer 模型的峰值DRAM 需求分別降低 9.7% 與 6.2%。
Neural Processing Units (NPUs) offer high energy efficiency for edge inference, but their value depends on a compiler that can map fast-evolving model architectures onto fixed hardware. This thesis takes Swin Transformer---a hierarchical Vision Transformer whose windowed attention is representative of the complex tensor reshaping and batched matrix multiplication found in modern large models---as a driving case study, and delivers its end-to-end deployment on the Novella-NPU (Brown). The work spans three layers. First, the TVM compiler middle-end and back-end are systematically extended so that every operator of the model is legalized into the NPU subgraph, eliminating any fallback to the host CPU; new support covers batched matrix multiplication, spatial transposition, and cyclic shift. Second, the Memory Layout Unit (MLU) data-movement engine is extended from three to four dimensions by adding the ALen, SrcZStride, and DstZStride fields, so that an outer loop of many similar transfers collapses into a single MacroOp instead of being unrolled into many separate ones; this extension is co-implemented consistently in the compiler back-end, the instruction-set simulator, and the RTL hardware, with backward compatibility to the original three-dimensional behavior. Third, correctness is established through a three-level, bit-exact verification chain (AlgoSim / ISS / RTL) that checks every transformation at the operator, instruction, and circuit levels. On the ImageNet-1k validation set, BF16 + msfp16-block8 inference reaches 81.11% top-1 accuracy, only 0.08 points below the FP32 baseline, while top-5 accuracy is essentially unchanged. The 4D extension reduces total instruction count by 25-35% across four models---MobileNetV2, DeiT-tiny, Swin Transformer, and a layer-fused DeiT---for only a 0.66% NPU area overhead from full RTL synthesis, and lowers peak DRAM demand by up to 9.7%. A fallback-cost analysis further shows that, without this operator support, the operators would instead execute on a host CPU and end-to-end inference would be roughly 104 times slower, quantifying the practical value of keeping the entire model inside the NPU subgraph.
[1] Y. Tai, “Design of a kernel-agnostic compute core for convolution and gemm,” Master’s thesis, National Cheng Kung University, 2024.
[2] P. H. Lee, “Memory layout unit design of cnn/transformer unified accelerator and memory subsystem analysis,” Master’s thesis, National Cheng Kung University, 2024.
[3] E. Y. Pong, “Non-linear function approximations and their npu operator legalization in tvm compiler,” Master’s thesis, National Cheng Kung University, 2025.
[4] C. H. Wang, “Custom compiler instruction generation and scheduling optimization for novella-npu with n:m sparse matrix operations support,” Master’s thesis, National Cheng Kung University, 2025.
[5] W.-L. Lo, “Compiler-hardware co-optimization and msfp-supported compute core design for novella-npu,” Master’s thesis, National Cheng Kung University, 2025.
[6] T. Chen, T. Moreau, Z. Jiang, L. Zheng, E. Yan, H. Shen, M. Cowan, L. Wang, Y. Hu, L. Ceze, et al., “TVM: An automated End-to-End optimizing compiler for deep learning,” in 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18), pp. 578–594, 2018.
[7] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “MobileNetV2: Inverted residuals and linear bottlenecks,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4510–4520, 2018.
[8] H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in International Conference on Machine Learning (ICML), pp. 10347–10357, 2021.
[9] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in IEEE International Conference on Computer Vision (ICCV), pp. 10012–10022, 2021.
[10] A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems (NIPS), pp. 1097–1105, 2012.
[11] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems (NeurIPS), pp. 5998–6008, 2017.
[12] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations (ICLR), 2021.
[13] O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei, “ImageNet large scale visual recognition challenge,” International Journal of Computer Vision (IJCV), vol. 115, no. 3, pp. 211–252, 2015.