簡易檢索 / 詳目顯示

研究生: 凃韋帆
Tu, Wei-Fan
論文名稱: 基於記憶體佈局單元增強之編譯器與硬體協同優化實現 Swin Transformer 於 Novella NPU 之端對端部署
End-to-End Swin Transformer Deployment on Novella NPU via Compiler-Hardware Co-Optimization with Memory Layout Unit Enhancement
指導教授: 陳中和
Chen, Chung-Ho
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 電機工程學系
Department of Electrical Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 109
中文關鍵詞: 神經網路加速器Swin TransformerTVM 編譯器硬體軟體協同設計
外文關鍵詞: Neural Processing Unit, Swin Transformer, TVM Compiler, Hardware-Software Co-design
相關次數: 點閱:39下載:2
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本研究以 Swin Transformer 部署至 Novella Neural Processing Unit(NPU)(Brown 版本)為核心,完成涵蓋編譯器中端、後端與硬體的端對端協同設計。

    Swin Transformer 相較於 DeiT 等標準 Vision Transformer,帶來兩類核心挑戰。其一,Window Partition 與 Cyclic Shift 等產生大量複雜的資料重排搬運,涉及四維張量操作,現有以三維 DMA 為主的指令集無法有效表達;其二,Shifted Window Multi-Head Attention 因視窗切割使 Batch 維度大於 1,相較於 DeiT 的 Multi-Head Attention,需要硬體支援 Batch Matrix Multiply(BMM),同時編譯器亦須對應擴充,以正確映射此類多批次矩陣運算。

    針對上述問題,本研究對 TVM 編譯器中端進行系統性擴充:Cleanup 階段負責計算圖清理,將 Swin 特有的複雜子圖化簡為硬體可處理或編譯器後端友善的形式;Partition 階段將計算圖切割為 NPU 子圖與 CPU 子圖,目標是使 Swin 的核心運算可以完全由 NPU 完整執行,降低對 CPU 的依賴;Legalization 階段則將子圖中的原生算子合法化為 NPU 硬體真正支援的操作;在編譯器後端,DRAM Allocator 修正處理 Reshape 的策略,解決 Swin 中大量 Reshape 帶來的記憶體搬運問題;Tiling 與 Lower 階段針對新增的 Spatial Transpose 與 BMM 算子設計對應的切分策略與硬體映射,並修改 dependency tag 排程以支援新算子的功能單元相依性。

    針對大量四維資料搬運造成的指令膨脹問題,本研究在 MLU 搬運單元的指令格式中新增 ALen、SrcZStride 與 DstZStride 三個欄位,將三維 MLU 擴充為四維,使相鄰的多筆搬運得以合併為單筆 MacroOp,同時大幅減少對應的 SRC_BASE 與 DST_BASE 等差指令寫入次數。此四維擴充同步實作於 TVM 編譯器後端、ISS 指令集模擬器與 RTL 硬體,並驗證三層行為完全一致。

    驗證方面,以 AlgoSim / ISS / RTL 三層架構確認模型端對端推論正確性,針對 Swin Transformer 模型,三層級之間的輸出均達到 bit-exact 一致。在 ImageNet-1k 驗證集(50,000 張)上,BF16 + msfp16-block8 量化推論的 top-1 精度為 81.11%,相較 FP32 原始模型基準(81.19%)僅下降 0.08 個百分點,驗證了編譯器轉換的數值正確性與量化格式的有效性。四維搬運優化在 MobileNet V2、DeiT-tiny、Swin Transformer 與 DeiT+Fusion四個模型上,使總指令數減少 25-35%,其中 LOAD/STORE MacroOp 減少約 50%,SRC_BASE/DST_BASE 等差指令減少 62-72%。硬體代價方面,RTL 合成結果顯示此擴充僅使整體 NPU 面積增加 0.66%,以極小的硬體成本換取顯著的指令效率提升;就記憶體佔用而言,指令壓縮亦使 DeiT 與 Swin 等 Transformer 模型的峰值DRAM 需求分別降低 9.7% 與 6.2%。

    Neural Processing Units (NPUs) offer high energy efficiency for edge inference, but their value depends on a compiler that can map fast-evolving model architectures onto fixed hardware. This thesis takes Swin Transformer---a hierarchical Vision Transformer whose windowed attention is representative of the complex tensor reshaping and batched matrix multiplication found in modern large models---as a driving case study, and delivers its end-to-end deployment on the Novella-NPU (Brown). The work spans three layers. First, the TVM compiler middle-end and back-end are systematically extended so that every operator of the model is legalized into the NPU subgraph, eliminating any fallback to the host CPU; new support covers batched matrix multiplication, spatial transposition, and cyclic shift. Second, the Memory Layout Unit (MLU) data-movement engine is extended from three to four dimensions by adding the ALen, SrcZStride, and DstZStride fields, so that an outer loop of many similar transfers collapses into a single MacroOp instead of being unrolled into many separate ones; this extension is co-implemented consistently in the compiler back-end, the instruction-set simulator, and the RTL hardware, with backward compatibility to the original three-dimensional behavior. Third, correctness is established through a three-level, bit-exact verification chain (AlgoSim / ISS / RTL) that checks every transformation at the operator, instruction, and circuit levels. On the ImageNet-1k validation set, BF16 + msfp16-block8 inference reaches 81.11% top-1 accuracy, only 0.08 points below the FP32 baseline, while top-5 accuracy is essentially unchanged. The 4D extension reduces total instruction count by 25-35% across four models---MobileNetV2, DeiT-tiny, Swin Transformer, and a layer-fused DeiT---for only a 0.66% NPU area overhead from full RTL synthesis, and lowers peak DRAM demand by up to 9.7%. A fallback-cost analysis further shows that, without this operator support, the operators would instead execute on a host CPU and end-to-end inference would be roughly 104 times slower, quantifying the practical value of keeping the entire model inside the NPU subgraph.

    摘要 i 英文延伸摘要 ii 誌謝 xv 目錄 xvi 表格 xix 圖片 xx 演算法 xxi Chapter 1. 緒論 1 1.1. 論文動機 1 1.2. 論文貢獻 2 1.3. 論文架構 2 Chapter 2. 背景知識與相關研究 3 2.1. 卷積神經網路 (Convolution Neural Network) 3 2.1.1. 卷積層 (Convolution Layer) 3 2.1.2. 分組卷積 (Group Convolution) 4 2.1.3. 點卷積 (Pointwise Convolution) 5 2.1.4. 池化層 (Pooling Layer) 6 2.1.5. MobileNet V2 6 2.2. Transformer Network 6 2.2.1. 自注意力機制 (Self-Attention) 7 2.2.2. Transformer Encoder 結構 8 2.2.3. 位置編碼 (Positional Encoding) 8 2.3. Vision Transformer 9 2.3.1. 輸入前處理 9 2.3.2. ViT 架構與原始 Transformer 的差異 10 2.3.3. DeiT 10 2.4. Swin Transformer 10 2.4.1. 階層式特徵圖與 Patch Merging 11 2.4.2. 視窗自注意力 (Window Multi-Head Self-Attention) 12 2.4.3. 移位視窗注意力 (Shifted Window Multi-Head Self-Attention) 12 2.4.4. 相對位置偏置 (Relative Position Bias) 13 2.5. Novella-NPU 系統架構概覽 13 2.5.1. Apache TVM 模型編譯器 15 2.5.2. Novella-NPU 編譯流程 15 2.5.3. 模擬與驗證工具鏈 17 2.6. Novella-NPU Brown 硬體架構 17 2.6.1. 計算單元(CTU 與 PPU) 18 2.6.2. 資料搬移(MLU) 19 2.6.3. 控制與記憶體(CC 與 UB) 20 2.6.4. 指令集架構 20 Chapter 3. Swin Transformer 編譯器實作 23 3.1. 編譯器中端 23 3.1.1. Cleanup:算子等價化簡 24 3.1.2. Partition:Relay 子圖切割 27 3.1.3. Legalization:算子合法化 28 3.2. 編譯器後端 31 3.2.1. 特徵圖記憶體管理 32 3.2.2. 計算映射: Batch Matrix Multiplication 34 3.2.3. 計算映射: Spatial Reduction 36 3.2.4. 資料洗牌: Channel Transposition 37 3.2.5. 資料洗牌: Cyclic Shift 39 Chapter 4. 四維 Memory Layout Unit 優化 42 4.1. 原有 MLU 架構回顧 42 4.1.1. Scatter-Gather 搬運機制 42 4.1.2. Emitter 差分發射機制 43 4.1.3. 3D 搬運的侷限 44 4.1.4. 指令組成分析 44 4.2. 四維搬運設計與更動 45 4.2.1. 演算法 45 4.2.2. 指令擴充 46 4.3. 四維搬運實作 47 4.3.1. ALen 合併策略 47 4.3.2. MacroOp 合併與 ALen 決策 48 4.3.3. ZStride 決策 49 4.3.4. Dependency Tag 修正 51 4.4. 應用實例:Spatial Transposition 51 4.4.1. Stride 交換策略 53 4.4.2. Tiling 策略 56 Chapter 5. AlgoSim、ISS 與 RTL 改動及驗證 58 5.1. 驗證架構概觀 58 5.2. 演算法模擬器(AlgoSim) 58 5.2.1. 算子模擬擴充 60 5.3. 指令集模擬器(ISS) 61 5.3.1. 四維 DMA 擴充 63 5.3.2. 基址擴充 63 5.4. RTL 硬體驗證 64 5.4.1. ISA 擴充 66 5.4.2. MLU 四維搬運 66 5.4.3. PPU Transpose Unit 修正 67 5.5. 三層驗證結論 68 Chapter 6. 實驗結果與分析 69 6.1. 實驗設置 69 6.1.1. 硬體平台 69 6.1.2. 評估模型與資料集 69 6.2. MLU 指令優化分析 70 6.2.1. 指令組成佔比 70 6.2.2. 四維搬運指令壓縮效益 73 6.2.3. 合理性分析 74 6.2.4. 硬體面積開銷 76 6.2.5. 指令記憶體佔比分析 77 6.3. 精度驗證 78 6.4. CPU Fallback 開銷分析 79 Chapter 7. 結論與未來展望 82 7.1. 研究總結 82 7.1.1. 端對端部署 82 7.1.2. MLU 四維搬運擴充 82 7.2. 未來展望 83 7.2.1. 大型語言模型部署 83 7.2.2. Layer Fusion 擴充 83 7.2.3. MLU Dispatch 動態展開 83 7.2.4. MLU Queue 寬度縮減 83 參考文獻 85

    [1] Y. Tai, “Design of a kernel-agnostic compute core for convolution and gemm,” Master’s thesis, National Cheng Kung University, 2024.
    [2] P. H. Lee, “Memory layout unit design of cnn/transformer unified accelerator and memory subsystem analysis,” Master’s thesis, National Cheng Kung University, 2024.
    [3] E. Y. Pong, “Non-linear function approximations and their npu operator legalization in tvm compiler,” Master’s thesis, National Cheng Kung University, 2025.
    [4] C. H. Wang, “Custom compiler instruction generation and scheduling optimization for novella-npu with n:m sparse matrix operations support,” Master’s thesis, National Cheng Kung University, 2025.
    [5] W.-L. Lo, “Compiler-hardware co-optimization and msfp-supported compute core design for novella-npu,” Master’s thesis, National Cheng Kung University, 2025.
    [6] T. Chen, T. Moreau, Z. Jiang, L. Zheng, E. Yan, H. Shen, M. Cowan, L. Wang, Y. Hu, L. Ceze, et al., “TVM: An automated End-to-End optimizing compiler for deep learning,” in 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18), pp. 578–594, 2018.
    [7] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “MobileNetV2: Inverted residuals and linear bottlenecks,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4510–4520, 2018.
    [8] H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in International Conference on Machine Learning (ICML), pp. 10347–10357, 2021.
    [9] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in IEEE International Conference on Computer Vision (ICCV), pp. 10012–10022, 2021.
    [10] A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems (NIPS), pp. 1097–1105, 2012.
    [11] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems (NeurIPS), pp. 5998–6008, 2017.
    [12] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations (ICLR), 2021.
    [13] O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei, “ImageNet large scale visual recognition challenge,” International Journal of Computer Vision (IJCV), vol. 115, no. 3, pp. 211–252, 2015.

    下載圖示
    校外:立即公開
    QR CODE