簡易檢索 / 詳目顯示

研究生: 涂維淳
Tu, Wei-Chun
論文名稱: 自回歸模型之 TVM 編譯器動態形狀支援與 JIT 技術研究
Dynamic Shape and JustInTime Support in TVM Compiler for Auto-regressive Models
指導教授: 陳中和
Chen, Chung-Ho
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 電機工程學系
Department of Electrical Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 86
中文關鍵詞: 模型編譯器TVM動態形狀Just-In-TimeJIT
外文關鍵詞: Model Compiler, TVM, Dynamic Shape, Just-In-Time, JIT
相關次數: 點閱:21下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 近年來,自回歸生成模型廣泛應用於自然語言處理、語音辨識與大型語言模型等領域。此類模型在推論過程中常具有輸入長度隨執行階段變化的特性,並可能搭配 KV Cache 保存過去 token 的中間結果。然而,傳統模型編譯流程多以靜態形狀為假設,張量大小、記憶體配置、tiling 結果與指令流通常皆於編譯階段決定,因此難以直接支援自回歸模型中的動態輸入與執行階段狀態管理。
    本論文以 Novella-NPU 之 TVM 編譯器為研究對象,提出動態形狀支援與 Just-In-Time(JIT)編譯機制。在動態形狀支援方面,本論文將中後端編譯流程逐步遷移至 Relax IR,並透過 Relay-to-Relax 轉換與 Shape Mutator,將會隨 token length 改變的維度表示為 symbolic shape,使編譯器能保留並推導動態維度與張量形狀之間的關係。同時,本論文亦修改後端 Legalize、Allocation、Tiling、MacroOp Gen 與 Inst Gen 等階段,使其能依據不同輸入長度產生對應執行內容,並由 runtime 選擇正確的指令流執行。
    此外,本論文提出 JIT 編譯機制,將可重複使用的指令樣板與每個 tile 不同的位址、offset 及 dependency 資訊分離,分別以 MacroMeta 與 TileMeta 表示,並於執行階段動態產生實際送入 Novella-NPU 的指令。此設計能降低大量 tile 與不同輸入形狀所造成的重複指令負擔,提升編譯器面對大型模型與動態形狀模型時的可擴展性。
    本論文之貢獻亦為後續 Shape Bucketing 與 KV Cache 支援奠定基礎。Shape Bucketing 需要編譯器能管理不同 token length 的執行流程,而本論文完成的動態形狀支援與 runtime instruction selection 正是其前提。KV Cache 則需要 runtime 依據生成步驟動態調整 Key 與 Value 的記憶體位址,而本論文提出的 JIT 機制可作為執行階段注入位址與 offset 的基礎。整體而言,本論文使 Novella-NPU TVM 編譯器具備處理動態形狀模型的能力,並為未來支援長序列與大型自回歸模型推論提供關鍵架構。

    This thesis investigates dynamic shape support and Just-In-Time (JIT) compilation in the TVM compiler for Novella-NPU, with the objective of improving support for auto-regressive models whose input lengths and execution states may change during runtime. Such models introduce challenges to conventional static-shape compilation flows, since tensor shapes, memory allocation, tiling strategies, and instruction streams are usually determined at compile time.
    To address these challenges, this work extends the Novella-NPU compiler flow by migrating the mid-end and back-end toward Relax IR and introducing symbolic shape handling through Relay-to-Relax conversion and a Shape Mutator. The back-end compilation stages, including Legalize, Allocation, Tiling, MacroOp generation, and instruction generation, are modified to generate executable contents for different input lengths and allow the runtime to select the proper instruction stream. In addition, a JIT mechanism is proposed to separate reusable instruction patterns from tile-specific information. MacroMeta and TileMeta are generated at compile time, while the runtime uses them to produce the final instructions with correct addresses, offsets, and dependency information.
    The proposed dynamic shape support is validated on multiple Novella operations with different token lengths, and the JIT flow is verified using Whisper Encoder. The results show that the compiler can correctly execute dynamic-shape workloads and reduce repeated instruction storage.
    In conclusion, this thesis enables Novella-NPU's TVM compiler to support dynamic-shape auto-regressive workloads and provides a foundation for future Shape Bucketing and KV Cache optimizations.

    摘要 i 英文延伸摘要 ii 誌謝 xvii 目錄 xviii 表格 xx 圖片 xxi Chapter 1. Introduction 1 1.1. 論文動機 1 1.2. 論文貢獻 2 1.3. 論文架構 2 Chapter 2. Background 3 2.1. 自注意力模型Self­Attention and Transformer 3 2.1.1. Encoder 3 2.1.2. Decoder 7 2.1.3. Output Linear Projection and Softmax 9 2.2. 自回歸推論與KV 快取Auto­Regression & KV­Cache 10 2.2.1. 自回歸生成 11 2.2.2. KV­Cache 13 2.3. TVM 模型編譯器及動態形狀支援 16 2.3.1. 模型編譯器Model Compiler 16 2.3.2. 中介表示(IR) 18 2.3.3. 新版IR(Relax IR) 19 2.4. Novella­NPU 後端編譯流程 23 Chapter 3. 動態形狀(Dynamic Shape) 之支援 30 3.1. 中端(Mid­end) 30 3.1.1. Relay­to­Relax 31 3.1.2. Shape Mutator 31 3.2. 後端(Back­end) 32 3.2.1. Legalize 33 3.2.2. Const Gen 35 3.2.3. Allocation 35 3.2.4. Tiling 36 3.2.5. MacroOp Gen 37 3.2.6. Inst Gen(Emitter) 37 3.3. 執行階段(Runtime) 38 Chapter 4. JIT(Just In Time) 編譯與執行階段支援 39 4.1. 後端(Back­end) 39 4.1.1. TileMeta 41 4.1.2. MacroMeta 42 4.1.3. 編譯流程 44 4.2. 執行階段(Runtime) 45 4.2.1. Share Bucket 45 4.2.2. Unique Bucket 46 Chapter 5. 驗證與評估 48 5.1. 動態形狀支援之功能正確性驗證 48 5.2. JIT 編譯與執行結果之一致性驗證 49 5.3. JIT 對指令大小壓縮效果之評估 50 5.3.1. Binary Elementwise 與Unary Elementwise Operation 50 5.3.2. Reduce Operation 51 5.3.3. Conv2D Operation 51 5.3.4. Whisper Encoder 54 Chapter 6. 未來展望 57 6.1. Shape Bucketing 57 6.1.1. 後端Back­end 58 6.1.2. Runtime 58 6.2. 利用KV Cache 優化執行速度 59 6.2.1. 後端Backend 59 6.2.2. Runtime 60 Chapter 7. Conclusion 61 7.1. 動態形狀張量之支援 61 7.2. Just­In­Time 編譯機制 61 7.3. 未來展望 62 References 63

    [1] Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, et al. {TVM}: An automated {End­to­End} optimizing compiler for deep learning. In 13th USENIX symposium on operating systems design and implementation (OSDI 18), pages 578–594, 2018.
    [2] Zhi Chen, Cody Hao Yu, Trevor Morris, Jorn Tuyls, Yi­Hsiang Lai, Jared Roesch, Elliott Delaye, Vin Sharma, and Yida Wang. Bring your own codegen to deep learning compiler. arXiv preprint arXiv:2105.03215, 2021.
    [3] Ruihang Lai, Junru Shao, Siyuan Feng, Steven S. Lyubomirsky, Bohan Hou, Wuwei Lin, Zihao Ye, Hongyi Jin, Yuchen Jin, Jiawei Liu, Lesheng Jin, Yaxing Cai, Ziheng Jiang, Yong Wu, Sunghyun Park, Prakalp Srivastava, Jared G. Roesch, Todd C. Mowry, and Tianqi Chen. Relax: Composable abstractions for end­to­end dynamic machine learning. arXiv preprint arXiv:2311.02103, 2023.
    [4] Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. Robust speech recognition via large­scale weak supervision. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 28492–28518. PMLR, 2023.
    [5] Jared Roesch, Steven Lyubomirsky, Marisa Kirisame, Josh Pollock, Logan Weber, Luis Vega, Ziheng Jiang, Tianqi Chen, Thierry Moreau, and Zachary Tatlock. Relay: A high­ level compiler for deep learning. arXiv preprint arXiv:1904.08368, 2019.
    [6] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.

    下載圖示
    校外:立即公開
    QR CODE