簡易檢索 / 詳目顯示

研究生: 洪維辰
Hung, Wei-Chen
論文名稱: 多NPU運算下深度神經網路之切點分析
Best Partition Point Analysis for Deep Neural Networks on Multi-NPU Systems
指導教授: 侯庭偉
Hou, Ting-Wei
學位類別: 碩士
Master
系所名稱: 工學院 - 工程科學系
Department of Engineering Science
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 62
中文關鍵詞: 分割式運算邊緣運算多 NPU 協同推論推論卸載
外文關鍵詞: Split Computing, Edge Computing, Multi-NPU Collaborative Inference, Inference Offloading
相關次數: 點閱:4下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 為克服單一硬體之資源限制,本研究提出並建構了一套基於主控端微處理器與多神經網路處理單元(Neural Processing Unit, NPU)之分散式管線化協同推論系統。本研究之核心貢獻在於開發一套最佳切點選擇機制。此分析工具僅需將「完整版之 AI 模型」先行執行一次,並透過 Docker 環境提取模型各階段之硬體運算週期(Cycles),藉此推估出各階段所需之執行時間。同時,該工具利用各切點產生之特徵圖大小計算網路傳輸時間。在無須預先切割模型與個別編譯的情況下,系統即可直接綜合比較「前端運算」、「網路傳輸」與「後端運算」三者之時間成本,進而精確找出管線化流程中之效能瓶頸與最佳平衡點。 有別於傳統依賴人工試錯之切分方式,本系統透過追蹤計算圖中之「特徵圖空間解析度縮減點」,有效規避了深度學習編譯器(如 TPU-MLIR)因運算子融合(Operator Fusion)所導致之切分失敗風險。同時,本研究基於靜態分析提出「記憶體超載預防機制」,在編譯前精準估算各子模型之靜態權重與動態特徵圖暫存需求,提前過濾超出硬體安全閾值之無效切點,大幅降低了實機部署與除錯之時間成本。
    在系統執行架構上,本研究將主控端與加速工作端之任務,透過作業系統層級之管線化平行處理機制,使網路張量傳輸與 NPU矩陣運算在時間軸上達到最大程度之重疊。此設計不僅維持了中間特徵圖的無損完整性,更有效隱藏了跨節點之通訊延遲。本系統所推測之最佳切分點在不同影像推論下皆具備高度一致性。系統所推估之各切點效能排名,與實機高壓連續推論下之實際 FPS 排序展現出正面的對應關係。此一高度吻合之現象,意味著本系統成功徹底免除傳統試錯法(Trial-and-Error)高昂的實機編譯與測試成本。開發者僅需仰賴編譯前之靜態資源精算與效能推估,即可達成部署。此特性確保了本分散式管線化架構在面對未來多樣化之模型與未知邊緣網路環境時,具備極強之自適應能力與免除重複調校之大優勢。
    以處理 5000 張影像之壓力測試為例,ResNet-18 之最佳切點 s3p0 吞吐量可達 10.89 FPS,相較於其他切分位置(如 s1p0 之 5.67 FPS)展現出近兩倍之效能領先;在 ResNet-50 模型中,最佳切點 S4P0 達到 6.03 FPS,亦大幅優於高通訊代價之 S2P3(2.93 FPS);而在輕量化網路 MobileNetV2 中,最佳切點 layer6 吞吐量達 28.51 FPS,對比 layer1 的 4.46 FPS,更凸顯了正確選擇切點之重要性。

    To overcome edge device resource limits, this research proposes a distributed pipelined inference system using a host microprocessor and multiple Neural Processing Units (NPUs). The core contribution is a best split-point selection mechanism. By profiling the full model once via Docker to extract hardware cycles and calculating transmission latency from feature map sizes, it identifies bottlenecks without pre-partitioning or individual compilation. Tracking "dimensionality reduction points" prevents compiler operator fusion failures, while static memory protection net pre-filters invalid points, eliminating manual trial-and-error.
    During execution, OS-level pipelined parallelism maximally overlaps network transmission with NPU computations. Because cross-node bandwidth remains the primary bottleneck, this dual-NPU setup achieves the best compute-communication balance; scaling to three or more NPUs is currently unsuitable. Stress tests on ResNet-18, ResNet-50, and MobileNetV2 demonstrate that estimated performance rankings highly align with actual high-stress FPS. This demonstrates the architecture can rapidly and adaptively deploy models in diverse edge environments without repetitive physical testing. The best splits achieved nearly 2x throughput gains, reaching 10.89 FPS for ResNet-18, 6.03 FPS for ResNet-50, and 28.51 FPS for MobileNetV2 compared to other configurations.

    摘要 I Extended Abstract II 致謝 VIII 表目錄 XI 圖目錄 XII 第一章 緒論 1 1.1 研究動機 1 1.2 研究目的 1 1.3 研究貢獻 2 1.4 研究架構 2 第二章 文獻探討 3 2.1 邊緣運算設備架構 3 2.2 分割式運算 6 2.2.1 模型切分 6 2.2.2 深度學習編譯器 7 2.2.3 切點選擇 9 2.3 管線化架構之優勢 10 2.4 任務調度 13 第三章 系統設計與實作 16 3.1 系統架構 17 3.1.1 系統軟體架構 18 3.1.2 OOM機制 22 3.2 分割式模型 23 3.2.1 深度神經網路切分機制 23 3.2.2 管線化平行處理 25 3.3 切點選擇 27 3.3.1 特徵圖降維點搜尋 27 3.3.2 OOM 記憶體超載預防機制 28 3.4 效能推測 29 第四章 研究成果與討論 31 4.1 實驗規格與環境 31 4.2 切點推測之效能評估 36 4.3 分割式管線化架構評估 40 第五章 結論與未來展望 45 5.1 結論 45 5.2 未來展望 46 參考文獻 47

    [1] M. G. S. Murshed, C. Murphy, D. Hou, N. Khan, G. Ananthanarayanan, and F. Hussain, “Machine Learning at the Network Edge: A Survey,” ACM Comput. Surv., vol. 54, no. 8, pp. 1–37, Nov. 2022, doi: 10.1145/3469029.
    [2] X. Wang, Y. Han, V. C. M. Leung, D. Niyato, X. Yan, and X. Chen, “Convergence of Edge Computing and Deep Learning: A Comprehensive Survey,” IEEE Commun. Surv. Tutor., vol. 22, no. 2, pp. 869–904, 2020, doi: 10.1109/COMST.2020.2970550.
    [3] J. Lin, W.-M. Chen, Y. Lin, J. Cohn, C. Gan, and S. Han, “MCUNet: tiny deep learning on IoT devices,” in Proceedings of the 34th International Conference on Neural Information Processing Systems, in NIPS ’20. Red Hook, NY, USA: Curran Associates Inc., 6 2020, pp. 11711–11722. Accessed: Aug. 10, 2026. [Online]. Available: https://dl.acm.org/doi/10.5555/3495724.3496706
    [4] A. Ignatov, R. Timofte, A. Kulik, et al.,“AI Benchmark: All About Deep Learning on Smartphones in 2019,” in 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), Oct. 2019, pp. 3617–3635. doi: 10.1109/ICCVW.2019.00447.
    [5] B. Jacob, S. Kligys, B. Chen, et al., “Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT: IEEE, Jun. 2018, pp. 2704–2713. doi: 10.1109/CVPR.2018.00286.
    [6] X. Peng, X. Shi, H. Dai, et al.,“Capuchin: Tensor-based GPU Memory Management for Deep Learning,” in Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems, in ASPLOS ’20. New York, NY, USA: Association for Computing Machinery, 13 2020, pp. 891–905. doi: 10.1145/3373376.3378505.
    [7] P. Hu, M. Lu, L. Wang, and G. Jiang, “TPU-MLIR: A Compiler For TPU Using MLIR,” Feb. 09, 2023, arXiv: arXiv:2210.15016. doi: 10.48550/arXiv.2210.15016.
    [8] Y.-S. Li, “Design of an Inference Offloading System with Multiple NPUs on an Edge Device,” M.S. Thesis, National Cheng Kung University. Accessed: Jul. 04, 2026. [Online]. Available: https://ndltd.ncl.edu.tw/cgi-bin/gs32/gsweb.cgi?randomimg=wZhzrk_1783156126&validpath=%2Ftmp%2F%5Enclcdr__doschk%2FwZhzrk_1783156126__MzY0MzIx&validinput=364321&check=%E7%A2%BA%E5%AE%9A
    [9] Y. Kang, J. Hauswald, C. Gao, et al., “Neurosurgeon: Collaborative Intelligence Between the Cloud and Mobile Edge,” in Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems, Xi’an China: ACM, Apr. 2017, pp. 615–629. doi: 10.1145/3037697.3037698.
    [10] E. Li, Z. Zhou, and X. Chen, “Edge Intelligence: On-Demand Deep Learning Model Co-Inference with Device-Edge Synergy,” in Proceedings of the 2018 Workshop on Mobile Edge Communications, in MECOMM’18. New York, NY, USA: Association for Computing Machinery, 7 2018, pp. 31–36. doi: 10.1145/3229556.3229562.
    [11] C.-T. Cheng, “Design and Implementation of Split Computing on Edge Devices,” National Cheng Kung University. Accessed: Jul. 04, 2026. [Online]. Available: https://ndltd.ncl.edu.tw/cgi-bin/gs32/gsweb.cgi?randomimg=TA6Ar__1783157718&validpath=%2Ftmp%2F%5Enclcdr__doschk%2FTA6Ar__1783157718__MTEzMTU0&validinput=113154&check=%E7%A2%BA%E5%AE%9A
    [12] D. Narayanan, A. Phanishayee, K. Shi, X. Chen, and M. Zaharia, “Memory-Efficient Pipeline-Parallel DNN Training,” in Proceedings of the 38th International Conference on Machine Learning, PMLR, Jul. 2021, pp. 7937–7947. Accessed: Aug. 10, 2026. [Online]. Available: https://proceedings.mlr.press/v139/narayanan21a.html
    [13] J. Karjee, P. Naik S, K. Anand, and V. N. Bhargav, “Split computing: DNN inference partition with load balancing in IoT-edge platform for beyond 5G,” Meas. Sens., vol. 23, p. 100409, Oct. 2022, doi: 10.1016/j.measen.2022.100409.
    [14] J. Kwon, J. Lee, and H. Kim, “Pipelining of a Mobile SoC and an External NPU for Accelerating CNN Inference,” IEEE Embed. Syst. Lett., vol. 16, no. 2, pp. 150–153, Jun. 2024, doi: 10.1109/LES.2023.3305016.
    [15] “resnet18 — Torchvision main documentation.” Accessed: Jul. 03, 2026. [Online]. Available: https://docs.pytorch.org/vision/main/models/generated/torchvision.models.resnet18.html
    [16] K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2016, pp. 770–778. doi: 10.1109/CVPR.2016.90.
    [17] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “MobileNetV2: Inverted Residuals and Linear Bottlenecks,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2018, pp. 4510–4520. doi: 10.1109/CVPR.2018.00474.

    QR CODE