簡易檢索 / 詳目顯示

研究生: 黃晟瑋
Huang, Chen-Wei
論文名稱: 基於 LLVM OpenMP 之網路附加型運算儲存異質卸載框架
An LLVM OpenMP-Based Heterogeneous Offloading Framework for Network-Attached Computational Storage
指導教授: 侯廷偉
Hou, Ting-Wei
學位類別: 碩士
Master
系所名稱: 工學院 - 工程科學系
Department of Engineering Science
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 135
中文關鍵詞: 運算儲存裝置近資料處理分散式異質系統LLVMOpenMP
外文關鍵詞: Computational Storage Device (CSD), Near-Data Processing (NDP), Heterogeneous Distributed Systems, LLVM, OpenMP
相關次數: 點閱:65下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 隨著邊緣運算與物件辨識技術的普及,如何於資源受限的嵌入式環境中高效部署深度學習模型,已成為分散式運算領域的重要課題。本研究針對運算儲存裝置(Computational Storage Device, CSD)之層次化分散式異質叢集,探討軟體並行調度與影像讀取策略對整體推論吞吐量的影響。
    原基礎卸載實作在處理 AI 推論任務時,因設計偏向單一遠端加速裝置的驗證,在面對橫向擴充的多節點硬體環境時,容易受限於主從端傳輸通道的實體頻寬,且執行期全域鎖(Global Lock)之限制常引發通訊阻塞,導致強延展性(Strong Scaling)效能飽和。為突破此一軟體調度限制,本研究深入 LLVM OpenMP 執行期軟體層進行架構重構,導入適用於多控制器的獨立互斥鎖陣列,以實現主機端多執行緒與遠端多核心微處理器(MPU)間的非同步並行通訊。同時,本研究藉由優化近資料影像讀取策略,將實體檔案讀取與影像解碼分流至控制節點本地端執行,免除主從端間之大量影像資料搬移,從而規避通訊頻寬之限制。
    本研究在由新唐 MA35D1 與多台 Milk-V Duo 構成的實體分散式硬體平台上,以YOLOv8 物件辨識模型進行多維度的基準測試。實驗結果顯示,在雙 MA35D1 微處理器(MPU)控制器配置下,本研究框架能將端到端吞吐量由 11.45 FPS 提升至35.73 FPS。此外,透過結合 Amdahl's Law 與 Universal Scalability Law (USL) 建立效能分析方法,能有效量化並定位系統在實體 I/O 存取與跨節點通訊上的效能瓶頸,並在排除本地讀圖開銷下達成 64.54 FPS 的純運算表現,說明了軟體框架自身的低通訊與建立成本。本研究之成果呈現了多裝置異質卸載調度在邊緣叢集下的延展表現,亦為未來 CSD 異質運算調度與通用化執行期平行框架之建構,提供具實驗數據佐證的實體設計範例。

    Computational Storage Devices (CSDs) improve overall system efficiency by processing data near storage, thereby reducing unnecessary data movement and communication overhead. This thesis proposes a parallel runtime framework for a heterogeneous MPU-NPU edge cluster, designed to scale object detection workloads. To overcome bandwidth bottlenecks and resource contention in multi-device environments, the LLVM OpenMP runtime is extended and optimized to support MPUs as offload targets. The proposed framework incorporates an independent mutex array for concurrent device management and adopts near-data processing principles to perform image loading and decoding directly on MPU controllers. The framework is evaluated using the YOLOv8 model on a distributed platform comprising Nuvoton MA35D1 MPUs and Milk-V Duo accelerators. In addition, the Universal Scalability Law (USL) is adopted to characterize system scalability and analyze performance limitations under different data processing scenarios. Experimental results demonstrate that a dual-MPU configuration achieves an end-to-end throughput of 35.73 FPS, representing a 3.12× speedup over the baseline, and reaches a pure computation throughput of 64.54 FPS when excluding local file I/O overheads.

    摘要 I EXTENDED ABSTRACT II 致謝 XI 目錄 XIII 表目錄 XVI 圖目錄 XVII 第一章 緒論 1 1.1 研究動機 1 1.2 研究目的 2 1.3 研究貢獻 3 1.4 研究架構 4 第二章 文獻探討 5 2.1 近資料運算 (NDP) 與運算儲存裝置 (CSD) 5 2.2 異質運算並行模型:OpenMP 與指令式開發 7 2.3 LLVM 編譯器架構 10 2.4 LLVM OpenMP Offloading 執行架構支援 13 2.5 邊緣分散式通訊機制:MPI/OpenMP 混合式平行架構 18 2.6 異質卸載架構演進與優化方向 21 第三章 系統架構設計與實作 23 3.1 層次化異質叢集之硬體架構設計 24 3.2 系統軟體分層設計 28 3.2.1 主機端應用程式設計與控制流程 30 3.2.2 中間層控制器(MPU)軟體組件與調度流程 34 3.2.3 底層加速單元(NPU)軟體架構與推論流程 39 3.3 基於 LLVM OpenMP 之卸載工具鏈基礎 41 3.4 執行時並行優化與層次化通訊函式庫實作 47 3.4.1 中間層 MA35D1 控制器之 OpenMP 運行時環境補全 47 3.4.2 主機端通訊外掛程式之獨立裝置鎖重構 51 3.4.3 跨裝置通訊橋接函式庫封裝與 API 實作 54 3.4.4 裝置端近資料影像讀取函式庫封裝與 API 實作 56 3.5 通用開發介面與應用程式執行方法 60 3.5.1 通用指令式開發範式與應用範例 61 3.5.2 執行期系統生命週期與非同步通訊時序 62 3.5.3 框架實作小結 65 第四章 研究成果與討論 67 4.1 實驗規格與環境 67 4.1.1 異質叢集硬體規格與作業系統環境 68 4.1.2 軟體環境與編譯工具鏈配置 69 4.1.3 基準測試模型與資料集配置 69 4.2 基準測試案例與影像前處理效能分析 70 4.2.1 主機端影像讀取之多 NPU 叢集效能表現 71 4.2.2 裝置端近資料影像讀取之並行效能優化 75 4.2.3 雙控制器架構下主機端影像讀取之並行效能表現 80 4.2.4 雙控制器架構下裝置端近資料影像讀取之效能優化 83 4.3 系統效能預測模型與實測對比 88 4.3.1 Amdahl's Law之並行比例擬合與驗證(主機端讀圖) 88 4.3.2 USL之多前處理策略與多控制器參數擬合驗證 92 4.3.3 系統效能擴充邊界與極限分析 98 4.4 綜合討論與小結 103 第五章 結論與未來展望 108 5.1 結論 108 5.2 未來展望 109 參考文獻 111

    [1] D. Patterson et al., “A Case for Intelligent RAM,” IEEE Micro, vol. 17, no. 2, pp. 34–44, Mar. 1997, doi: 10.1109/40.592312.
    [2] D. Tiwari et al., “Active Flash: Towards Energy-Efficient, In-Situ Data Analytics on Extreme-Scale Machines,” in 11th USENIX Conference on File and Storage Technologies (FAST 13), 2013, pp. 119–132. Accessed: Jun. 22, 2026. [Online]. Available: https://www.usenix.org/conference/fast13/technical-sessions/presentation/tiwari
    [3] D. Fakhry, M. Abdelsalam, M. W. El-Kharashi, and M. Safar, “A Review on Computational Storage Devices and Near Memory Computing for High Performance Applications,” Mem.-Mater. Devices Circuits Syst., vol. 4, p. 100051, 2023.
    [4] B. Gu et al., “Biscuit: A Framework for Near-Data Processing of Big Data Workloads,” ACM SIGARCH Comput. Archit. News, vol. 44, pp. 153–165, Jun. 2016, doi: 10.1145/3007787.3001154.
    [5] J. H. Lee, H. Zhang, V. Lagrange, P. Krishnamoorthy, X. Zhao, and Y. S. Ki, “SmartSSD: FPGA Accelerated Near-Storage Data Analytics on SSD,” IEEE Comput. Archit. Lett., vol. 19, no. 2, pp. 110–113, Jul. 2020, doi: 10.1109/LCA.2020.3009347.
    [6] Z. Ruan, T. He, and J. Cong, “INSIDER: Designing In-Storage Computing System for Emerging High-Performance Drive,” presented at the 2019 USENIX Annual Technical Conference (USENIX ATC 19), 2019, pp. 379–394. Accessed: Jun. 22, 2026. [Online]. Available: https://www.usenix.org/conference/atc19/presentation/ruan
    [7] S. Lee and R. Eigenmann, “OpenMPC: Extended OpenMP Programming and Tuning for GPUs,” in Proceedings of the 2010 ACM/IEEE International Conference for High Performance Computing, Networking, Storage and Analysis, in SC ’10. USA: IEEE Computer Society, 13 2010, pp. 1–11. doi: 10.1109/SC.2010.36.
    [8] S. J. Pennycook, S. D. Hammond, S. A. Wright, J. A. Herdman, I. Miller, and S. A. Jarvis, “An Investigation of the Performance Portability of OpenCL,” J. Parallel Distrib. Comput., vol. 73, no. 11, pp. 1439–1450, Nov. 2013, doi: 10.1016/j.jpdc.2012.07.005.
    [9] B. Chapman, G. Jost, and R. van der Pas, Using OpenMP: Portable Shared Memory Parallel Programming (Scientific and Engineering Computation). The MIT Press, 2007.
    [10] C. Piñeiro and J. C. Pichel, “OMP4Py: A Pure Python Implementation of OpenMP,” May 15, 2025, arXiv: arXiv:2411.14887. doi: 10.48550/arXiv.2411.14887.
    [11] N. I. Jr, “Mixing C and Java for High Performance Computing,” Sep. 2013, Accessed: Jun. 22, 2026. [Online]. Available: https://www.mitre.org/news-insights/publication/mixing-c-and-java-high-performance-computing
    [12] OpenMP, “Specifications,” OpenMP. Accessed: Jun. 22, 2026. [Online]. Available: https://www.openmp.org/specifications/
    [13] B. Amy and W. Greg, “The Architecture of Open Source Applications (Volume 1) LLVM.” Accessed: Jun. 22, 2026. [Online]. Available: https://aosabook.org/en/v1/llvm.html
    [14] B. Shan, M. Araya-Polo, A. M. Malik, and B. Chapman, “MPI-based Remote OpenMP Offloading: A More Efficient and Easy-to-use Implementation,” in Proceedings of the 14th International Workshop on Programming Models and Applications for Multicores and Manycores, in PMAM’23. New York, NY, USA: Association for Computing Machinery, 25 2023, pp. 50–59. doi: 10.1145/3582514.3582519.
    [15] B. Shan, M. Araya-Polo, and B. Chapman, “DiOMP-Offloading: Toward Portable Distributed Heterogeneous OpenMP,” in Proceedings of the SC ’25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis, in SC Workshops ’25. New York, NY, USA: Association for Computing Machinery, 15 2025, pp. 1289–1301. doi: 10.1145/3731599.3767505.
    [16] R. Rabenseifner, G. Hager, and G. Jost, “Hybrid MPI/OpenMP Parallel Programming on Clusters of Multi-Core SMP Nodes,” in 2009 17th Euromicro International Conference on Parallel, Distributed and Network-based Processing, Feb. 2009, pp. 427–436. doi: 10.1109/PDP.2009.43.
    [17] H. Jin, D. Jespersen, P. Mehrotra, R. Biswas, L. Huang, and B. Chapman, “High Performance Computing Using MPI and OpenMP on Multi-Core Parallel Systems,” Parallel Comput., vol. 37, no. 9, pp. 562–575, Sep. 2011, doi: 10.1016/j.parco.2011.02.002.
    [18] Chen, Jin-Lin, Design of a Storage-Oriented Heterogeneous Edge Computing Platform Extended with LLVM OpenMP Offloading. Tainan, Taiwan: M.S. thesis, National Cheng Kung University, 2025. [Online]. Available: https://hdl.handle.net/11296/5b386s
    [19] G. M. Amdahl, “Validity of the Single Processor Approach to Achieving Large Scale Computing Capabilities,” in Proceedings of the April 18-20, 1967, spring joint computer conference, in AFIPS ’67 (Spring). New York, NY, USA: Association for Computing Machinery, 18 1967, pp. 483–485. doi: 10.1145/1465482.1465560.
    [20] A. Alasandagutti, P. G. Bridges, and T. Estrada, “Grey-Box Machine Learning Prediction of Parallel Application Scaling,” in 2025 IEEE 32nd International Conference on High Performance Computing, Data, and Analytics (HiPC), Feb. 2025, pp. 33–43. doi: 10.1109/HiPC66333.2025.00013.
    [21] N. J. Gunther, “A General Theory of Computational Scalability Based on Rational Functions,” arXiv.org. Accessed: Jul. 06, 2026. [Online]. Available: https://arxiv.org/abs/0808.1431v2
    [22] Neil J. Gunther, Guerrilla Capacity Planning: A Tactical Approach to Planning for Highly Scalable Applications and Services. Berlin, Heidelberg: Springer, 2007. doi: 10.1007/978-3-540-31010-5.

    下載圖示
    校外:立即公開
    QR CODE