| 研究生: |
黃晟瑋 Huang, Chen-Wei |
|---|---|
| 論文名稱: |
基於 LLVM OpenMP 之網路附加型運算儲存異質卸載框架 An LLVM OpenMP-Based Heterogeneous Offloading Framework for Network-Attached Computational Storage |
| 指導教授: |
侯廷偉
Hou, Ting-Wei |
| 學位類別: |
碩士 Master |
| 系所名稱: |
工學院 - 工程科學系 Department of Engineering Science |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 135 |
| 中文關鍵詞: | 運算儲存裝置 、近資料處理 、分散式異質系統 、LLVM 、OpenMP |
| 外文關鍵詞: | Computational Storage Device (CSD), Near-Data Processing (NDP), Heterogeneous Distributed Systems, LLVM, OpenMP |
| 相關次數: | 點閱:65 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
隨著邊緣運算與物件辨識技術的普及,如何於資源受限的嵌入式環境中高效部署深度學習模型,已成為分散式運算領域的重要課題。本研究針對運算儲存裝置(Computational Storage Device, CSD)之層次化分散式異質叢集,探討軟體並行調度與影像讀取策略對整體推論吞吐量的影響。
原基礎卸載實作在處理 AI 推論任務時,因設計偏向單一遠端加速裝置的驗證,在面對橫向擴充的多節點硬體環境時,容易受限於主從端傳輸通道的實體頻寬,且執行期全域鎖(Global Lock)之限制常引發通訊阻塞,導致強延展性(Strong Scaling)效能飽和。為突破此一軟體調度限制,本研究深入 LLVM OpenMP 執行期軟體層進行架構重構,導入適用於多控制器的獨立互斥鎖陣列,以實現主機端多執行緒與遠端多核心微處理器(MPU)間的非同步並行通訊。同時,本研究藉由優化近資料影像讀取策略,將實體檔案讀取與影像解碼分流至控制節點本地端執行,免除主從端間之大量影像資料搬移,從而規避通訊頻寬之限制。
本研究在由新唐 MA35D1 與多台 Milk-V Duo 構成的實體分散式硬體平台上,以YOLOv8 物件辨識模型進行多維度的基準測試。實驗結果顯示,在雙 MA35D1 微處理器(MPU)控制器配置下,本研究框架能將端到端吞吐量由 11.45 FPS 提升至35.73 FPS。此外,透過結合 Amdahl's Law 與 Universal Scalability Law (USL) 建立效能分析方法,能有效量化並定位系統在實體 I/O 存取與跨節點通訊上的效能瓶頸,並在排除本地讀圖開銷下達成 64.54 FPS 的純運算表現,說明了軟體框架自身的低通訊與建立成本。本研究之成果呈現了多裝置異質卸載調度在邊緣叢集下的延展表現,亦為未來 CSD 異質運算調度與通用化執行期平行框架之建構,提供具實驗數據佐證的實體設計範例。
Computational Storage Devices (CSDs) improve overall system efficiency by processing data near storage, thereby reducing unnecessary data movement and communication overhead. This thesis proposes a parallel runtime framework for a heterogeneous MPU-NPU edge cluster, designed to scale object detection workloads. To overcome bandwidth bottlenecks and resource contention in multi-device environments, the LLVM OpenMP runtime is extended and optimized to support MPUs as offload targets. The proposed framework incorporates an independent mutex array for concurrent device management and adopts near-data processing principles to perform image loading and decoding directly on MPU controllers. The framework is evaluated using the YOLOv8 model on a distributed platform comprising Nuvoton MA35D1 MPUs and Milk-V Duo accelerators. In addition, the Universal Scalability Law (USL) is adopted to characterize system scalability and analyze performance limitations under different data processing scenarios. Experimental results demonstrate that a dual-MPU configuration achieves an end-to-end throughput of 35.73 FPS, representing a 3.12× speedup over the baseline, and reaches a pure computation throughput of 64.54 FPS when excluding local file I/O overheads.
[1] D. Patterson et al., “A Case for Intelligent RAM,” IEEE Micro, vol. 17, no. 2, pp. 34–44, Mar. 1997, doi: 10.1109/40.592312.
[2] D. Tiwari et al., “Active Flash: Towards Energy-Efficient, In-Situ Data Analytics on Extreme-Scale Machines,” in 11th USENIX Conference on File and Storage Technologies (FAST 13), 2013, pp. 119–132. Accessed: Jun. 22, 2026. [Online]. Available: https://www.usenix.org/conference/fast13/technical-sessions/presentation/tiwari
[3] D. Fakhry, M. Abdelsalam, M. W. El-Kharashi, and M. Safar, “A Review on Computational Storage Devices and Near Memory Computing for High Performance Applications,” Mem.-Mater. Devices Circuits Syst., vol. 4, p. 100051, 2023.
[4] B. Gu et al., “Biscuit: A Framework for Near-Data Processing of Big Data Workloads,” ACM SIGARCH Comput. Archit. News, vol. 44, pp. 153–165, Jun. 2016, doi: 10.1145/3007787.3001154.
[5] J. H. Lee, H. Zhang, V. Lagrange, P. Krishnamoorthy, X. Zhao, and Y. S. Ki, “SmartSSD: FPGA Accelerated Near-Storage Data Analytics on SSD,” IEEE Comput. Archit. Lett., vol. 19, no. 2, pp. 110–113, Jul. 2020, doi: 10.1109/LCA.2020.3009347.
[6] Z. Ruan, T. He, and J. Cong, “INSIDER: Designing In-Storage Computing System for Emerging High-Performance Drive,” presented at the 2019 USENIX Annual Technical Conference (USENIX ATC 19), 2019, pp. 379–394. Accessed: Jun. 22, 2026. [Online]. Available: https://www.usenix.org/conference/atc19/presentation/ruan
[7] S. Lee and R. Eigenmann, “OpenMPC: Extended OpenMP Programming and Tuning for GPUs,” in Proceedings of the 2010 ACM/IEEE International Conference for High Performance Computing, Networking, Storage and Analysis, in SC ’10. USA: IEEE Computer Society, 13 2010, pp. 1–11. doi: 10.1109/SC.2010.36.
[8] S. J. Pennycook, S. D. Hammond, S. A. Wright, J. A. Herdman, I. Miller, and S. A. Jarvis, “An Investigation of the Performance Portability of OpenCL,” J. Parallel Distrib. Comput., vol. 73, no. 11, pp. 1439–1450, Nov. 2013, doi: 10.1016/j.jpdc.2012.07.005.
[9] B. Chapman, G. Jost, and R. van der Pas, Using OpenMP: Portable Shared Memory Parallel Programming (Scientific and Engineering Computation). The MIT Press, 2007.
[10] C. Piñeiro and J. C. Pichel, “OMP4Py: A Pure Python Implementation of OpenMP,” May 15, 2025, arXiv: arXiv:2411.14887. doi: 10.48550/arXiv.2411.14887.
[11] N. I. Jr, “Mixing C and Java for High Performance Computing,” Sep. 2013, Accessed: Jun. 22, 2026. [Online]. Available: https://www.mitre.org/news-insights/publication/mixing-c-and-java-high-performance-computing
[12] OpenMP, “Specifications,” OpenMP. Accessed: Jun. 22, 2026. [Online]. Available: https://www.openmp.org/specifications/
[13] B. Amy and W. Greg, “The Architecture of Open Source Applications (Volume 1) LLVM.” Accessed: Jun. 22, 2026. [Online]. Available: https://aosabook.org/en/v1/llvm.html
[14] B. Shan, M. Araya-Polo, A. M. Malik, and B. Chapman, “MPI-based Remote OpenMP Offloading: A More Efficient and Easy-to-use Implementation,” in Proceedings of the 14th International Workshop on Programming Models and Applications for Multicores and Manycores, in PMAM’23. New York, NY, USA: Association for Computing Machinery, 25 2023, pp. 50–59. doi: 10.1145/3582514.3582519.
[15] B. Shan, M. Araya-Polo, and B. Chapman, “DiOMP-Offloading: Toward Portable Distributed Heterogeneous OpenMP,” in Proceedings of the SC ’25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis, in SC Workshops ’25. New York, NY, USA: Association for Computing Machinery, 15 2025, pp. 1289–1301. doi: 10.1145/3731599.3767505.
[16] R. Rabenseifner, G. Hager, and G. Jost, “Hybrid MPI/OpenMP Parallel Programming on Clusters of Multi-Core SMP Nodes,” in 2009 17th Euromicro International Conference on Parallel, Distributed and Network-based Processing, Feb. 2009, pp. 427–436. doi: 10.1109/PDP.2009.43.
[17] H. Jin, D. Jespersen, P. Mehrotra, R. Biswas, L. Huang, and B. Chapman, “High Performance Computing Using MPI and OpenMP on Multi-Core Parallel Systems,” Parallel Comput., vol. 37, no. 9, pp. 562–575, Sep. 2011, doi: 10.1016/j.parco.2011.02.002.
[18] Chen, Jin-Lin, Design of a Storage-Oriented Heterogeneous Edge Computing Platform Extended with LLVM OpenMP Offloading. Tainan, Taiwan: M.S. thesis, National Cheng Kung University, 2025. [Online]. Available: https://hdl.handle.net/11296/5b386s
[19] G. M. Amdahl, “Validity of the Single Processor Approach to Achieving Large Scale Computing Capabilities,” in Proceedings of the April 18-20, 1967, spring joint computer conference, in AFIPS ’67 (Spring). New York, NY, USA: Association for Computing Machinery, 18 1967, pp. 483–485. doi: 10.1145/1465482.1465560.
[20] A. Alasandagutti, P. G. Bridges, and T. Estrada, “Grey-Box Machine Learning Prediction of Parallel Application Scaling,” in 2025 IEEE 32nd International Conference on High Performance Computing, Data, and Analytics (HiPC), Feb. 2025, pp. 33–43. doi: 10.1109/HiPC66333.2025.00013.
[21] N. J. Gunther, “A General Theory of Computational Scalability Based on Rational Functions,” arXiv.org. Accessed: Jul. 06, 2026. [Online]. Available: https://arxiv.org/abs/0808.1431v2
[22] Neil J. Gunther, Guerrilla Capacity Planning: A Tactical Approach to Planning for Highly Scalable Applications and Services. Berlin, Heidelberg: Springer, 2007. doi: 10.1007/978-3-540-31010-5.