| 研究生: |
余季芸 Yu, Chi-Yun |
|---|---|
| 論文名稱: |
基於儲存裝置內運算之平行處理介面設計 A Parallel Programming Interface for In-storage Processing |
| 指導教授: |
侯廷偉
Hou, Ting-Wei |
| 學位類別: |
碩士 Master |
| 系所名稱: |
工學院 - 工程科學系 Department of Engineering Science |
| 論文出版年: | 2024 |
| 畢業學年度: | 112 |
| 語文別: | 中文 |
| 論文頁數: | 64 |
| 中文關鍵詞: | 平行處理介面 、儲存裝置內運算 、人臉辨識 、Open MPI 、OpenMP |
| 外文關鍵詞: | Parallel Programming Interface, In-storage Processing, Facial Recognition, Open MPI, OpenMP |
| 相關次數: | 點閱:148 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
儲存裝置內運算(In-Storage Processing, ISP)通過在儲存設備內部執行計算,減少資料搬移、節省時間和能源,從而提高系統效率,成為解決大量資料搬移問題的一種方法。但若能一次使多台儲存裝置內運算設備同時平行運算,整體運行速率即可再次提高。本論文提出的平行處理介面即是同時驅動多台儲存裝置內運算設備,並且以平行計算的角度將運算分配到多台ISP單元執行。再以人臉辨識應用為例,展現本研究之可行性。除此之外,還進一步利用Open MPI和OpenMP驗證本研究之介面可以執行平行處理程式。本研究在資料搬移方面,所選用的人臉辨識應用的測試程式,確實可以大量減少資料搬移量。運算方面,根據實驗結果顯示,使用8台ISP裝置進行100張照片辨識時,Open MPI和OpenMP都得到5.8倍加速。另外本論文所提出的設計在與主機連接市售USB 隨身碟比較中,8台ISP裝置效率最多可提升 13.2倍。
In-Storage Processing (ISP) enhances system efficiency by performing computations within storage devices, reducing data movement, saving time, and conserving energy. This approach has emerged as a solution to the challenges of large-scale data movement. The overall operational speed can be further increased by enabling multiple ISP devices to compute simultaneously. This thesis proposes a parallel programming interface that drives multiple ISP devices simultaneously, distributing computations across multiple ISP units from a parallel computing perspective. The feasibility of this thesis is demonstrated using a facial recognition application as a case study. Additionally, the interface's ability to execute parallel processing programs is validated using Open MPI and OpenMP. Regarding data transfer, the chosen facial recognition test program significantly reduces data movement. Regarding computation, experimental results show that using 8 ISP devices for recognizing 100 photos achieves a speedup of 5.8 times with both Open MPI and OpenMP. Compared to a commercially available USB flash drive connected to the host, the proposed design with 8 ISP devices achieves an efficiency improvement of up to 13.2 times.
[1] "Open MPI," [Online]. Available: https://www.open-mpi.org/. [Accessed 15 July 2024].
[2] O. A. R. Board, "OpenMP," [Online]. Available: https://www.openmp.org/. [Accessed 15 July 2024].
[3] H. Jin, D. Jespersen, P. Mehrotra, R. Biswas, L. Huang and B. Chapman, "High performance computing using MPI and OpenMP on multi-core parallel systems," Parallel Computing, Volume 37, Issue 9, pp. 562-575, September 2011.
[4] N. Stankovic and K. Zhang, "A distributed parallel programming framework," IEEE Transactions on Software Engineering, pp. 478-493, 2002, doi: 10.1109/TSE.2002.1000451.
[5] C. Jiang , G. Su and X. Liu , "A distributed parallel system for face recognition," Proceedings of the Fourth International Conference on Parallel and Distributed Computing, Applications and Technologies, Chengdu, China, pp. 797-800, August 2003, doi: 10.1109/PDCAT.2003.1236417.
[6] T. Rauber and G. Rünger, Parallel Programming: for Multicore and Cluster Systems, New York City: Springer International Publishing, 2010, p. 450.
[7] A. Ali and K. S. Syed, "An outlook of high performance computing infrastructures for scientific computing," in Advances in Computers, Amsterdam, Netherlands, Elsevier, 2013, pp. 87-118.
[8] E. Gabriel, G. E. Fagg, G. Bosilca, T. Angskun , . J. J. Dongarra, J. M. Squyres, . V. Sahay, P. Kambadur, . B. Barrett , . A. Lumsdaine, R. H. Castain, . D. J. Daniel, R. L. Graham and T. S. Woodall, "Open MPI: Goals, Concept, and Design of a Next Generation MPI Implementation," in Recent Advances in Parallel Virtual Machine and Message Passing Interface, Berlin Heidelberg, Springer, 2004, doi: https://doi.org/10.1007/978-3-540-30218-6_19, pp. 97-104.
[9] G. Shipman, T. Woodall, R. Graham, A. Maccabe and P. Bridges, "Infiniband scalability in Open MPI," Proceedings 20th IEEE International Parallel & Distributed Processing Symposium, Rhodes, Greece, p. 100, April 2006, doi: 10.1109/IPDPS.2006.1639335.
[10] B. R. de Supinski, T. R. W. Scogland, A. Duran, M. Klemm, S. M. Bellido, S. L. Olivier, C. Terboven and T. G. Mattson, "The Ongoing Evolution of OpenMP," Proceedings of the IEEE Volume: 106, Issue: 11, November 2018, pp. 2004 - 2019, August 2018, doi: 10.1109/JPROC.2018.2853600.
[11] R. Chandra, R. Menon, L. Dagum, D. Kohr, D. Maydan and J. McDonald, Parallel Programming in OpenMP, San Francisco, California : Morgan Kaufmann, 2001, October 2000.
[12] L. Dagum and R. Menon, "OpenMP: an industry standard API for shared-memory programming," IEEE Computational Science and Engineering ( Volume: 5, Issue: 1, Jan.-March 1998), pp. 46-55, Jan.-March 1998, doi: 10.1109/99.660313.
[13] M. Torabzadehkashi, S. Rezaei, V. Alves and N. Bagherzadeh, "CompStor: An In-storage Computation Platform for Scalable Distributed Processing," 2018 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), Vancouver, BC, Canada, pp. 1260-1267, May 2018, doi: 10.1109/IPDPSW.2018.00195.
[14] A. HeydariGorji, M. Torabzadehkashi, S. Rezaei, H. Bobarshad, V. Alves and P. H. Chou, "Stannis: Low-Power Acceleration of DNN Training Using Computational Storage Devices," 2020 57th ACM/IEEE Design Automation Conference (DAC), San Francisco, CA, USA, pp. 1-6, July 2020, doi: 10.1109/DAC18072.2020.9218687.
[15] M. Torabzadehkashi, S. Rezaei, A. Heydarigorji, H. Bobarshad, V. Alves and N. Bagherzadeh, "Catalina: In-Storage Processing Acceleration for Scalable Big Data Analytics," 2019 27th Euromicro International Conference on Parallel, Distributed and Network-Based Processing (PDP), Pavia, Italy, pp. 430-437, February 2019, doi: 10.1109/EMPDP.2019.8671589.
[16] A. HeydariGorji, M. Torabzadehkashi, S. Rezaei, H. Bobarshad, V. Alves and P. H. Chou, "In-storage Processing of I/O Intensive Applications on Computational Storage Drives," 2022 23rd International Symposium on Quality Electronic Design (ISQED), Santa Clara, CA, USA, pp. 1-6, April 2022, doi: 10.1109/ISQED54688.2022.9806270.
[17] Y. Zheng, J. Fixelle, N. Challapalle, P. Huo, Z. Shen, Z. Shao, M. Stan and V. Narayanan, "ISKEVA: in-SSD key-value database engine for video analytics applications," LCTES 2022: Proceedings of the 23rd ACM SIGPLAN/SIGBED International Conference on Languages, Compilers, and Tools for Embedded Systems, CA, San Diego, USA, pp. 50 - 60, June 2022, doi:10.1145/3519941.
[18] Y. Zheng, J. Fixelle, P. Huo, M. R. Stan, M. P. Mesnier and V. Narayanan, "ISVABI: In-Storage Video Analytics Engine with Block Interface," LCTES 2023: Proceedings of the 24th ACM SIGPLAN/SIGBED International Conference on Languages, Compilers, and Tools for Embedded Systems, FL, Orlando, USA, pp. 111 - 121, June 2023, doi:10.1145/3589610.
[19] V. S. Mailthody, Z. Qureshi, W. Liang, Z. Feng, S. G. de Gonzalo, Y. Li, H. Franke, J. Xiong, J. Huang and W.-m. Hwu, "DeepStore: In-Storage Acceleration for Intelligent Queries," MICRO '52: Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture, OH, Columbus, USA, pp. 224 - 238, October 2019, doi:10.1145/3352460.
[20] S. Kang, J. Kim, G. Lee, J. Lee, J. Seo, H. Jung, Y. H. Song and Y. Park, "ISP Agent: A Generalized In-storage-processing Workload Offloading Framework by Providing Multiple Optimization Opportunities," ACM Transactions on Architecture and Code Optimization, Volume 21, Issue 1, New York, NY, United States, pp. 1-24, 2024, doi:10.1145/3613496.
[21] M. Torabzadehkashi, A. Heydarigorji, S. Rezaei, H. Bobarshad, V. Alves and N. Bagherzadeh, "Accelerating HPC Applications Using Computational Storage Devices," 2019 IEEE 21st International Conference on High Performance Computing and Communications; IEEE 17th International Conference on Smart City; IEEE 5th International Conference on Data Science and Systems (HPCC/SmartCity/DSS), Zhangjiajie, China, pp. 1878-1885, August 2019, doi: 10.1109/HPCC/SmartCity/DSS.2019.00259.
[22] B. Antonio and D. Jaeyoung, "Computational Storage: Where Are We Today?," in Innovative Data Systems Research, Virtual Conference, 2021.
[23] Y. Kang, Y.-s. Kee, E. L. Miller and C. Park, "Enabling cost-effective data processing with smart SSD," 2013 IEEE 29th Symposium on Mass Storage Systems and Technologies (MSST), Long Beach, CA, USA, pp. 1-12, 2013, doi: 10.1109/MSST.2013.6558444.
[24] Y.-D. Su, Design of an In-storage Processing Architecture, M.S. thesis, Department of Engineering Science, National Cheng Kung University, Tainan, Taiwan, 2024.
[25] T.-Y. Lai, Design and Implementation of an Embedded System-Based In-Storage Processing Prototype, M.S. thesis, Department of Engineering Science, National Cheng Kung University, Tainan, Taiwan, 2024.
[26] C.-H. Wu, Design and implement a face recognition system combining embedded device and mobile phone application, M.S. thesis, Department of Engineering Science, National Cheng Kung University, Tainan, Taiwan, 2022.
[27] OpenCV, "haarcascade_frontalface_default.xml," [Online]. Available: https://github.com/opencv/opencv/blob/4.x/data/haarcascades/haarcascade_frontalface_default.xml. [Accessed 15 July 2024].
[28] Nuvoton, "MA35D1 Overview," [Online]. Available: https://www.nuvoton.com/products/microprocessors/arm-cortex-a35-mpus/ma35d1-high-performance-edge-iiot-series/. [Accessed 15 July 2024].
[29] Winbond, "W25N04KV Datasheet," [Online]. Available: https://www.winbond.com/hq/product/code-storage-flash-memory/qspinand-flash/?__locale=zh_TW&partNo=W25N04KV. [Accessed 15 July 2024].
[30] ADATA, "UV320," [Online]. Available: https://www.adata.com/hk/faq/506. [Accessed 15 July 2024].
[31] R. P. Weicker, "benchmark-dhrystone," Sifive, [Online]. Available: https://github.com/sifive/benchmark-dhrystone. [Accessed 15 July 2024].
[32] kaggle, "aisanfaces Dataset," [Online]. Available: https://www.kaggle.com/datasets/lukexng/aisanfaces?resource=download. [Accessed 15 July 2024].