簡易檢索 / 詳目顯示

研究生: 韓文斌
HON, WEN BIN
論文名稱: 結合IBM之多 GPU 跨節點CFD求解器開發
Development of a Multi-GPU CFD Solver Combining Immersed Boundary Methods
指導教授: 李崇綱
Li, Chung-Gang
學位類別: 碩士
Master
系所名稱: 工學院 - 機械工程學系
Department of Mechanical Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 89
中文關鍵詞: GPUMPIBCMLUSGSRunge-Kutta
外文關鍵詞: GPU, MPI, BCM, LUSGS, Runge-Kutta
相關次數: 點閱:42下載:1
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本研究發展一套基於 GPU 加速之 Building Cube Method(BCM)計算流體力學(Computational Fluid Dynamics, CFD)求解器,並以 Taylor-Green vortex 與圓柱繞流作為驗證案例,評估數值方法之正確性、能量耗散特性與平行計算效能。空間離散採用 Roe flux 計算數值通量,時間積分則分別實作顯式 Runge-Kutta 方法與隱式 LU-SGS(Lower-Upper Symmetric Gauss-Seidel)方法。
    在顯式方法方面,本研究以 16×16×16 cells 作為基本 cube 進行 GPU 平行化,並透過與文獻之動能衰減曲線比較,驗證程式於 Taylor-Green vortex 問題中的數值準確性與能量耗散行為。針對隱式 LU-SGS 方法,由於前後向掃描具有資料相依特性,限制 GPU 平行效率,因此進一步將計算域拆分為 4×4×4 之 small-block,以提升計算粒度與平行並行度。
    為降低 cube 細分後所增加之 halo communication 開銷,本研究設計專用資料陣列儲存所有 cube 資訊,使各 cube 在 halo region 計算時可直接存取陣列資料,而無需頻繁進行 GPU 或 MPI 通訊。此方法雖增加 GPU 記憶體使用量,但可有效降低通訊成本,提升 GPU 資源利用率與整體平行效率,使隱式 LU-SGS 求解器GPU、多節點環境下仍具良好擴充性。
    平行架構採用 MPI 進行跨節點運算,每節點配置 8 張 GPU,目前已擴展至 32 GPU。單一 GPU 約負責 71,991,296個網格,整體模擬規模可達 2,303,721,472 個網格。研究進一步進行 strong scaling 與 weak scaling 測試,並透過時間分解分析量化計算與通訊比例。結果顯示,本研究之 GPU BCM CFD 平台於 weak scaling 下可維持良好平行效率,而 strong scaling 效率下降則主要來自 halo exchange 通訊負擔增加,且於多階段顯式 Runge-Kutta 方法中特別明顯。整體而言,small-block LUSGS 策略可有效提升隱式求解器之平行效率與擴充能力,證明本研究所發展之 GPU BCM CFD 平台具備大規模高解析流場模擬之能力與可擴展性。

    This study develops a GPU-accelerated Building Cube Method (BCM) based Computational Fluid Dynamics solver and validates its numerical accuracy and parallel performance using Taylor–Green vortex and circular cylinder flow cases. The numerical scheme employs Roe flux for spatial discretization, while both explicit Runge–Kutta method and implicit LU-SGS methods are implemented for time integration.

    For the explicit solver, GPU parallelization is performed using 16×16×16-cell cubes, and numerical accuracy is verified through kinetic energy decay comparisons with reference data. For the implicit LU-SGS solver, a 4×4×4 small-block strategy is introduced to im-prove parallelism and computational efficiency by reducing data dependency during for-ward and backward sweeps.

    To minimize halo communication overhead, a dedicated data array is designed to store cube information, allowing direct data access during halo-region computations and reducing communication costs. This approach improves GPU resource utilization and overall parallel efficiency.

    The solver is implemented on a multi-GPU cluster using MPI, scaling up to 32 GPUs with a total problem size of 2,303,721,472 cells. Strong and weak scaling tests show good scalability, particularly for weak scaling. The proposed small-block LU-SGS strategy significantly enhances the scalability of the implicit solver, demonstrating the capability of the developed GPU BCM platform for large-scale, high-resolution flow simulations.

    摘要 I Abstract III 致謝 IX 目錄 X 表目錄 XIII 圖目錄 XIV 符號說明 XVI 1 第一章 緒論 1 1.1 研究背景與動機 1 1.2相關文獻回顧 2 2 第二章 數值方法 4 2.1 BCM (Building Cube Method) 4 2.2 沉浸邊界法 (Immersed Boundary Method) 6 2.3 全域統一解法 8 2.4 時間步數模組 10 2.5 預處理法Preconditioning 11 2.6 Roe scheme 14 2.7 黏性項離散 18 2.8 Runge-Kutta 3階顯示方法 20 2.9 LUSGS隱式方法 21 2.10 MUSCL 高階空間離散法 25 2.11 全域非反射性邊界 26 3 第三章 平行化架構 30 3.1 前處理 30 3.1.1 立方體鄰接關係與資料結構 30 3.1.2 負載平衡 31 3.1.3 GPU間通訊關係分類與通訊表建構 31 3.1.4 邊界拓樸分類 33 3.1.5 跨級網格處理 34 3.1.6 各 GPU 專屬通訊清單之建構與資料重組 35 3.1.7 沉浸邊界法處理 36 3.2 時間迭代與平行化計算 40 3.2.1 系統開發架構與異質計算策略 40 3.2.2 系統初始化與異質記憶體配置 40 3.2.3 四維邏輯坐標的反線性化重建 41 3.2.4 資料交換 42 3.2.5 平行化歸約(Parallel Reduction) 44 3.2.6 原位探針與數據輸出優化 45 3.2.7 Runge-Kutta平行化 46 3.2.8 LUSGS平行化 46 4 第四章 結果與討論 48 4.1 Taylor Green Vortex 48 4.2 Circular Cylinder Flow 55 4.3 計算效能與擴展性評估 63 4.3.1 效能測試環境與配置 63 4.3.2 強擴展性分析 (Strong Scaling Analysis) 63 4.3.3 弱擴展性分析 (Weak Scaling Analysis) 64 5 第五章 結論與未來展望 67 5.1 結論 67 5.2 未來展望 68 參考文獻 69

    [1] W.S. Fu, et al.,“ Roe scheme with preconditioning method for large eddy simulation of compressible turbulent channel flow.” INTERNATIONAL JOURNAL FOR NUMERICAL METHODS IN FLUIDS. 61(8): p. 888-910. 2009.
    [2] C.G. Li, et al.,“ Compressible direct numerical simulation with a hybrid boundary condition of transitional phenomena in natural convection.” INTERNATIONAL JOURNAL OF HEAT AND MASS TRANSFER. 90: p. 654-664. 2015.
    [3] C.G. Li, et al.,“ A sharp interface immersed boundary method for thin-walled geometries in viscous compressible flows.” INTERNATIONAL JOURNAL OF MECHANICAL SCIENCES. 253. 2023.
    [4] K. Komatsu, et al.,“ Parallel processing of the Building-Cube Method on a GPU platform.” COMPUTERS & FLUIDS. 45(1): p. 122-128. 2011.
    [5] S. Yoon and A. Jameson,“ Lower-upper Symmetric-Gauss-Seidel method for the Euler and Navier-Stokes equations.” AIAA journal. 26(9): p. 1025-1026. 1988.
    [6] G.V. Candler, M.J. Wright, and J.D. McDonald,“ Data-parallel lower-upper relaxation method for reacting flows.” AIAA journal. 32(12): p. 2380-2386. 1994.
    [7] J.M.Weiss, W.A.Smith, Preconditioning applied to variable and constant density flows, AIAA J. 33 (1995) 2050–2057.
    [8] Felix Rieper, “A low-Mach number fix for Roe’s approximate Riemann solver.” JOURNAL OF COMPUTATIONAL PHYSICS. 230(13): p. 5263-5287. 2011.
    [9] M.J. Wright, G.V. Candler, and M. Prampolini, “Data-Parallel Lower-Upper Relaxation Method for the Navier-Stokes Equations.” AIAA journal. 34(7): p. 1371-1377. 1996.
    [10] K. Ekici and A.S. Lyrintzis, “Parallelization of Rotorcraft Aerodynamics Navier–Stokes Codes” AIAA journal. 40(5): p. 887-896. 2002.
    [11] G. Giangaspero, E. van der Weide, M. Svärd, M.H. Carpenter, and K. Mattsson, “Case C3.3: Taylor–Green Vortex,” in Proceedings of the 1st International Workshop on High-Order CFD Methods, 2012.
    [12] Kuiju Xue, Qinling Li, and Liangyu Zhao, “Compressibility Effect on Flow Characteristics over a Circular Cylinder at Reynolds Number of 3900,” Physics of Fluids, Vol. 36, No. 8, Art. no. 085165, 2024.
    [13] J. B. Freund, “Proposed Inflow/Outflow Boundary Condition for Direct Computation of Aerodynamic Sound”, AIAA journal. 35(4): p. 740-742. 1997
    [14] X. S. Li, C. W. Gu, “An All-Speed Roe-type scheme and its asymptotic analysis of low Mach number behaviour”, JOURNAL OF COMPUTATIONAL PHYSICS. 230 (2008), p. 5144-5159.
    [15] K. Oßwald, A. Siegmund, P. Birken, V. Hannemann, A. Meister, “A low-dissipation version of Roe’s approximate Riemann solver for low Mach numbers”.
    [16] S. Venkateswaran and C. Merkle, "Dual time-stepping and preconditioning for unsteady computations," in 33rd Aerospace Sciences Meeting and Exhibit, 1995, p.78

    下載圖示
    校外:立即公開
    QR CODE