簡易檢索 / 詳目顯示

研究生: 藍鴻毅
Lan, Hung-Yi
論文名稱: 結合串流資料傳輸與部分重組之 FPGA 推論加速器
FPGA-Based Inference Accelerator with Streaming Data Transfer and Partial Reconfiguration
指導教授: 侯廷偉
Hou, Ting-Wei
學位類別: 碩士
Master
系所名稱: 工學院 - 工程科學系
Department of Engineering Science
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 65
中文關鍵詞: 現場可程式化邏輯閘陣列部分重組串流架構
外文關鍵詞: Field-Programmable Gate Array, Partial Reconfiguration, Streaming Architecture
相關次數: 點閱:4下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本研究提出一套結合串流資料傳輸與部分重組之 FPGA 推論加速架構,透過有限數量之硬體加速器串聯與可重組區域的重複使用,達成節省資源並提升推論效率之目的。部分重組可於執行期間動態載入對應的硬體加速模組,使資源受限之FPGA具有得以支援規模較大的神經網路模型之適應性,然而需將中間特徵資料寫回外部記憶體。另一方面,串流式架構雖能使相鄰運算模組直接傳遞資料,但須同時配置模型所需之硬體模組,容易受到有限資源限制。因此本研究將串流資料傳輸與部分重組機制加以整合,在資源受限環境下兼顧資料傳輸效率與架構適應能力。
    本系統具備三項主要設計。首先,將多個硬體加速器組成串流運算群組,使同一群組內前一硬體模組之輸出可直接傳遞至下一硬體模組;其次,透過部分重組機制動態替換不同推論階段所需之硬體模組,並搭配管線化執行機制設計,藉此降低載入模組所產生的等待時間;最後,建立一套推論時間估計模型,依據各階段之配置時間、執行時間與相依關係,估計不同批次大小、加速器串聯數量及群組配置下之推論時間與資料吞吐量,作為架構配置與效能評估之參考。
    本研究以 AMD Kria KV260 作為驗證平台,並採用MobileNetV1 模型測試。實驗結果顯示,相較於未串聯硬體加速器之部分重組配置,本研究提出的架構可降低 19.4% 與28.1% 的總推論時間。此外,推論時間估計模型與實際量測結果呈現相近的效能變化趨勢,各配置之平均相對誤差介於 1.0% 至 8.4% 之間。本研究所提出之架構能在有限 FPGA 資源下改善推論效率,並保留依據模型運算階段動態調整硬體功能之能力,證明結合串流資料傳輸與部分重組技術應用於 FPGA 推論加速器之可行性。

    This study presents an FPGA-based neural-network inference accelerator that integrates streaming data transfer with partial reconfiguration to address the conflicting requirements of low data-movement overhead and limited programmable-logic resources. Accelerator modules are organized into streaming groups so that intermediate features pass directly between adjacent layers, while reconfigurable partitions are reused across inference stages to support networks whose total resource requirements exceeds the available FPGA capacity. An asynchronous pipeline overlaps configuration and execution when resource and data dependencies permit. A recursive timing model estimates inference throughput for different batch sizes and hardware configurations, characterized by the number of accelerators per group and the number of groups. The architecture is evaluated on an AMD Kria KV260 using an INT8-quantized MobileNetV1. Relative to a partial-reconfiguration baseline without accelerator chaining, the proposed reduce total inference latency by 19.4% to 28.1%, respectively. Predicted throughput follows the measured trends, with mean relative errors of 1.0 to 8.4%.

    摘要 I Extended Abstract II 致謝 X 目錄 XI 表目錄 XIII 圖目錄 XIV 第一章 緒論 1 1.1 研究動機 1 1.2 研究目的 1 1.3 研究貢獻 2 1.4 研究架構 2 第二章 文獻探討 3 2.1 邊緣運算平台與架構 3 2.2 FPGA現場可程式化邏輯閘陣列 4 2.2.1 High-Level Synthesis高階合成 5 2.2.2 自動化框架 6 2.3 串流式架構 6 2.4 部分重組 8 2.5 綜合討論 9 第三章 系統設計與實作 12 3.1 串流傳輸與部分重組系統架構 12 3.2 硬體架構設計 13 3.2.1 硬體加速模組 15 3.2.2 硬體系統建置 17 3.3 軟體設計 21 3.3.1 軟體控制流程 21 3.3.2 管線化執行機制設計 22 3.4 推論時間估計 25 第四章 實驗結果與討論 29 4.1 實驗規格與環境 29 4.2 串流式架構與部分重組之效能評估 31 4.3 與時間估計模型之比較 34 4.4 跨平台效能比較 37 4.5 問題與討論 39 第五章 結論與未來展望 42 5.1 結論 42 5.2 未來展望 42 參考文獻 44

    [1] D. Ngo, H.-C. Park, and B. Kang, “Edge Intelligence: A Review of Deep Neural Network Inference in Resource-Limited Environments,” Electronics, vol. 14, no. 12, p. 2495, Jan. 2025, doi: 10.3390/electronics14122495.
    [2] C. Silvano, D. Ielmini, F. Ferrandi, et al., “A Survey on Deep Learning Hardware Accelerators for Heterogeneous HPC Platforms,” ACM Comput. Surv., vol. 57, no. 11, p. 286:1-286:39, Jun. 2025, doi: 10.1145/3729215.
    [3] L. Du, Y. Du, Y. Li, et al., “A Reconfigurable Streaming Deep Convolutional Neural Network Accelerator for Internet of Things,” IEEE Trans. Circuits Syst. Regul. Pap., vol. 65, no. 1, pp. 198–208, Jan. 2018, doi: 10.1109/TCSI.2017.2735490.
    [4] S. Alam, C. Yakopcic, Q. Wu, M. Barnell, S. Khan, and T. M. Taha, “Survey of Deep Learning Accelerators for Edge and Emerging Computing,” Electronics, vol. 13, no. 15, p. 2988, Jan. 2024, doi: 10.3390/electronics13152988.
    [5] A. Boutros, A. Arora, and V. Betz, “Field-Programmable Gate Array Architecture for Deep Learning: Survey and Future Directions,” Proc. IEEE, vol. 113, no. 7, pp. 613–639, Jul. 2025, doi: 10.1109/JPROC.2025.3623023.
    [6] L. Liu, J. Zhu, Z. Li, et al., “A Survey of Coarse-Grained Reconfigurable Architecture and Design: Taxonomy, Challenges, and Applications,” ACM Comput. Surv., vol. 52, no. 6, pp. 1–39, Nov. 2020, doi: 10.1145/3357375.
    [7] F. Yan, A. Koch, and O. Sinnen, “A survey on FPGA-based accelerator for ML models,” Dec. 20, 2024, arXiv: arXiv:2412.15666. doi: 10.48550/arXiv.2412.15666.
    [8] S. Lahti and T. D. Hämäläinen, “High-Level Synthesis for FPGAs—A Hardware Engineer’s Perspective,” IEEE Access, vol. 13, pp. 28574–28593, 2025, doi: 10.1109/ACCESS.2025.3540320.
    [9] Y. Bai, A. Sohrabizadeh, Z. Qin, Z. Hu, Y. Sun, and J. Cong, “Towards a Comprehensive Benchmark for High-Level Synthesis Targeted to FPGAs,” Adv. Neural Inf. Process. Syst., vol. 36, pp. 45288–45299, 2023, doi: 10.52202/075280-1962.
    [10] F. Hamanaka, T. Odan, K. Kise, and T. V. Chu, “An Exploration of State-of-the-Art Automation Frameworks for FPGA-Based DNN Acceleration,” IEEE Access, vol. 11, pp. 5701–5713, 2023, doi: 10.1109/ACCESS.2023.3236974.
    [11] Y. Umuroglu, N. J. Fraser, G. Gambardella, et al., “FINN: A Framework for Fast, Scalable Binarized Neural Network Inference,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays, Feb. 2017, pp. 65–74. doi: 10.1145/3020078.3021744.
    [12] Xilinx/Vitis-AI. (Jul. 05, 2026). Accessed: Jul. 05, 2026. [Online]. Available: https://github.com/Xilinx/Vitis-AI
    [13] J. Duarte, S. Han, P. Harris, et al., “Fast Inference of Deep Neural Networks in FPGAs for Particle Physics,” J. Instrum., vol. 13, no. 07, pp. P07027–P07027, Jul. 2018, doi: 10.1088/1748-0221/13/07/P07027.
    [14] F. Fahim, B. Hawks, C. Herwig, et al., “hls4ml: An Open-Source Codesign Workflow to Empower Scientific Low-Power Machine Learning Devices,” in Proceedings of TinyML Research Symposium, ACM, Mar. 2021, pp. 1–10. doi: 10.48550/arXiv.2103.05579.
    [15] C. Baskin, N. Liss, E. Zheltonozhskii, A. M. Bronstein, and A. Mendelson, “Streaming Architecture for Large-Scale Quantized Neural Networks on an FPGA-Based Dataflow Platform,” in 2018 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), May 2018, pp. 162–169. doi: 10.1109/IPDPSW.2018.00032.
    [16] C. Shao-Yi, T. Xin-Yi, W. Jian, H. Jing-Si, and L. Zheng, “An Ultra-efficient Streaming-based FPGA Accelerator for Infrared Target Detection,” J. Infrared Millim. Waves, vol. 41, no. 5, p. 914, 2022, doi: 10.11972/j.issn.1001-9014.2022.05.016.
    [17] R. Hou, J. Zhai, Y. Wang, Z. Lin, and K. Zhao, “Array Partitioning Method for Streaming Dataflow Optimization in High-level Synthesis,” in 2024 2nd International Symposium of Electronics Design Automation (ISEDA), May 2024, pp. 278–282. doi: 10.1109/ISEDA62518.2024.10618042.
    [18] A. Maclellan, L. H. Crockett, and R. W. Stewart, “RFSoC Modulation Classification With Streaming CNN: Data Set Generation & Quantized-Aware Training,” IEEE Open J. Circuits Syst., vol. 6, pp. 38–49, 2025, doi: 10.1109/OJCAS.2024.3509627.
    [19] N. K. Shaydyuk and E. B. John, “FPGA Implementation of MobileNetV2 CNN Model Using Semi-Streaming Architecture for Low Power Inference Applications,” in 2020 IEEE Intl Conf on Parallel & Distributed Processing with Applications, Big Data & Cloud Computing, Sustainable Computing & Communications, Social Computing & Networking (ISPA/BDCloud/SocialCom/SustainCom), Feb. 2020, pp. 160–167. doi: 10.1109/ISPA-BDCloud-SocialCom-SustainCom51426.2020.00046.
    [20] Q. Tang, B. Guo, and Z. Wang, “Sw/Hw Partitioning and Scheduling on Region-Based Dynamic Partial Reconfigurable System-on-Chip,” Electronics, vol. 9, no. 9, p. 1362, Aug. 2020, doi: 10.3390/electronics9091362.
    [21] M. Nguyen and J. C. Hoe, “Time-Shared Execution of Realtime Computer Vision Pipelines by Dynamic Partial Reconfiguration,” in 2018 28th International Conference on Field Programmable Logic and Applications (FPL), Aug. 2018, pp. 230–2304. doi: 10.1109/FPL.2018.00046.
    [22] C.-H. Huang, S.-W. Tang, and P.-A. Hsiung, “ACNNE: An Adaptive Convolution Engine for CNNs Acceleration Exploiting Partial Reconfiguration on FPGAs,” in 2024 IEEE International Symposium on Circuits and Systems (ISCAS), IEEE, May 2024, pp. 1–5. doi: 10.1109/ISCAS58744.2024.10558457.
    [23] J. Boudjadar, S. U. Islam, and R. Buyya, “Dynamic FPGA Reconfiguration for Scalable Embedded Artificial Intelligence (AI): A Co-design Methodology for Convolutional Neural Networks (CNN) Acceleration,” Future Gener. Comput. Syst., vol. 169, p. 107777, Aug. 2025, doi: 10.1016/j.future.2025.107777.
    [24] L. Deng, “The MNIST Database of Handwritten Digit Images for Machine Learning Research [Best of the Web],” IEEE Signal Process. Mag., vol. 29, no. 6, pp. 141–142, Jan. 2012, doi: 10.1109/MSP.2012.2211477.
    [25] A. Lengyel, S. Garg, M. Milford, and J. C. van Gemert, “Zero-Shot Day-Night Domain Adaptation with a Physics Prior,” Oct. 11, 2021, arXiv: arXiv:2108.05137. doi: 10.48550/arXiv.2108.05137.
    [26] S. R. Bommana, S. Veeramachaneni, S. Ershad, and M. B. Srinivas, “Mitigating Side Channel Attacks on FPGA Through Deep Learning and Dynamic Partial Reconfiguration,” Sci. Rep., vol. 15, no. 1, p. 13745, Apr. 2025, doi: 10.1038/s41598-025-98473-3.
    [27] M. Nguyen, N. Serafin, and J. C. Hoe, “Partial Reconfiguration for Design Optimization,” in 2020 30th International Conference on Field-Programmable Logic and Applications (FPL), Aug. 2020, pp. 328–334. doi: 10.1109/FPL50879.2020.00061.
    [28] Advanced Micro Devices, Inc., “Vitis High-Level Synthesis User Guide (UG1399).” Accessed: Jul. 05, 2026. [Online]. Available: https://docs.amd.com/r/en-US/ug1399-vitis-hls
    [29] Advanced Micro Devices, Inc., “Vivado Design Suite Tutorial: Design Flows Overview (UG888).” Accessed: Jul. 05, 2026. [Online]. Available: https://docs.amd.com/r/2024.2-English/ug888-vivado-design-flows-overview-tutorial
    [30] Advanced Micro Devices, Inc., “AXI DMA LogiCORE IP Product Guide (PG021).” Accessed: Jul. 05, 2026. [Online]. Available: https://docs.amd.com/r/en-US/pg021_axi_dma
    [31] Advanced Micro Devices, Inc., “AXI4-Stream Infrastructure IP Suite LogiCORE IP Product Guide (PG085).” Accessed: Jul. 05, 2026. [Online]. Available: https://docs.amd.com/r/en-US/pg085-axi4stream-infrastructure
    [32] Advanced Micro Devices, Inc., “Vivado Design Suite User Guide: Dynamic Function eXchange (UG909).” Accessed: Jul. 05, 2026. [Online]. Available: https://docs.amd.com/r/en-US/ug909-vivado-partial-reconfiguration/Introduction
    [33] Advanced Micro Devices, Inc., “Dynamic Function eXchange Decoupler LogiCORE IP Product Guide (PG375).” Accessed: Jul. 05, 2026. [Online]. Available: https://docs.amd.com/r/en-US/pg375-dfx-decoupler
    [34] Advanced Micro Devices, Inc., “AXI GPIO LogiCORE IP Product Guide (PG144).” Accessed: Jul. 05, 2026. [Online]. Available: https://docs.amd.com/r/en-US/pg144-axi-gpio
    [35] Advanced Micro Devices, Inc., “PYNQ.” Accessed: Jul. 05, 2026. [Online]. Available: http://www.pynq.io/
    [36] A. G. Howard, M. Zhu, B. Chen, et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications,” Apr. 17, 2017, arXiv: arXiv:1704.04861. doi: 10.48550/arXiv.1704.04861.
    [37] O. Russakovsky, J. Deng, H. Su, et al., “ImageNet Large Scale Visual Recognition Challenge,” Int. J. Comput. Vis., vol. 115, no. 3, pp. 211–252, 2015, doi: 10.1007/s11263-015-0816-y.
    [38] H. Hong, D. Choi, N. Kim, and H. Kim, “Mobile-X: Dedicated FPGA Implementation of the MobileNet Accelerator Optimizing Depthwise Separable Convolution,” IEEE Trans. Circuits Syst. II Express Briefs, vol. 71, no. 11, pp. 4668–4672, Jan. 2024, doi: 10.1109/TCSII.2024.3440884.
    [39] J. Liao, L. Cai, Y. Xu, and M. He, “Design of Accelerator for MobileNet Convolutional Neural Network Based on FPGA,” in 2019 IEEE 4th Advanced Information Technology, Electronic and Automation Control Conference (IAEAC), Feb. 2019, pp. 1392–1396. doi: 10.1109/IAEAC47372.2019.8997842.
    [40] W. Liu, Y. Li, Y. Yang, J. Zhu, and L. Liu, “Design an Efficient DNN Inference Framework with PS-PL Synergies in FPGA for Edge Computing,” in 2022 China Automation Congress (CAC), Jan. 2022, pp. 4186–4190. doi: 10.1109/CAC57257.2022.10055526.

    QR CODE