| 研究生: |
藍鴻毅 Lan, Hung-Yi |
|---|---|
| 論文名稱: |
結合串流資料傳輸與部分重組之 FPGA 推論加速器 FPGA-Based Inference Accelerator with Streaming Data Transfer and Partial Reconfiguration |
| 指導教授: |
侯廷偉
Hou, Ting-Wei |
| 學位類別: |
碩士 Master |
| 系所名稱: |
工學院 - 工程科學系 Department of Engineering Science |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 65 |
| 中文關鍵詞: | 現場可程式化邏輯閘陣列 、部分重組 、串流架構 |
| 外文關鍵詞: | Field-Programmable Gate Array, Partial Reconfiguration, Streaming Architecture |
| 相關次數: | 點閱:4 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
本研究提出一套結合串流資料傳輸與部分重組之 FPGA 推論加速架構,透過有限數量之硬體加速器串聯與可重組區域的重複使用,達成節省資源並提升推論效率之目的。部分重組可於執行期間動態載入對應的硬體加速模組,使資源受限之FPGA具有得以支援規模較大的神經網路模型之適應性,然而需將中間特徵資料寫回外部記憶體。另一方面,串流式架構雖能使相鄰運算模組直接傳遞資料,但須同時配置模型所需之硬體模組,容易受到有限資源限制。因此本研究將串流資料傳輸與部分重組機制加以整合,在資源受限環境下兼顧資料傳輸效率與架構適應能力。
本系統具備三項主要設計。首先,將多個硬體加速器組成串流運算群組,使同一群組內前一硬體模組之輸出可直接傳遞至下一硬體模組;其次,透過部分重組機制動態替換不同推論階段所需之硬體模組,並搭配管線化執行機制設計,藉此降低載入模組所產生的等待時間;最後,建立一套推論時間估計模型,依據各階段之配置時間、執行時間與相依關係,估計不同批次大小、加速器串聯數量及群組配置下之推論時間與資料吞吐量,作為架構配置與效能評估之參考。
本研究以 AMD Kria KV260 作為驗證平台,並採用MobileNetV1 模型測試。實驗結果顯示,相較於未串聯硬體加速器之部分重組配置,本研究提出的架構可降低 19.4% 與28.1% 的總推論時間。此外,推論時間估計模型與實際量測結果呈現相近的效能變化趨勢,各配置之平均相對誤差介於 1.0% 至 8.4% 之間。本研究所提出之架構能在有限 FPGA 資源下改善推論效率,並保留依據模型運算階段動態調整硬體功能之能力,證明結合串流資料傳輸與部分重組技術應用於 FPGA 推論加速器之可行性。
This study presents an FPGA-based neural-network inference accelerator that integrates streaming data transfer with partial reconfiguration to address the conflicting requirements of low data-movement overhead and limited programmable-logic resources. Accelerator modules are organized into streaming groups so that intermediate features pass directly between adjacent layers, while reconfigurable partitions are reused across inference stages to support networks whose total resource requirements exceeds the available FPGA capacity. An asynchronous pipeline overlaps configuration and execution when resource and data dependencies permit. A recursive timing model estimates inference throughput for different batch sizes and hardware configurations, characterized by the number of accelerators per group and the number of groups. The architecture is evaluated on an AMD Kria KV260 using an INT8-quantized MobileNetV1. Relative to a partial-reconfiguration baseline without accelerator chaining, the proposed reduce total inference latency by 19.4% to 28.1%, respectively. Predicted throughput follows the measured trends, with mean relative errors of 1.0 to 8.4%.
[1] D. Ngo, H.-C. Park, and B. Kang, “Edge Intelligence: A Review of Deep Neural Network Inference in Resource-Limited Environments,” Electronics, vol. 14, no. 12, p. 2495, Jan. 2025, doi: 10.3390/electronics14122495.
[2] C. Silvano, D. Ielmini, F. Ferrandi, et al., “A Survey on Deep Learning Hardware Accelerators for Heterogeneous HPC Platforms,” ACM Comput. Surv., vol. 57, no. 11, p. 286:1-286:39, Jun. 2025, doi: 10.1145/3729215.
[3] L. Du, Y. Du, Y. Li, et al., “A Reconfigurable Streaming Deep Convolutional Neural Network Accelerator for Internet of Things,” IEEE Trans. Circuits Syst. Regul. Pap., vol. 65, no. 1, pp. 198–208, Jan. 2018, doi: 10.1109/TCSI.2017.2735490.
[4] S. Alam, C. Yakopcic, Q. Wu, M. Barnell, S. Khan, and T. M. Taha, “Survey of Deep Learning Accelerators for Edge and Emerging Computing,” Electronics, vol. 13, no. 15, p. 2988, Jan. 2024, doi: 10.3390/electronics13152988.
[5] A. Boutros, A. Arora, and V. Betz, “Field-Programmable Gate Array Architecture for Deep Learning: Survey and Future Directions,” Proc. IEEE, vol. 113, no. 7, pp. 613–639, Jul. 2025, doi: 10.1109/JPROC.2025.3623023.
[6] L. Liu, J. Zhu, Z. Li, et al., “A Survey of Coarse-Grained Reconfigurable Architecture and Design: Taxonomy, Challenges, and Applications,” ACM Comput. Surv., vol. 52, no. 6, pp. 1–39, Nov. 2020, doi: 10.1145/3357375.
[7] F. Yan, A. Koch, and O. Sinnen, “A survey on FPGA-based accelerator for ML models,” Dec. 20, 2024, arXiv: arXiv:2412.15666. doi: 10.48550/arXiv.2412.15666.
[8] S. Lahti and T. D. Hämäläinen, “High-Level Synthesis for FPGAs—A Hardware Engineer’s Perspective,” IEEE Access, vol. 13, pp. 28574–28593, 2025, doi: 10.1109/ACCESS.2025.3540320.
[9] Y. Bai, A. Sohrabizadeh, Z. Qin, Z. Hu, Y. Sun, and J. Cong, “Towards a Comprehensive Benchmark for High-Level Synthesis Targeted to FPGAs,” Adv. Neural Inf. Process. Syst., vol. 36, pp. 45288–45299, 2023, doi: 10.52202/075280-1962.
[10] F. Hamanaka, T. Odan, K. Kise, and T. V. Chu, “An Exploration of State-of-the-Art Automation Frameworks for FPGA-Based DNN Acceleration,” IEEE Access, vol. 11, pp. 5701–5713, 2023, doi: 10.1109/ACCESS.2023.3236974.
[11] Y. Umuroglu, N. J. Fraser, G. Gambardella, et al., “FINN: A Framework for Fast, Scalable Binarized Neural Network Inference,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays, Feb. 2017, pp. 65–74. doi: 10.1145/3020078.3021744.
[12] Xilinx/Vitis-AI. (Jul. 05, 2026). Accessed: Jul. 05, 2026. [Online]. Available: https://github.com/Xilinx/Vitis-AI
[13] J. Duarte, S. Han, P. Harris, et al., “Fast Inference of Deep Neural Networks in FPGAs for Particle Physics,” J. Instrum., vol. 13, no. 07, pp. P07027–P07027, Jul. 2018, doi: 10.1088/1748-0221/13/07/P07027.
[14] F. Fahim, B. Hawks, C. Herwig, et al., “hls4ml: An Open-Source Codesign Workflow to Empower Scientific Low-Power Machine Learning Devices,” in Proceedings of TinyML Research Symposium, ACM, Mar. 2021, pp. 1–10. doi: 10.48550/arXiv.2103.05579.
[15] C. Baskin, N. Liss, E. Zheltonozhskii, A. M. Bronstein, and A. Mendelson, “Streaming Architecture for Large-Scale Quantized Neural Networks on an FPGA-Based Dataflow Platform,” in 2018 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), May 2018, pp. 162–169. doi: 10.1109/IPDPSW.2018.00032.
[16] C. Shao-Yi, T. Xin-Yi, W. Jian, H. Jing-Si, and L. Zheng, “An Ultra-efficient Streaming-based FPGA Accelerator for Infrared Target Detection,” J. Infrared Millim. Waves, vol. 41, no. 5, p. 914, 2022, doi: 10.11972/j.issn.1001-9014.2022.05.016.
[17] R. Hou, J. Zhai, Y. Wang, Z. Lin, and K. Zhao, “Array Partitioning Method for Streaming Dataflow Optimization in High-level Synthesis,” in 2024 2nd International Symposium of Electronics Design Automation (ISEDA), May 2024, pp. 278–282. doi: 10.1109/ISEDA62518.2024.10618042.
[18] A. Maclellan, L. H. Crockett, and R. W. Stewart, “RFSoC Modulation Classification With Streaming CNN: Data Set Generation & Quantized-Aware Training,” IEEE Open J. Circuits Syst., vol. 6, pp. 38–49, 2025, doi: 10.1109/OJCAS.2024.3509627.
[19] N. K. Shaydyuk and E. B. John, “FPGA Implementation of MobileNetV2 CNN Model Using Semi-Streaming Architecture for Low Power Inference Applications,” in 2020 IEEE Intl Conf on Parallel & Distributed Processing with Applications, Big Data & Cloud Computing, Sustainable Computing & Communications, Social Computing & Networking (ISPA/BDCloud/SocialCom/SustainCom), Feb. 2020, pp. 160–167. doi: 10.1109/ISPA-BDCloud-SocialCom-SustainCom51426.2020.00046.
[20] Q. Tang, B. Guo, and Z. Wang, “Sw/Hw Partitioning and Scheduling on Region-Based Dynamic Partial Reconfigurable System-on-Chip,” Electronics, vol. 9, no. 9, p. 1362, Aug. 2020, doi: 10.3390/electronics9091362.
[21] M. Nguyen and J. C. Hoe, “Time-Shared Execution of Realtime Computer Vision Pipelines by Dynamic Partial Reconfiguration,” in 2018 28th International Conference on Field Programmable Logic and Applications (FPL), Aug. 2018, pp. 230–2304. doi: 10.1109/FPL.2018.00046.
[22] C.-H. Huang, S.-W. Tang, and P.-A. Hsiung, “ACNNE: An Adaptive Convolution Engine for CNNs Acceleration Exploiting Partial Reconfiguration on FPGAs,” in 2024 IEEE International Symposium on Circuits and Systems (ISCAS), IEEE, May 2024, pp. 1–5. doi: 10.1109/ISCAS58744.2024.10558457.
[23] J. Boudjadar, S. U. Islam, and R. Buyya, “Dynamic FPGA Reconfiguration for Scalable Embedded Artificial Intelligence (AI): A Co-design Methodology for Convolutional Neural Networks (CNN) Acceleration,” Future Gener. Comput. Syst., vol. 169, p. 107777, Aug. 2025, doi: 10.1016/j.future.2025.107777.
[24] L. Deng, “The MNIST Database of Handwritten Digit Images for Machine Learning Research [Best of the Web],” IEEE Signal Process. Mag., vol. 29, no. 6, pp. 141–142, Jan. 2012, doi: 10.1109/MSP.2012.2211477.
[25] A. Lengyel, S. Garg, M. Milford, and J. C. van Gemert, “Zero-Shot Day-Night Domain Adaptation with a Physics Prior,” Oct. 11, 2021, arXiv: arXiv:2108.05137. doi: 10.48550/arXiv.2108.05137.
[26] S. R. Bommana, S. Veeramachaneni, S. Ershad, and M. B. Srinivas, “Mitigating Side Channel Attacks on FPGA Through Deep Learning and Dynamic Partial Reconfiguration,” Sci. Rep., vol. 15, no. 1, p. 13745, Apr. 2025, doi: 10.1038/s41598-025-98473-3.
[27] M. Nguyen, N. Serafin, and J. C. Hoe, “Partial Reconfiguration for Design Optimization,” in 2020 30th International Conference on Field-Programmable Logic and Applications (FPL), Aug. 2020, pp. 328–334. doi: 10.1109/FPL50879.2020.00061.
[28] Advanced Micro Devices, Inc., “Vitis High-Level Synthesis User Guide (UG1399).” Accessed: Jul. 05, 2026. [Online]. Available: https://docs.amd.com/r/en-US/ug1399-vitis-hls
[29] Advanced Micro Devices, Inc., “Vivado Design Suite Tutorial: Design Flows Overview (UG888).” Accessed: Jul. 05, 2026. [Online]. Available: https://docs.amd.com/r/2024.2-English/ug888-vivado-design-flows-overview-tutorial
[30] Advanced Micro Devices, Inc., “AXI DMA LogiCORE IP Product Guide (PG021).” Accessed: Jul. 05, 2026. [Online]. Available: https://docs.amd.com/r/en-US/pg021_axi_dma
[31] Advanced Micro Devices, Inc., “AXI4-Stream Infrastructure IP Suite LogiCORE IP Product Guide (PG085).” Accessed: Jul. 05, 2026. [Online]. Available: https://docs.amd.com/r/en-US/pg085-axi4stream-infrastructure
[32] Advanced Micro Devices, Inc., “Vivado Design Suite User Guide: Dynamic Function eXchange (UG909).” Accessed: Jul. 05, 2026. [Online]. Available: https://docs.amd.com/r/en-US/ug909-vivado-partial-reconfiguration/Introduction
[33] Advanced Micro Devices, Inc., “Dynamic Function eXchange Decoupler LogiCORE IP Product Guide (PG375).” Accessed: Jul. 05, 2026. [Online]. Available: https://docs.amd.com/r/en-US/pg375-dfx-decoupler
[34] Advanced Micro Devices, Inc., “AXI GPIO LogiCORE IP Product Guide (PG144).” Accessed: Jul. 05, 2026. [Online]. Available: https://docs.amd.com/r/en-US/pg144-axi-gpio
[35] Advanced Micro Devices, Inc., “PYNQ.” Accessed: Jul. 05, 2026. [Online]. Available: http://www.pynq.io/
[36] A. G. Howard, M. Zhu, B. Chen, et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications,” Apr. 17, 2017, arXiv: arXiv:1704.04861. doi: 10.48550/arXiv.1704.04861.
[37] O. Russakovsky, J. Deng, H. Su, et al., “ImageNet Large Scale Visual Recognition Challenge,” Int. J. Comput. Vis., vol. 115, no. 3, pp. 211–252, 2015, doi: 10.1007/s11263-015-0816-y.
[38] H. Hong, D. Choi, N. Kim, and H. Kim, “Mobile-X: Dedicated FPGA Implementation of the MobileNet Accelerator Optimizing Depthwise Separable Convolution,” IEEE Trans. Circuits Syst. II Express Briefs, vol. 71, no. 11, pp. 4668–4672, Jan. 2024, doi: 10.1109/TCSII.2024.3440884.
[39] J. Liao, L. Cai, Y. Xu, and M. He, “Design of Accelerator for MobileNet Convolutional Neural Network Based on FPGA,” in 2019 IEEE 4th Advanced Information Technology, Electronic and Automation Control Conference (IAEAC), Feb. 2019, pp. 1392–1396. doi: 10.1109/IAEAC47372.2019.8997842.
[40] W. Liu, Y. Li, Y. Yang, J. Zhu, and L. Liu, “Design an Efficient DNN Inference Framework with PS-PL Synergies in FPGA for Edge Computing,” in 2022 China Automation Congress (CAC), Jan. 2022, pp. 4186–4190. doi: 10.1109/CAC57257.2022.10055526.