簡易檢索 / 詳目顯示

研究生: 王中辰
Wang, Chung-Chen
論文名稱: 基於神經網路模型壓縮與 FPGA 加速之行人偵測研究
The Study on Neural Network Compression and FPGA Acceleration for Pedestrian Detection
指導教授: 田思齊
Tien, Szu-Chi
學位類別: 碩士
Master
系所名稱: 工學院 - 機械工程學系
Department of Mechanical Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 125
中文關鍵詞: FPGA 加速 、行人偵測 、模型壓縮
外文關鍵詞: FPGA Acceleration, Pedestrian Detection, Model Compression
相關次數: 點閱:79  下載:0 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本研究透過神經網路模型壓縮與 FPGA 硬體加速設計,建立一套低功耗即時行人偵測系統。模型壓縮方面,使用 YOLOv3 Tiny 作為基準模型進行架構壓縮,並且將標準卷積替換為深度可分離卷積。隨後使用通道層級知識蒸餾技術提升壓縮後模型的特徵表達能力與準確度。接續採用結構化迭代剪枝移除逐點卷積中的冗餘濾波器,大幅減少推論計算負擔與硬體乘法器及記憶體資源消耗。最後透過訓練後量化將模型量化至 8 位元整數格式,進一步縮減硬體儲存需求。FPGA 硬體加速設計方面,使用管線串流與資源複用的複合式硬體架構,在有限的硬體資源與推論速度之間取得良好平衡。實驗結果顯示,YOLOv3 Tiny 經本研究之模型壓縮流程後,參數量與計算量分別縮減 99.00% 及 97.34%,mAP@50 僅下降 5.05%。結合本研究所提出之硬體架構,於約十萬個邏輯閘等級的 FPGA 開發板上,在 50 MHz的工作時脈下,僅需 13.18 毫秒的低延遲即可完成行人偵測。

    This study integrates neural network model compression and FPGA hardware acceleration to develop a low-power, real-time pedestrian detection system. YOLOv3 Tiny is adopted as the baseline model and compressed through architectural simplification, depthwise separable convolutions, channel-wise knowledge distillation, structured iterative pruning, and 8-bit post-training quantization. These techniques significantly reduce the model size and computational complexity while maintaining satisfactory detection performance. For hardware acceleration, a hybrid architecture combining pipeline streaming and resource reuse is designed to balance FPGA resource utilization and inference speed. Experimental results show that the proposed compression process reduces the number of parameters and computational complexity by 99.00% and 97.34%, respectively, while the mAP@50 decreases by only 5.05 percentage points. The proposed system achieves an inference latency of approximately 13.18 milliseconds per image on the Altera DE2-115 FPGA at an operating frequency of 50 MHz, demonstrating its feasibility for real-time embedded pedestrian detection.

    摘要 i 目錄 viii 圖目錄 x 表目錄 xiv 符號表 xv 第一章 緒論 1 第二章 模型壓縮方法與原理 6 2.1 深度可分離卷積 6 2.2 知識蒸餾 8 2.3 模型剪枝 12 2.4 模型量化 16 第三章 模型壓縮流程 21 3.1 基準模型 22 3.2 模型架構壓縮 25 3.3 知識蒸餾 28 3.4 模型剪枝 32 3.5 模型量化 37 第四章 FPGA 硬體加速設計 43 4.1 硬體加速架構 43 4.2 運算單元設計 47 第五章 系統整合與實現 60 5.1 系統硬體配置 60 5.2 FPGA 系統設計 66 第六章 實驗結果與討論 75 6.1 模型訓練環境與資料集 75 6.2 模型壓縮實驗結果 78 6.3 硬體設計驗證 86 第七章 結論與未來展望 97 7.1 結論 97 7.2 未來展望 98 參考文獻99

    [1] J P Gou et. al. Knowledge distillation: A survey. International journal of computer vision, 129(6):1789--1819, 2021.

    [2] J Redmon and A Farhadi. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767, 2018.

    [3] G Zheng et. al. Yolox: Exceeding yolo series in 2021. arXiv preprint arXiv:2107.08430, 2021.

    [4] C Y Shu et. al. Channel-wise knowledge distillation for dense prediction. In Proceedings of the IEEE/CVF international conference on computer vision, pages 5311--5320, 2021.

    [5] H Li et. al. Pruning filters for efficient convnets. arXiv preprint arXiv:1608.08710, 2016.

    [6] C H Tu et. al. Pruning depthwise separable convolutions for mobilenet compression. In 2020 international joint conference on neural networks (IJCNN)}, pages 1--8. IEEE, 2020.

    [7] N Dalal and B Triggs. Histograms of oriented gradients for human detection. In 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR'05), volume 1, pages 886--893. Ieee, 2005.

    [8] J Glenn et. al. Ultralytics yolo, 2023.

    [9] A Krizhevsky et. al. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 2012.

    [10] R Girshick et. al. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 580--587, 2014.

    [11] J Redmon et. al. You only look once: Unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition}, pages 779--788, 2016.

    [12] O Ronneberger et. al. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pages 234--241. Springer, 2015.

    [13] K M He et. al. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770--778, 2016.

    [14] V Sze et. al. Efficient processing of deep neural networks: A tutorial and survey. Proceedings of the IEEE}, 105(12):2295--2329, 2017.

    [15] C Y Chen et. al. Deepdriving: Learning affordance for direct perception in autonomous driving. In Proceedings of the IEEE international conference on computer vision, pages 2722--2730, 2015.

    [16] G Sreenu and M A Saleem Durai. Intelligent video surveillance: a review through deep learning techniques for crowd analysis. Journal of Big Data, 6(1):1--27, 2019.

    [17] M Saqib et. al. A study on detecting drones using deep convolutional neural networks. In 2017 14th IEEE international conference on advanced video and signal based surveillance (AVSS), pages 1--5. IEEE, 2017.

    [18] Y Cheng et. al. A survey of model compression and acceleration for deep neural networks. arXiv preprint arXiv:1710.09282, 2017.

    [19] Andrew G Howard et. al. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017.

    [20] G Hinton et. al. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015.

    [21] S Han et. al. Learning both weights and connections for efficient neural network. Advances in neural information processing systems, 28, 2015.

    [22] B Jacob et. al. Quantization and training of neural networks for efficient integer-arithmetic-only inference. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2704--2713, 2018.

    [23] L Deng et. al. Model compression and hardware acceleration for neural networks: A comprehensive survey. Proceedings of the IEEE, 108(4):485--532, 2020.

    [24] K Y Guo et. al. [dl] a survey of fpga-based neural network inference accelerators. ACM Transactions on Reconfigurable Technology and Systems (TRETS), 12(1):1--26, 2019.

    [25] C Zhang et. al. Optimizing fpga-based accelerator design for deep convolutional neural networks. In Proceedings of the 2015 ACM/SIGDA international symposium on field-programmable gate arrays, pages 161--170, 2015.

    [26] T Yi Lin et. al. Focal loss for dense object detection. In Proceedings of the IEEE international conference on computer vision, pages 2980--2988, 2017.

    [27] R Girshick. Fast r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 1440--1448, 2015.

    [28] S Q Ren et. al. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 28, 2015.

    [29] W Liu et. al. Ssd: Single shot multibox detector. In European conference on computer vision, pages 21--37. Springer, 2016.

    [30] M X Tan and Q V Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In International conference on machine learning, pages 6105--6114. PMLR, 2019.

    [31] F N. Iandola et. al. Squeezenet: Alexnet-level accuracy with 50x fewer parameters and< 0.5 mb model size. arXiv preprint arXiv:1602.07360, 2016.

    [32] S Zagoruyko and N Komodakis. Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer. arXiv preprint arXiv:1612.03928, 2016.

    [33] J Frankle and M Carbin. The lottery ticket hypothesis: Finding sparse, trainable neural networks. arXiv preprint arXiv:1803.03635, 2018.

    [34] M Nagel et. al. A white paper on neural network quantization. arXiv preprint arXiv:2106.08295, 2021.

    [35] L C Wang et. al. Design of fpga-based reconfigurable hardware acceleration system. In 2024 6th International Conference on Electronics and Communication, Network and Computer Technology (ECNCT), pages 561--566. IEEE, 2024.

    [36] A Romero et. al. Fitnets: Hints for thin deep nets. In International Conference on Learning Representations (ICLR), 2015.

    [37] W Parket. al. Relational knowledge distillation.In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3967--3976, 2019.

    [38] P Molchanov et. al. Pruning convolutional neural networks for resource efficient inference. arXiv preprint arXiv:1611.06440, 2016.

    [39] H Y Hu et. al. Network trimming: A data-driven neuron pruning approach towards efficient deep architectures. arXiv preprint arXiv:1607.03250, 2016.

    [40] A Paszke et. al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32, 2019.

    [41] M Abadi et. al. TensorFlow: a system for Large-Scale machine learning. In 12th USENIX symposium on operating systems design and implementation (OSDI 16), pages 265--283, 2016.

    [42] Y Bengio et. al. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432, 2013.

    [43] T Y Lin et. al. Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2117--2125, 2017.

    [44] S Tao et. al. A quantization-friendly separable convolution for mobilenets. arXiv preprint arXiv:1803.08607, 2018.

    [45] J H Cho and B Hariharan.On the efficacy of knowledge distillation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2019.

    [46] S H Hung. Fpga-based real-time moving target detection with a moving camera. Master's thesis, NCKU, 2021.

    [47] T Adiono et. al. Low latency yolov3-tiny accelerator for low-cost fpga using general matrix multiplication principle. IEEE access, 9:141890--141913, 2021.

    [48] M S Kimand et. al. A low-latency fpga accelerator for yolov3-tiny with flexible layerwise mapping and dataflow. IEEE Transactions on Circuits and Systems I: Regular Papers, 71(3):1158--1171, 2023.

    下載圖示
    校外:立即公開
    QR CODE