簡易檢索 / 詳目顯示

研究生: 邱靖云
Chiu, Ching-Yun
論文名稱: 一個輕量且有效之動態頻域神經網路用於可見光紅外線行人重識別
A Lightweight and Effective Dynamic Frequency Domain Neural Network for Visible-Infrared Person Re-Identification
指導教授: 戴顯權
Tai, Shen-Chuan
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 電機工程學系
Department of Electrical Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 71
中文關鍵詞: 可見光紅外線行人重識別注意力機制輕量化神經網路
外文關鍵詞: Visible-Infrared Person Re-Identification, Attention Mechanism, Lightweight Neural Network
相關次數: 點閱:2下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 可見光紅外線行人重識別(Visible-Infrared Person Re-Identification, VI-ReID)旨在判斷不同攝影機拍攝之可見光與紅外線行人影像是否屬於同一身分,是全天候智慧監控的重要技術。然而,由於可見光與紅外線影像在光照、色彩及成像特性上的差異,以及視角、姿態與遮擋造成的外觀變化,使跨模態特徵學習與身分匹配仍具挑戰性。此外現有頻域方法多採用固定的頻率分割方式,缺乏根據影像內容進行自適應調整的能力,且辨識效能與模型複雜度之間仍存在權衡。
    為改善上述問題,本論文提出一種輕量且有效的動態頻域神經網路。其中,自適應高低頻調制模組透過內容感知的動態頻率分割與低頻特徵調制,提升跨模態特徵學習能力;其高頻路徑則結合可變形多尺度卷積,自適應提取高頻特徵中的局部結構資訊。同時重新設計輕量化最終殘差階段,以兼顧辨識效能與模型效率。實驗結果顯示,所提出的方法在SYSU-MM01與LLCM兩個公開VI-ReID資料集上皆展現具競爭力的辨識效能。整體而言,本方法在辨識準確度、模型參數量與計算成本之間取得良好的平衡,具備應用於智慧監控系統之潛力。

    Visible-Infrared Person Re-Identification (VI-ReID) aims to determine whether visible and infrared pedestrian images captured by different cameras belong to the same identity and is an important technology for round-the-clock intelligent surveillance. However, differences between visible and infrared images in illumination, color, and imaging characteristics, together with appearance variations caused by viewpoint, pose, and occlusion, make cross modality feature learning and identity matching challenging. Furthermore, most existing frequency-domain methods employ fixed frequency partitioning schemes and lack the ability to adapt the partitioning to image content, while a trade-off remains between recognition performance and model complexity.
    To address these issues, this Thesis proposes a lightweight and effective dynamic frequency-domain neural network. The proposed Adaptive HiLo-Frequency Modulation (Adaptive HiLo-FM) module improves cross-modality feature learning through content aware dynamic frequency partitioning and low-frequency feature modulation. Within its high-frequency path, deformable multi-scale convolutions adaptively extract local structural information from high-frequency features. The lightweight final residual stage is also redesigned to balance recognition performance and model efficiency. Experimental results demonstrate that the proposed method achieves competitive recognition performance on two public VI-ReID datasets, SYSU-MM01 and LLCM. Overall, the proposed method strikes a favorable balance among recognition accuracy, model parameter count, and computational cost, demonstrating its potential for application in intelligent surveillance systems.

    中文摘要 i Abstract ii Acknowledgements iv Contents v List of Tables viii List of Figures ix Chapter 1 Introduction 1 Chapter 2 Related Works 5 2.1 Visible-Infrared Person Re-Identification 5 2.1.1 Image-Level Methods 6 2.1.2 Feature-Level Methods 7 2.2 Feature Extraction and Enhancement Methods 8 2.2.1 Residual Network 8 2.2.2 Omni-Scale Feature Learning 9 2.2.3 Squeeze-and-Excitation 10 2.2.4 Coordinate Attention 11 2.2.5 Deformable Convolution 12 2.3 Frequency Domain Analysis in Deep Learning 13 2.4 MFENet 14 Chapter 3 The Proposed Method 18 3.1 Overall System Architecture 18 3.2 Adaptive HiLo-Frequency Modulation 21 3.2.1 Motivation for Adaptive HiLo-Frequency Modulation 21 3.2.2 Adaptive Gate Soft Mask Generation 22 3.2.3 Spatial-Driven Dynamic Filter 23 3.3 Deformable Multiscale Convolutions 25 3.3.1 Bottleneck of Multi-Scale Feature Extraction 26 3.3.2 Introducing Large-Kernel Convolution 26 3.3.3 Deformable Convolution for Geometric Variation Modeling 27 3.4 Lightweight Reconstruction of the Final Residual Stage 29 3.4.1 Motivation and Replacement Strategy 29 3.4.2 Lightweight Residual Bottleneck Design for Layer 4 31 3.4.3 Synergistic Mechanism of IBN, DWConv, and CoordAtt 33 3.5 Loss Function Settings 35 Chapter 4 Performance Evaluation 39 4.1 Experimental Setup 39 4.1.1 Datasets 39 4.1.2 Evaluation Metrics 40 4.1.3 Implementation Details 42 4.1.4 Hardware and Software Environment 43 4.2 Comparison with State-of-the-Art Methods 43 4.2.1 Quantitative Results and Discussion 43 4.2.2 Computational Cost Comparison 47 4.3 Ablation Study 48 4.3.1 Progressive Ablation Design 49 Chapter 5 Conclusions 52 5.1 Research Summary 52 5.2 Future Work 54 References 55

    [1] M. Ye, J. Shen, G. Lin, T. Xiang, L. Shao, and S. C. H. Hoi, “Deep learning for person re-identification: A survey and outlook,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 6, pp. 2872–2893, 2022.
    [2] L. Zheng, L. Shen, L. Tian, S. Wang, J. Wang, and Q. Tian, “Scalable person re-identification: A benchmark,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp. 1116–1124, 2015.
    [3] A. Wu, W. Zheng, H. Yu, S. Gong, and J. Lai, “RGB-infrared cross-modality person re-identification,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp. 5380–5389, 2017.
    [4] Z. Wang, Z. Wang, Y. Zheng, Y. Chuang, and S. Satoh, “Learning to reduce dual-level discrepancy for infrared-visible person re-identification,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 618–626, 2019.
    [5] Z. Zhao, R. Sun, Z. Yang, and J. Gao, “Visible-infrared person re-identification based on frequency-domain simulated multispectral modality for dual-mode cameras,” IEEE Sensors Journal, vol. 22, no. 1, pp. 989–1002, 2022.
    [6] G. Zhang, Y. Zhang, T. Zhang, B. Li, and S. Pu, “PHA: Patch-wise high-frequency augmentation for transformer-based person re-identification,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 14 133–14 142, 2023.
    [7] M. Ye, J. Shen, D. J. Crandall, L. Shao, and J. Luo, “Dynamic dual-attentive aggregation learning for visible-infrared person re-identification,” in Proceedings of the European Conference on Computer Vision (ECCV), pp. 229–247, 2020.
    [8] Y. Zhang and H. Wang, “Diverse embedding expansion network and low-light cross-modality benchmark for visible-infrared person re-identification,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2153–2162, 2023.
    [9] Z. Chai, Y. Ling, Z. Luo, D. Lin, M. Jiang, and S. Li, “Dual-stream transformer with distribution alignment for visible-infrared person re-identification,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 11, pp. 6764–6776, 2023.
    [10] X. Yu, N. Dong, L. Zhu, H. Peng, and D. Tao, “CLIP-driven semantic discovery network for visible-infrared person re-identification,” IEEE Transactions on Multimedia, vol. 27, pp. 4137–4151, 2025.
    [11] N. Dong, S. Yan, L. Zhang, and J. Tang, “Diverse semantics-guided feature alignment and decoupling for visible-infrared person re-identification,” IEEE Transactions on Information Forensics and Security, vol. 20, pp. 12 245–12 259, 2025.
    [12] H. Gu, X. Yang, R. Lu, L. Pu, S. Han, and M. Wu, “Discovering multi-frequency embedding for visible-infrared person re-identification,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 36, no. 2, pp. 1766–1780, 2026.
    [13] P. Dai, R. Ji, H. Wang, Q. Wu, and Y. Huang, “Cross-modality person re-identification with generative adversarial training,” in Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence (IJCAI), pp. 677–683, 2018.
    [14] G. Wang, T. Zhang, J. Cheng, S. Liu, Y. Yang, and Z. Hou, “RGB-infrared cross-modality person re-identification via joint pixel and feature alignment,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 3623–3632, 2019.
    [15] S. Choi, S. Lee, Y. Kim, T. Kim, and C. Kim, “Hi-CMD: Hierarchical cross-modality disentanglement for visible-infrared person re-identification,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10 257–10 266, 2020.
    [16] L. Qiu, S. Chen, J.-H. Xue, D.-H. Wang, S. Zhu, and Y. Yan, “HOH-Net: High-order hierarchical middle-feature learning network for visible-infrared person re-identification,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 36, no. 2, pp. 2607–2622, 2026.
    [17] Y. Zhang, H. Wang, Y. Lu, Y. Yan, and X. Li, “Frequency domain nuances mining for visible-infrared person re-identification,” IEEE Transactions on Information Forensics and Security, vol. 20, pp. 5411–5424, 2025.
    [18] R. Sun, X. Wang, G. Huang, L. Chen, L. Qian, and J. Gao, “Robust visible-infrared person re-identification via frequency-space joint disentanglement and fusion network,” International Journal of Machine Learning and Cybernetics, vol. 16, no. 11, pp. 8893–8906, 2025.
    [19] X. Cao, P. Ding, J. Li, and M. Chen, “BiFFN: Bi-frequency guided feature fusion network for visible-infrared person re-identification,” Sensors, vol. 25, no. 5, p. 1298, 2025.
    [20] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770–778, 2016.
    [21] Y. Sun, L. Zheng, Y. Yang, Q. Tian, and S. Wang, “Beyond part models: Person retrieval with refined part pooling (and a strong convolutional baseline),” in Proceedings of the European Conference on Computer Vision (ECCV), pp. 501–518, 2018.
    [22] G. Wang, Y. Yuan, X. Chen, J. Li, and X. Zhou, “Learning discriminative features with multiple granularities for person re-identification,” in Proceedings of the 26th ACM International Conference on Multimedia (ACM MM), pp. 274–282, 2018.
    [23] H. Luo, Y. Gu, X. Liao, S. Lai, and W. Jiang, “Bag of tricks and a strong baseline for deep person re-identification,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 1487–1495, 2019.
    [24] K. Zhou, Y. Yang, A. Cavallaro, and T. Xiang, “Omni-scale feature learning for person re-identification,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 3702–3712, 2019.
    [25] J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7132–7141, 2018.
    [26] Q. Hou, D. Zhou, and J. Feng, “Coordinate attention for efficient mobile network design,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 13 713–13 722, 2021.
    [27] J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei, “Deformable convolutional networks,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp. 764–773, 2017.
    [28] X. Zhu, H. Hu, S. Lin, and J. Dai, “Deformable ConvNets v2: More deformable, better results,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 9308–9316, 2019.
    [29] Y. Rao, W. Zhao, Z. Zhu, J. Lu, and J. Zhou, “Global filter networks for image classification,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 34, pp. 980–993, 2021.
    [30] Z. Li, N. Kovachki, K. Azizzadenesheli, B. Liu, K. Bhattacharya, A. Stuart, and A. Anandkumar, “Fourier neural operator for parametric partial differential equations,” in International Conference on Learning Representations (ICLR), 2021.
    [31] G. Zhang, Y. Zhang, and Z. Tan, “ProtoHPE: Prototype-guided high-frequency patch enhancement for visible-infrared person re-identification,” in Proceedings of the 31st ACM International Conference on Multimedia, pp. 944–954, 2023.
    [32] J. Xiao, Z. Lyu, C. Zhang, Y. Ju, C. Shui, and K.-M. Lam, “Towards progressive multi-frequency representation for image warping,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2995–3004, 2024.
    [33] P. Goyal, P. Dollar, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He, “Accurate, large minibatch SGD: Training ImageNet in 1 hour,” arXiv preprint arXiv:1706.02677, 2017.
    [34] M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” in Proceedings of the European Conference on Computer Vision (ECCV), pp. 818–833, 2014.
    [35] J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “How transferable are features in deep neural networks?” in Advances in Neural Information Processing Systems, vol. 27, pp. 3320–3328, 2014.
    [36] Y. Chen, S. Xia, J. Zhao, Y. Zhou, Q. Niu, R. Yao, D. Zhu, and D. Liu, “ResT-ReID: Transformer block-based residual learning for person re-identification,” Pattern Recognition Letters, vol. 157, pp. 90–96, 2022.
    [37] A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,” arXiv preprint arXiv:1704.04861, 2017.
    [38] X. Pan, P. Luo, J. Shi, and X. Tang, “Two at once: Enhancing learning and generalization capacities via IBN-Net,” in Proceedings of the European Conference on Computer Vision (ECCV), pp. 464–479, 2018.
    [39] M. Blondel, O. Teboul, Q. Berthet, and J. Djolonga, “Fast differentiable sorting and ranking,” in Proceedings of the 37th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 119, pp. 950–959, 2020.
    [40] D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proceedings of the 3rd International Conference on Learning Representations (ICLR), 2015.

    QR CODE