| 研究生: |
邱靖云 Chiu, Ching-Yun |
|---|---|
| 論文名稱: |
一個輕量且有效之動態頻域神經網路用於可見光紅外線行人重識別 A Lightweight and Effective Dynamic Frequency Domain Neural Network for Visible-Infrared Person Re-Identification |
| 指導教授: |
戴顯權
Tai, Shen-Chuan |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 電機工程學系 Department of Electrical Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 71 |
| 中文關鍵詞: | 可見光紅外線行人重識別 、注意力機制 、輕量化神經網路 |
| 外文關鍵詞: | Visible-Infrared Person Re-Identification, Attention Mechanism, Lightweight Neural Network |
| 相關次數: | 點閱:2 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
可見光紅外線行人重識別(Visible-Infrared Person Re-Identification, VI-ReID)旨在判斷不同攝影機拍攝之可見光與紅外線行人影像是否屬於同一身分,是全天候智慧監控的重要技術。然而,由於可見光與紅外線影像在光照、色彩及成像特性上的差異,以及視角、姿態與遮擋造成的外觀變化,使跨模態特徵學習與身分匹配仍具挑戰性。此外現有頻域方法多採用固定的頻率分割方式,缺乏根據影像內容進行自適應調整的能力,且辨識效能與模型複雜度之間仍存在權衡。
為改善上述問題,本論文提出一種輕量且有效的動態頻域神經網路。其中,自適應高低頻調制模組透過內容感知的動態頻率分割與低頻特徵調制,提升跨模態特徵學習能力;其高頻路徑則結合可變形多尺度卷積,自適應提取高頻特徵中的局部結構資訊。同時重新設計輕量化最終殘差階段,以兼顧辨識效能與模型效率。實驗結果顯示,所提出的方法在SYSU-MM01與LLCM兩個公開VI-ReID資料集上皆展現具競爭力的辨識效能。整體而言,本方法在辨識準確度、模型參數量與計算成本之間取得良好的平衡,具備應用於智慧監控系統之潛力。
Visible-Infrared Person Re-Identification (VI-ReID) aims to determine whether visible and infrared pedestrian images captured by different cameras belong to the same identity and is an important technology for round-the-clock intelligent surveillance. However, differences between visible and infrared images in illumination, color, and imaging characteristics, together with appearance variations caused by viewpoint, pose, and occlusion, make cross modality feature learning and identity matching challenging. Furthermore, most existing frequency-domain methods employ fixed frequency partitioning schemes and lack the ability to adapt the partitioning to image content, while a trade-off remains between recognition performance and model complexity.
To address these issues, this Thesis proposes a lightweight and effective dynamic frequency-domain neural network. The proposed Adaptive HiLo-Frequency Modulation (Adaptive HiLo-FM) module improves cross-modality feature learning through content aware dynamic frequency partitioning and low-frequency feature modulation. Within its high-frequency path, deformable multi-scale convolutions adaptively extract local structural information from high-frequency features. The lightweight final residual stage is also redesigned to balance recognition performance and model efficiency. Experimental results demonstrate that the proposed method achieves competitive recognition performance on two public VI-ReID datasets, SYSU-MM01 and LLCM. Overall, the proposed method strikes a favorable balance among recognition accuracy, model parameter count, and computational cost, demonstrating its potential for application in intelligent surveillance systems.
[1] M. Ye, J. Shen, G. Lin, T. Xiang, L. Shao, and S. C. H. Hoi, “Deep learning for person re-identification: A survey and outlook,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 6, pp. 2872–2893, 2022.
[2] L. Zheng, L. Shen, L. Tian, S. Wang, J. Wang, and Q. Tian, “Scalable person re-identification: A benchmark,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp. 1116–1124, 2015.
[3] A. Wu, W. Zheng, H. Yu, S. Gong, and J. Lai, “RGB-infrared cross-modality person re-identification,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp. 5380–5389, 2017.
[4] Z. Wang, Z. Wang, Y. Zheng, Y. Chuang, and S. Satoh, “Learning to reduce dual-level discrepancy for infrared-visible person re-identification,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 618–626, 2019.
[5] Z. Zhao, R. Sun, Z. Yang, and J. Gao, “Visible-infrared person re-identification based on frequency-domain simulated multispectral modality for dual-mode cameras,” IEEE Sensors Journal, vol. 22, no. 1, pp. 989–1002, 2022.
[6] G. Zhang, Y. Zhang, T. Zhang, B. Li, and S. Pu, “PHA: Patch-wise high-frequency augmentation for transformer-based person re-identification,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 14 133–14 142, 2023.
[7] M. Ye, J. Shen, D. J. Crandall, L. Shao, and J. Luo, “Dynamic dual-attentive aggregation learning for visible-infrared person re-identification,” in Proceedings of the European Conference on Computer Vision (ECCV), pp. 229–247, 2020.
[8] Y. Zhang and H. Wang, “Diverse embedding expansion network and low-light cross-modality benchmark for visible-infrared person re-identification,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2153–2162, 2023.
[9] Z. Chai, Y. Ling, Z. Luo, D. Lin, M. Jiang, and S. Li, “Dual-stream transformer with distribution alignment for visible-infrared person re-identification,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 11, pp. 6764–6776, 2023.
[10] X. Yu, N. Dong, L. Zhu, H. Peng, and D. Tao, “CLIP-driven semantic discovery network for visible-infrared person re-identification,” IEEE Transactions on Multimedia, vol. 27, pp. 4137–4151, 2025.
[11] N. Dong, S. Yan, L. Zhang, and J. Tang, “Diverse semantics-guided feature alignment and decoupling for visible-infrared person re-identification,” IEEE Transactions on Information Forensics and Security, vol. 20, pp. 12 245–12 259, 2025.
[12] H. Gu, X. Yang, R. Lu, L. Pu, S. Han, and M. Wu, “Discovering multi-frequency embedding for visible-infrared person re-identification,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 36, no. 2, pp. 1766–1780, 2026.
[13] P. Dai, R. Ji, H. Wang, Q. Wu, and Y. Huang, “Cross-modality person re-identification with generative adversarial training,” in Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence (IJCAI), pp. 677–683, 2018.
[14] G. Wang, T. Zhang, J. Cheng, S. Liu, Y. Yang, and Z. Hou, “RGB-infrared cross-modality person re-identification via joint pixel and feature alignment,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 3623–3632, 2019.
[15] S. Choi, S. Lee, Y. Kim, T. Kim, and C. Kim, “Hi-CMD: Hierarchical cross-modality disentanglement for visible-infrared person re-identification,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10 257–10 266, 2020.
[16] L. Qiu, S. Chen, J.-H. Xue, D.-H. Wang, S. Zhu, and Y. Yan, “HOH-Net: High-order hierarchical middle-feature learning network for visible-infrared person re-identification,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 36, no. 2, pp. 2607–2622, 2026.
[17] Y. Zhang, H. Wang, Y. Lu, Y. Yan, and X. Li, “Frequency domain nuances mining for visible-infrared person re-identification,” IEEE Transactions on Information Forensics and Security, vol. 20, pp. 5411–5424, 2025.
[18] R. Sun, X. Wang, G. Huang, L. Chen, L. Qian, and J. Gao, “Robust visible-infrared person re-identification via frequency-space joint disentanglement and fusion network,” International Journal of Machine Learning and Cybernetics, vol. 16, no. 11, pp. 8893–8906, 2025.
[19] X. Cao, P. Ding, J. Li, and M. Chen, “BiFFN: Bi-frequency guided feature fusion network for visible-infrared person re-identification,” Sensors, vol. 25, no. 5, p. 1298, 2025.
[20] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770–778, 2016.
[21] Y. Sun, L. Zheng, Y. Yang, Q. Tian, and S. Wang, “Beyond part models: Person retrieval with refined part pooling (and a strong convolutional baseline),” in Proceedings of the European Conference on Computer Vision (ECCV), pp. 501–518, 2018.
[22] G. Wang, Y. Yuan, X. Chen, J. Li, and X. Zhou, “Learning discriminative features with multiple granularities for person re-identification,” in Proceedings of the 26th ACM International Conference on Multimedia (ACM MM), pp. 274–282, 2018.
[23] H. Luo, Y. Gu, X. Liao, S. Lai, and W. Jiang, “Bag of tricks and a strong baseline for deep person re-identification,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 1487–1495, 2019.
[24] K. Zhou, Y. Yang, A. Cavallaro, and T. Xiang, “Omni-scale feature learning for person re-identification,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 3702–3712, 2019.
[25] J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7132–7141, 2018.
[26] Q. Hou, D. Zhou, and J. Feng, “Coordinate attention for efficient mobile network design,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 13 713–13 722, 2021.
[27] J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei, “Deformable convolutional networks,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp. 764–773, 2017.
[28] X. Zhu, H. Hu, S. Lin, and J. Dai, “Deformable ConvNets v2: More deformable, better results,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 9308–9316, 2019.
[29] Y. Rao, W. Zhao, Z. Zhu, J. Lu, and J. Zhou, “Global filter networks for image classification,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 34, pp. 980–993, 2021.
[30] Z. Li, N. Kovachki, K. Azizzadenesheli, B. Liu, K. Bhattacharya, A. Stuart, and A. Anandkumar, “Fourier neural operator for parametric partial differential equations,” in International Conference on Learning Representations (ICLR), 2021.
[31] G. Zhang, Y. Zhang, and Z. Tan, “ProtoHPE: Prototype-guided high-frequency patch enhancement for visible-infrared person re-identification,” in Proceedings of the 31st ACM International Conference on Multimedia, pp. 944–954, 2023.
[32] J. Xiao, Z. Lyu, C. Zhang, Y. Ju, C. Shui, and K.-M. Lam, “Towards progressive multi-frequency representation for image warping,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2995–3004, 2024.
[33] P. Goyal, P. Dollar, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He, “Accurate, large minibatch SGD: Training ImageNet in 1 hour,” arXiv preprint arXiv:1706.02677, 2017.
[34] M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” in Proceedings of the European Conference on Computer Vision (ECCV), pp. 818–833, 2014.
[35] J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “How transferable are features in deep neural networks?” in Advances in Neural Information Processing Systems, vol. 27, pp. 3320–3328, 2014.
[36] Y. Chen, S. Xia, J. Zhao, Y. Zhou, Q. Niu, R. Yao, D. Zhu, and D. Liu, “ResT-ReID: Transformer block-based residual learning for person re-identification,” Pattern Recognition Letters, vol. 157, pp. 90–96, 2022.
[37] A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,” arXiv preprint arXiv:1704.04861, 2017.
[38] X. Pan, P. Luo, J. Shi, and X. Tang, “Two at once: Enhancing learning and generalization capacities via IBN-Net,” in Proceedings of the European Conference on Computer Vision (ECCV), pp. 464–479, 2018.
[39] M. Blondel, O. Teboul, Q. Berthet, and J. Djolonga, “Fast differentiable sorting and ranking,” in Proceedings of the 37th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 119, pp. 950–959, 2020.
[40] D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proceedings of the 3rd International Conference on Learning Representations (ICLR), 2015.