| 研究生: |
陳宏仲 Chen, Hong-Zhong |
|---|---|
| 論文名稱: |
無監督領域自適應物件偵測之雙模式教師蒸餾 Dual-Teacher Distillation for Unsupervised Domain Adaptive Object Detection |
| 指導教授: |
莊坤達
Chuang, Kun-Ta |
| 共同指導: |
高宏宇
Kao, Hung-Yu |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 人工智慧科技碩士學位學程 Graduate Program of Artificial Intelligence |
| 論文出版年: | 2025 |
| 畢業學年度: | 113 |
| 語文別: | 英文 |
| 論文頁數: | 55 |
| 中文關鍵詞: | 領域自適應 、物體偵測 、偽標籤學習 、多教師蒸餾 、EMA 教師模型 、無監督學習 |
| 外文關鍵詞: | Domain Adaptation, Object Detection, Pseudo-Label Learning, Multi-Teacher Distillation, EMA Teacher, Unsupervised Learning |
| 相關次數: | 點閱:178 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
無監督域自適應物體檢測(Unsupervised Domain-Adaptive Object Detection, UDAOD)[25] 的目標是將在已標註的來源域上訓練的檢測器,適應到具有巨大外觀差異的未標註目標域。
為了應對ALDI++ [11] 中單一EMA 設計對雜訊的敏感性,我們提出了兩種互補的蒸餾策略:離線教師蒸餾(Offline-Teacher Distillation, OTD),透過固定教師提供穩定的logits;以及雙教師蒸餾(Dual-Teacher Distillation, DTD),將離線教師與EMA 教師融合以兼具適應性。融合策略包含:均值(mean)、熵(entropy)、硬投票(hard-vote)、以及置信度加權(confidence weighting)。這兩種方法皆結合了強增強與proposal-level 一致性。
我們在四種設定下進行評估:(i) ALDI++、(ii) 離線head 微調、(iii) DTD、以及(iv)OTD。
在Cityscapes [4] → Foggy Cityscapes [22] (CS → FCS)的任務中,我們重新實作的ALDI++ 僅達到65.96 AP50(相比論文報導的66.8)。在相同設定下,DTD 將分數提升至66.21(+0.25),而OTD 則達到66.07(+0.11),明顯超越我們的基準結果,並接近報導數值
在我們自行建立的撲克牌基準(Card-Std → Card-3D 任務)中,ALDI++ 在目標域(Card-3D)僅達到22.55 / 37.84(AP / AP50)。在相同設定下,DTD 提升至23.19 /39.30,而OTD 進一步提升至24.48 / 41.82,展現了其在嚴重風格與光照變化下的有效性。
Unsupervised Domain-Adaptive Object Detection (UDAOD) [25] adapts detectors from a labeled source domain to an unlabeled target domain with large appearance gaps.
To address the noise sensitivity of the single-EMA design in ALDI++ [11], we propose two complementary distillation strategies: Offline-Teacher Distillation (OTD) with a fixed teacher for stable logits, and Dual-Teacher Distillation (DTD) which fuses the offline teacher with an EMA teacher for adaptability. Fusion is performed via mean, entropy, hard-vote, or confidence weighting. Both are combined with strong augmentation and proposal-level consistency.
We evaluate four settings: (i) ALDI++, (ii) offline head fine-tuning, (iii) DTD, and (iv) OTD.
In the Cityscapes [4] to Foggy Cityscapes [22] (CS to FCS) task, our re-implementation of ALDI++ reaches only 65.96 AP50 (compared to the reported 66.8). Under the same setting,DTD improves the score to 66.21 (+0.25), while OTD achieves 66.07 (+0.11), clearly surpassing our baseline and approaching the reported figure.
In our in-house card benchmark (Card-Std to Card-3D task), ALDI++ achieves only 22.55 /37.84 (AP / AP50) on the target domain (Card-3D). Under the same setting, DTD improves performance to 23.19 / 39.30, and OTD further boosts it to 24.48 / 41.82, demonstrating effectiveness under severe style and lighting shifts.
[1] Konstantinos Bousmalis, Nathan Silberman, David Dohan, Dumitru Erhan, and Dilip Krishnan. Unsupervised pixel-level domain adaptation with generative adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition,pages 3722–3731, 2017.
[2] Shengcao Cao, Dhiraj Joshi, Liang-Yan Gui, and Yu-Xiong Wang. Contrastive mean teacher for domain adaptive object detectors. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 23839–23848, 2023.
[3] Yuhua Chen, Wen Li, Christos Sakaridis, Dengxin Dai, and Luc Van Gool. Domain adaptive faster r-cnn for object detection in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3339–3348, 2018.
[4] Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler,Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3213–3223, 2016.
[5] Jinhong Deng, Wen Li, Yuhua Chen, and Lixin Duan. Unbiased mean teacher for crossdomain object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4091–4101, 2021.
[6] Tommaso Furlanello, Zachary Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar. Born again neural networks. In International conference on machine learning, pages 1607–1616. PMLR, 2018.
[7] Ross Girshick. Fast r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 1440–1448, 2015.
[8] Zhenwei He and Lei Zhang. Multi-adversarial faster-rcnn for unrestricted object detection. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6668–6677, 2019.
[9] Judy Hoffman, Eric Tzeng, Taesung Park, Jun-Yan Zhu, Phillip Isola, Kate Saenko, Alexei Efros, and Trevor Darrell. Cycada: Cycle-consistent adversarial domain adaptation. In International conference on machine learning, pages 1989–1998. Pmlr, 2018.
[10] Naoto Inoue, Ryosuke Furuta, Toshihiko Yamasaki, and Kiyoharu Aizawa. Crossdomain weakly-supervised object detection through progressive domain adaptation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5001–5009, 2018.
[11] Justin Kay, Timm Haucke, Suzanne Stathatos, Siqi Deng, Erik Young, Pietro Perona, Sara Beery, and Grant Van Horn. Align and distill: Unifying and improving domain adaptive object detection. Transactions on Machine Learning Research, 2025. Featured Certification.
[12] Kyungmin Kwon, Hojin Na, Hwangbeom Lee, and Nam Soo Kim. Adaptive knowledge distillation based on entropy. In ICASSP 2020 – IEEE International Conference on Acoustics, Speech and Signal Processing, pages 7409–7413. IEEE, 2020.
[13] Yu-Jhe Li, Xiaoliang Dai, Chih-Yao Ma, Yen-Cheng Liu, Kan Chen, Bichen Wu, Zijian He, Kris Kitani, and Peter Vajda. Cross-domain adaptive teacher for object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 7581–7590, 2022.
[14] Jian Liang, Dapeng Hu, Yunbo Wang, Ran He, and Jiashi Feng. Source data-absent unsupervised domain adaptation through hypothesis transfer and labeling transfer. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11):8602–8617, 2021.
[15] Feng Liu, Xiaosong Zhang, Fang Wan, Xiangyang Ji, and Qixiang Ye. Domain contrast for domain adaptive object detection. IEEE Transactions on Circuits and Systems for Video Technology, 32(12):8227–8237, 2021.
[16] Sangwoo Mo, Minsu Cho, and Jinwoo Shin. Instagan: Instance-aware image-to-image translation. arXiv preprint arXiv:1812.10889, 2018.
[17] Hyeonseob Nam and Bohyung Han. Learning multi-domain convolutional neural networks for visual tracking. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4293–4302, 2016.
[18] Taesung Park, Alexei A. Efros, Richard Zhang, and Jun-Yan Zhu. Contrastive learning for unpaired image-to-image translation. In Computer Vision—ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IX, pages 319–345. Springer, 2020.
[19] Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards realtime object detection with region proposal networks. In Advances in Neural Information Processing Systems (NeurIPS), 2015.
[20] Aruni RoyChowdhury, Prithvijit Chakrabarty, Ashish Singh, SouYoung Jin, Huaizu Jiang, Liangliang Cao, and Erik Learned-Miller. Automatic adaptation of object detectors to new domains using self-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 780–790, 2019.
[21] Kuniaki Saito, Yoshitaka Ushiku, Tatsuya Harada, and Kate Saenko. Strong-weak distribution alignment for adaptive object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6956–6965, 2019.
[22] Christos Sakaridis, Dengxin Dai, Simon Hecker, and Luc Van Gool. Model adaptation with synthetic and real data for semantic dense foggy scene understanding. In Proceedings of the european conference on computer vision (ECCV), pages 687–704, 2018.
[23] Antti Tarvainen and Harri Valpola. Mean teachers are better role models: Weightaveraged consistency targets improve semi-supervised deep learning results. In NeurIPS, 2017.
[24] T. Vu and H. Jain. Dada: Depth-aware domain adaptation in semantic segmentation. In CVPR, 2020.
[25] Siqi Yang, Lin Wu, Arnold Wiliem, and Brian C Lovell. Unsupervised domain adaptive object detection using forward-backward cyclic adaptation. In Proceedings of the Asian Conference on Computer Vision, 2020.
[26] Shan You, Chen Xu, Chaojian Xu, and Dacheng Tao. Learning from multiple teacher networks. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 1285–1294. ACM, 2017.
[27] Hailin Zhang, Defang Chen, and Can Wang. Confidence-aware multi-teacher knowledge distillation. In ICASSP 2022 – IEEE International Conference on Acoustics, Speech and Signal Processing, pages 4498–4502. IEEE, 2022.
[28] Zhi-Hua Zhou and Ming Li. Tri-training: Exploiting unlabeled data using three classifiers. IEEE Transactions on Knowledge and Data Engineering, 17(11):1529–1541, 2005.
[29] Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros. Unpaired image-toimage translation using cycle-consistent adversarial networks. In Proceedings of the IEEE international conference on computer vision, pages 2223–2232, 2017.