| 研究生: |
林仕騏 Lin, Shi-Qi |
|---|---|
| 論文名稱: |
防禦對抗式攻擊的穩定性感知多模態 Deepfake 偵測方法 SAD : A Stability-Aware Multimodal Deepfake Detection Method for Defending Against Adversarial Attacks |
| 指導教授: |
郭耀煌
Kuo, Yau-Hwang |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 醫學資訊研究所 Institute of Medical Informatics |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 110 |
| 中文關鍵詞: | 多模態 、深偽 、對抗式攻擊 |
| 外文關鍵詞: | Multimodal, Deepfake, Adversarial Attack |
| 相關次數: | 點閱:74 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
隨著生成式人工智慧技術的快速進步,影像合成、語音轉換與影音偽造技術持續發展,使偽造內容與真實內容之間的差異愈來愈難以辨識。早期深偽多集中於單一模態,例如臉部影像替換、生成或操控,或語音內容轉換與合成。然而,隨著多媒體偽造技術的進步,深偽已逐漸發展為同時涉及視覺與音訊的多模態影音偽造,可能包含偽造臉部影像、修改語音內容,或音訊與影像不匹配的影音資料。若此類偽造內容被惡意散播,將可能對個人隱私、社群平台可信度與數位媒體安全造成嚴重影響。
為了因應深偽技術所帶來的威脅,過去已有許多研究提出不同的偵測方法,並在一般測試情境下達到良好表現。然而,現有方法多數著重於乾淨樣本上的偵測能力,較少考量多模態深偽偵測模型在對抗式攻擊下的安全性與穩健性。當攻擊者在偽造影音內容中加入人類不易感知的對抗式擾動時,模型原本依賴的影像偽造痕跡、語音偽造特徵,或音訊-視覺之間的跨模態一致性線索可能遭到破壞,進而導致偵測模型產生錯誤判斷。
因此,本研究提出可抵禦對抗式攻擊的多模態深偽偵測方法,Stability-Aware Multimodal Deepfake Detection(SAD)。SAD 的目標是在維持乾淨影音樣本偵測準確率的同時,提升模型面對對抗式擾動時的穩健性。SAD 主要包含三個模組:多視角轉換(MVT)、多視角穩定性估計器(MSE),以及可靠度引導融合(RGF)。首先,MVT 針對同一筆影音輸入產生多個固定轉換後的視角,用以觀察模型在不同轉換條件下對視覺、音訊與跨模態關係的反應是否穩定。接著,MSE 透過比較不同視角下的預測變化、特徵偏移與跨模態相似度變化,估計視覺、音訊與跨模態資訊的穩定性證據。最後,RGF 根據這些穩定性相關資訊產生可靠度權重,並引導影音特徵進行融合,使最終判斷能同時考量單一模態內部的偽造線索,以及不同模態之間的關聯性與一致性。
實驗結果顯示,相較於現有深偽偵測方法,SAD 在面對不同對抗式攻擊時能維持較穩定的偵測表現,並有效降低對抗式擾動對多模態偵測模型造成的影響。這說明所提出的穩定性感知設計能提升 SAD 在攻擊情境下的穩健性,並證明結合視覺、音訊與跨模態一致性分析,對於建構安全且可靠的多模態深偽偵測方法具有重要價值。
With the rapid advancement of generative artificial intelligence, image synthesis, voice conversion, and audio-visual forgery technologies have continued to develop, making it increasingly difficult to distinguish forged content from real content. Early Deepfake content mainly focused on a single modality, such as facial image replacement, generation, or manipulation, or speech content conversion and synthesis. However, with the progress of multimedia forgery technologies, Deepfake has gradually evolved into multimodal audio-visual forgery involving both visual and audio modalities. Such forged content may include manipulated facial images, modified speech content, or audio-visual data in which the audio and visual information do not match. If such forged content is maliciously disseminated, it may cause serious impacts on personal privacy, the credibility of social media platforms, and digital media security.
To address the threats brought by Deepfake technologies, many studies have proposed different detection methods and achieved promising performance under general testing scenarios. However, most existing methods focus on detection capability on clean samples and pay less attention to the security and robustness of multimodal Deepfake detection models under adversarial attacks. When attackers introduce human-imperceptible adversarial perturbations into forged audio-visual content, the visual forgery traces, speech forgery features, or cross-modal consistency cues between audio and visual information that the model originally relies on may be disrupted, thereby causing the detection model to make incorrect predictions.
Therefore, this study proposes a multimodal Deepfake detection method for defending against adversarial attacks, called Stability-Aware Multimodal Deepfake Detection (SAD). The goal of SAD is to improve the robustness of the model against adversarial perturbations while maintaining the detection accuracy of clean audio-visual samples. SAD mainly consists of three modules: Multi-View Transformation (MVT), Multi-View Stability Estimator (MSE), and Reliability-Guided Fusion (RGF). First, MVT generates multiple fixed transformed views from the same audio-visual input to observe whether the model responses to the visual modality, audio modality, and cross-modal relationship remain stable under different transformation conditions. Next, MSE estimates stability-related evidence for visual, audio, and cross-modal information by comparing prediction changes, feature deviations, and cross-modal similarity variations across different views. Finally, RGF generates reliability weights based on these stability-related cues and uses them to guide audio-visual feature fusion, enabling the final decision to consider both forgery cues within each individual modality and the relationships and consistency between different modalities.
Experimental results show that, compared with existing Deepfake detection methods, SAD can maintain more stable detection performance under different adversarial attacks and effectively reduce the impact of adversarial perturbations on multimodal detection models. These results demonstrate that the proposed stability-aware design improves the robustness of SAD under attack scenarios, and further show that integrating visual, audio, and cross-modal consistency analysis is valuable for building secure and reliable multimodal Deepfake detection methods.
[1] R. Tolosana, R. Vera-Rodriguez, J. Fierrez, A. Morales, and J. Ortega-Garcia, "Deepfakes and beyond: A survey of face manipulation and fake detection," Information fusion, vol. 64, pp. 131–148, 2020.
[2] Y. Mirsky and W. Lee, "The creation and detection of deepfakes: A survey," ACM computing surveys (CSUR), vol. 54, no. 1, pp. 1–41, 2021.
[3] Y. Nirkin, Y. Keller, and T. Hassner, "Fsgan: Subject agnostic face swapping and reenactment," in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 7184–7193.
[4] K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa-Johnson, "Autovc: Zero-shot voice style transfer with only autoencoder loss," in International Conference on Machine Learning, 2019: PMLR, pp. 5210–5219.
[5] J. Shen et al., "Natural tts synthesis by conditioning wavenet on mel spectrogram predictions," in 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP), 2018: IEEE, pp. 4779–4783.
[6] H. Khalid, S. Tariq, M. Kim, and S. S. Woo, "FakeAVCeleb: A novel audio-video multimodal deepfake dataset," arXiv preprint arXiv:2108.05080, 2021.
[7] Y. Hou, H. Fu, C. Chen, Z. Li, H. Zhang, and J. Zhao, "Polyglotfake: A novel multilingual and multimodal deepfake dataset," in International conference on pattern recognition, 2024: Springer, pp. 180–193.
[8] A. Hashmi, S. A. Shahzad, C. W. Lin, Y. Tsao, and H.-M. Wang, "AVTENet: A human-cognition-inspired audio-visual transformer-based ensemble network for video deepfake detection," IEEE Transactions on Cognitive and Developmental Systems, vol. 17, no. 6, pp. 1360–1376, 2025.
[9] M. Astrid, E. Ghorbel, and D. Aouada, "Statistics-aware audio-visual deepfake detector," in 2024 IEEE International Conference on Image Processing (ICIP), 2024: IEEE, pp. 2557–2563.
[10] C. Koutlis and S. Papadopoulos, "DiMoDif: Discourse modality-information differentiation for audio-visual deepfake detection and localization," arXiv preprint arXiv:2411.10193, 2024.
[11] S. Smeu, D.-A. Boldisor, D. Oneata, and E. Oneata, "Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning," in 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025: IEEE, pp. 18815–18825.
[12] N. Carlini and H. Farid, "Evading deepfake-image detectors with white-and black-box attacks," in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, 2020, pp. 658–659.
[13] P. Kawa, M. Plata, and P. Syga, "Defense against adversarial attacks on audio deepfake detection," arXiv preprint arXiv:2212.14597, 2022.
[14] J. Thies, M. Zollhofer, M. Stamminger, C. Theobalt, and M. Nießner, "Face2face: Real-time face capture and reenactment of rgb videos," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 2387–2395.
[15] A. Rossler, D. Cozzolino, L. Verdoliva, C. Riess, J. Thies, and M. Nießner, "Faceforensics++: Learning to detect manipulated facial images," in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 1–11.
[16] J.-w. Jung et al., "Aasist: Audio anti-spoofing using integrated spectro-temporal graph attention networks," in ICASSP 2022-2022 IEEE international conference on acoustics, speech and signal processing (ICASSP), 2022: IEEE, pp. 6367–6371.
[17] L. Li et al., "Face x-ray for more general face forgery detection," in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 5001–5010.
[18] Y. Luo, Y. Zhang, J. Yan, and W. Liu, "Generalizing face forgery detection with high-frequency features," in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 16317–16326.
[19] H. Tak, J.-w. Jung, J. Patino, M. Kamble, M. Todisco, and N. Evans, "End-to-end spectro-temporal graph attention networks for speaker verification anti-spoofing and speech deepfake detection," arXiv preprint arXiv:2107.12710, 2021.
[20] X. Liu, M. Sahidullah, K. A. Lee, and T. Kinnunen, "Speaker-aware anti-spoofing," arXiv preprint arXiv:2303.01126, 2023.
[21] K. Prajwal, R. Mukhopadhyay, V. P. Namboodiri, and C. Jawahar, "A lip sync expert is all you need for speech to lip generation in the wild," in Proceedings of the 28th ACM international conference on multimedia, 2020, pp. 484–492.
[22] K. Cheng et al., "Videoretalking: Audio-based lip synchronization for talking head video editing in the wild," in SIGGRAPH Asia 2022 Conference Papers, 2022, pp. 1–9.
[23] Y. Zhou and S.-N. Lim, "Joint audio-visual deepfake detection," in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 14800–14809.
[24] T. Oorloff et al., "Avff: Audio-visual feature fusion for video deepfake detection," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 27102–27112.
[25] S. Hussain, P. Neekhara, M. Jere, F. Koushanfar, and J. McAuley, "Adversarial deepfakes: Evaluating vulnerability of deepfake detectors to adversarial examples," in Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2021, pp. 3348–3357.
[26] I. J. Goodfellow, J. Shlens, and C. Szegedy, "Explaining and harnessing adversarial examples," arXiv preprint arXiv:1412.6572, 2014.
[27] A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu, "Towards deep learning models resistant to adversarial attacks," arXiv preprint arXiv:1706.06083, 2017.
[28] J. Wen, X. Wu, S. Zhao, Y. Jia, and Y. Li, "Investigating vulnerabilities and defenses against audio-visual attacks: A comprehensive survey emphasizing multimodal models," arXiv preprint arXiv:2506.11521, 2025.
[29] Y. Tian and C. Xu, "Can audio-visual integration strengthen robustness under multimodal attacks?," in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 5601–5611.
[30] M. U. Farooq, A. Khan, K. Uddin, and K. M. Malik, "Transferable adversarial attacks on audio deepfake detection," in Proceedings of the Winter Conference on Applications of Computer Vision, 2025, pp. 1640–1649.
[31] Z. Zhang, S. Liang, D. Shimada, and C. Xu, "Rethinking audio-visual adversarial vulnerability from temporal and modality perspectives," arXiv preprint arXiv:2502.11858, 2025.
[32] K. Yang, W.-Y. Lin, M. Barman, F. Condessa, and Z. Kolter, "Defending multimodal fusion models against single-source adversaries," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 3340–3349.
[33] W. Yang et al., "Avoid-df: Audio-visual joint learning for detecting deepfake," IEEE Transactions on Information Forensics and Security, vol. 18, pp. 2015–2029, 2023.
[34] D. Cozzolino, A. Pianese, M. Nießner, and L. Verdoliva, "Audio-visual person-of-interest deepfake detection," in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 943–952.
[35] C. Feng, Z. Chen, and A. Owens, "Self-supervised video forensics by audio-visual anomaly detection," in proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 10491–10503.
[36] P. Xie et al., "Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial Attacks," in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 14679–14689.
[37] J. Zhang et al., "Anyattack: Towards large-scale self-supervised adversarial attacks on vision-language models," in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 19900–19909.
[38] X. Xu, X. Chen, C. Liu, A. Rohrbach, T. Darrell, and D. Song, "Fooling vision and language models despite localization and attention mechanism," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 4951–4961.
[39] A. Kurakin, I. J. Goodfellow, and S. Bengio, "Adversarial examples in the physical world," in Artificial intelligence safety and security: Chapman and Hall/CRC, 2018, pp. 99–112.
[40] Y. Zheng, J. Bao, D. Chen, M. Zeng, and F. Wen, "Exploring temporal coherence for more general video face forgery detection," in 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 2021: IEEE, pp. 15024–15034.
[41] H. Tak, M. Todisco, X. Wang, J.-w. Jung, J. Yamagishi, and N. Evans, "Automatic speaker verification spoofing and deepfake detection using wav2vec 2.0 and data augmentation," arXiv preprint arXiv:2202.12233, 2022.
[42] Y. Gao, T. Vuong, M. Elyasi, G. Bharaj, and R. Singh, "Generalized spoofing detection inspired from audio generation artifacts," arXiv preprint arXiv:2104.04111, 2021.
[43] J. Xue et al., "Audio deepfake detection based on a combination of f0 information and real plus imaginary spectrogram features," in Proceedings of the 1st international workshop on deepfake detection for audio multimedia, 2022, pp. 19–26.
[44] B. Shi, W.-N. Hsu, K. Lakhotia, and A. Mohamed, "Learning audio-visual speech representation by masked multimodal cluster prediction," arXiv preprint arXiv:2201.02184, 2022.
[45] P. Ma, S. Petridis, and M. Pantic, "Visual speech recognition for multiple languages in the wild," Nature Machine Intelligence, vol. 4, no. 11, pp. 930–939, 2022.
[46] J. S. Chung, A. Nagrani, and A. Zisserman, "Voxceleb2: Deep speaker recognition," arXiv preprint arXiv:1806.05622, 2018.