簡易檢索 / 詳目顯示

研究生: 陳宏恩
Chen, Hung-En
論文名稱: 對抗乾淨標籤後門攻擊之無先驗知識資料清洗方法
BlindClean: A Prior-Knowledge-Free Data Sanitization Method against Clean-Label Backdoor Attacks
指導教授: 郭耀煌
Kuo, Yau-Hwang
莊宜勲
Chuang, I-Hsun
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 資訊工程學系
Department of Computer Science and Information Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 113
中文關鍵詞: 資料清洗乾淨標籤後門攻擊無先驗知識
外文關鍵詞: Data sanitization, Clean-label backdoor attack, No prior knowledge
相關次數: 點閱:3下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 今日的模型開發者經常使用第三方提供的資料集來訓練其模型,藉此提升訓練資料的多樣性並降低資料收集與標記的成本,卻也使得其訓練的模型極易遭受各種後門攻擊的威脅,因此亟需一資料清洗演算法來降低資安風險。在現有多種模型後門攻擊中,最具隱蔽性的當屬乾淨標籤後門攻擊。此類攻擊由於其中毒樣本與正常樣本擁有相同標籤,導致識別並清洗此類中毒樣本變得極為困難。然而,現有資料清洗的方法應對乾淨標籤後門攻擊往往需要大幅犧牲模型準確率才能抑制攻擊成功率,或需仰賴乾淨資料與攻擊相關資訊的先驗知識才能正常執行,大大限制了他們的可用性。
    為此,本論文提出一種對抗乾淨標籤後門攻擊的資料清洗方法(BlindClean),能在不依賴先驗知識且不顯著降低模型準確率的情況下,有效辨識並移除中毒樣本。首先,透過多尺度輸入擾動評估樣本的跨層特徵穩定性;接著,利用中毒樣本在輸入擾動下具有較高特徵穩定性的特性,以各類別中的低穩定性樣本建立乾淨參考分布;再衡量高穩定性樣本相對於參考分布的跨層特徵偏離程度,藉此鎖定可能遭受攻擊的目標類別;最後,於目標類別內評估各樣本的特徵分布異常程度,完成中毒樣本的分離。透過先定位目標類別,再進行類別內中毒樣本辨識,BlindClean 能在有效辨識中毒樣本的同時,降低正常樣本的誤判率。
    實驗結果顯示,BlindClean 在三種主流乾淨標籤後門攻擊下皆展現出優異的資料清洗能力。在上游中毒樣本辨識任務方面,相較於現有最先進的資料清洗方法,BlindClean 不僅將真陽性率提升最高超過 90%,更同時將假陽性率降低最多60%。而在下游模型訓練任務方面,BlindClean 是唯一能將所有攻擊的平均攻擊成功率壓制於 1% 以下,且最高可將模型準確率提升 44.4% 的方法。上述結果顯示,本論文所提出之 BlindClean 能在不需先驗知識的情況下,有效兼顧資料清洗能力與下游模型效能,為對抗乾淨標籤後門攻擊提供一套實用的資料清洗方法。

    With the rapid development of artificial intelligence, model developers increasingly rely on third-party datasets to train machine learning models. This practice increases the diversity of training data and reduces the cost of data collection and annotation. However, it also makes trained models more vulnerable to various backdoor attacks, creating an urgent need for effective data sanitization methods to reduce security risks. Among existing backdoor attacks, clean-label backdoor attacks are particularly stealthy because poisoned samples retain labels consistent with their original classes, making them difficult to distinguish from clean samples and remove from the training dataset. In addition, when defending against clean-label backdoor attacks, existing data sanitization methods often suppress the attack success rate at the cost of a significant decrease in model accuracy or rely on prior knowledge, such as clean reference data or attack-related information. These limitations greatly reduce their applicability.
    To address these challenges, this research proposes BlindClean, a data sanitization method against clean-label backdoor attacks. BlindClean can effectively identify and remove poisoned samples without relying on prior knowledge or significantly reducing model accuracy. First, multi-scale input perturbations are applied to evaluate the cross-layer feature stability of each sample. Then, based on the observation that poisoned samples exhibit higher feature stability under input perturbations, low stability samples from each class are used to construct clean reference distributions. Next, the cross-layer feature distribution deviation of high stability samples from the corresponding reference distributions is aggregated to determine the attacked target class. Finally, each sample within the identified target class is evaluated according to its deviation from the clean reference distribution to distinguish poisoned samples from clean samples.
    Experimental results demonstrate that BlindClean achieves strong data sanitization performance against three representative clean-label backdoor attacks. For poisoned sample detection, compared with existing state-of-the-art data sanitization methods, BlindClean achieves an improvement of up to 90% in the true positive rate while reducing the false positive rate by up to 60%. For downstream model training, BlindClean is the only method that reduces the average attack success rate across all evaluated attacks to less than 1% while improving model accuracy by up to 44.4%. These results demonstrate that, without relying on prior knowledge, BlindClean achieves strong data sanitization performance while maintaining downstream model performance. Therefore, BlindClean provides a practical method for defending against clean-label backdoor attacks.

    CHAPTER 1 INTRODUCTION 1 1.1 Background 2 1.2 Motivation 9 1.3 Comparison and Contributions 15 1.4 Organization 19 CHAPTER 2 RELATED WORK 20 2.1 Clean-Label Backdoor Attacks 21 2.2 Data Sanitization Methods 26 CHAPTER 3 BLINDCLEAN: A PRIOR-KNOWLEDGE-FREE DATA SANITIZATION METHOD AGAINST CLEAN-LABEL BACKDOOR ATTACKS 34 3.1 System Model and Threat Model 35 3.2 Framework of BlindClean 38 3.3 Multi-scale Perturbed Sample Generator 41 3.4 Perturbation-based Feature Stability Evaluation 45 3.5 Layer-based Target Class Determination 49 3.6 Target-class-aware Poison Sample Detection 58 3.7 Algorithm 63 CHAPTER 4 EXPERIMENTAL RESULTS 64 4.1 Experimental Settings 65 4.2 Overall Performance Analysis 71 4.3 Performance under Different Conditions 76 4.4 Ablation Studies 83 4.5 Computational Cost Analysis 93 CHAPTER 5 CONCLUSION 95 CHAPTER 6 FUTURE WORK 98 REFERENCES 99

    [1] A. E. Cinà et al., "Wild patterns reloaded: A survey of machine learning security against training data poisoning," ACM Computing Surveys, vol. 55, no. 13s, pp. 1–39, 2023.
    [2] S. Zhao, X. Ma, X. Zheng, J. Bailey, J. Chen, and Y.-G. Jiang, "Clean-label backdoor attacks on video recognition models," in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 14443–14452.
    [3] M. Goldblum et al., "Dataset security for machine learning: Data poisoning, backdoor attacks, and defenses," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 2, pp. 1563–1580, 2022.
    [4] Z. Wang, J. Ma, X. Wang, J. Hu, Z. Qin, and K. Ren, "Threats to training: A survey of poisoning attacks and defenses on machine learning systems," ACM Computing Surveys, vol. 55, no. 7, pp. 1–36, 2022.
    [5] P. Zhao, W. Zhu, P. Jiao, D. Gao, and O. Wu, "Data poisoning in deep learning: A survey," arXiv preprint arXiv:2503.22759, 2025.
    [6] Y. Hu et al., "Artificial intelligence security: Threats and countermeasures," ACM Computing Surveys (CSUR), vol. 55, no. 1, pp. 1–36, 2021.
    [7] A. E. Cinà, K. Grosse, A. Demontis, B. Biggio, F. Roli, and M. Pelillo, "Machine learning security against data poisoning: Are we there yet?," Computer, vol. 57, no. 3, pp. 26–34, 2024.
    [8] Y. Li, Y. Jiang, Z. Li, and S.-T. Xia, "Backdoor learning: A survey," IEEE transactions on neural networks and learning systems, vol. 35, no. 1, pp. 5–22, 2022.
    [9] Y. Bai et al., "Backdoor attack and defense on deep learning: A survey," IEEE Transactions on Computational Social Systems, vol. 12, no. 1, pp. 404–434, 2024.
    [10] T. Gu, K. Liu, B. Dolan-Gavitt, and S. Garg, "Badnets: Evaluating backdooring attacks on deep neural networks," Ieee Access, vol. 7, pp. 47230–47244, 2019.
    [11] X. Chen, C. Liu, B. Li, K. Lu, and D. Song, "Targeted backdoor attacks on deep learning systems using data poisoning," arXiv preprint arXiv:1712.05526, 2017.
    [12] A. Turner, D. Tsipras, and A. Madry, "Label-consistent backdoor attacks," arXiv preprint arXiv:1912.02771, 2019.
    [13] H. Souri, L. Fowl, R. Chellappa, M. Goldblum, and T. Goldstein, "Sleeper agent: Scalable hidden trigger backdoors for neural networks trained from scratch," Advances in Neural Information Processing Systems, vol. 35, pp. 19165–19178, 2022.
    [14] Y. Zeng, M. Pan, H. A. Just, L. Lyu, M. Qiu, and R. Jia, "Narcissus: A practical clean-label backdoor attack with limited information," in Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security, 2023, pp. 771–785.
    [15] W. Ma, D. Wang, R. Sun, M. Xue, S. Wen, and Y. Xiang, "The" Beatrix''Resurrections: Robust Backdoor Detection via Gram Matrices," in The Network and Distributed System Security Symposium (NDSS), 2023.
    [16] L. Hou, W. Luo, Z. Hua, S. Chen, L. Y. Zhang, and Y. Li, "Flare: Towards universal dataset purification against backdoor attacks," IEEE Transactions on Information Forensics and Security, 2025.
    [17] Y. Li, X. Lyu, N. Koren, L. Lyu, B. Li, and X. Ma, "Anti-backdoor learning: Training clean models on poisoned data," Advances in Neural Information Processing Systems, vol. 34, pp. 14900–14912, 2021.
    [18] M. Pan, Y. Zeng, L. Lyu, X. Lin, and R. Jia, "{ASSET}: Robust backdoor data detection across a multiplicity of deep learning paradigms," in 32nd USeNIX security symposium (USeNIX security 23), 2023, pp. 2725–2742.
    [19] S. Pal, Y. Yao, R. Wang, B. Shen, and S. Liu, "Backdoor secrets unveiled: Identifying backdoor data with optimized scaled prediction consistency," in International conference on learning representations, 2024, vol. 2024, pp. 14061–14083.
    [20] L. Hou, R. Feng, Z. Hua, W. Luo, L. Y. Zhang, and Y. Li, "Ibd-psc: Input-level backdoor detection via parameter-oriented scaling consistency," in International Conference on Machine Learning, 2024, vol. 235, pp. 18992–19022.
    [21] L. Yasur, T. Lederer, L. Rokach, and Y. Mirsky, "Clean-Label Backdoor Attacks: A Survey," IEEE Access, 2026.
    [22] I. J. Goodfellow et al., "Generative adversarial nets," Advances in neural information processing systems, vol. 27, 2014.
    [23] A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu, "Towards deep learning models resistant to adversarial attacks," in International conference on learning representations, 2018.
    [24] J. Geiping et al., "Witches' brew: Industrial scale data poisoning via gradient matching," in International conference on learning representations, 2021.
    [25] A. Saha, A. Subramanya, and H. Pirsiavash, "Hidden trigger backdoor attacks," in Proceedings of the AAAI conference on artificial intelligence, 2020, vol. 34, no. 07, pp. 11957–11965.
    [26] M. A. Hanif, N. Chattopadhyay, B. Ouni, and M. Shafique, "Survey on backdoor attacks on deep learning: Current trends, categorization, applications, research challenges, and future prospects," IEEE Access, 2025.
    [27] O. Mengara, A. Avila, and T. H. Falk, "Backdoor attacks to deep neural networks: A survey of the literature, challenges, and future research directions," IEEE Access, vol. 12, pp. 29004–29023, 2024.
    [28] B. Chen et al., "Detecting backdoor attacks on deep neural networks by activation clustering," arXiv preprint arXiv:1811.03728, 2018.
    [29] J. Hayase, W. Kong, R. Somani, and S. Oh, "Spectre: Defending against backdoor attacks using robust statistics," in International Conference on Machine Learning, 2021: PMLR, pp. 4129–4139.
    [30] B. Tran, J. Li, and A. Madry, "Spectral signatures in backdoor attacks," Advances in neural information processing systems, vol. 31, 2018.
    [31] X. Qi, T. Xie, J. T. Wang, T. Wu, S. Mahloujifar, and P. Mittal, "Towards a proactive {ML} approach for detecting backdoor poison samples," in 32nd USENIX Security Symposium (USENIX Security 23), 2023, pp. 1685–1702.
    [32] J. Guo, Y. Li, X. Chen, H. Guo, L. Sun, and C. Liu, "Scale-up: An efficient black-box input-level backdoor detection via analyzing scaled prediction consistency," in International conference on learning representations, 2023.
    [33] M. Hu, Z. Guan, J. Guo, Z. Zhou, J. Zhang, and S. Li, "BBCaL: Black-box Backdoor Detection under the Causality Lens," Transactions on Machine Learning Research, 2024.
    [34] S. Danafar, P. Rancoita, T. Glasmachers, K. Whittingstall, and J. Schmidhuber, "Testing hypotheses by regularized maximum mean discrepancy," International Journal of Computer and Information Technology, vol. 3, no. 2, pp. 223–232, 2014.
    [35] L. McInnes, J. Healy, and J. Melville, "Umap: Uniform manifold approximation and projection for dimension reduction," arXiv preprint arXiv:1802.03426, 2018.
    [36] R. J. Campello, D. Moulavi, A. Zimek, and J. Sander, "Hierarchical density estimates for data clustering, visualization, and outlier detection," ACM Transactions on Knowledge Discovery from Data (TKDD), vol. 10, no. 1, pp. 1–51, 2015.
    [37] Y. Roh, G. Heo, and S. E. Whang, "A survey on data collection for machine learning: a big data-ai integration perspective," IEEE Transactions on Knowledge and Data Engineering, vol. 33, no. 4, pp. 1328–1347, 2019.
    [38] A. Krizhevsky and G. Hinton, "Learning multiple layers of features from tiny images," 2009.
    [39] Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng, "Reading digits in natural images with unsupervised feature learning," in NIPS workshop on deep learning and unsupervised feature learning, 2011, vol. 2011, no. 2: Granada, p. 4.
    [40] B. Wu et al., "Backdoorbench: A comprehensive benchmark of backdoor learning," Advances in Neural Information Processing Systems, vol. 35, pp. 10546–10559, 2022.
    [41] Y. Li, M. Ya, Y. Bai, Y. Jiang, and S.-T. Xia, "Backdoorbox: A python toolbox for backdoor learning," arXiv preprint arXiv:2302.01762, 2023.

    QR CODE