簡易檢索 / 詳目顯示

研究生: 游家達
Yu, Chia-Ta
論文名稱: 基於提示引導與空間交叉注意力對齊之異常感知 擴散模型用於多樣化少樣本工業異常生成
Prompt-Guided Anomaly-Aware Diffusion with Spatial Cross Attention Alignment for Diverse Few-Shot Industrial Anomaly Generation
指導教授: 蔡家齊
Tsai, Chia-Chi
學位類別: 碩士
Master
系所名稱: 智慧半導體及永續製造學院 - 晶片設計學位學程
Program on Integrated Circuit Design
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 129
中文關鍵詞: 工業異常檢測 、少樣本異常影像生成 、擴散模型 、機器學習
外文關鍵詞: Industrial Anomaly Detection, Few-Shot Anomaly Image Generation, Diffusion Models, Machine Learning
相關次數: 點閱:92  下載:2 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 工業異常檢測在現代製造流程中扮演重要角色,其目標在於自動辨識產品瑕疵並定位異常區域。然而,在實際工業場景中,異常樣本通常具有數量稀少、型態多變且標註成本高等特性,使得監督式異常檢測與定位模型難以取得足夠的訓練資料。為了緩解異常資料不足的問題,近年來許多研究開始透過生成模型合成異常影像與對應遮罩,以提供下游異常檢測、定位與分類任務所需的額外監督資料。
    近年擴散模型在影像生成任務中展現出優異的生成品質與文字控制能力,並逐漸被應用於少樣本工業異常影像生成。然而,在文字條件式擴散模型中,文字 token 與影像空間區域之間的對應關係通常是由模型隱式學習而來。在少樣本異常生成情境下,由於瑕疵區域往往具有小範圍、稀疏且局部化的特性,異常相關 token 可能無法穩定對齊真實瑕疵區域,進而造成生成異常影像與遮罩之間的空間不一致,並降低生成結果的可控性。
    為了解決上述問題,本研究提出一種 anomaly-aware diffusion fine-tuning framework,以強化文字語意與空間異常區域之間的對應關係。首先,本研究根據異常遮罩將輸入異常影像分解為完整異常影像表示與異常區域聚焦表示,使模型能同時學習全域外觀與局部瑕疵特徵。其次,本研究設計可控制的語意提示詞,將正常內容與異常內容分別對應至不同的 learnable tokens,並使用多個異常 token 以表示多樣化的瑕疵外觀。最後,本研究引入 token-specific cross-attention alignment loss,直接約束異常相關 token 的 cross-attention map 對齊瑕疵遮罩,同時使正常相關 token 聚焦於非異常區域,以降低正常與異常語意之間的混淆。
    在模型訓練上,本研究採用 LoRA 進行參數高效微調,凍結預訓練 U-Net 與 text encoder 的原始權重,僅更新插入的 LoRA adapters 與 learnable prompt embeddings,以降低少樣本情境下的訓練成本與過擬合風險。實驗部分,本研究於 MVTec AD [28]資料集上進行評估,並透過生成品質、異常定位效能與消融實驗分析所提出之 prompt design 與 attention alignment supervision 的有效性。實驗結果顯示,本研究方法能改善異常語意與瑕疵區域之間的空間對應關係,並提升生成異常影像與遮罩的對齊品質。

    Industrial defect detection and localization plays an important role in modern manufacturing systems by automatically identifying defective products and localizing anomalous regions. However, anomalous samples in real-world industrial scenarios are typically scarce, visually diverse, and expensive to annotate, which limits the training of supervised anomaly inspection models, particularly when pixel-level annotations are required. To alleviate this limitation, recent studies have explored anomaly image generation to synthesize anomalous images together with associated masks for downstream training.
    Although diffusion models provide strong image generation quality and flexible text controllability, the correspondence between textual tokens and spatial image regions is usually learned implicitly. Under few-shot anomaly generation settings, defect regions are often small, sparse, and highly localized, making anomaly-related tokens difficult to consistently associate with actual defect regions. This may result in semantic leakage and spatial inconsistency between generated anomalies and their associated masks.
    To address this issue, this thesis proposes an anomaly-aware diffusion fine-tuning framework that explicitly strengthens the semantic correspondence between textual tokens and spatial anomaly regions. For each anomalous sample and its mask, the image is decomposed into a global anomalous representation and an anomaly-focused representation, allowing the model to jointly learn overall appearance and localized defect characteristics. A controllable semantic prompt design is further introduced to assign normal-content and anomaly-related concepts to separate learnable tokens, while multiple anomaly tokens are employed to represent diverse defect appearances. In addition, a token-specific cross-attention alignment loss explicitly guides anomaly-related attention maps toward defect regions and normal-related tokens toward non-defective regions, thereby reducing semantic leakage and improving spatial controllability. For parameter-efficient adaptation, the pretrained U-Net and text encoder are frozen, and only LoRA adapters and learnable prompt embeddings are optimized.
    Experiments on the MVTec AD dataset demonstrate the effectiveness of the proposed framework. The method obtains a mean Inception Score of 1.97 and an IC-LPIPS score of 0.37. For downstream anomaly inspection, it obtains an average image-level average precision of 99.51%, a pixel-level average precision of 84.2%, and a pixel-level F1-max score of 78.5%. These results demonstrate that explicit cross-attention alignment improves the spatial correspondence between anomaly semantics and defect regions and produces more reliable synthesized anomaly images and their associated masks for downstream Industrial defect detection and localization.

    摘要 iii ABSTRACT v 致謝 vii Content viii List of Tables xi List of Figures xii Chapter 1 Introduction 1 1.1 Motivation 1 1.2 Thesis Contributions 4 1.3 Thesis Organization 5 Chapter 2 Background and Related Work 7 2.1 Industrial defect detection and localization 8 2.2 Diffusion Models 13 2.2.1 Denoising Diffusion Models 14 2.2.2 Latent Diffusion Models 16 2.2.3 Text-conditioned Diffusion and Cross-attention 19 2.3 Anomaly Image Generation 22 2.3.1 Hand-crafted and Self-supervised Anomaly Synthesis 23 2.3.2 GAN-based Anomaly Generation 26 2.3.3 Diffusion-based Anomaly Generation 28 2.4 Prompt-based Diffusion Customization 33 2.5 Summary 36 Chapter 3 Problems Analysis and Design Methodology 37 3.1 Preliminaries 37 3.2 Overview of the Proposed Framework 40 3.3 Anomaly-aware Data Decomposition 43 3.4 Controllable Semantic Prompt Design 45 3.5 Paired Feature Interaction 47 3.6 Attention Alignment Loss 50 3.7 Training Objective and Inference 54 Chapter 4 Experimental Results Comparison and Evaluation 56 4.1 Experiment Setup 56 4.1.1 Datasets 57 4.1.2 Implementation Details 59 4.1.3 Evaluation Protocols 60 4.1.4 Downstream Models and Evaluation Metrics 61 4.2 Generation Quality Evaluation 62 4.3 Generated-only Evaluation 63 4.3.1 Anomaly Classification 64 4.3.2 Anomaly Detection and Localization 66 4.4 Mixed Training with Real Defects 73 4.4.1 Anomaly Classification 74 4.4.2 Anomaly Detection and Localization 75 4.5 Ablation Study 78 4.5.1 Effect of Prompt Design 78 4.5.2 Effect of Attention Alignment Loss 80 4.5.3 Discussion 82 4.6 Qualitative Analysis 83 Chapter 5 Conclusion and Future Work 86 5.1 Conclusion 86 5.2 Future Work 88 References 90 Appendix 95 A.1 Background Refinement Exploration 95 A.1.1 Motivation 95 A.1.2 Refinement Method 96 A.1.3 Qualitative Results 99 A.2 Experiment on different dataset 101 A.3 Experiment on different few-shot setting 103 A.4 MVTec AD Generated Results 104

    [1] T. Schlegl, P. Seebock, S. M. Waldstein, U. Schmidt-Erfurth, and G. Langs, “Unsupervised anomaly detection with generative adversarial networks to guide marker discovery,” in Proceedings of the International Conference on Information Processing in Medical Imaging, 2017.
    [2] T. Schlegl, P. Seebock, S. M. Waldstein, G. Langs, and U. Schmidt-Erfurth, “f-anogan: Fast unsupervised anomaly detection with generative adversarial networks,” Medical Image Analysis, vol. 54, pp. 30–44, 2019.
    [3] V. Zavrtanik, M. Kristan, and D. Skocaj, “Draem–a discriminatively trained reconstruction embedding for surface anomaly detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 8330–8339.
    [4] K. Roth, L. Pemula, J. Zepeda, B. Scholkopf, T. Brox, and P. Gehler, “Towards total recall in industrial anomaly detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
    [5] H. Deng and X. Li, “Anomaly detection via reverse distillation from one-class embedding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
    [6] S. Lee, S. Lee, and B. C. Song, “Cfa: Coupled-hypersphere-based feature adaptation for target-oriented anomaly localization,” IEEE Access, vol. 10, pp. 78 446–78 454, 2022.
    [7] H. M. Schluter, J. Tan, B. Hou, and B. Kainz, “Natural synthetic anomalies for self-supervised anomaly detection and localization,” in Proceedings of the European Conference on Computer Vision. Springer, 2022, pp. 474–489.
    [8] C.-L. Li, K. Sohn, J. Yoon, and T. Pfister, “Cutpaste: Self-supervised learning for anomaly detection and localization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 9664–9674.
    [9] D. Lin, Y. Cao, W. Zhu, and Y. Li, “Few-shot defect segmentation leveraging abundant defect-free training samples through normal background regularization and crop-and-paste operation,” in Proceedings of the IEEE International Conference on Multimedia and Expo, 2021, pp. 1–6.
    [10] S. Niu, B. Li, X. Wang, and H. Lin, “Defect image sample generation with gan for improving defect recognition,” IEEE Transactions on Automation Science and Engineering, vol. 17, no. 3, pp. 1611–1622, 2020.
    [11] G. Zhang, K. Cui, T.-Y. Hung, and S. Lu, “Defectgan: High-fidelity defect synthesis for automated defect inspection,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2021, pp. 2524–2534.
    [12] Y. Duan, Y. Hong, L. Niu, and L. Zhang, “Few-shot defect image generation via defect-aware feature manipulation,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, 2023, pp. 571–578.
    [13] T. Hu, J. Zhang, R. Yi, Y. Du, X. Chen, L. Liu, Y. Wang, and C. Wang, “Anomalydiffusion: Few-shot anomaly image generation with diffusion model,” in Proceedings of the AAAI Conference on Artificial Intelligence, 2024.
    [14] Y. Jin, J. Peng, Q. He, T. Hu, J. Wu, H. Chen, H. Wang, W. Zhu, M. Chi, J. Liu, and Y. Wang, “Dual-interrelated diffusion model for few-shot anomaly image generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025.
    [15] Z. Dai, S. Zeng, H. Liu, X. Li, F. Xue, and Y. Zhou, “Seas: Few-shot industrial anomaly image generation with separation and sharing finetuning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025.
    [16] J. Choi, M. Kim, and J. H. Hong, “Magic: Few-shot mask-guided anomaly inpainting with prompt perturbation, spatially adaptive guidance, and context awareness,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026, cVPR 2026 Findings.
    [17] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 10 684–10 695.
    [18] R. Gal, Y. Alaluf, Y. Atzmon, O. Patashnik, A. H. Bermano, G. Chechik, and D. Cohen-Or, “An image is worth one word: Personalizing text-to-image generation using textual inversion,” in Proceedings of the International Conference on Learning Representations, 2023.
    [19] N. Ruiz, Y. Li, V. Jampani, Y. Pritch, M. Rubinstein, and K. Aberman, “Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 22 500–22 510.
    [20] J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Advances in Neural Information Processing Systems, 2020.
    [21] J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in Proceedings of the International Conference on Learning Representations, 2021.
    [22] D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” arXiv preprint arXiv:1312.6114, 2013.
    [23] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in Proceedings of the International Conference on Machine Learning. PMLR, 2021, pp. 8748–8763.
    [24] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” arXiv preprint arXiv:2106.09685, 2021.
    [25] A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollar, and R. Girshick, “Segment anything,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 4015–4026.
    [26] X. Qin, Z. Zhang, C. Huang, M. Dehghan, O. R. Zaiane, and M. Jagersand, “U2-net: Going deeper with nested u-structure for salient object detection,” Pattern Recognition, vol. 106, p. 107404, 2020.
    [27] P. Bergmann, M. Fauser, D. Sattlegger, and C. Steger, “Mvtec ad–a comprehensive real-world dataset for unsupervised anomaly detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 9592–9600.
    [28] R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 586–595.
    [29] S. Barratt and R. Sharma, “A note on the inception score,” arXiv preprint arXiv:1801.01973, 2018.
    [30] O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015. Springer, 2015, pp. 234–241.
    [31] T. Defard, A. Setkov, A. Loesch, and R. Audigier, “PaDiM: A patch distribution modeling framework for anomaly detection and localization,” in Proceedings of the International Conference on Pattern Recognition Workshops, 2021, pp. 475–489.
    [32] N. Cohen and Y. Hoshen, “Sub-image anomaly detection with deep pyramid correspondences,” arXiv preprint arXiv:2005.02357, 2020.
    [33] Y. Yu, Y. Shin, M. Lee, and S. Lee, “FastFlow: Unsupervised anomaly detection and localization via 2D normalizing flows,” arXiv preprint arXiv:2111.07677, 2021.
    [34] Z. Liu, Y. Zhou, Y. Xu, and Z. Wang, “SimpleNet: A simple network for image anomaly detection and localization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 20402–20411.
    [35] A. Hertz, R. Mokady, J. Tenenbaum, K. Aberman, Y. Pritch, and D. Cohen-Or, “Prompt-to-Prompt image editing with cross-attention control,” in Proceedings of the International Conference on Learning Representations, 2023.
    [36] H. Chefer, Y. Alaluf, Y. Vinker, L. Wolf, and D. Cohen-Or, “Attend-and-Excite: Attention-based semantic guidance for text-to-image diffusion models,” ACM Transactions on Graphics, vol. 42, no. 4, pp. 1–10, 2023.
    [37] R. Tang, L. Liu, A. Pandey, Z. Jiang, G. Yang, K. Kumar, P. Stenetorp, J. Lin, and F. Ture, “What the DAAM: Interpreting Stable Diffusion using cross attention,” in Proceedings of the Annual Meeting of the Association for Computational Linguistics, 2023, pp. 5644–5659.
    [38] Y. Li, M. Keuper, D. Zhang, and A. Khoreva, “Divide & Bind your attention for improved generative semantic nursing,” in Proceedings of the British Machine Vision Conference, 2023.
    [39] N. Kumari, B. Zhang, R. Zhang, E. Shechtman, and J.-Y. Zhu, “Multi-concept customization of text-to-image diffusion,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 1931–1941.
    [40] L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 3836–3847.
    [41] C. Mou, X. Wang, L. Xie, Y. Wu, J. Zhang, Z. Qi, Y. Shan, and X. Qie, “T2I-Adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 5, 2024, pp. 4296–4304.
    [42] O. Avrahami, O. Fried, and D. Lischinski, “Blended latent diffusion,” ACM Transactions on Graphics, vol. 42, no. 4, pp. 1–11, 2023.

    下載圖示
    校外:立即公開
    QR CODE