簡易檢索 / 詳目顯示

研究生: 陳麒為
Chen, Chi-Wei
論文名稱: 殘差特徵對齊與陰影增強於文字驅動之擴散式物件移除
Residual Feature Alignment and Shadow Enhancement for Text-Driven Diffusion-Based Object Removal
指導教授: 楊家輝
Yang, Jar-Ferr
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 電腦與通信工程研究所
Institute of Computer & Communication Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 67
中文關鍵詞: 擴散模型文字驅動物件消除穩定擴散殘差特徵對齊陰影增強
外文關鍵詞: Diffusion Models, Text-Driven Object Removal, Stable Diffusion, Residual MLP, Shadow Enhancement
相關次數: 點閱:43下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 近年來,擴散模型(Diffusion Models)於影像生成與影像修復領域展現優異表現,其中物件移除(Object Removal)已成為電腦視覺的重要研究方向。然而,現有方法多仰賴人工標註遮罩,限制了實際應用的便利性。此外,戶外場景中物件陰影常於移除後仍殘留於影像,降低修復結果的自然度與視覺品質。
    本研究提出一套基於擴散模型之文字驅動物件移除方法,以 CLIPAway [13] 為基礎架構,結合 CLIPSeg [11]、AlphaCLIP [10]、IP-Adapter [12] 與 Stable Diffusion Inpainting,實現僅透過文字即可完成物件移除。首先,利用 CLIPSeg [11] 根據文字提示自動產生物件遮罩,取代人工標註流程。為解決戶外場景陰影殘留問題,本研究提出陰影強化模組,利用亮度分析、區域生長及幾何限制,自動擴展遮罩範圍以涵蓋與目標物件相連之投影陰影,提升修復影像的一致性。此外,考量 AlphaCLIP [10] 與 IP-Adapter [12] 特徵維度差異,本研究提出殘差多層次映射(Residual MLP)模組,在凍結預訓練模型權重下,僅訓練特徵映射模組,以提升特徵對齊能力,並降低訓練成本與 GPU 記憶體需求。

    In recent years, diffusion models have demonstrated remarkable performance in image generation and image restoration, making object removal an important research topic in computer vision. However, most existing methods rely on manually annotated masks, limiting their practicality in real-world applications. Moreover, in outdoor scenes, shadows cast by target objects often remain after object removal, degrading the visual quality of restored images.
    This thesis proposes a text-driven diffusion-based object removal framework built upon CLIPAway [13] by integrating CLIPSeg [11], AlphaCLIP [10], IP-Adapter [12], and Stable Diffusion Inpainting. The proposed framework enables object removal using only natural language descriptions without manually annotated masks. Specifically, CLIPSeg [11] automatically generates object masks from text prompts. To address shadow artifacts in outdoor scenes, a shadow enhancement module expands the object mask by incorporating adjacent shadow regions through luminance analysis, region growing, and geometric constraints, improving the completeness and visual consistency of restored images. Furthermore, to bridge the feature space discrepancy between AlphaCLIP [10] and IP-Adapter [12], a residual multi-layer perceptron (MLP) feature alignment module is introduced. By freezing the pre-trained models and optimizing only the feature mapping network, the proposed method improves feature alignment while reducing training cost and GPU memory consumption.

    摘要 I Abstract II 誌謝 III Contents IV List of Tables VII List of Figures VIII Chapter 1 Introduction 1 1.1 Research Background 1 1.2 Motivations 2 1.3 Research Contributions 3 Chapter 2 Related Work 5 2.1 Diffusion-based Image Inpainting 5 2.1.1 Denoising Diffusion Probabilistic Models (DDPM) 6 2.1.2 Latent Diffusion Models (LDM) 7 2.1.3 Diffusion-based Image Inpainting 8 2.2 Object Removal 9 2.3 Vision-Language Models 9 2.4 Text-Driven Object Removal 11 2.4.1 InstructPix2Pix 11 2.4.2 Inst-inpaint 11 2.4.3 CLIPAway 12 Chapter 3 The Proposed Method 13 3.1 Overview Framework 14 3.2 Baseline Architecture 16 3.2.1 AlphaCLIP 17 3.2.2 Projection Block 18 3.2.3 IP-Adapter 21 3.3 Propose Method 23 3.3.1 Text-Guided Mask Generation 23 3.3.2 Shadow Enhancement 26 3.3.3 Residual MLP 31 Chapter 4 Experiment Results 36 4.1 Environment Setup and Datasets 37 4.1.1 Experimental Environment 37 4.1.2 Datasets 37 4.2 Evaluation Metrics 38 4.3 Comparison with Other Methods 40 4.4 Ablation Study 45 4.4.1 Shadow Enhancement 45 4.5 Visualization of Prediction Results 47 4.5.1 GQA-Inpaint Dataset 47 4.5.2 OBER-Test and RORD 50 4.5.3 OBER-Wild 51 Chapter 5 Conclusions 53 Chapter 6 Future Work 54 References 55

    [1] Ho, J., Jain, A., Abbeel, P. Denoising Diffusion Probabilistic Models. Advances in Neural Information Processing Systems (NeurIPS), Vol. 33, pp. 6840–6851, 2020.
    [2] Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B. High-Resolution Image Synthesis with Latent Diffusion Models. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10684–10695, 2022.
    [3] Nichol, A. Q., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., Chen, M. GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models. Proceedings of the International Conference on Machine Learning (ICML), pp. 16784–16804, 2022.
    [4] Lugmayr, A., Danelljan, M., Romero, A., Yu, F., Van Gool, L., Timofte, R. RePaint: Inpainting Using Denoising Diffusion Probabilistic Models. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 11461–11471, 2022.
    [5] Sun, W., Cui, B., Dong, X.-M., Tang, J. Attentive Eraser: Unleashing Diffusion Model's Object Removal Potential via Self-Attention Redirection Guidance. Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), 2025.
    [6] Chen, Z., Wang, W., Yang, Z., Yuan, Z., Chen, H., Shen, C. FreeCompose: Generic Zero-Shot Image Composition with Diffusion Prior. Proceedings of the European Conference on Computer Vision (ECCV), 2024.
    [7] Zhao, J., Wang, Z., Yang, P., Zhou, S. ObjectClear: Precise Object and Effect Removal with Adaptive Target-Aware Attention. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026.
    [8] [8] Yildirim, A. B., Baday, V., Erdem, E., Erdem, A., Dundar, A. Inst-Inpaint: Instruction-Guided Image Inpainting. arXiv preprint, 2023.
    [9] Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I. Learning Transferable Visual Models From Natural Language Supervision. Proceedings of the 38th International Conference on Machine Learning (ICML), Vol. 139, pp. 8748–8763, 2021.
    [10] Sun, X., Zhang, Q., Shi, Y., Wei, Y., Gao, Y., et al. Alpha-CLIP: A CLIP Model Focusing on Wherever You Want. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7867–7876, 2024.
    [11] Lüddecke, T., Ecker, A. S. Image Segmentation Using Text and Image Prompts. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7086–7096, 2022.
    [12] Ye, H., Zhang, J., Liu, S., Han, X., Yang, W. IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models. arXiv preprint arXiv:2308.06721, 2023.
    [13] Yildirim, A. B., Ekin, Y., Caglar, E. E., Erdem, A., Erdem, E., Dundar, A. CLIPAway: Harmonizing Focused Embeddings for Removing Objects via Diffusion Models. Advances in Neural Information Processing Systems (NeurIPS), Vol. 37, 2024.
    [14] Brooks, T., Holynski, A., Efros, A. A. InstructPix2Pix: Learning to Follow Image Editing Instructions. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 18392–18402, 2023.
    [15] Zou, X., Yang, J., Zhang, Z., Li, Y., Liu, H., Huang, J.-B., Loy, C. C. X-Decoder: Generalized Decoding for Pixel, Image, and Language. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 12786–12797, 2023.
    [16] Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., Rombach, R. SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis. arXiv preprint arXiv:2307.01952, 2023.
    [17] Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C. L. Microsoft COCO: Common Objects in Context. Proceedings of the European Conference on Computer Vision (ECCV), pp. 740–755, 2014.
    [18] Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al. Segment Anything. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 4015–4026, 2023.
    [19] Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., Hochreiter, S. GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. Advances in Neural Information Processing Systems (NeurIPS), Vol. 30, 2017.
    [20] Hessel, J., Holtzman, A., Forbes, M., Le Bras, R., Choi, Y. CLIPScore: A Reference-Free Evaluation Metric for Image Captioning. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 7514–7528, 2021.
    [21] Zhang, R., Isola, P., Efros, A. A., Shechtman, E., Wang, O. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 586–595, 2018.
    [22] Chandrasekar, A., Chakrabarty, G., Bardhan, J., Hebbalaguppe, R., Prathosh, A. P. ReMOVE: A Reference-Free Metric for Object Erasure. arXiv preprint arXiv:2409.00707, 2024.
    [23] Sagong, M. C., Yeo, Y. J., Jung, S. W., Ko, S. J. RORD: A Real-world Object Removal Dataset. Proceedings of the 33rd British Machine Vision Conference (BMVC), 2022.

    下載圖示
    校外:立即公開
    QR CODE