簡易檢索 / 詳目顯示

研究生: 林晏緯
Lin, Yen-Wei
論文名稱: 基於條件式生成對抗網路之模糊圖像分割模型
An Image Segmentation Model Based on Conditional Generative Adversarial Networks for Blurry Image
指導教授: 賴槿峰
Lai, Chin-Feng
學位類別: 碩士
Master
系所名稱: 工學院 - 工程科學系
Department of Engineering Science
論文出版年: 2021
畢業學年度: 109
語文別: 中文
論文頁數: 52
中文關鍵詞: 條件式生成對抗網路 、圖像分割 、模糊圖像
外文關鍵詞: Conditional Generative Adversarial Networks, Image Segmentation, Blurry Image
相關次數: 點閱:195  下載:0 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 圖像分割為機器視覺中重要的基礎之一,可做為其他研究領域的前處理過程。但當
    圖像因為擷取設備晃動或是老舊問題導致圖像模糊時,有可能造成目標物件的輪廓
    容於背景導致分割不準確的問題。
    在本論文當中,我們提出一種編碼器解碼器架構作為圖像分割生成器,期望可以在
    不經過任何清晰化的前處理方法下為模糊的圖像進行分割。相較於傳統使用編碼器
    解碼器架構,我們搭配 Concatentation 及 Summation 加強編碼跟解碼的過程的特徵傳
    遞。資料集採用現有公開的資料集進行均值模糊、水平/垂直動態模糊,模擬相機失
    焦或是快門過慢導致動態模糊發生的情形,並搭配提出的生成器架構進行條件式生
    成對抗網路訓練。
    在實驗結果當中,本研究使用 IoU 作為圖像分割結果的評估指標。於實驗一的部分,
    本論文擷取數張測試集中的圖像,在模糊情況下皆保持跟清晰原圖的分割結果有相
    同水準。並於實驗二當中將本研究的實驗模型與 U-Net 生成器及 DeepLab 模型做比
    較並利用長條圖呈現結果分布,從實驗結果可以發現本研究實驗模型在模糊程度的
    增長下,低 IoU 的數量並沒有劇烈增加。而且本實驗模型在測試集的測試下沒有發
    生過度擬合的情形,因此可以證明此實驗模型的通用性。

    In this paper, we propose a image segmentation method based on Conditional Generative Adversarial Networks to improve the problem of segmenting inaccurately due to blurry image, which mean the outline of target object in the image are similar to the background or other object. This issue leads to errors in mainstream image segmentation model. In order to reduce this problem, we use an encoder-decoder architecture to expect image segmentation without any pre-processing methods. Compared with the traditional encoder and decoder architecture, we use Concatentation and Summation to enhance the feature transfer of the encoding and decoding process. In the experimental results, this paper use the Intersection over Union(IoU) as an criteria to evaluate the performance of the image segmentation model. We compare the experimental results with U-net and DeepLab, and use the bar graph to show the distribution of IoU. The number of low IoU in our experimental model does not increase drastically with the degree of blur. Moreover, there is no overfitting problem in the testing dataset, it can prove the universality of our experimental model.

    摘要i 英文延伸摘要ii 誌謝vii 目錄viii 表格x 圖片xi Chapter 1. 簡介1 1.1. 研究動機 1 1.2. 研究方向與貢獻 2 1.3. 章節提要 2 Chapter 2. 研究背景與相關文獻3 2.1. 研究背景 3 2.1.1. 生成對抗網路架構研究 3 2.2. 圖像分割研究 6 2.2.1. 資料集與評估指標 7 2.2.2. 圖像分割研究概述 8 2.3. 模糊圖像分割 11 2.3.1. 模糊圖像之因與其分割問題 11 2.4. 近年相關研究 12 Chapter 3. 研究方法14 3.1. 網路模型 14 3.1.1. 模型訓練方法 15 3.1.2. 編解碼器架構 15 3.1.3. 判別器 18 3.2. Skip Connection 19 3.2.1. Res-Net 20 3.3. 模型架構 21 3.4. 圖像模糊化 22 Chapter 4. 研究結果與討論26 4.1. 實驗設計 26 4.1.1. 實驗環境 26 4.1.2. 實驗流程 27 4.1.3. 實驗資料集 28 4.2. 實驗結果 31 4.2.1. 模糊圖像分割模型實驗結果 31 4.2.2. 模糊圖像分割結果比較 36 4.2.3. 結果討論 45 Chapter 5. 結論與未來展望46 5.1. 研究結論 46 5.2. 未來展望 46 References 48 Chapter 6. Appendix 52

    [1] Shervin Minaee, Yuri Y Boykov, Fatih Porikli, Antonio J Plaza, Nasser Kehtarnavaz, and Demetri Terzopoulos. Image segmentation using deep learning: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021.
    [2] Qing Li, Weidong Cai, Xiaogang Wang, Yun Zhou, David Dagan Feng, and Mei Chen.Medical image classification with convolutional neural network. In 2014 13th international conference on control automation robotics & vision (ICARCV), pages 844–848.IEEE, 2014.
    [3] Ahmed Ali Mohammed Al-Saffar, Hai Tao, and Mohammed Ahmed Talab.Review of deep convolution neural network in image classification.
    In 2017 International Conference on Radar, Antenna, Microwave, Electronics, and Telecommunications (ICRAMET), pages 26–31. IEEE, 2017.
    [4] Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3431–3440, 2015.
    [5] Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Semantic image segmentation with deep convolutional nets and fully connected crfs. arXiv preprint arXiv:1412.7062, 2014.
    [6] Ian J Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley,Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks. arXiv preprint arXiv:1406.2661, 2014.
    [7] Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
    [8] Hoang Thanh-Tung and Truyen Tran. Catastrophic forgetting and mode collapse in gans. In 2020 International Joint Conference on Neural Networks (IJCNN), pages 1–10. IEEE, 2020.
    [9] Mehdi Mirza and Simon Osindero.Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784, 2014.
    [10] Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE international conference on computer vision, pages 2223–2232, 2017.
    [11] Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1125–1134, 2017.
    [12] Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pages 234–241. Springer, 2015.48
    [13] Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang,et al. End to end learning for self-driving cars. arXiv preprint arXiv:1604.07316, 2016.
    [14] Ashish Shrivastava, Tomas Pfister, Oncel Tuzel, Joshua Susskind, Wenda Wang, and Russell Webb. Learning from simulated and unsupervised images through adversarial training. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2107–2116, 2017.
    [15] Orest Kupyn, Volodymyr Budzan, Mykola Mykhailych, Dmytro Mishkin, and Jiri Matas. Deblurgan: Blind motion deblurring using conditional adversarial networks, 2018.
    [16] Seungjun Nah, Tae Hyun Kim, and Kyoung Mu Lee. Deep multi-scale convolutional neural network for dynamic scene deblurring. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3883–3891, 2017.
    [17] Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In European conference on computer vision, pages 694–711. Springer, 2016.
    [18] Martin Arjovsky, Soumith Chintala, and Léon Bottou. Wasserstein gan, 2017.
    [19] Han Zhang, Ian Goodfellow, Dimitris Metaxas, and Augustus Odena. Self-attention generative adversarial networks, 2019.
    [20] Augustus Odena, Jacob Buckman, Catherine Olsson, Tom Brown, Christopher Olah, Colin Raffel, and Ian Goodfellow. Is generator conditioning causally related to gan performance? In International conference on machine learning, pages 3849–3858.
    PMLR, 2018.
    [21] Tamar Rott Shaham, Tali Dekel, and Tomer Michaeli. Singan: Learning a generative model from a single natural image. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4570–4580, 2019.
    [22] Chuan Li and Michael Wand. Precomputed real-time texture synthesis with markovian generative adversarial networks. In European conference on computer vision, pages 702–716. Springer, 2016.
    [23] Xinming Zhang, Xiaobin Zhu, 3rd Xiao–Yu Zhang, Naiguang Zhang, Peng Li, and Lei Wang. Seggan: Semantic segmentation with generative adversarial network. In 2018 IEEE Fourth International Conference on Multimedia Big Data (BigMM), pages 1–5,2018.
    [24] ZengShun Zhaoa, Yulong Wang, Ke Liu, Haoran Yang, Qian Sun, and Heng Qiao. Semantic segmentation by improved generative adversarial networks, 2021.
    [25] Michael Treml, José Arjona-Medina, Thomas Unterthiner, Rupesh Durgesh, Felix Friedmann, Peter Schuberth, Andreas Mayr, Martin Heusel, Markus Hofmarcher, Michael Widrich, et al. Speeding up semantic segmentation for autonomous driving. In MLITS, NIPS Workshop, volume 2, 2016.49
    [26] Sunder Ali Khowaja and Seok-Lyong Lee. Semantic image networks for human action recognition. International Journal of Computer Vision, 128(2):393–419, 2020.
    [27] Nima Tajbakhsh, Laura Jeyaseelan, Qian Li, Jeffrey N Chiang, Zhihao Wu, and Xiaowei Ding. Embracing imperfect datasets: A review of deep learning solutions for medical image segmentation. Medical Image Analysis, 63:101693, 2020.
    [28] Li Liu, Wanli Ouyang, Xiaogang Wang, Paul Fieguth, Jie Chen, Xinwang Liu, and Matti Pietikäinen. Deep learning for generic object detection: A survey. International journal of computer vision, 128(2):261–318, 2020.
    [29] Derek Hoiem, Santosh K Divvala, and James H Hays. Pascal voc 2008 challenge. World Literature Today, 2009.
    [30] Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 580–587, 2014.
    [31] Alexey Bochkovskiy, Chien-Yao Wang, and Hong-Yuan Mark Liao. Yolov4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:2004.10934, 2020.
    [32] Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg. Ssd: Single shot multibox detector. In European conference on computer vision, pages 21–37. Springer, 2016.
    [33] Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014.
    [34] Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Scharwächter, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele.The cityscapes dataset. In CVPR Workshop on the Future of Datasets in Vision, volume 2,2015.
    [35] Nameirakpam Dhanachandra, Khumanthem Manglem, and Yambem Jina Chanu. Image segmentation using k-means clustering algorithm and subtractive clustering algorithm. Procedia Computer Science, 54:764–771, 2015.
    [36] Salem Saleh Al-Amri, NV Kalyankar, and SD Khamitkar. Image segmentation by using edge detection. International journal on computer science and engineering, 2(3):804–807, 2010.
    [37] Orlando José Tobias and Rui Seara. Image segmentation by histogram thresholding using fuzzy sets. IEEE transactions on Image Processing, 11(12):1457–1465, 2002.
    [38] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25:1097–1105, 2012.
    [39] Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1–9, 2015.50
    [40] Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
    [41] Vijay Badrinarayanan, Alex Kendall, and Roberto Cipolla. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE transactions on pattern analysis and machine intelligence, 39(12):2481–2495, 2017.
    [42] Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for largescale image recognition, 2015.
    [43] Léon Bottou. Large-scale machine learning with stochastic gradient descent. In Proceedings of COMPSTAT’2010, pages 177–186. Springer, 2010.
    [44] Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia. Pyramid scene parsing network. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2881–2890, 2017.
    [45] G Lin, A Milan, C Shen, and I Reid. Refinenet: Multi-path refinement networks with identity mappings for high-resolution semantic segmentation. arxiv 2016. arXiv preprint arXiv:1611.06612.
    [46] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
    [47] Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017.
    [48] Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. arXiv preprint arXiv:1506.01497, 2015.
    [49] Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L. Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs, 2017.
    [50] Liang-Chieh Chen, George Papandreou, Florian Schroff, and Hartwig Adam. Rethinking atrous convolution for semantic image segmentation. arXiv preprintarXiv:1706.05587, 2017.
    [51] Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and HartwigAdam. Encoder-decoder with atrous separable convolution for semantic image seg-mentation. In Proceedings of the European conference on computer vision (ECCV),pages 801–818, 2018.
    [52] Yuan Xue, Tao Xu, Han Zhang, L. Rodney Long, and Xiaolei Huang. Segan: Adversarial network with multi-scale l1 loss for medical image segmentation. Neuroinformatics,16(3-4):383–392, May 2018.
    [53] Pierre Baldi. Autoencoders, unsupervised learning, and deep architectures. In Proceedings of ICML workshop on unsupervised and transfer learning, pages 37–49. JMLR Workshop and Conference Proceedings, 2012.

    下載圖示
    2026-09-06公開
    QR CODE