簡易檢索 / 詳目顯示

研究生: 鄭永正
Cheng, Yung-Cheng
論文名稱: 運用複合縮放與深度分離卷積於輕量化影像修復模型設計與研究
Lightweight Image Inpainting Model Using Compound Scaling and Depthwise Separable Convolution
指導教授: 賴槿峰
Lai, Chin-­Feng
學位類別: 碩士
Master
系所名稱: 工學院 - 工程科學系
Department of Engineering Science
論文出版年: 2021
畢業學年度: 109
語文別: 中文
論文頁數: 68
中文關鍵詞: 影像修復 、輕量化模型 、對抗生成神經網路
外文關鍵詞: Image Inpainting, Light Weight Model, GAN
相關次數: 點閱:142  下載:0 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本研究旨在探討使用於影像修復之對抗生成網路架構,有鑑於當今的深度網路模型偏向增加網路複雜度以換取模型精確度而造成模型缺乏效率的現象。在本研究中試圖利用深度分離捲積與複合縮放法所搭建之輕量化影像修復模型加以解決。
    研究中所進行的實驗一共分為三階段,第一階段為基礎模型探尋,透過實驗不同的網路模型架構來找尋一適合進行縮放之網路模型作為基礎模型。第二階段為縮放係數探尋,透過實驗多組有限制的縮放係數組合,觀察其中最具縮放效率的係數組作為後續實驗所需使用的縮放係數。第三階段為縮放階段,利用第二階段所得到的縮放係數對第一階段的基礎模型進行數次的縮放直到達成目標效果。
    本研究中最後所提出的輕量化影像修復模型,透過使用深度分離捲積與複合縮放法,使模型在建構時能有效率的複雜化。與當今影像修復模型相比,在保持相同修復能力的情況下,約能將模型參數量下降至現有模型的四分之一左右。
    關鍵字:影像修復、輕量化模型、對抗生成神經網路

    The purpose of this study is to investigate the use of GAN architecture for image inpainting. In view of the lack of efficiency of today’s deep network models due to their tendency to increase network complexity in exchange for model accuracy. In this study, we attempt to solve this problem by using a lightweight image inpainting model built with depthwise separable convolution and compound scaling.
    The final lightweight image inpainting model proposed in this study uses depthwise separable convolution and compound scaling to efficiently complicate the model construction. Compared with the current image inpainting model, the number of parameter can be reduced to about onefourth of the current model while maintaining the same inpainting capability.

    中文摘要 ii 英文摘要 iii 誌謝 x 目錄 xi 表目錄 xiv 圖目錄 xvii 縮寫表 xviii 符號說明 xix 1 緒論 1 1.1 研究背景與動機 1 1.2 研究目標 1 1.3 章節提要 2 2 文獻探討 3 2.1 深度學習研究 3 2.1.1 神經網路架構研究 3 2.1.2 生成對抗神經網路研究 4 2.2 影像修復研究 7 2.2.1 影像修復文獻探討 7 2.2.2 評估指標與資料集 9 2.3 模型輕量化研究 11 2.3.1 模型輕量化文獻探討 11 3 研究方法 13 3.1 影像修復 13 3.1.1 Gated Convolution 13 3.1.2 SpectralNormalized Patch GAN 15 3.2 輕量化模型 17 3.2.1 複合縮放法(Compound Model Scaling) 17 3.2.2 深度分離卷積(Depthwise Separable Convolution) 19 3.3 影像修復模型重構方法 23 3.3.1 Depthwise Separable Gated Convolution 23 3.3.2 Compound Model Scaling with Deepfill 24 4 研究結果與討論 27 4.1 實驗設計 27 4.1.1 實驗環境 27 4.1.2 實驗流程 28 4.2 實驗結果 31 4.2.1 基礎模型探索實驗結果 32 4.2.2 縮放係數探索實驗結果 44 4.2.3 模型縮放實驗 45 4.3 結果討論 58 5 結論與未來展望 60 5.1 研究結論 60 5.2 未來展望 61 參考文獻 62

    [1] Hebb, D. O. (2005). The organization of behavior: A neuropsychological theory. Psychology Press.
    [2] Farley, B. W. A. C., & Clark, W. (1954). Simulation of selforganizing systems by digital computer. Transactions of the IRE Professional Group on Information Theory, 4(4), 76-84.
    [3] Rochester, N., Holland, J., Haibt, L., & Duda, W. (1956). Tests on a cell assembly theory of the action of the brain, using a large digital computer. IRE Transactions on information Theory, 2(3), 80-93.
    [4] Rosenblatt, F. (1958). The perceptron: a probabilistic model for information storage and organization in the brain. Psychological review, 65(6), 386.
    [5] Minsky, M., & Papert, S. A. (2017). Perceptrons: An introduction to computational geometry.
    [6] Werbos, P. J. (1982). Applications of advances in nonlinear sensitivity analysis. System modeling and optimization, 762-770.
    [7] McClelland, J. L., Rumelhart, D. E., & PDP Research Group. (1986). Parallel distributed processing (Vol. 2, pp. 2021). Cambridge, MA: MIT press.
    [8] Hinton, G. E., & Salakhutdinov, R. R. (2006). Reducing the dimensionality of data with neural networks. science, 313(5786), 504-507.
    [9] Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 1097-1105.
    [10] LeCun, Y., Bottou, L., Bengio, Y., & Haffner, P. (1998). Gradientbased learning applied to document recognition. Proceedings of the IEEE, 86(11), 2278-2324.
    [11] Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for largescale image recognition. arXiv preprint arXiv:1409.1556.
    [12] Lin, M., Chen, Q., & Yan, S. (2013). Network in network. arXiv preprint arXiv:1312.4400.
    [13] Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., ... & Rabinovich,A. (2015). Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1-9).
    [14] Ioffe, S., & Szegedy, C. (2015, June). Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning (pp. 448-456). PMLR.
    [15] Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., & Wojna, Z. (2016). Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 2818-2826).
    [16] Szegedy, C., Ioffe, S., Vanhoucke, V., & Alemi, A. (2017, February). Inceptionv4, inceptionresnet and the impact of residual connections on learning. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 31, No. 1).
    [17] He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770-778).
    [18] Tan, M., & Le, Q. (2019, May). Efficientnet: Rethinking model scaling for convolutional neural networks. In International Conference on Machine Learning (pp. 6105-6114).PMLR.
    [19] Girshick, R., Donahue, J., Darrell, T., & Malik, J. (2014). Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 580-587).
    [20] Girshick, R. (2015). Fast rcnn. In Proceedings of the IEEE international conference on computer vision (pp. 1440-1448).
    [21] Ren, S., He, K., Girshick, R., & Sun, J. (2015). Faster rcnn: Towards realtime object detection with region proposal networks. arXiv preprint arXiv:1506.01497.
    [22] He, K., Gkioxari, G., Dollár, P., & Girshick, R. (2017). Mask rcnn. In Proceedings of the IEEE international conference on computer vision (pp. 2961-2969).
    [23] Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You only look once: Unified, realtime object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 779-788).
    [24] Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by backpropagating errors. nature, 323(6088), 533-536.
    [25] Cho, K., Van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., & Bengio, Y. (2014). Learning phrase representations using RNN encoderdecoder for statistical machine translation. arXiv preprint arXiv:1406.1078.
    [26] Hochreiter, S., & Schmidhuber, J. (1997). Long shortterm memory. Neural computation, 9(8), 1735-1780.
    [27] Howard, J., & Ruder, S. (2018). Universal language model finetuning for text classification. arXiv preprint arXiv:1801.06146.
    [28] Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. arXiv preprint arXiv:1706.03762.
    [29] Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2018). Bert: Pretraining of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
    [30] Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., & Salakhutdinov, R. (2019). Transformerxl: Attentive language models beyond a fixedlengt context. arXiv preprint arXiv:1901.02860.
    [31] Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., ... & Amodei, D. (2020). Language models are fewshot learners. arXiv preprint arXiv:2005.14165.
    [32] Goodfellow, I. J., PougetAbadie, J., Mirza, M., Xu, B., WardeFarley, D., Ozair, S., ... & Bengio, Y. (2014). Generative adversarial networks. arXiv preprint arXiv:1406.2661.
    [33] Karras, T., Aila, T., Laine, S., & Lehtinen, J. (2017). Progressive growing of gans for improved quality, stability, and variation. arXiv preprint arXiv:1710.10196.
    [34] Zhang, H., Xu, T., Li, H., Zhang, S., Wang, X., Huang, X., & Metaxas, D. N. (2017). Stackgan: Text to photorealistic image synthesis with stacked generative adversarial networks. In Proceedings of the IEEE international conference on computer vision (pp. 5907-5915).
    [35] Zhu, J. Y., Park, T., Isola, P., & Efros, A. A. (2017). Unpaired imagetoimage translation using cycleconsistent adversarial networks. In Proceedings of the IEEE international conference on computer vision (pp. 2223-2232).
    [36] [237] Akcay, S., AtapourAbarghouei, A., & Breckon, T. P. (2018, December). Ganomaly: Semisupervised anomaly detection via adversarial training. In Asian conference on computer vision (pp. 622-637). Springer, Cham.
    [37] [238] Arjovsky, M., & Bottou, L. (2017). Towards principled methods for training generative adversarial networks. arXiv preprint arXiv:1701.04862.
    [38] Arjovsky, M., Chintala, S., & Bottou, L. (2017, July). Wasserstein generative adversarial networks. In International conference on machine learning (pp. 214-223). PMLR.
    [39] [237] Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., & Courville, A. (2017). Improved training of wasserstein gans. arXiv preprint arXiv:1704.00028.
    [40] [238] Yoshida, Y., & Miyato, T. (2017). Spectral norm regularization for improving the generalizability of deep learning. arXiv preprint arXiv:1705.10941.
    [41] Miyato, T., Kataoka, T., Koyama, M., & Yoshida, Y. (2018). Spectral normalization for generative adversarial networks. arXiv preprint arXiv:1802.05957.
    [42] Radford, A., Metz, L., & Chintala, S. (2015). Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv preprint arXiv:1511.06434.
    [43] Brock, A., Donahue, J., & Simonyan, K. (2018). Large scale GAN training for high fidelity natural image synthesis. arXiv preprint arXiv:1809.11096.
    [44] Deng, J., Dong, W., Socher, R., Li, L. J., Li, K., & FeiFei,L. (2009, June). Imagenet: A largescale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition (pp. 248-255). Ieee.
    [45] Reed, S., Akata, Z., Lee, H., & Schiele, B. (2016). Learning deep representations of finegrained visual descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 49-58).
    [46] Mirza, M., & Osindero, S. (2014). Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784. [47] Pathak, D., Krahenbuhl, P., Donahue, J., Darrell, T., & Efros, A. A. (2016). Context encoders: Feature learning by inpainting. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 2536-2544).
    [47] Pathak, D., Krahenbuhl, P., Donahue, J., Darrell, T., & Efros, A. A. (2016). Context encoders: Feature learning by inpainting. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 2536-2544).
    [48] Yang, C., Lu, X., Lin, Z., Shechtman, E., Wang, O., & Li, H. (2017). Highresolution image inpainting using multiscale neural patch synthesis. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 6721-6729).
    [49] Yan, Z., Li, X., Li, M., Zuo, W., & Shan, S. (2018). Shiftnet: Image inpainting via deep feature rearrangement. In Proceedings of the European conference on computer vision (ECCV) (pp. 1-17).
    [50] Liu, G., Reda, F. A., Shih, K. J., Wang, T. C., Tao, A., & Catanzaro, B. (2018). Image inpainting for irregular holes using partial convolutions. In Proceedings of the European Conference on Computer Vision (ECCV) (pp. 85-100).
    [51] Nazeri, K., Ng, E., Joseph, T., Qureshi, F. Z., & Ebrahimi, M. (2019). Edgeconnect: Generative image inpainting with adversarial edge learning. arXiv preprint arXiv:1901.00212.
    [52] Yu, J., Lin, Z., Yang, J., Shen, X., Lu, X., & Huang, T. S. (2018). Generative image inpainting with contextual attention. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 5505-5514).
    [53] Yu, J., Lin, Z., Yang, J., Shen, X., Lu, X., & Huang, T. S. (2019). Freeform image inpainting with gated convolution. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 4471-4480).
    [54] Lin, T. Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., ... & Zitnick, C. L. (2014, September). Microsoft coco: Common objects in context. In European conference on computer vision (pp. 740-755). Springer, Cham.
    [55] Zhou, B., Khosla, A., Lapedriza, A., Torralba, A., & Oliva, A. (2016). Places: An image database for deep scene understanding. arXiv preprint arXiv:1610.02055. [56] Wang, Z., Bovik, A. C., Sheikh, H. R., Simoncelli, E. P. (2004). Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4), 600-612.
    [56] Wang, Z., Bovik, A. C., Sheikh, H. R., Simoncelli, E. P. (2004). Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4), 600-612.
    [57] Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., Hochreiter, S. (2017). Gans trained by a two timescale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30.
    [58] Howard, A. G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., ... & Adam, H. (2017). Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861.
    [59] Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., & Chen, L. C. (2018). Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 4510-4520).
    [60] Howard, A., Sandler, M., Chu, G., Chen, L. C., Chen, B., Tan, M., ... & Adam, H. (2019). Searching for mobilenetv3. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 1314-1324).
    [61] Zhang, X., Zhou, X., Lin, M., & Sun, J. (2018). Shufflenet: An extremely efficient convolutional neural network for mobile devices. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 6848-6856).
    [62] Ma, N., Zhang, X., Zheng, H. T., & Sun, J. (2018). Shufflenet v2: Practical guidelines for efficient cnn architecture design. In Proceedings of the European conference on computer vision (ECCV) (pp. 116-131).
    [63] Huang, G., Liu, Z., Van Der Maaten, L., & Weinberger, K. Q. (2017). Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 4700-4708).
    [64] Isola, P., Zhu, J. Y., Zhou, T., & Efros, A. A. (2017). Imagetoimage translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1125-1134).

    下載圖示
    2026-08-09公開
    QR CODE