簡易檢索 / 詳目顯示

研究生: 廖祐德
Liao, You-De
論文名稱: 適用於多類多件試衣情境之端到端混搭虛擬試衣網路
An End-to-End Mix-and-Match Virtual Try-On Network for Multi-Type Multi-Garment Scenario
指導教授: 莊坤達
Chuang, Kun-Ta
共同指導: 郭耀煌
Kuo, Yau-Hwang
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 資訊工程學系
Department of Computer Science and Information Engineering
論文出版年: 2021
畢業學年度: 109
語文別: 中文
論文頁數: 86
中文關鍵詞: 影像生成虛擬試衣多類多件試衣
外文關鍵詞: Image Generation, Virtual Try-On, Multi-Type Multi-Garment Virtual Try On
相關次數: 點閱:143下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 近年來人工智慧技術蓬勃發展,許多大型研究機構及產業界投入了大量資源來發展生技醫療、工業4.0、時尚領域等人工智慧相關應用領域的產品與服務,並陸續開發出多種先進技術。虛擬試衣是時尚領域中一種能模擬參考人物試穿目標衣物的技術,過去的方法僅能模擬單件衣物的試穿結果,且能夠模擬的衣物類型也較少。然而衣物的種類五花八門,人們通常想要看到多件不同類型衣物的搭配情形,若使用現有方法進行模擬多件衣物試穿,則需要針對不同衣物類型訓練多個模型,花費很久的時間來依序使用這多個模型進行衣服模擬試穿,生成結果也會因此更容易出現不真實的瑕疵,也無法考慮到衣物之間的關聯性。
    在本論文中,我們提出端到端混搭虛擬試衣網路,僅需訓練端到端深度學習網路就能讓參考人物一次換上多件目標衣物,並且適用於各式各樣類型的衣物,不但節省時間,也能夠考慮到各種衣物搭配的情況。我們的方法可分為三個步驟。第ㄧ,使用條件式的外觀流使得各類型衣物能夠更正確地摺疊匹配至合適的位置。第二,使用生成模型將摺疊後的各類衣物以及人體資訊合成出初步的生成結果,並且利用多遮罩的方式確保衣物之間沒有重疊混合的情形。第三,使用精煉模型修正初步生成結果的衣物接縫處及陰影等瑕疵,將圖片變得更加自然且清晰。
    此外,我們還提出了兩種評估虛擬試衣困難度的方法,分別為姿勢困難度估計方法和衣物困難度估計方法,透過將資料集運算評估指標,能夠使得虛擬試衣實驗結果評估更加客觀,而非各自表述。
    最後,我們的方法不論是在單件衣物虛擬試衣還是在多件衣物虛擬試衣都擁有比過去方法更好的結果,能夠避免衣物之間的重疊且生成出精細的衣物紋理細節,並在無論衣物及姿勢難度為何的狀況下都能生成出非常傑出的結果。我們的方法在人類測驗中取得了百分之八十六的優異成果。

    With the rapid development of artificial intelligence (AI), many research institutes and companies invest lots of resources to develop AI-enabled technologies for biotechnology, industry 4.0, the fashion field, and other fields. Virtual try-on is a technology in the fashion field that can simulate a target person to try on the desired clothing. However, the existing methods cannot deal with multi-garment try-on at the same time, and are limited to one type of clothing items. Since there are various types of clothing in the real world, people usually want to mix and match several types of clothing together before making a buying decision. Previous methods need to train multiple models for different clothing types to simulate multi-garment virtual try-on. It takes a long time to use these multiple models to simulate multi-garment virtual try-on in sequence. Hence, the generated results are more prone to unreal flaws and cannot take the correlation between the clothing items into account.
    In this thesis, we propose an End-to-End Mix-and-Match Virtual Try-On Network (EMVTON). It only needs to train an end-to-end deep learning network, the target person can try on multiple desired clothing items at a time, and it is suitable for all kinds of garments. It can not only save time but also consider the matching of garments. Our method can be seen as three steps. First, the conditional appearance flow method makes various types of clothing to be warped and matched to a suitable position more accurately. Second, the generative model synthesizes the warped garments and body information into a preliminary result, and we propose the multi-mask method to ensure that there is no overlap and mixing between the garments. Third, the refined model corrects the blemishes such as the seams and shadows of the preliminary result to make the picture more natural and clearer.
    Moreover, we also proposed two metrics for evaluating the difficulty of virtual try-on, the pose-angle-based difficulty estimation method, and the clothing-based difficulty estimation method. By calculating the evaluation index from the dataset, the evaluation of the virtual try-on experiment results can be more objective.
    Finally, our method has better results than the previous methods, whether it is the single-garment virtual try-on task or the multi-garment virtual try-on task. It can avoid overlapping between clothing and generate fine clothing texture details. Moreover, the proposed method can produce outstanding results regardless of the difficulty of the garment and posture. Notably, EMVTON has achieved 86% excellent results in human tests.

    CHAPTER 1 INTRODUCTION 1 1.1 Background 1 1.2 Problem Description 6 1.3 Motivation 9 1.4 Contribution 12 1.5 Organization 14 CHAPTER 2 RELATED WORK 15 2.1 Generative Model 15 2.2 Virtual Try-On 19 CHAPTER 3 END-TO-END MIX-AND-MATCH VIRTUAL TRY-ON NETWORK 35 3.1 Research Process 36 3.2 Input Data 38 3.3 Warping Model 40 3.4 Generator 47 3.5 Refiner 53 CHAPTER 4 EXPERIMENTS AND DISCUSSION 56 4.1 Dataset 56 4.2 Experimental Environment 58 4.3 Difficulty Estimation of Virtual Try-On 59 4.4 Experimental Results 66 CHAPTER 5 CONCLUSION 81 CHAPTER 6 FUTURE WORK 82 REFERENCES 83

    [ARJ 17] M. Arjovsky, S. Chintala, and L.Bottou, “Wasserstein gan”, arXiv preprint arXiv:1701.07875, 2017.
    [BAR 18] Barratt, Shane, and Rishi Sharma. "A note on the inception score." arXiv preprint arXiv:1801.01973 (2018).
    [DEN 09] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. FeiFei, “ImageNet: A
    large-scale hierarchical image database”, IEEE Conference on Computer
    Vision and Pattern Recognition, 2009.
    [DEN 12] J. Deng, A. Berg, S. Satheesh, H.Su, A. Khosla, and L. FeiFei, ”ImageNet Large Scale Visual Recognition Competition”, 2012.
    [FAD 16] Facebook AI Director Yann LeCun on the comment to answer the question in quora
    https://www.quora.com/What-are-some-recent-and-potentially-upcoming-breakthroughs-in-deep-learning
    [GE 21] Ge, Yuying, et al. "Parser-Free Virtual Try-on via Distilling Appearance Flows." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021.
    [GOO 14] I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets”, Conference and Workshop on Neural Information Processing Systems, 2014.
    [GUA 12] P. Guan, L. Reiss, D. A. Hirshberg, A. Weiss, and M. J. Black, “DRAPE: Dressing Any PErson”, ACM Transactions on Graphics, 2012.
    [HAN 17] X. Han, Z. Wu, Z. Wu, R. Yu, L. S. Davis, “Viton: An image-based virtual try-on network”, arXiv preprint arXiv:1711.08447, 2017.
    [HAN 19] Han, Xintong, et al. "Clothflow: A flow-based model for clothed person generation." Proceedings of the IEEE/CVF International Conference on Computer Vision. 2019.
    [HE 16] He, Kaiming, et al. "Deep residual learning for image recognition." Proceedings of the IEEE conference on computer vision and pattern recognition. 2016.
    [HYV 00] Hyvarinen and E. Oja, “Independent component analysis: algorithms and applications”, Neural Network, 2000.
    [KAR 17] Karras, Tero, et al. "Progressive growing of gans for improved quality, stability, and variation." arXiv preprint arXiv:1710.10196 (2017).
    [KAR 19] Karras, Tero, Samuli Laine, and Timo Aila. "A style-based generator architecture for generative adversarial networks." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019.
    [KAR 20] Karras, Tero, et al. "Analyzing and improving the image quality of stylegan." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020.
    [KIN 14] D. P. Kingma and M. Welling, “Auto-encoding variational bayes”, International Conference on Learning Representations, 2014.
    [LEC 98] Y. LeCun, L. Bottou, and Y. Bengio, “Gradient-based learning applied to document recognition”, Proceedings of the IEEE, 1998.
    [MIN 20] Minar, Matiur Rahman, et al. "CP-VTON+: Clothing shape and texture preserving image-based virtual try-on." CVPR Workshops. 2020.
    [MNI 10] V. Mnih, G. E. Hinton, et al. “Generating more realistic images using gated mrf’s”, Advances in Neural Information Processing Systems, 2010.
    [MRO 17] Y. Mroueh, T. Sercu, and V. Goel, “Mcgan: Mean and covariance feature matching gan”, arXiv preprint arXiv:1702.08398, 2017.
    [OOR 17] Oord, Aaron van den, Oriol Vinyals, and Koray Kavukcuoglu. "Neural discrete representation learning." arXiv preprint arXiv:1711.00937 (2017).
    [PER 03] H. Permuter, J. Francos, and I. H. Jermyn. “Gaussian mixture models of texture and colour for image database retrieval”, International Conference on Acoustics, Speech and Signal Processing, 2003.
    [QI 17] G. -J. Qi, “Loss-sensitive generative adversarial networks on lipschitz densities”, arXiv preprint arXiv:1701.06264, 2017.
    [QIN 20] Qin, Xuebin, et al. "U2-Net: Going deeper with nested U-structure for salient object detection." Pattern Recognition 106 (2020): 107404.
    [RAZ 19] Razavi, Ali, Aaron van den Oord, and Oriol Vinyals. "Generating diverse high-fidelity images with vq-vae-2." Advances in neural information processing systems. 2019.
    [RON 15] Ronneberger, Olaf, Philipp Fischer, and Thomas Brox. "U-net: Convolutional networks for biomedical image segmentation." International Conference on Medical image computing and computer-assisted intervention. Springer, Cham, 2015.
    [SAI 13] T. N. Sainath, A. R. Mohamed, B. Kingsbury, and B. Ramabhadran, “Deep
    convolutional neural networks for LVCSR”, International Conference on
    Acoustics, Speech and Signal Processing, 2013.
    [SAL 09] R. Salakhutdinov and G. E. Hinton. “Deep Boltzmann machines”, International Conference on Artificial Intelligence and Statistics, 2009.
    [STA 97] T. Starner and A. Pentland. “Real-time American sign language recognition form video using hidden markov models”, Motion-Based Recognition, 1997.
    [TAN 20] Tang, Hao, et al. "Xinggan for person image generation." European Conference on Computer Vision. Springer, Cham, 2020.
    [TU 07] Z. Tu. “Learning generative models via discriminative approaches”, IEEE Conference on Computer Vision and Pattern Recognition, 2007.
    [TUR 91] M. A. Turk and A. P. Pentland. “Face recognition using eigenfaces”, Conference on Computer Vision and Pattern Recognition, 1991.
    [XU 96] L. Xu and M. I. Jordan. “On convergence properties of the em algorithm for gaussian mixtures”, Neural computation, 1996.
    [YAN 20] Yang, Han, et al. "Towards photo-realistic virtual try-on by adaptively generating-preserving image content." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020.
    [YU 19] Yu, Jiahui, et al. "Free-form image inpainting with gated convolution." Proceedings of the IEEE/CVF International Conference on Computer Vision. 2019.
    [ZHU 19] Zhu, Zhen, et al. "Progressive pose attention transfer for person image generation." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019.

    下載圖示
    2026-06-30公開
    QR CODE