簡易檢索 / 詳目顯示

研究生: 徐雍軒
Hsu, Yung-Hsuan
論文名稱: 基於WACGAN-GP 資料增強之有限資料影像辨識改善方法
Improving Image Recognition with Limited Data via WACGAN-GP-Based Data Augmentation
指導教授: 李坤洲
Lee, Kun-Chou
學位類別: 碩士
Master
系所名稱: 工學院 - 系統及船舶機電工程學系
Department of Systems and Naval Mechatronic Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 69
中文關鍵詞: 深度學習 、類神經網路 、生成對抗網路 、資料增強
外文關鍵詞: Deep Learning, Neural Networks, Generative Adversarial Network, Data Augmentation
相關次數: 點閱:83  下載:1 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 標註訓練影像不足仍是影像辨識任務中的主要限制。深度學習模型通常需要大量訓練資料,才能學習具有代表性的特徵。當可用樣本數量有限時,訓練後的分類器容易產生過度擬合,並降低泛化能力。為了解決此問題,本論文探討使用 WACGAN-GP 進行資料增強。此模型結合輔助類別資訊、Wasserstein 距離與梯度懲罰。輔助分類器透過類別標籤引導生成過程,而 Wasserstein 距離與梯度懲罰則有助於提升對抗訓練的穩定性。模型訓練完成後,生成器會針對各類別產生合成影像。接著,這些影像會加入原始訓練資料集中,以在分類器訓練前增加資料多樣性。本研究於多個影像資料集上進行實驗,以評估所提出資料增強方法在不同資料條件下的有效性。實驗中同時考慮平衡與不平衡的訓練資料設定。此外,本論文也分析不同訓練設定對結果的影響,並在實驗中使用 ACGAN 作為比較模型。實驗結果顯示,WACGAN-GP 在多數情況下能提升分類準確率,尤其是在原始訓練影像數量較少時,其改善效果較為明顯。結果也說明,生成影像能為分類器訓練提供有用的補充資訊。與 ACGAN 及傳統資料增強方法相比,WACGAN-GP 在測試設定中具有較穩定的表現。綜合以上,本論文證實 WACGAN-GP 是一種有效的資料增強方法,可用於改善訓練資料有限情況下的影像辨識表現。

    Insufficient labeled training images remain a major limitation in image recognition tasks. Deep learning models usually require a large amount of training data to learn useful features. When the available samples are limited, the trained classifier may suffer from overfitting and poor generalization. To address this problem, this thesis investigates data augmentation using WACGAN-GP. The model combines auxiliary class information with Wasserstein distance and gradient penalty. The auxiliary classifier guides the generation process by using class labels, while the Wasserstein distance and gradient penalty help improve the stability of adversarial training. After training, the generator produces synthetic images for each class. These images are then added to the original training set to increase data diversity before classifier training. Experiments are conducted on several image datasets to evaluate the effectiveness of the proposed augmentation method under different data conditions. Both balanced and imbalanced training sets are considered. In addition, the effect of different training settings is examined, and ACGAN is used as a comparison model in part of the experiments. The experimental results show that WACGAN-GP can improve classification accuracy in most cases, especially when the number of original training images is small. The results also indicate that the generated images provide useful supplementary information for classifier training. Compared with ACGAN and conventional augmentation methods, WACGAN-GP gives more stable performance in the tested settings. Based on these findings, this thesis shows that WACGAN-GP is an effective data augmentation method for image recognition tasks with limited training data.

    中文摘要 i Abstract ii Contents iii List of Tables v List of Figures vi 1 Introduction 1 1.1 Motivations and Objectives 1 1.2 Related Works 4 1.3 Organization 6 2 Preliminaries 7 2.1 Neural Networks 7 2.2 Convolutional Neural Networks 9 2.3 Generative Adversarial Networks 9 2.4 Wasserstein Generative Adversarial Networks 12 2.5 Wasserstein GAN with Gradient Penalty 13 2.6 Conditional Generative Adversarial Networks 14 2.7 Auxiliary Classifier Generative Adversarial Networks 15 2.8 WACGAN-GP 16 3 Experimental Results 20 3.1 Dataset and Experimental Setup 20 3.2 Fashion-MNIST Experiments 24 3.3 CIFAR-100 Experiments 29 3.4 MSTAR Experiments 35 3.5 STL-10 Experiments 41 3.6 EuroSAT Experiments 45 3.7 Summary of Experimental Results Across Datasets 48 4 Conclusions 52 4.1 Conclusions 52 4.2 Future Work 53 References 55

    [1] Connor Shorten and Taghi M Khoshgoftaar. A survey on image data augmentation for deep learning. Journal of Big Data, 6(1):1-48, 2019.
    [2] Yaqing Wang Quanming Yao, James T Kwok, and Lionel M Ni. Generalizing from a few examples: A survey on few-shot learning. ACM Computing Surveys (CSUR), 53(3):1-34, 2020.
    [3] Kiran Maharana, Surajit Mondal, and Bhushankumar Nemade. A review: Data pre-processing and data augmentation techniques. Global Transitions Proceedings, 3(1):91-99, 2022.
    [4] Brandon Trabucco, Kyle Doherty, Max Gurinas, and Ruslan Salakhutdinov. Effective data augmentation with diffusion models. ArXiv Preprint arXiv:2302.07944, 2023.
    [5] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Communications of the ACM, 60(6):84-90, 2017.
    [6] Antonia Creswell, Tom White, Vincent Dumoulin, Kai Arulkumaran, Biswa Sengupta, and Anil A Bharath. Generative adversarial networks: An overview. IEEE Signal Processing Magazine, 35(1):53-65, 2018.
    [7] Martin Arjovsky, Soumith Chintala, and Léon Bottou. Wasserstein generative adversarial networks. In International Conference on Machine Learning, pages 214-223. Pmlr, 2017.
    [8] Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville. Improved training of wasserstein gans. Advances in Neural Information Processing Systems, 30, 2017.
    [9] Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training gans. Advances in Neural Information Processing Systems, 29, 2016.
    [10] Kun-Chou Lee and Yung-Hsuan Hsu. Improving image recognition with limited data via wacgan-gp-based data augmentation. Applied Sciences, 16(6):2805, 2026.
    [11] Shu-Xiang Jiang. Application of generative adversarial network to recognition of ship targets. Master's thesis, National Cheng Kung University, Tainan, Taiwan, 2020.
    [12] Andrés Anaya-Isaza and Leonel Mera-Jiménez. Data augmentation and transfer learning for brain tumor detection in magnetic resonance imaging. IEEE Access, 10:23217-23233, 2022.
    [13] Yuhang Li, Youngeun Kim, Hyoungseob Park, Tamar Geller, and Priyadarshini Panda. Neuromorphic data augmentation for training spiking neural networks. In European Conference on Computer Vision, pages 631-649. Springer, 2022.
    [14] Barret Zoph, Ekin D Cubuk, Golnaz Ghiasi, Tsung-Yi Lin, Jonathon Shlens, and Quoc V Le. Learning data augmentation strategies for object detection. In European Conference on Computer Vision, pages 566-583. Springer, 2020.
    [15] Aghiles Kebaili, Jérôme Lapuyade-Lahorgue, and Su Ruan. Deep learning approaches for data augmentation in medical imaging: a review. Journal of Imaging, 9(4):81, 2023.
    [16] Teerath Kumar, Rob Brennan, Alessandra Mileo, and Malika Bendechache. Image data augmentation approaches: A comprehensive survey and future directions. IEEE Access, 12:187536-187571, 2024.
    [17] Kushankur Ghosh, Colin Bellinger, Roberto Corizzo, Paula Branco, Bartosz Krawczyk, and Nathalie Japkowicz. The class imbalance problem in deep learning. Machine Learning, 113(7):4845-4901, 2024.
    [18] Nitesh V Chawla, Kevin W Bowyer, Lawrence O Hall, and W Philip Kegelmeyer. Smote: synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16:321-357, 2002.
    [19] Xiaotian Han, Zhimeng Jiang, Ninghao Liu, and Xia Hu. G-mixup: Graph data augmentation for graph classification. In International Conference on Machine Learning, pages 8230-8248. PMLR, 2022.
    [20] Manisha Sharma, Alka Verma, and Uma Rani. Optimizing efficientnetv2 model with randaugment data augmentation for detecting wheat diseases in smart farming. Scalable Computing: Practice and Experience, 26(5):2087-2104, 2025.
    [21] Agnieszka Mikołajczyk and Michał Grochowski. Data augmentation for improving deep learning in image classification problem. In 2018 International Interdisciplinary PhD Workshop (IIPhDW), pages 117-122. IEEE, 2018.
    [22] Jun-Hyung Kim and Youngbae Hwang. Gan-based synthetic data augmentation for infrared small target detection. IEEE Transactions on Geoscience and Remote Sensing, 60:1-12, 2022.
    [23] Jonathan Ho, Chitwan Saharia, William Chan, David J Fleet, Mohammad Norouzi, and Tim Salimans. Cascaded diffusion models for high fidelity image generation. Journal of Machine Learning Research, 23(47):1-33, 2022.
    [24] Zhongxuan Mei. Advances in fidelity-preserving gans for small datasets: Focus on stylegan-ada and its variants. In ITM Web of Conferences, volume 80, page 01010. EDP Sciences, 2025.
    [25] Haoyang Li, Wei Chen, and Xiaojin Zhang. Fed-augmix: Balancing privacy and utility via data augmentation. ArXiv Preprint arXiv:2412.13818, 2024.
    [26] Runlian Zhang, Yu Mo, Zhaoxuan Pan, Hailong Zhang, Yongzhuang Wei, and Xiaonian Wu. Intra-class cutmix data augmentation based deep learning side channel attacks. Integration, 102:102373, 2025.
    [27] Samuel G Müller and Frank Hutter. Trivialaugment: Tuning-free yet state-of-the-art data augmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 774-782, 2021.
    [28] Swaminathan Gurumurthy, Ravi Kiran Sarvadevabhatla, and R Venkatesh Babu. Deligan: Generative adversarial networks for diverse and limited data. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 166-174, 2017.
    [29] Kaleb E Smith and Anthony O Smith. Conditional gan for timeseries generation. ArXiv Preprint arXiv:2006.16477, 2020.
    [30] Abdul Waheed, Muskan Goyal, Deepak Gupta, Ashish Khanna, Fadi Al-Turjman, and Plácido Rogerio Pinheiro. Covidgan: data augmentation using auxiliary classifier gan for improved covid-19 detection. IEEE Access, 8:91916-91923, 2020.
    [31] Zhaoyu Zhang, Mengyan Li, and Jun Yu. On the convergence and mode collapse of gan. In SIGGRAPH Asia 2018 Technical Briefs, pages 1-4. 2018.
    [32] Monica Welfert, Gowtham R Kurri, Kyle Otstot, and Lalitha Sankar. Addressing gan training instabilities via tunable classification losses. IEEE Journal on Selected Areas in Information Theory, 5:534-553, 2024.
    [33] Saurabh Vijay Parhad, Krishna K Warhade, and Sanjay S Shitole. Speckle noise reduction in sar images using improved filtering and supervised classification. Multimedia Tools and Applications, 83(18):54615-54636, 2024.
    [34] Steven M Scarborough, LeRoy Gorham, Michael J Minardi, Uttam K Majumder, Matthew G Judge, Linda Moore, Leslie Novak, Steven Jaroszewksi, Laura Spoldi, and Alan Pieramico. A challenge problem for sar change detection and data compression. In Algorithms for Synthetic Aperture Radar Imagery XVII, volume 7699, pages 287-291. SPIE, 2010.
    [35] Lu Wang, Yuhang Qi, P Takis Mathiopoulos, Chunhui Zhao, and Suleman Mazhar. An improved sar ship classification method using text-to-image generation-based data augmentation and squeeze and excitation. Remote Sensing, 16(7):1299, 2024.
    [36] Enze Liu, Zhiguang Chu, and Xing Zhang. Wasserstein gan for moving differential privacy protection. Scientific Reports, 15(1):19634, 2025.
    [37] Tristan Milne and Adrian I Nachman. Wasserstein gans with gradient penalty compute congested transport. In Conference on Learning Theory, pages 103-129. PMLR, 2022.

    下載圖示
    校外:立即公開
    QR CODE