簡易檢索 / 詳目顯示

研究生: 謝承恩
Hsieh, Cheng-En
論文名稱: 資料平衡式生成對抗網路惡意資料處理機制
Data Balanced Algorithm Based on Generative Adversarial Network
指導教授: 李忠憲
Li, Jung-Shian
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 電腦與通信工程研究所
Institute of Computer & Communication Engineering
論文出版年: 2021
畢業學年度: 109
語文別: 中文
論文頁數: 63
中文關鍵詞: 惡意流量偵測 、生成對抗網路 、入侵偵測系統資料集 、機器學習
外文關鍵詞: Anomaly Traffic Detection, Machine Learning, IDS Dataset, Generate Adversarial Network
相關次數: 點閱:160  下載:0 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 近年來科技與網路技術的普及與進步,網路駭客的攻擊手法也日益精進,入侵偵測系統為了能夠阻擋惡意攻擊,因而引入機器學習作為防護策略,然而機器學習演算法與資料集對模型的成效有很大的影響,因此本研究以CNN、LSTM、BAT、SVM、NB五種機器學習演算法針對NSL-KDD、UNSW-NB15、CICIDS 2017三種資料集進行模型訓練,透過演算法對資料集的評估指數表現來分析演算法對資料的敏感性,並且我們設計用於改善資料集資料平衡性問題的生成對抗網路演算法,本研究將其應用在CICIDS 2017資料集上,其結果在CNN的召回率提高了20%,準確率提升4%; LSTM的召回率提高了10%,準確率提升2%;BAT的召回率提高了16%,準確率提升3%;SVM的召回率提高了15%,準確率提升4%,並且我們利用該演算法生成含有標籤的混合資料集,用以解釋非監督式模型的分群效果,透過此方法來解釋Xgboost、Gradient Boosting和隨機森林三種非監督式機器學習演算法生成的標籤結果。

    To defend against malicious attacks, intrusion detection systems have introduced machine learning as a protection strategy. However, machine learning algorithms and datasets have a great influence on the effectiveness of the machine learning model. This study uses five algorithms which are Naïve Bayes, CNN, LSTM, BAT, and SVMto train the IDS machine learning model. We use three datasets which are NSL-KDD, UNSW-NB15, and CICIDS 2017 to train and evaluate the model performance. We design a data-balanced method based on the GAN algorithm to improve the data imbalance problem of the IDS dataset. We apply the method to the CICIDS 2017 dataset. As a result, the recall rate and the accuracy rate of maching learning models have increased. Also, we use the method to generate labels, these labels are used to explain the clustering effect of the unsupervised model. We set three unsupervised machine learning algorithms as Xgboost, Gradient Boosting, and Random Forest as the target, and use our label to explain the clustering effect of the above-unsupervised model.

    摘要 I EXTENDED ABSTRACT II 誌謝 XII 目錄 XIV 表目錄 XVI 圖目錄 XVII 一、 緒論 1 1.1 研究背景 1 1.2 研究動機 2 1.3 研究貢獻 4 1.4 全文架構 5 二、 相關研究 6 2.1 惡意流量偵測 6 2.2 網路流量資料集 8 2.2.1 NSL-KDD 8 2.2.2 UBSW-NB15 10 2.2.3 CICIDS 2017 12 2.3 網路攻擊概述 14 2.4 常見入侵偵測系統之機器學習演算法 16 2.4.1 單純貝氏分類器(Naïve Bayes, NB) 16 2.4.2 支持向量機(Support Vector Machine, SVM) 17 2.4.3 卷積神經網路(Convolution Neural Network, CNN) 18 2.4.4 長短期記憶模型(Long Short-term Memory, LSTM) 20 2.4.5 BAT演算法 22 三、 系統架構 23 3.1 資料預處理 24 3.2 敏感性分析 29 3.2.1 混淆矩陣 29 3.2.2 敏感性與資料集缺陷探討 31 3.3 基於生成對抗網路之資料平衡演算法 32 3.3.1 資料平衡演算法 33 3.3.2 非監督式機器學習驗證方法探討 37 四、 實驗結果 39 4.1 敏感性分析結果 40 4.2 基於生成對抗網路的資料平衡演算法結果呈現 44 4.2.1 監督式機器學習演算法評估結果 44 4.2.2 驗證非監督式機器學習演算法分群效果 51 4.3 改良資料集結果呈現 55 4.3.1 惡意流量重現 55 4.3.2 改良資料集與公開資料集比較 57 五、 結論與未來展望 59 5.1 結論 59 5.2 未來展望 60 參考資料 61

    NSA, CISA, FBI, and NCSC, “Russian GRU Conducting Global Brute Force Campaign to Compromise Enterprise and Cloud Environments,” 15 4 2021. [Online]. Available: https://media.defense.gov/2021/Jul/01/2002753896/-1/-1/1/CSA_GRU_GLOBAL_BRUTE_FORCE_CAMPAIGN_UOO158036-21.PDF. [Accessed 2 7 2021].
    林妍溱, “比利時政府被DDoS攻陷,” 5 5 2021. [Online]. Available: https://www.ithome.com.tw/news/144202. [Accessed 1 7 2021].
    CBINSIGHTS, “AI 100: The Artificial Intelligence Startups Redefining Industries,” 3 3 2020. [Online]. Available: https://www.cbinsights.com/research/artificial-intelligence-top-startups/?utm_source=CB+Insights+Newsletter&utm_campaign=a4d41a607e-TuesNL_11_28_2017&utm_medium=email&utm_term=0_9dc0513989-a4d41a607e-89402301. [Accessed 1 5 2021].
    R. Tolido, G. V. D. Linden, L. Delabarre, and J. Theisler, “Reinventing Cybersecurity with Artificial Intelligence,” Capgemini., USA, 2019.
    CISCO, “Cisco Annual Internet Report(2018-2023),” 9 3 2020. [Online]. Available: https://www.cisco.com/c/en/us/solutions/collateral/executive-perspectives/annual-internet-report/white-paper-c11-741490.html. [Accessed 1 5 2021].
    Y. Hu, Y. Lu.,S. Wang, M. Zhang, X. Qu, and B. Niu, “Application of Machine Learning Approaches for the Design and Study of Anticancer Drugs,” Current Drug Targets, vol. 20, no. 5, pp. 488-500, 2016.
    J. Schmidt, M. R. G. Marques, S. Botti2, and M. A. L. Marques, “Recent advances and applications of machine learning in solidstate materials science,” npj Computational Materials volume, vol. 5, no. 83, 2019.
    V. Ganganwar, “An overview of classification algorithms for imbalanced datasets,” International Journal of Emerging Technology and Advanced Engineering, vol. 2, no. 4, 2012.
    A. Sundaram, “An Introduction to Intrusion Detection,” XRDS: Crossroads, The ACM Magazine for Students, vol. 2, no. 4, pp. 3-7, 1996.
    S. A. R. Shah, and B. Issac, “Performance comparison of intrusion detection systems and application of,” Future Generation Computer Systems, vol. 80, pp. 157-170, Mar. 2018.
    I. Ahmad, M. Basheri, M. J. Iqbal, and A. Rahim, “Performance Comparison of Support Vector Machine, Random Forest, and Extreme Learning Machine for Intrusion Detection,” IEEE Access, vol. 6, pp. 33789 - 33795, 30 May 2018.
    K. Singh and Dr.K. J. Mathai, “Performance Comparison of Intrusion Detection System Between Deep Belief Network (DBN)Algorithm and State Preserving Extreme Learning Machine (SPELM) Algorithm,” 2019 IEEE International Conference on Electrical, Computer and Communication Technologies (ICECCT 2019), Coimbatore, India, 20-22 Feb., 2019.
    R. Vinayakumar, M. Alazab, P. Poornachandran, A. Al-Nemrat, and S. Venkatraman, “Deep Learning Approach for Intelligent Intrusion Detection System,” IEEE Access, vol. 7, pp. 41525 - 41550, 3 3 2019.
    M. Ring, S. Wunderlich, D. Scheuring, D. Landes, and A. Hotho, “A survey of network-based intrusion detection data sets,” Computers & Security, vol. 86, pp. 147-167, Sep. 2019.
    LINCOLN LABORATORY, “1998 DARPA INTRUSION DETECTION EVALUATION DATASET,” 2 1998. [Online]. Available: https://www.ll.mit.edu/r-d/datasets/1998-darpa-intrusion-detection-evaluation-dataset. [Accessed 1 6 2021].
    The UCI KDD Archive, “KDD Cup 1999 Data,” 28 10 1999. [Online]. Available: https://kdd.ics.uci.edu/databases/kddcup99/kddcup99.html. [Accessed 1 6 2021].
    M. Tavallaee, E. Bagheri, W. Lu, and A. A. Ghorbani, “A Detailed Analysis of the KDD CUP 99 Data Set,” 2009 IEEE Symposium on Computational Intelligence for Security and Defense Applications, Ottawa, Canada, 8-10 July, 2009.
    N. Moustafa, IEEE student Member, and J. Slay, “UNSW-NB15: A Comprehensive Data set for Network Intrusion Detection systems,” 2015 Military Communications and Information Systems Conference (MilCIS), Canberra, ACT, Australia, 10-12 Nov., 2015.
    A. Gharib, I. Sharafaldin, A. H. Lashkari, and A. A. Ghorbani, “An Evaluation Framework for Intrusion Detection Dataset,” 2016 International Conference on Information Science and Security (ICISS), Pattaya, Thailand, 19-22 Dec., 2016.
    I. Sharafaldin, A. H. Lashkari, and A. A. Ghorbani, “Toward Generating a New Intrusion Detection Dataset and Intrusion Traffic Characterization,” 4th International Conference on Information Systems Security and Privacy (ICISSP 2018), Funchal, Portugal, 22-24 Jan., 2018.
    J. Seidl, “GoldenEye,” 15 10 2014. [Online]. Available: https://github.com/jseidl/GoldenEye. [Accessed 1 6 2021].
    A. I.Grafov, “Hulk DoS tool,” 17 10 2018. [Online]. Available: https://github.com/grafov/hulk/pulls. [Accessed 1 6 2021].
    T. K. Oo, “udpstorm,” 9 10 2015. [Online]. Available: https://github.com/toekhaing/udpstorm. [Accessed 1 6 2021].
    F. Gumus, C. O. Sakar, Z. Erdem, O. Kursun, “Online Naive Bayes classification for network intrusion detection,” 2014 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM 2014), Beijing, China, 17-20 Aug., 2014.
    V. N. Vapnik, “An overview of statistical learning theory,” IEEE Transactions on Neural Networks, vol. 10, no. 5, pp. 988 - 999, 1999.
    H. Zhang, Y. Li, Z. Lv, A. K. Sangaiah, and T. Huang, “A real-time and ubiquitous network attack detection based on deep belief network and support vector machine,” IEEE/CAA Journal of Automatica Sinica, vol. 7, no. 3, pp. 790-799, 2020.
    J. Kim, J. Kim, H. Kim, M. Shim, and E. Choi, “CNN-Based Network Intrusion Detection against Denial-of-Service Attacks,” Electronics 2020, vol. 9, no. 6, 2020.
    S. Hochreiter and J. Schmidhuber, “Long Short-Term Memory,” Neural Computation (1997), vol. 9, no. 8, pp. 1735-1780, 1997.
    J. Kim, J. Kim, H. L. T. Thu, and H. Kim, “Long Short Term Memory Recurrent Neural Network Classifier for Intrusion Detection,” 2016 International Conference on Platform Technology and Service (PlatCon), Jeju, Korea, 15-17 Feb, 2016.
    T. Su, H. Sun, J. Zhu, S. Wang, Y. Li, “BAT: Deep Learning Methods on Network Intrusion Detection Using NSL-KDD Dataset,” IEEE Access, vol. 8, pp. 29575 - 29585, 2020.
    莊易叡, “使用seq2seq、R-Transformer和TCN-BiLSTM方法的入侵檢測系統,” 國立成功大學工程科學系碩士在職專班, 台南, 2021.
    University of New Brunswick, “CICIDS 2017 dataset,” University of New Brunswick, 1 7 2017. [Online]. Available: https://www.unb.ca/cic/datasets/ids-2017.html. [Accessed 15 6 2021].
    P. Korshunov and S. Marcel, “Vulnerability assessment and detection of Deepfake videos,” 2019 International Conference on Biometrics (ICB 2019), Crete, Greece, 4-7 June, 2019.
    S. Huang and K. Lei, “IGAN-IDS: An imbalanced generative adversarial network towards intrusion detection system in ad-hoc networks,” Ad Hoc Networks, vol. 105, p. 102177, 2020.
    B. A. Tama, L. Nkenyereye, S.M. R. Islam, and K.-S. Kwak, “An Enhanced Anomaly Detection in Web Traffic Using a Stack of Classifier Ensemble,” IEEE Access, vol. 8, pp. 24120 - 24134, 4 2 2020.
    Torch Contributors, “Pytorch,” Torch Contributors., 2019. [Online]. Available: https://pytorch.org/. [Accessed 1 6 2021].
    S. Yegulalp, “Facebook brings GPU-powered machine learning to Python,” InfoWorld, 19 1 2017. [Online]. Available: https://www.infoworld.com/article/3159120/facebook-brings-gpu-powered-machine-learning-to-python.html. [Accessed 1 6 2021].
    M. H. Shahriar, N. I. Haque, M. A. Rahman, and M. Alonso, “G-IDS: Generative Adversarial Networks Assisted Intrusion Detection System,” 2020 IEEE 44th Annual Computers, Software, and Applications Conference (COMPSAC), Madrid, Spain, 13-17 July, 2020.

    下載圖示
    2026-09-17公開
    QR CODE