簡易檢索 / 詳目顯示

研究生: 陳源志
Chen, Yuan-Chih
論文名稱: 空間-時間轉譯與自我注意卷積神經網路於步步為營遊戲之理解與實現
Comprehension and Implementation of Quoridor by Spatial-Temporal Transformer and Self-Attention Convolutional Neural Network
指導教授: 李祖聖
Li, Tzuu-Hseng
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 電機工程學系
Department of Electrical Engineering
論文出版年: 2021
畢業學年度: 109
語文別: 英文
論文頁數: 85
中文關鍵詞: 卷積神經網路自我注意機制自我監督學習自然語言處理空間時間學習自動編碼模型代表學習
外文關鍵詞: Convolutional Neural Network, Self-Attention Mechanism, Self-Supervised Learning, Natural Language Processing, Spatial-Temporal Learning, Auto-Encoding Model, Representation Learning
相關次數: 點閱:153下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 機器人的思考模式可藉由學習人類的邏輯概念來達成,結合現今深度學習網路與自然語言學習對資料的擬合,將可能讓機器人了解是如何來思考。本論文主旨為透過多種神經網路學習步步為營遊戲來了解機器人如何學習一個遊戲。本論文第一部分透過圖像的方式,提出卷積網路結合空間基底自我注意機制與自我對弈的架構來學習空間上的概念。第二部分透過自然語言的訓練流程來理解整個遊戲,使用自然語言的概念結合空間與時間的資訊來學習遊戲。訓練的第一階段利用大量的隨機資訊來進行預訓練使模型理解遊戲規則,第二階段則透過自我對弈的結果進行模型的微調來學習獲勝的策略。在模擬實驗中可以看出使用自然語言的訓練流程,透過預訓練能大幅度提升模型的策略。第三部分提出全新的架構競技轉換器(Game Transformer, GaT)來詮釋空間與時間,藉由空間嵌入(embedding)與時間嵌入將空間與時間的輸入資訊分離,使神經網路能分別學習時間與空間的資訊。此架構與原本的轉換器基底的架構相比,能去除輸入空間與時間資訊結合可能帶來的稀疏性與干擾,進而提升模型的表現力。由模擬實驗可以看出空間嵌入的重要性,並且GaT成功地利用預訓練學習到時間嵌入。

    Nowadays, the intelligence of a robot can be mimicked by learning the human logic. By combining deep learning and natural language (NLP), we have the chance to make robot comprehend the text and rule of a game. The main purpose of this thesis is to understand how a robot learns to play board games by learning to play the Quoridor game with several kinds of neural networks. We can observe the distinct results from different aspects. The first network learns the Quoridor by spatial domain by using Convolutional Neural Network (CNN) with spatial self-attention mechanism and self-play. The second network learns the board game by natural language, where one can consider the rule of the board game as a kind of language. To this end, the learning approach is using a Bidirectional Encoder Representations from Transformers (BERT) based model with spatial-temporal input by self-play and self-supervised pretraining. In self-supervised pretraining, the network learns the rule of the Quoridor and the policy of playing Quoridor by finetuning with the self-play results. In the experiment, we can see the awesome result by applying rule pretraining to the language model. Finally, to improve the spatial and temporal representation, the Game Transformer (GaT) is purposed in this thesis. GaT separates the spatial domain and time domain information by position embedding and time embedding. Compared to general Transformers, GaT reduces the sparsity of the input and some potential confuse information in the input, which makes the network more informative and efficient. In the experiment, one can see the importance of the position embedding and the successful time embedding of the GaT.

    CONTENTS IV LIST OF FIGURES VII LIST OF TABLES X Chapter 1 Introduction 1 1.1 Motivation 1 1.2 Related Works 3 1.2.1 Residual based CNN 3 1.2.2 NLP 3 1.3 Thesis Organization 4 Chapter 2 Quoridor Game and Self-Play 6 2.1 Introduction 6 2.2 Quoridor 7 2.2.1 The rule of Quoridor 7 2.2.2 The improvement of the path finding algorithm 8 2.2.3 Quoridor setting in the experiment 10 2.3 Self-Play 11 2.3.1 AlphaZero 12 Chapter 3 Learning to Play Quoridor with Spatial Network 13 3.1 Introduction 13 3.2 Input Data 14 3.2.1 Image Representation 14 3.2.2 Image Representation with historical information 15 3.3 Spatial Networks 16 3.3.1 Introduction 16 3.3.2 Spatial Self-Attention 17 3.3.3 Squeeze and Excitation Block 19 3.3.4 Self-Attention + Squeeze and Excitation 20 3.4 Training Architecture 21 Chapter 4 Language Learning by BERT Based Language Model 23 4.1 Introduction 23 4.2 Word Embedding Representation 24 4.2.1 Word Representation Input 24 4.2.2 Word Embedding Output 24 4.2.3 Spatial Networks with Word Embedding Output 25 4.3 Related Work 26 4.3.1 BERT 26 4.3.2 ViT 29 4.4 Fine-tuning-Based Language Model 30 4.4.1 Game Rule Prediction (GRP) 30 4.4.2 Historical Action Replay (HAR) 31 4.5 Training Architecture 32 Chapter 5 Game Transformer (GaT) 34 5.1 Introduction 34 5.2 Independent Spatial-Temporal Input 35 5.2.1 Current Board State and Historical Actions 36 5.3 Game Transformer (GaT) 37 5.4 Time Embedding 37 5.5 Historical State Replay (HSR) 39 5.6 Training Architecture 39 Chapter 6 Experimental Results 42 6.1 Introduction 42 6.2 Experimental Setup 43 6.2.1 Representation 43 6.2.2 Self-play 44 6.2.3 Training 44 6.2.4 Optimization 46 6.3 Elo Rating 47 6.4 Experimental Results of Self-Attention 49 6.4.1 BertNet without Pretraining 50 6.4.2 BertNet with Pretraining 50 6.5 Experimental Results of Language Models 51 6.5.1 ViT without Pretraining 52 6.5.2 ViT with Pretraining 56 6.5.3 ViT_v2 with Pretraining 60 6.6 Experimental Results of GaT 64 6.6.1 Time Embedding of GaT 64 6.6.2 Visualize the Attention Map of GaT 64 6.6.3 Visualize the Temporal Attention Map of GaT 69 6.7 Summary 78 Chapter 7 Conclusions and Future Work 79 7.1 Conclusions 79 7.2 Future Work 80 REFERENCES 81

    [1] D. Silver et al., “Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm,” 2017, [Online]. Available: http://arxiv.org/abs/1712.01815.
    [2] J. Schrittwieser et al., “Mastering Atari, Go, chess and shogi by planning with a learned model,” Nature, vol. 588, no. 7839, pp. 604–609, 2020, doi: 10.1038/s41586-020-03051-4.
    [3] D. Silver, J. Schrittwieser, K. Simonyan, I. A.- Nature, and U. 2017, “Mastering the game of Go without human knowledge,” Nature, vol. 550, no. 7676, p. 354, 2016.
    [4] R. S. Sutton and A. G. Barto, “Reinforcement Learning: An Introduction,” IEEE Trans. Neural Networks, vol. 9, no. 5, pp. 1054–1054, 1998, doi: 10.1109/tnn.1998.712192.
    [5] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” Proc. IEEE Comput. Soc. Conf. Comput. Vis. Pattern Recognit., vol. 2016–Decem, pp. 770–778, 2016, doi: 10.1109/CVPR.2016.90.
    [6] C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the Inception Architecture for Computer Vision,” Proc. IEEE Comput. Soc. Conf. Comput. Vis. Pattern Recognit., vol. 2016–Decem, pp. 2818–2826, 2016, doi: 10.1109/CVPR.2016.308.
    [7] Simonyan.Karen and Zisserman.Andrew, “Very deep convolutional networks for large-scale image recognition,” 3rd Int. Conf. Learn. Represent. ICLR 2015 - Conf. Track Proc., 2015.
    [8] M. Längkvist, L. Karlsson, and A. Loutfi, “Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning,” Pattern Recognit. Lett., vol. 42, no. 1, pp. 11–24, 2014, [Online]. Available: http://arxiv.org/abs/1512.00567.
    [9] G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” Proc. - 30th IEEE Conf. Comput. Vis. Pattern Recognition, CVPR 2017, vol. 2017–Janua, pp. 2261–2269, 2017, doi: 10.1109/CVPR.2017.243.
    [10] H. Zhang et al., “ResNeSt: Split-Attention Networks,” 2020, [Online]. Available: http://arxiv.org/abs/2004.08955.
    [11] S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” Proc. - 30th IEEE Conf. Comput. Vis. Pattern Recognition, CVPR 2017, vol. 2017–Janua, pp. 5987–5995, 2017, doi: 10.1109/CVPR.2017.634.
    [12] H. A. Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications,” Comput. Vis. Pattern Recognit., vol. 14, no. 2, pp. 53–57, 2009, [Online]. Available: http://arxiv.org/abs/1704.04861.
    [13] M. Tan and Q. V. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” 36th Int. Conf. Mach. Learn. ICML 2019, vol. 2019–June, pp. 10691–10700, 2019.
    [14] A. Torfi, R. A. Shirvani, Y. Keneshloo, N. Tavaf, and E. A. Fox, “Natural Language Processing Advancements By Deep Learning: A Survey,” 2020, [Online]. Available: http://arxiv.org/abs/2003.01200.
    [15] J. Devlin, M.-W. Chang, K. Lee, K. T. Google, and A. I. Language, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” Naacl-Hlt 2019, 2018, [Online]. Available: https://github.com/tensorflow/tensor2tensor.
    [16] Y. Liu et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” 2019, [Online]. Available: http://arxiv.org/abs/1907.11692.
    [17] K. Clark, M.-T. Luong, Q. V. Le, and C. D. Manning, “ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators,” 2020, [Online]. Available: http://arxiv.org/abs/2003.10555.
    [18] Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut, “ALBERT: A Lite BERT for Self-supervised Learning of Language Representations,” 2019, [Online]. Available: http://arxiv.org/abs/1909.11942.
    [19] A. Vaswani et al., “Attention is all you need,” Adv. Neural Inf. Process. Syst., vol. 2017–Decem, pp. 5999–6009, 2017.
    [20] I. J. Goodfellow et al., “Generative adversarial nets,” Adv. Neural Inf. Process. Syst., vol. 3, no. January, pp. 2672–2680, 2014, doi: 10.3156/jsoft.29.5_177_2.
    [21] Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. Salakhutdinov, and Q. V. Le, “XLNet: Generalized autoregressive pretraining for language understanding,” Adv. Neural Inf. Process. Syst., vol. 32, 2019.
    [22] Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. V. Le, and R. Salakhutdinov, “Transformer-XL: Attentive language models beyond a fixed-length context,” ACL 2019 - 57th Annu. Meet. Assoc. Comput. Linguist. Proc. Conf., pp. 2978–2988, 2020, doi: 10.18653/v1/p19-1285.
    [23] V. M. Respall, J. A. Brown, and H. Aslam, “Monte carlo tree search for quoridor,” 19th Int. Conf. Intell. Games Simulation, GAME-ON 2018, pp. 5–9, 2018.
    [24] P. Mertens, “A Quoridor-playing Agent,” Bachelor Thesis, Dep. Knowl. …, 2006, [Online]. Available: http://www.unimaas.nl/games/files/bsc/Mertens_BSc-paper.pdf.
    [25] M. Campbell, A. J. Hoane, and F. H. Hsu, “Deep Blue,” Artif. Intell., vol. 134, no. 1–2, pp. 57–83, 2002, doi: 10.1016/S0004-3702(01)00129-1.
    [26] R. Coulom, “Efficient selectivity and backup operators in Monte-Carlo tree search,” Lect. Notes Comput. Sci. (including Subser. Lect. Notes Artif. Intell. Lect. Notes Bioinformatics), vol. 4630 LNCS, pp. 72–83, 2007, doi: 10.1007/978-3-540-75538-8_7.
    [27] H. Zhang, I. Goodfellow, D. Metaxas, and A. Odena, “Self-attention generative adversarial networks,” 36th Int. Conf. Mach. Learn. ICML 2019, vol. 2019–June, pp. 12744–12753, 2019.
    [28] J. Hu, L. Shen, S. Albanie, G. Sun, and E. Wu, “Squeeze-and-Excitation Networks,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 42, no. 8, pp. 2011–2023, 2020, doi: 10.1109/TPAMI.2019.2913372.
    [29] A. Howard et al., “Searching for mobileNetV3,” Proc. IEEE Int. Conf. Comput. Vis., vol. 2019–Octob, pp. 1314–1324, 2019, doi: 10.1109/ICCV.2019.00140.
    [30] S. Woo, J. Park, J. Y. Lee, and I. S. Kweon, “CBAM: Convolutional block attention module,” Lect. Notes Comput. Sci. (including Subser. Lect. Notes Artif. Intell. Lect. Notes Bioinformatics), vol. 11211 LNCS, pp. 3–19, 2018, doi: 10.1007/978-3-030-01234-2_1.
    [31] A. Dosovitskiy et al., “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” 2020, [Online]. Available: http://arxiv.org/abs/2010.11929.
    [32] L. Glendenning, “Mastering Quoridor,” 2005.
    [33] S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” 32nd Int. Conf. Mach. Learn. ICML 2015, vol. 1, pp. 448–456, 2015.
    [34] S. Panigrahi, A. Nanda, and T. Swarnkar, “A Survey on Transfer Learning,” Smart Innov. Syst. Technol., vol. 194, pp. 781–789, 2021, doi: 10.1007/978-981-15-5971-6_83.
    [35] G. I. Parisi, R. Kemker, J. L. Part, C. Kanan, and S. Wermter, “Continual lifelong learning with neural networks: A review,” Neural Networks, vol. 113, pp. 54–71, 2019, doi: 10.1016/j.neunet.2019.01.012.
    [36] R. Coulom, “Whole-history rating: A Bayesian rating system for players of time-varying strength,” Int. Conf. Artif. Neural Networks, vol. 5131 LNCS, pp. 113–124, 2008, doi: 10.1007/978-3-540-87608-3_11.

    下載圖示
    2026-08-27公開
    QR CODE