| 研究生: |
張弘遠 Chang, Hung-Yuan |
|---|---|
| 論文名稱: |
多智能體增強式學習之自私生成對抗網路應用於多自走車排程與路徑規劃研究 Generative Adversarial Selfish Networks Based on Multi-Agent Reinforcement Learning for Multi-AGV Path Planning and Scheduling |
| 指導教授: |
李祖聖
Li, Tzuu-Hseng |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 電機工程學系 Department of Electrical Engineering |
| 論文出版年: | 2021 |
| 畢業學年度: | 109 |
| 語文別: | 英文 |
| 論文頁數: | 69 |
| 中文關鍵詞: | 自私對抗網路 、多智能體增強式學習 、自走車 、路經規劃 、排程 |
| 外文關鍵詞: | GASNet, Multi-Agent Reinforcement Learning, AGV, Path-planning, Scheduling |
| 相關次數: | 點閱:196 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
在多智能體增強式學習中,自私與合作被廣泛地討論。而在眾多議題當中,「如何建立強健且高效的溝通方式」是本篇討論的重點。本論文提出了自私生成對抗網路(Generative Adversarial Selfish Networks, GASNet),其中透過判斷網路判斷在自私或是合作下的表現來進行自我對抗進而不斷更新生成網路,因此生成網路能產生出更好的關係圖。GASNet包含生成網路、判斷網路以及圖神經網路,其中生成網路能將各個智能體的想法去建構出關係圖,而判斷網路會根據該關係圖去產生各智能體之自私程度,以用來判斷是否採取自私或合作行為,最後,圖神經網路會不斷地去整合關係圖、自私程度以及各智能體想法以決定動作。在本論文所設計之模擬環境中,包含簡單、困難兩種程度中的實驗結果可以佐證,所設計之GASNet演算法表現遠超過目前最新演算法,並且也經由實驗顯示出智能體在群聚時,會與關鍵智能體或鄰近智能體進行溝通。本論文並將此演算法應用在真實多自走車系統上,達到可以同時解決多自走車之路徑規劃與排程之成效。另外,該系統透過Kubernetes讓多自走車系統能夠更簡單地去擴展、自動化佈署以及管理,並藉由真實的實驗中了解到,多智能體增強式學習不僅學到路徑規劃,並且還自主地學到排程、分群、工作分配的能力與成功通過測試,因此多智能體增強式學習可優化整體系統。
By Multi-Agent Reinforcement Learning (MARL), selfish and cooperative agents have been widely researched for last decades. Among all of issues, how to construct the robust and high-performance communication is considered in this thesis. We proposed a self-competition algorithm by a discriminator to judge the performance on the behavior of selfishness and cooperation, so the generator is able to create more and more robust relationship graphs, called Generative Adversarial Selfish Networks (GASNet). GASNet comprises a Generator Adversarial Networks (GAN) and a Graph Neural Networks (GNN). The generator figures out the graphs which are passed to discriminator for a selfish extent and it is evaluated which selfish or cooperative behavior is better. Eventually, GNN estimates their actions based on the graphs, selfish extent and aggregation of all the thoughts. Empirically, the simulation results, including two levels simple and hard, show GASNet outperforms the state-of-the-art algorithms and also indicate that agents only coordinate and communicate with dominant agents or close agents. Furthermore, we implement the developed algorithm into real multiple Autonomous Guided Vehicles (Multi-AGV) system. The AI Multi-AGV system is established to complete path-planning and scheduling simultaneously. Moreover, the AI Multi-AGV system is based on Kubernetes, which makes itself easily scaling, automatic deployment and management. Experiments demonstrate MARL autonomously learn scheduling, clustering and workload balancing and a huge success of alpha testing.
[1]. V. M. Janik, “Pitfalls in the categorization of behaviour: a comparison of dolphin whistle classification methods,” Animal Behaviour, vol. 57, pp. 133-143, 1999.
[2]. C. Berner, G. Brockman, B. Chan, V. Cheung, P. Dębiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, R. Józefowicz, S. Gray, C. Olsson, J. Pachocki, M. Petrov, H. P. d.O. Pinto, J. Raiman, T. Salimans, J. Schlatter, J. Schneider, S. Sidor, I. Sutskever, J. Tang, F. Wolski, and S. Zhang, “Dota 2 with large scale deep reinforcement learning,” arXiv:1912.06680, 2019.
[3]. M. Samvelyan, T. Rashid, C. S. de Witt, G. Farquhar, N. Nardelli, T. G. J. Rudner, C. M. Hung, P. HS Torr, J. Foerster, and S. Whiteson, “The starcraft multi-agent challenge,” in Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems, pp. 2186-2188, 2019.
[4]. H. Zhang, S. Feng, C. Liu, Y. Ding, Y. Zhu, Z. Zhou, W. Zhang, Y. Yu, H. Jin, and Z. Li, “Cityflow: A multi-agent reinforcement learning environment for large scale city traffic scenario,” in The World Wide Web Conference, pp. 3620-3624, 2019.
[5]. W. Fan, Y. Ma, Q. Li, Y. He, E. Zhao, J. Tang, and D. Yin, “Graph neural networks for social recommendation,” arXiv:1902.07243, 2019.
[6]. X. Wang, X. He, M. Wang, F. Feng, and T. S. Chua, “Neural graph collaborative filtering,” in Proceedings of the 42nd international ACM SIGIR conference on Research and development in Information Retrieval. pp. 165-174, 2019.
[7]. H. Wang, M. Zhao, X. Xie, W. Li, and M. Guo, “Knowledge graph convolutional networks for recommender systems,” in The World Wide Web Conference, pp. 3307-3313, 2019.
[8]. K. Huang, C. Xiao, L. Glass, M. Zitnik, and J. Sun, “SkipGNN: Predicting molecular interactions with Skip-Graph Networks,” arXiv:2004.14949, 2020.
[9]. H. Ma, Y. Bian, Y. Rong, W. Huang, T. Xu, W. Xie, G. Ye, and J. Huang “Multi-view graph neural networks for molecular property prediction,” arXiv preprint arXiv:2005.13607, 2020.
[10]. M. Hu, W. Liu, K. Peng, X. Ma, W. Cheng, J. Liu, and B. Li “Joint routing and scheduling for vehicle-assisted multidrone surveillance,” in IEEE Internet of Things Journal, vol. 6, no. 2, pp. 1781-1790, April 2019.
[11]. C. Yin, Z. Xiao, X. Cao, X. Xi, P. Yang, and D. Wu, “Offline and online search: UAV multiobjective path planning under dynamic urban environment,” in IEEE Internet of Things Journal, vol. 5, no. 2, pp. 546-558, April 2018.
[12]. J. Li, Y. Xiong, J. She, and M. Wu, “A path planning method for sweep coverage with multiple UAVs,” in IEEE Internet of Things Journal, vol. 7, no. 9, pp. 8967-8978, Sept. 2020.
[13]. D. Liu, S. Wang, Z. Wen, L. Cheng, M. Wen, and Y. -C. Wu, “Edge learning with Unmanned Ground Vehicle: joint path, energy, and sample size planning,” in IEEE Internet of Things Journal, vol. 8, no. 4, pp. 2959-2975, 15 Feb.15, 2021.
[14]. L. -L. Wang, J. -S. Gui, X. -H. Deng, F. Zeng, and Z. -F. Kuang, “Routing algorithm based on vehicle position analysis for internet of vehicles,” in IEEE Internet of Things Journal, vol. 7, no. 12, pp. 11701-11712, Dec. 2020.
[15]. Y. Zheng, J. Wang, and K. Li, “Smoothing traffic flow via control of Autonomous Vehicles,” in IEEE Internet of Things Journal, vol. 7, no. 5, pp. 3882-3896, May 2020.
[16]. S. Sainbayar, A. Szlam, and R. Fergus, “Learning multiagent communication with backpropagation,” arXiv:1605.07736, 2016.
[17]. S. Iqbal and F. Sha, “Actor-Attention-Critic for Multi-Agent Reinforcement Learning,” arXiv:1810.02912, 2018.
[18]. Y. Liu, W. Wang, Y. Hu, J. Hao, X. Chen, and Y. Gao, “Multi-agent game abstraction via graph attention neural network,” arXiv:1911.10715, 2019.
[19]. K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio, “Show, Attend and Tell: Neural Image Caption Generation with Visual Attention,” in International Conf. on Machine Learning, pp. 2048-2057, 2015.
[20]. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems, pp. 5998-6008, 2017.
[21]. Z. Han, D. Wang, F. Liu, and Z. Zhao, “Multi-AGV path planning with double-path constraints by using an improved genetic algorithm,” PLoS ONE, vol. 12, no. 7, 2017.
[22]. C. Wang, L. Wang, J. Qin, Z. Wu, L. Duan, Z. Li, M. Caom, X. Ou, X. Su, W. Li, Z. Lu, M. Li, Y. Wang, J. Long, M. Huang, Y. Li, and Q. Wang, “Path planning of automated guided vehicles based on improved A-Star algorithm,” in 2015 IEEE International Conference on Information and Automation, pp. 2071-2086, 2015.
[23]. B. Shirazi, H. Fazlollahtabar, and I. Mahdavi, “A six sigma based multi-objective optimization for machine grouping control in flexible cellular manufacturing systems with guide-path flexibility,” Advances in Engineering Software, vol. 41, pp. 865-873, 2010.
[24]. J. Jian and X. H. Zhang, “Multi AGV scheduling problem in automated container terminal,” Journal of Marine Science and Technology, vol. 24, no. 1, pp. 32-38 2016.
[25]. B. S. P. Reddy and C. S. P Rao, “A hybrid multi-objective GA for simultaneous scheduling of machines and AGVs in FMS,” The International Journal of Advanced Manufacturing Technology, vol. 31, pp. 602-613, 2006.
[26]. S. S. Sankar, S. G. Ponnambalam, and M. Gurumarimuthu, “Scheduling flexible manufacturing systems using parallelization of multi-objective evolutionary algorithms,” The International Journal of Advanced Manufacturing Technology, vol. 30, no. 3, pp. 279-285, 2006.
[27]. M. L. Littman, “Markov games as a framework for multi-agent reinforcement learning,” in Proceedings of the Eleventh International Conf. on Machine Learning, vol. 157, pp. 157-163, 1994.
[28]. T. N. Kipf and M. Welling, “Semi-supervised classification with graph convolutional networks,” arXiv preprint arXiv:1609.02907, 2016.
[29]. D. K. Hammond, P. Vandergheynst, and R. Gribonval, “Wavelets on graphs via spectral graph theory,” Applied and Computational Harmonic Analysis, vol. 30, no.2, pp. 129-150, 2011.
[30]. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” arXiv preprint arXiv:1706.03762, 2017.
[31]. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative Adversarial Networks,” arXiv:1406.2661, 2014.
[32]. J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proceedings of the IEEE Conf. on Computer Vision and Pattern Recognition, pp. 7132-7141, 2018.
[33]. T. Karras, T. Aila, S. Laine, and J. Lehtinen, “Progressive growing of GANs for improved quality, stability, and variation,” arXiv:1710.10196v3, 2017.
[34]. P. Isola, J. Y. Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation with conditional adversarial networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1125-1134, 2017.
[35]. B. Baker, “Emergent Reciprocity and Team Formation from Randomized Uncertain Social Preferences,” arXiv preprint arXiv:2011.05373, 2020.
[36]. H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 30, no. 1, 2016.
[37]. V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller, “Playing ATARI with deep reinforcement learning,” arXiv preprint arXiv:1312.5602, 2013
[38]. The Amazon Robotics Family: Kiva, Pegasus, Xanthus, and more… [Online] Available: https://www.allaboutlean.com/amazon-robotics-family/
[39]. Tugger AGVs: Workhorse AGV for Multi-Trailer Automatic Material Transport [Online] Available: https://www.dematic.com/en/products/products-overview/agv-systems/tugger-agvs/
[40]. Robotis e-Manual MX-106R, MX-106T (Protocol 2.0) [Online]. Available: https://emanual.robotis.com/docs/en/dxl/mx/mx-106-2/
[41]. Slamtec Mapper M1M1 ToF Laser Scanner Kit - 20M Range [Online]. Available: http://bucket.download.slamtec.com/975678f44875b6cb4db5ecf30202163c8910519b/SLAMTEC_mapper_datasheet_M1M1_v1.2_en.pdf
[42]. S. Su, X. Zeng, S. Song, et al., “Positioning accuracy improvement of automated guided vehicles based on a novel magnetic tracking approach,” IEEE Intelligent Transportation Systems Magazine, vol. 12, no. 4, pp. 138-148, winter 2020.
[43]. Nvidia Nano [Online]. Available: https://www.nvidia.com/zh-tw/autonomous-machines/embedded-systems/jetson-nano/
[44]. Docker Introduction — What You Need To Know To Start Creating Containers [Online]. Available: https://medium.com/zero-equals-false/docker-introduction-what-you-need-to-know-to-start-creating-containers-8ffaf064930a
[45]. Z. Zhang, Q. Guo, J. Chen, and P. Yuan, “Collision-free route planning for multiple AGVs in an automated warehouse based on collision classification,” IEEE Access, vol. 6, pp. 26022-26035, 2018.