| 研究生: |
鍾肇沅 Chung, Chao-Yuan |
|---|---|
| 論文名稱: |
基於即時視覺關係偵測與三維地圖重建之AIoT環境安全監控系統 Real-time Visual Relationship Detection and 3D Map Reconstruction based AIoT Environment Surveillance-oriented Security System |
| 指導教授: |
李祖聖
Li, Tzuu-Hseng S. |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 電機工程學系 Department of Electrical Engineering |
| 論文出版年: | 2021 |
| 畢業學年度: | 109 |
| 語文別: | 英文 |
| 論文頁數: | 100 |
| 中文關鍵詞: | 卷積神經網路 、視覺關係偵測 、三維地圖重建 、物聯網 |
| 外文關鍵詞: | Deep Convolutional Neural Network, Visual Relationship Detection, 3D Map Reconstruction, Internet of Things |
| 相關次數: | 點閱:178 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
本論文提出一應用於倉儲物流環境的AIoT(Artificial Intelligence of Things)監控系統。物聯網建立主伺服器端與客戶端之間的關係,深度攝影機結合重建技術建立即時的三維地圖,結合兩項技術讓系統掌握工作環境中機器人的狀態與位置,接著整合視覺關係偵測系統,讓監控系統辨別環境中物體之間的相對關係並適時給予指令。首先建立物件辨識系統讓主伺服器端瞭解場景中的物體類別、位置以及相對關係,取得物體的類別以及邊界框。根據邊界框預先訓練一個卷積神經網路,用以獲取圖像中的物體特徵,將邊界框取得兩個具相對關係的物體特徵進行串接以及分類,即可獲得相對物體在環境中的相對關係。同時本論文使用ROS系統以及無線感測器系統建立物聯網,讓主伺服器端能夠即時取得環境中機器人的狀態,例如:電池消耗、馬達運轉、機械手運轉等等,進行即時的資訊傳遞以及命令控制,讓監控者能即時掌握機器人的狀態。最後為了降低監控者監控時的負擔,將許多二維的攝影機圖像轉換成三維地圖,讓監視者能夠一目了然整個工作環境。首先將相機圖像共同取特徵進行剛性轉換,找出環境四周深度攝影機之間的轉移矩陣,將其導入三維地圖重建,讓大量點雲資訊即時顯示,讓監控者可以更輕鬆的監視工作中的機器人,能透過三維地圖在危機發生時,即時在地圖中顯示問題區域,也同時顯示區域位置於控制介面上。在通過實際實驗結果顯示,本論文所提方法,可讓環境中的監控能夠透過人工智慧的方法去達成,且效果準確,能降低高度的人力資源,也能夠讓整個產線更加有效率。
This thesis proposes the Artificial Intelligence of Things (AIoT) surveillance system for warehousing and logistic environments. In order to establish the contact between the server and the clients, the Internet of Things (IoT) is implemented. A reconstruction method is applied to create real-time 3D map with depth camera. With the combination of the previous two technologies, the server is able to get the status and position from the robots in the working environment. A deep learning system is proposed to understand the relationship among the robots and give the instructions. To make the server comprehend the object categories, positions and relative relationships of the objects, the object recognition system is established to obtain the classes and bounding boxes of the objects in the environment at first. According to the bounding boxes, pre-trained convolutional neural network (CNN) is used to obtain the features of the objects in the images and classify the relationships from the concatenated feature between object and subject. At the same time, this thesis utilizes the ROS system and Wireless Sensor Network (WSN) to establish IoT system. The server can control the working robots and obtain their status immediately, such as battery consumption, motor operation, manipulator operation, etc. Moreover, to reduce the burden of the monitor, 2D camera images are converted into 3D maps to monitor the entire working environment at a glance. To find the extrinsic transformation matrix from the cameras, the same features of different camera images are performed rigid transform and merged into 3D map reconstruction method, so that a large amount of point cloud will be displayed in real-time. Finally, the experimental results show that the method proposed in this thesis can achieve the goal of monitoring and surveillance, which can reduce the height of human resources, and can also make the entire production line more efficient.
[1] A. Gupta and L. S. Davis, “Beyond nouns: Exploiting prepositions and comparative adjectives for learning visual classifiers,” in Proceedings of European Conference on Computer Vision. Springer Berlin Heidelberg, pp. 16-29, 2008.
[2] Y. W. Chao, Z. Wang, Y. He, J. Wang and J. Deng, “Hico: A benchmark for recognizing human-object interactions in images,” in Proceedings of the IEEE International Conference on Computer Vision, pp. 1017-1025, 2015.
[3] A. Gupta, A. Kembhavi and L. S. Davis, “Observing Human-Object Interactions: Using Spatial and Functional Compatibility for Recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1775-1789, 2009.
[4] V. Ramanathan, C. Li, J. Denge, et al., “Learning semantic relationships for better action retrieval in images,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1100-1109, 2015.
[5] B. Yao and L. Fei-Fei, “Modeling mutual context of object and human pose in human-object interaction activities,” in Proceedings of 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pp. 17-24, 2010.
[6] Y. Atzmon, J. Berant, V. Kezami, A. Globerson and G. Chechik, “Learning to generalize to new compositions in image understanding,” arXiv preprint arXiv:1608.07639, 2016.
[7] M. A. Sadeghi and A. Farhadi, “Recognition using visual phrases,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1745-1752, 2011.
[8] H. Zhang, Z. Kyaw, S. F. Chang and T. S. Chua, “Visual translation embedding network for visual relation detection,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5532-5540, 2017.
[9] X. Yang, H. Zhang and J. Cai, “Shuffle-then-assemble: Learning object-agnostic visual relationship features,” in Proceedings of the European Conference on Computer Vision (ECCV), pp. 36-52, 2018.
[10] C. Lu, R. Krishna, M. Bernstein and L. Fei-Fei, “Visual relationship detection with language priors,” in Proceedings of the European Conference on Computer Vision, Springer Cham, pp. 852-869, 2016.
[11] S. Sharifzadeh, S. M. Baharlou, M. Berrendorf, R. Koner and V. Tresp, “Improving visual relation detection using depth maps,” in Proceedings of 2020 25th International Conference on Pattern Recognition (ICPR), pp. 3597-3604, 2021.
[12] R. Krishna, R. Krishna, Y. Zhu, et al., “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” International Journal of Computer Vision, pp. 32-73, 2017.
[13] M. Berger, A. Tagliasacchi, L. Seversky, et al., “A survey of surface reconstruction from point clouds,” Computer Graphics Forum, pp. 301-329, 2017.
[14] K. Chen, Y. K. Lai and S. M. Hu, “3D indoor scene modeling from RGB-D data: A survey,” Computational Visual Media, pp. 267-278, 2015.
[15] F. Calakli and G. Taubin, “SSD: Smooth signed distance surface reconstruction,” Computer Graph. Forum, pp. 1993-2002, 2011.
[16] A. L. Chauve, P. Labatut and J. P. Pons, “Robust piecewise-planar 3D reconstruction and completion from large-scale unstructured point data,” in Proceedings of the 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pp. 1261-1268, 2010.
[17] Ö. Çiçek, A. Abdulkadir, S. S. Lienkamp, T. Brox and O. Ronneberger, “3D U-net: Learning dense volumetric segmentation from sparse annotation,” in International Conference on Medical Image Computing and Computer-assisted Intervention, pp. 424-432, 2016.
[18] D. Shin, Z. Ren, E. B. Sudderth and C. C. Fowlkes, “3D scene reconstruction with multi-layer depth and epipolar transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 2172-2182, 2019.
[19] R. A. Newcombe, S. Izadi, O. Hilliges, et al., “Kinectfusion: Real-time dense surface mapping and tracking,” in Proceedings of the 2011 10th IEEE International Symposium on Mixed and Augmented Reality, pp. 127-136, 2011.
[20] H. Cai, B. Xu, L. Jiang and A. V. Vasilakos, “IoT-Based Big Data Storage Systems in Cloud Computing: Perspectives and Challenges,” IEEE Internet of Things Journal, pp. 75-87, 2017
[21] M. A. Al Faruque and K. Vatanparvar, “Energy Management-as-a-Service Over Fog Computing Platform,” IEEE Internet of Things Journal, pp. 161-169, 2016.
[22] J. E. Plazas, S. Bimonte, C. de Vaulx, et al., “A Conceptual Data Model and Its Automatic Implementation for IoT-Based Business Intelligence Applications,” IEEE Internet of Things Journal, pp. 10719-10732, 2020.
[23] A. Bochkovskiy, C. Y. Wang and H. Y. M. Liao, “YOLOv4: Optimal Speed and Accuracy of Object Detection,” arXiv preprint arXiv:2004.10934. 2020.
[24] J. Redmon and A. Farhadi, “YOLOv3: An incremental improvement,” arXiv preprint arXiv:1804.02767, 2018.
[25] C. Y. Wang, H.-Y. M. Liao, I-H. Yeh, Y.-H. Wu, P.-Y. Chen and J.-W. Hsieh, “CSPNET: A new backbone that can enhance learning capability of CNN,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 390-391, 2020.
[26] P. Purkait, C. Zhao and C. Zach, “SPP-Net: Deep absolute pose regression with synthetic views,” arXiv preprint arXiv:1712.03452, 2017.
[27] J. Yang, X. Fu, Y. Hu, Y .Huang, X. Ding and J. Paisley, “PanNet: A Deep Network Architecture for Pan-Sharpening,” in Proceedings of the IEEE International Conference on Computer Vision, pp. 1753-1761, 2017.
[28] F. Iandola, M. Moskewicz, S. Karayev, R. Girshick, T. Darrell and K. Keutzer, “DenseNet: Implementing Efficient ConvNet Descriptor Pyramids,” arXiv preprint arXiv: 1404.1869, pp. 1-11, 2014.
[29] T. Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan and S. Belongie, “Feature pyramid networks for object detection,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 936-944, 2017.
[30] S. Ren, K. He, R. Girshick and J. Sun, “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1137-1149, 2017.
[31] R. Girshick, J. Donahue, T. Darrell, J. Malik, U. C. Berkeley and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 580-587, 2014.
[32] K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” arXiv preprint arXiv: 1409.1556, 2014.
[33] V. Christlein, L. Spranger, M. Seuret, A. Nicolaou, P. Kral and A. Maier, “Deep generalized max pooling,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2473-2480, 2014.
[34] GitHub-tzutalin/labelImg. [Online] https://github.com/tzutalin/labelImg.
[35] GitHub-wkentaro/Labelme [Online] https://github.com/wkentaro/labelme.
[36] Depth Camera D435i–Intel® RealSenseTM Depth and Tracking Cameras. [Online] https://www.intelrealsense.com/depth-camera-d435i/.
[37] P. C. Ng and S. Henikoff, “SIFT: Predicting amino acid changes that affect protein function,” Nucleic Acids Research, pp. 3812-3814, 2003.
[38] E. Rublee, V. Rabaud, K. Konolige and G. Bradski, “ORB: An efficient alternative to SIFT or SURF,” in 2011 International Conference on Computer Vision, pp. 2564-2571, 2011.
[39] H. Bay, T. Tuytelaars and L. V. Gool, “Surf: Speeded up robust features,” in European Conference on Computer Vision, pp. 404-417, 2006.
[40] D. W. Eggert, A. Lorusso and R. B. Fisher, “Estimating 3-D rigid body transformations: A comparison of four major algorithms,” Machine Vision and Applications, pp. 272–290, 1997.
[41] M. Pick and Z. Šimon, “Closed formulae for transformation of the cartesian coordinate system into a system of geodetic coordinates,” Studia Geophysica et Geodaetica, pp. 112-119, 1985.
[42] V. C. Klema and A. J. Laub, “The Singular Value Decomposition: Its Computation and Some Applications,” IEEE Transactions on Automatic Control, pp. 164-176, 1980.
[43] D. Werner, A. Al-Hamadi and P. Werner, “Truncated signed distance function: experiments on voxel size,” in Proceedings of the International Conference Image Analysis and Recognition, pp. 357-364, 2014.
[44] CUDA Toolkit | NVIDIA Developer. [Online] https://developer.nvidia.com/cuda-toolkit.
[45] Pycuda-PyPI. [Online] https://pypi.org/project/pycuda/.
[46] rviz - ROS Wiki. [Online] http://wiki.ros.org/rviz.
[47] ROS.org. [Online]: https://www.ros.org/.
[48] NVIDIA Jetson Nano Developer Kit. [Online] https://developer.nvidia.com/embedded/jetson-nano-developer-kit.
[49] Jetson Xavier NX Developer Kit. [Online] https://developer.nvidia.com/embedded/jetson-xavier-nx-devkit.
[50] Arduino Uno. [Online]: https://store.arduino.cc/usa/arduino-uno-rev3.
[51] TAL220. [Online] https://www.robotics.org.za/TAL220-20KG.
[52] MX-106T/R(2.0). [Online] https://emanual.robotis.com/docs/en/dxl/mx/mx-106-2/.
[53] H54-200-S500-R(A). [Online] https://emanual.robotis.com/docs/en/dxl/pro/h54-200-s500-ra/.
[54] H54-100-S500-R(A). [Online] https://emanual.robotis.com/docs/en/dxl/pro/h54-100-s500-ra/.
[55] H42-20-S300-R(A). [Online] https://emanual.robotis.com/docs/en/dxl/pro/h42-20-s300-ra/.
[56] ROG Rapture GT-AX11000. [Online] https://rog.asus.com/tw/networking/rog-rapture-gt-ax11000-model/.
[57] L. Mi and Z. Chen, “Hierarchical Graph Attention Network for Visual Relationship Detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13883-13892, 2020.