| 研究生: |
向韋豪 Hsiang, Wei-Hao |
|---|---|
| 論文名稱: |
應用於自駕技術之空間注意力車道檢測網路 End-to-End Lane Detection Networks with Spatial Attention Mechanism for Autonomous Driving |
| 指導教授: |
楊家輝
Yang, Jar-Ferr |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 電腦與通信工程研究所 Institute of Computer & Communication Engineering |
| 論文出版年: | 2021 |
| 畢業學年度: | 109 |
| 語文別: | 英文 |
| 論文頁數: | 53 |
| 中文關鍵詞: | 深度學習 、自動駕駛 、車道線辨識 、影像分割 、空間注意力 、跨級融合 、卷積神經網路 |
| 外文關鍵詞: | deep learning, autonomous driving, lane detection, image segmentation, spatial attention, cross-level fusion, convolutional neural networks |
| 相關次數: | 點閱:111 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
隨著電腦視覺以及深度學習的發展,在自動駕駛領域中有越來越多的系統都引入了深度學習的技術。對於自駕車來說,如何有效率的去偵測道路駕駛場景中的車道位置是一項非常關鍵的任務。在本論文中,我們基於近幾年的影像分割網路和車道線分支網路的概念設計出更有效率的卷積神經網路偵測系統。藉由自編碼結構來達到逐像素的類別預測,並且透過兩個分支分別輸出二值化分割與實例分割的特徵圖,最後將兩分支的輸出合併來達到預測不同車道線的分割結果。此外,我們引入兩個輔助方法來針對細長型且沒有豐富的特徵資訊的車道線物體做加強。首先,我們在編碼器末端加入平行化空間注意力網路(PSAN),以強化細長型物體在空間關係上的特徵。接著我們通過跨級融合塊(CFB)來增強車道邊緣的特徵並做資訊的補償。跨級融合塊能過濾淺層特徵圖中非目標物體的特徵雜訊。從實驗結果可以發現,本論文在兩個車道線影像數據集中都能夠得到更精確的預測結果,並且同時減少了計算時間與成本。
With the development of computer vision and deep learning, more and more systems in the field of autonomous driving are introducing deep learning technology. It is a critical task to efficiently detect the lane positions in the road driving scenario. In this thesis, we design a more efficient detection system based on the recent concepts of image segmentation networks and lane branching networks. The auto-encoder structure is used to achieve pixel-wise class prediction, and the binary segmentation and instance segmentation feature maps are retrieved from two branches, and the outputs of the two branches are combined to predict the segmentation results of the lanes. In addition, we introduce two auxiliary methods to enhance lane objects that are long and thin and do not have rich feature information. First, we add a parallel spatial attention network at the end of the encoder to enhance the spatial relationship features of elongated objects. Then, we compensate the information by enhancing the lane edge features with cross-level fusion blocks. The cross-level fusion block can filter the feature noise of non-target objects in the low-level feature map. The experimental results show that the proposed network can obtain more accurate prediction results tested in two famous lane image datasets while it also reduces the computation time.
[1] D. Neven, B. D. Brabandere, S. Georgoulis, M. Proesmans and L. V. Gool, "Towards End-to-End Lane Detection: an Instance Segmentation Approach," Proc. of 2018 IEEE Intelligent Vehicles Symposium (IV), 2018, pp. 286-291, doi: 10.1109/IVS.2018.8500547.
[2] The tuSimple lane challange, http://benchmark.tusimple.ai/
[3] X. Pan, J. Shi, P. Luo, X. Wang, and X. Tang, “Spatial as Deep: Spatial CNN for Traffic Scene Understanding”, AAAI, vol. 32, no. 1, Apr. 2018.
[4] H. Jung, J. Min and J. Kim, "An efficient lane detection algorithm for lane departure detection," Proc. of 2013 IEEE Intelligent Vehicles Symposium (IV), 2013, pp. 976-981, doi: 10.1109/IVS.2013.6629593.
[5] G. Liu, F. Wörgötter and I. Markelić, "Combining Statistical Hough Transform and Particle Filter for robust lane detection and tracking," Proc. of 2010 IEEE Intelligent Vehicles Symposium, 2010, pp. 993-997, doi: 10.1109/IVS.2010.5548021.
[6] J. Long, E. Shelhamer, and T. Darrell. (2015). Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3431-3440).
[7] V. Badrinarayanan, A. Kendall and R. Cipolla, "SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation," in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 12, pp. 2481-2495, 1 Dec. 2017, doi: 10.1109/TPAMI.2016.2644615.
[8] A. Paszke, A. Chaurasia, S. Kim, and E. Culurciello, “Enet: A deep neural network architecture for real-time semantic segmentation,” arXiv preprint arXiv:1606.02147, 2016.
[9] K. He, X. Zhang, S. Ren, and J. Sun. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770-778).
[10] H. Noh, S. Hong, and B. Han. (2015). Learning deconvolution network for semantic segmentation. In Proceedings of the IEEE international conference on computer vision (pp. 1520-1528).
[11] S. Liu, L. Qi, H. Qin, J. Shi, and J. Jia. (2018). Path aggregation network for instance segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 8759-8768).
[12] T. Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie. (2017). Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 2117-2125).
[13] Q. Zhao, T. Sheng, Y. Wang, Z. Tang, Y. Chen, L. Cai, and H. Ling. (2019). M2Det: A Single-Shot Object Detector Based on Multi-Level Feature Pyramid Network. In Proceedings of the AAAI Conference on Artificial Intelligence, 33(01), 9259-9266.
[14] Y. Hou, Z. Ma, C. Liu, and C. C. Loy, “Learning lightweight lane detection cnns by self attention distillation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 1013–1021.
[15] T. Verelst, and T. Tuytelaars. (2020). Dynamic convolutions: Exploiting spatial sparsity for faster inference. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 2320-2329).
[16] Z. Zhang, C. Lan, W. Zeng, X. Jin, and Z. Chen. (2020). Relation-aware global attention for person re-identification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 3186-3195).
[17] R. Niu, X. Sun, Y. Tian, W. Diao, K. Chen and K. Fu, "Hybrid Multiple Attention Network for Semantic Segmentation in Aerial Images," in IEEE Transactions on Geoscience and Remote Sensing, doi: 10.1109/TGRS.2021.3065112.
[18] F. Chollet. (2017). Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1251-1258).
[19] A. Howard, M. Sandler, G. Chu, L. C. Chen, B. Chen, M. Tan, ... and H. Adam. (2019). Searching for mobilenetv3. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 1314-1324).
[20] X. Zhong, O. Gong, W. Huang, L. Li and H. Xia, "Squeeze-and-Excitation Wide Residual Networks in Image Classification," Proc. of 2019 IEEE International Conference on Image Processing (ICIP), Taipei, Taiwan, 2019, pp. 395-399, doi: 10.1109/ICIP.2019.8803000.
[21] X. Wang, R. Girshick, A. Gupta, & K. He (2018). "Non-local neural networks." In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 7794-7803).
[22] J. Fu, J. Liu, J. Jiang, Y. Li, Y. Bao and H. Lu, "Scene Segmentation With Dual Relation-Aware Attention Network," in IEEE Transactions on Neural Networks and Learning Systems, doi: 10.1109/TNNLS.2020.3006524.
[23] F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, and K. Keutzer. (2016). SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and< 0.5 MB model size. arXiv preprint arXiv:1602.07360.
[24] M. Sandler, A. Howard, M, Zhu, A. Zhmoginov, and L. C. Chen. (2018). Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 4510-4520).
[25] X. Zhang, X. Zhou, M. Lin, and J. Sun. (2018). Shufflenet: An extremely efficient convolutional neural network for mobile devices. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 6848-6856).
[26] G. Huang, S. Liu, L. Van der Maaten, and K. Q. Weinberger. (2018). Condensenet: An efficient densenet using learned group convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 2752-2761).
[27] Z. Yan, X. Li, M. Li, W. Zuo, and S. Shan. (2018). Shift-net: Image inpainting via deep feature rearrangement. In Proceedings of the European conference on computer vision (ECCV) (pp. 1-17).
[28] P. Ramachandran, B. Zoph, and Q. V. Le. (2017). Searching for activation functions. arXiv preprint arXiv:1710.05941.
[29] X. Li, W. Wang, X. Hu, J. Yang. (2019). Selective kernel networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 510-519).
[30] B. De Brabandere, D. Neven, and L. Van Gool. (2017). Semantic instance segmentation with a discriminative loss function. arXiv preprint arXiv:1708.02551.
[31] W. Yang, Y. Cheng and P. Chung, "Improved Lane Detection With Multilevel Features in Branch Convolutional Neural Networks," in IEEE Access, vol. 7, pp. 173148-173156, 2019, doi: 10.1109/ACCESS.2019.2957053.