| 研究生: |
曾丞平 Tseng, Cheng-Ping |
|---|---|
| 論文名稱: |
使用3D-CNN與Bi-GRU結合注意機制之WiFi通道狀態資訊人體骨幹姿態即時重建 Wifi Channel State Information Real-Time Reconstruction of Human Skeleton Pose Using 3D-CNN and Bi-GRU with Attention Mechanism |
| 指導教授: |
賴槿峰
Lai, Chin-Feng |
| 學位類別: |
碩士 Master |
| 系所名稱: |
工學院 - 工程科學系 Department of Engineering Science |
| 論文出版年: | 2021 |
| 畢業學年度: | 109 |
| 語文別: | 中文 |
| 論文頁數: | 60 |
| 中文關鍵詞: | 通道狀態資訊 、人體姿態重建 、深度學習 、注意機制 |
| 外文關鍵詞: | Channel State Information, Human Posture Reconstruction, Deep Learning, Attention Mechanism |
| 相關次數: | 點閱:194 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
未來智慧感測的需求會越來越多,然而以影像與音訊為主的感測器,由於環境因素、資訊安全與個人隱私,有的並不合適。使用無線設備 WiFi 通道狀態資訊(Channel State Information, CSI)的人體行為感測系統,已在室內定位、手勢辨識、活動檢測、人數統計有相關研究,然而人體的智慧感測並沒有一個全方位的解決辦法,無法對人類活動進行全面性的完整識別。
因此本研究透過深度神經網路,使 CSI 數據與 Kinect 裝置獲取的人體姿態資料相互匹配,進行更詳細的即時人體姿態重建。將 CSI 訊號的振幅和相位透過 3D-CNN 與 Bi-GRU 獲取空間與時間特徵,利用 Encoder-Decoder 架構結合注意機制,使特徵向量重建成連續的三維空間人體骨架資訊,以獲取更完整的人體動作,得以在智慧感測與互動領域有更完善的應用。同時針對本研究提出數種模型判斷指標,以檢視實驗的成效。
透過本研究的結果可發現,人體主幹的重建成果較佳,得以完整呈現人體在實驗環境的位置,四肢的重建成果雖說較差,但仍可大致看出人體活動的姿態。此套系統結合了視覺人體辨識的精細度及連續性,與 CSI 人體感測的方便性和安全性,勢必會隨著物聯網科技的發展而有極大的應用價值。
The demand for intelligent sensors will increase in the future, however, image and audio-based sensors are not suitable due to environmental factors, information security and personal privacy.
Therefore, this study matches the CSI data with the human pose data obtained from the Kinect device through a deep neural network to perform more detailed real-time human pose reconstruction. The spatial and temporal features of CSI signal amplitude and phase are obtained by 3D-CNN and Bi-GRU, and the feature vector is reconstructed into a continuous three-dimensional human skeleton information by combining the Encoder-Decoder architecture with the attention mechanism to obtain a more complete human body movement for better application in the field of intelligent sensing and interaction. Several model judgment metrics are proposed to examine the effectiveness of the experiment.
The results of this study show that the reconstruction of the human trunk is better, and the position of the human body in the experimental environment can be fully represented, while the reconstruction of the extremities is worse, but the posture of the human body can still be generally seen.
This system combines the precision and continuity of visual human recognition with the convenience and safety of CSI human body sensing, and will have great application value with the development of Internet of Things technology.
[1] C.-F. Lai, S.-Y. Chang, H.-C. Chao, and Y.-M. Huang, “Detection of cognitive injured body region using multiple triaxial accelerometers for elderly falling,” IEEE Sensors Journal, vol. 11, no. 3, pp. 763–770, 2010.
[2] Y.-S. Lu, C.-F. Lai, C.-C. Hu, Y.-M. Huang, and X.-H. Ge, “Path loss exponent estimation for indoor wireless sensor positioning,” KSII Transactions on Internet and Information Systems (TIIS), vol. 4, no. 3, pp. 243–257, 2010.
[3] D. R. Beddiar, B. Nini, M. Sabokrou, and A. Hadid, “Vision-based human activity recognition: a survey,” Multimedia Tools and Applications, vol. 79, no. 41, pp. 30509– 30555, 2020.
[4] C. Zhang, H. Li, X. Wang, and X. Yang, “Cross-scene crowd counting via deep convolutional neural networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 833–841, 2015.
[5] O. D. Lara and M. A. Labrador, “A survey on human activity recognition using wearable sensors,” IEEE Communications Surveys Tutorials, vol. 15, no. 3, pp. 1192–1209, 2013.
[6] Z. Iqbal, D. Luo, P. Henry, S. Kazemifar, T. Rozario, Y. Yan, K. Westover, W. Lu, D. Nguyen, T. Long, et al., “Accurate real time localization tracking in a clinical environment using bluetooth low energy and deep learning,” PloS one, vol. 13, no. 10, p. e0205392, 2018.
[7] C. Jiang, J. Shen, S. Chen, Y. Chen, D. Liu, and Y. Bo, “Uwb nlos/los classification using deep learning method,” IEEE Communications Letters, vol. 24, no. 10, pp. 2226– 2230, 2020.
[8] Y. Lee, J.-Y. Park, Y.-W. Choi, H.-K. Park, S.-H. Cho, S. H. Cho, and Y.-H. Lim, “A novel non-contact heart rate monitor using impulse-radio ultra-wideband (ir-uwb) radar technology,” Scientific reports, vol. 8, no. 1, pp. 1–10, 2018.
[9] A. R. Jiménez Ruiz and F. Seco Granja, “Comparing ubisense, bespoon, and decawave uwb location systems: Indoor performance analysis,” IEEE Transactions on Instrumentation and Measurement, vol. 66, no. 8, pp. 2106–2117, 2017.
[10] Z. Wang, K. Jiang, Y. Hou, W. Dou, C. Zhang, Z. Huang, and Y. Guo, “A survey on human behavior recognition using channel state information,” IEEE Access, vol. 7, pp. 155986–156024, 2019.
[11] Z. Jiang, S. Chen, A. F. Molisch, R. Vannithamby, S. Zhou, and Z. Niu, “Exploiting wireless channel state information structures beyond linear correlations: A deep learning approach,” IEEE Communications Magazine, vol. 57, no. 3, pp. 28–34, 2019.
[12] O. D. Lara and M. A. Labrador, “A survey on human activity recognition using wearable sensors,” IEEE communications surveys & tutorials, vol. 15, no. 3, pp. 1192–1209, 2012.
[13] Y. Xie, Z. Li, and M. Li, “Precise power delay profiling with commodity wifi,” in Proceedings of the 21st Annual International Conference on Mobile Computing and Networking, MobiCom ’15, (New York, NY, USA), p. 53–64, ACM, 2015.
[14] J. Liu, H. Liu, Y. Chen, Y. Wang, and C. Wang, “Wireless sensing for human activity: A survey,” IEEE Communications Surveys & Tutorials, vol. 22, no. 3, pp. 1629–1645, 2019.
[15] K. Wu, J. Xiao, Y. Yi, M. Gao, and L. M. Ni, “Fila: Fine-grained indoor localization,” in 2012 Proceedings IEEE INFOCOM, pp. 2210–2218, IEEE, 2012.
[16] H. Zhang, H. Du, Q. Ye, and C. Liu, “Utilizing csi and rssi to achieve high-precision outdoor positioning: A deep learning approach,” in ICC 2019 - 2019 IEEE International Conference on Communications (ICC), pp. 1–6, 2019.
[17] A. Sobehy, E. Renault, and P. Muhlethaler, “Ndr: Noise and dimensionality reduction of csi for indoor positioning using deep learning,” in 2019 IEEE Global Communications Conference (GLOBECOM), pp. 1–6, 2019.
[18] X. Wang, L. Gao, S. Mao, and S. Pandey, “Deepfi: Deep learning for indoor fingerprinting using channel state information,” in 2015 IEEE Wireless Communications and Networking Conference (WCNC), pp. 1666–1671, 2015.
[19] R. Zhou, M. Tang, Z. Gong, and M. Hao, “Freetrack: Device-free human tracking with deep neural networks and particle filtering,” IEEE Systems Journal, vol. 14, no. 2, pp. 2990–3000, 2020.
[20] S.-J. Liu, R. Y. Chang, and F.-T. Chien, “Analysis and visualization of deep neural networks in device-free wi-fi indoor localization,” IEEE Access, vol. 7, pp. 69379– 69392, 2019.
[21] H. Chen, Y. Zhang, W. Li, X. Tao, and P. Zhang, “Confi: Convolutional neural networks based indoor wi-fi localization using channel state information,” IEEE Access, vol. 5, pp. 18066–18074, 2017.
[22] C.-H. Hsieh, J.-Y. Chen, and B.-H. Nien, “Deep learning-based indoor localization using received signal strength and channel state information,” IEEE Access, vol. 7, pp. 33256–33267, 2019.
[23] W. Xun, L. Sun, C. Han, Z. Lin, and J. Guo, “Depthwise separable convolution based passive indoor localization using csi fingerprint,” in 2020 IEEE Wireless Communications and Networking Conference (WCNC), pp. 1–6, 2020.
[24] A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,” arXiv preprint arXiv:1704.04861, 2017.
[25] Y. Jing, J. Hao, and P. Li, “Learning spatiotemporal features of csi for indoor localization with dual-stream 3d convolutional neural networks,” IEEE Access, vol. 7, pp. 147571– 147585, 2019.
[26] N. Kostantinos, “Gaussian mixtures and their applications to signal processing,” Advanced signal processing handbook: theory and implementation for radar, sonar, and medical imaging real time systems, pp. 3–1, 2000.
[27] S. De Bast, A. P. Guevara, and S. Pollin, “Csi-based positioning in massive mimo systems using convolutional neural networks,” in 2020 IEEE 91st Vehicular Technology Conference (VTC2020-Spring), pp. 1–5, IEEE, 2020.
[28] H. Li, X. Zeng, Y. Li, S. Zhou, and J. Wang, “Convolutional neural networks based indoor wi-fi localization with a novel kind of csi images,” China Communications, vol. 16, no. 9, pp. 250–260, 2019.
[29] R. Ayyalasomayajula, A. Arun, C. Wu, S. Sharma, A. R. Sethi, D. Vasisht, and D. Bharadia, “Deep learning based wireless localization for indoor navigation,” in Proceedings of the 26th Annual International Conference on Mobile Computing and Networking, pp. 1–14, 2020.
[30] M. Nabati, H. Navidan, R. Shahbazian, S. A. Ghorashi, and D. Windridge, “Using synthetic data to enhance the accuracy of fingerprint-based localization: A deep learning approach,” IEEE Sensors Letters, vol. 4, no. 4, pp. 1–4, 2020.
[31] Q. Li, H. Qu, Z. Liu, N. Zhou, W. Sun, S. Sigg, and J. Li, “Af-dcgan: Amplitude feature deep convolutional gan for fingerprint construction in indoor localization systems,” IEEE Transactions on Emerging Topics in Computational Intelligence, vol. 5, no. 3, pp. 468–480, 2021.
[32] M. T. Hoang, B. Yuen, K. Ren, X. Dong, T. Lu, R. Westendorp, and K. Reddy, “A cnn-lstm quantifier for single access point csi indoor localization,” arXiv preprint arXiv:2005.06394, 2020.
[33] S. Fan, Y. Wu, C. Han, and X. Wang, “A structured bidirectional lstm deep learning method for 3d terahertz indoor localization,” in IEEE INFOCOM 2020 - IEEE Conference on Computer Communications, pp. 2381–2390, 2020.
[34] J.-Y. Chang, K.-Y. Lee, K. C.-J. Lin, and W. Hsu, “Wifi action recognition via visionbased methods,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 2782–2786, 2016.
[35] K. Wu, M. Yang, C. Ma, and J. Yan, “Csi-based wireless localization and activity recognition using support vector machine,” in 2019 IEEE International Conference on Signal Processing, Communications and Computing (ICSPCC), pp. 1–5, IEEE, 2019.
[36] J. Ding and Y. Wang, “Wifi csi-based human activity recognition using deep recurrent neural network,” IEEE Access, vol. 7, pp. 174257–174269, 2019.
[37] Z. Chen, L. Zhang, C. Jiang, Z. Cao, and W. Cui, “Wifi csi based passive human activity recognition using attention based blstm,” IEEE Transactions on Mobile Computing, vol. 18, no. 11, pp. 2714–2724, 2019.
[38] B. Sheng, F. Xiao, L. Sha, and L. Sun, “Deep spatial–temporal model based cross-scene action recognition using commodity wifi,” IEEE Internet of Things Journal, vol. 7, no. 4, pp. 3592–3601, 2020.
[39] X. Yang, R. Cao, M. Zhou, and L. Xie, “Temporal-frequency attention-based human activity recognition using commercial wifi devices,” IEEE Access, vol. 8, pp. 137758– 137769, 2020.
[40] Z. Shi, J. A. Zhang, R. Xu, and Q. Cheng, “Deep learning networks for human activity recognition with csi correlation feature extraction,” in ICC 2019 - 2019 IEEE International Conference on Communications (ICC), pp. 1–6, 2019.
[41] Z. Shi, J. A. Zhang, R. Xu, Q. Cheng, and A. Pearce, “Towards environmentindependent human activity recognition using deep learning and enhanced csi,” in GLOBECOM 2020 - 2020 IEEE Global Communications Conference, pp. 1–6, 2020.
[42] Z. Tang, A. Zhu, Z. Wang, K. Jiang, Y. Li, and F. Hu, “Human behavior recognition based on wifi channel state information,” in 2020 Chinese Automation Congress (CAC), pp. 1157–1162, 2020.
[43] j. zhang, f. wu, w. hu, q. zhang, w. xu, and j. cheng, “Wienhance: Towards data augmentation in human activity recognition using wifi signal,” in 2019 15th International Conference on Mobile Ad-Hoc and Sensor Networks (MSN), pp. 309–314, 2019.
[44] J. Zhang, F. Wu, B. Wei, Q. Zhang, H. Huang, S. W. Shah, and J. Cheng, “Data augmentation and dense-lstm for human activity recognition using wifi signal,” IEEE Internet of Things Journal, vol. 8, no. 6, pp. 4628–4641, 2021.
[45] F. Wang, J. Feng, Y. Zhao, X. Zhang, S. Zhang, and J. Han, “Joint activity recognition and indoor localization with wifi fingerprints,” IEEE Access, vol. 7, pp. 80058–80068, 2019.
[46] N. Damodaran and J. Schäfer, “Device free human activity recognition using wifi channel state information,” in 2019 IEEE SmartWorld, Ubiquitous Intelligence Computing, Advanced Trusted Computing, Scalable Computing Communications, Cloud Big Data Computing, Internet of People and Smart City Innovation (Smart- World/SCALCOM/UIC/ATC/CBDCom/IOP/SCI), pp. 1069–1074, 2019.
[47] J. Zong, B. Huang, L. He, B. Yang, and X. Cheng, “Device-free crowd counting based on the phase difference of channel state information,” in 2020 IEEE International Conference on Information Technology,Big Data and Artificial Intelligence (ICIBA), vol. 1, pp. 1343–1347, 2020.
[48] F. Wang, F. Zhang, C. Wu, B. Wang, and K. J. R. Liu, “Respiration tracking for people counting and recognition,” IEEE Internet of Things Journal, vol. 7, no. 6, pp. 5233– 5245, 2020.
[49] S. Liu, Y. Zhao, F. Xue, B. Chen, and X. Chen, “Deepcount: Crowd counting with wifi via deep learning,” arXiv preprint arXiv:1903.05316, 2019.
[50] D. Wang, Z. Zhou, X. Yu, and Y. Cao, “Csiid: Wifi-based human identification via deep learning,” in 2019 14th International Conference on Computer Science Education (ICCSE), pp. 326–330, 2019.
[51] Z. Zhou, C. Liu, X. Yu, C. Yang, P. Duan, and Y. Cao, “Deep-wiid: Wifi-based contactless human identification via deep learning,” in 2019 IEEE SmartWorld, Ubiquitous Intelligence Computing, Advanced Trusted Computing, Scalable Computing Communications, Cloud Big Data Computing, Internet of People and Smart City Innovation (SmartWorld/SCALCOM/UIC/ATC/CBDCom/IOP/SCI), pp. 877–884, 2019.
[52] M. R. Bastwesy, N. M. ElShennawy, and M. T. F. Saidahmed, “Deep learning sign language recognition system based on wi-fi csi.,” International Journal of Intelligent Systems & Applications, vol. 12, no. 6, 2020.
[53] H. F. T. Ahmed, H. Ahmad, and C. Aravind, “Device free human gesture recognition using wi-fi csi: A survey,” Engineering Applications of Artificial Intelligence, vol. 87, p. 103281, 2020.
[54] W. Liu, Q. Cheng, Z. Deng, H. Chen, X. Fu, X. Zheng, S. Zheng, C. Chen, and S. Wang, “Survey on csi-based indoor positioning systems and recent advances,” in 2019 International Conference on Indoor Positioning and Indoor Navigation (IPIN), pp. 1–8, 2019.
[55] Y. He, Y. Chen, Y. Hu, and B. Zeng, “Wifi vision: Sensing, recognition, and detection with commodity mimo-ofdm wifi,” IEEE Internet of Things Journal, vol. 7, no. 9, pp. 8296–8317, 2020.
[56] L. Guo, L. Wang, C. Lin, J. Liu, B. Lu, J. Fang, Z. Liu, Z. Shan, J. Yang, and S. Guo,“Wiar: A public dataset for wifi-based activity recognition,” IEEE Access, vol. 7, pp. 154935–154945, 2019.
[57] M. H. Kefayati, V. Pourahmadi, and H. Aghaeinia, “Wi2vi: Generating video frames from wifi csi samples,” IEEE Sensors Journal, vol. 20, no. 19, pp. 11463–11473, 2020.
[58] J. G. Rohra, B. Perumal, S. J. Narayanan, P. Thakur, and R. B. Bhatt, “User localization in an indoor environment using fuzzy hybrid of particle swarm optimization & gravitational search algorithm with neural networks,” in Proceedings of Sixth International Conference on Soft Computing for Problem Solving, pp. 286–295, Springer, 2017.
[59] M. Muaaz, A. Chelli, A. A. Abdelgawwad, A. C. Mallofré, and M. Pätzold, “Wiwehar: Multimodal human activity recognition using wi-fi and wearable sensing modalities,” IEEE Access, vol. 8, pp. 164453–164470, 2020.
[60] R. Saini, P. Kumar, P. P. Roy, and D. P. Dogra, “A novel framework of continuous human-activity recognition using kinect,” Neurocomputing, vol. 311, pp. 99–111, 2018.
[61] G. T. Papadopoulos, A. Axenopoulos, and P. Daras, “Real-time skeleton-tracking-based human action recognition using kinect data,” in International Conference on Multimedia Modeling, pp. 473–483, Springer, 2014.
[62] A. F. Elaraby, A. Hamdy, and M. Rehan, “A kinect-based 3d object detection and recognition system with enhanced depth estimation algorithm,” in 2018 IEEE 9th Annual Information Technology, Electronics and Mobile Communication Conference (IEMCON), pp. 247–252, 2018.
[63] M. Chen, Y. Li, X. Luo, W. Wang, L. Wang, and W. Zhao, “A novel human activity recognition scheme for smart health using multilayer extreme learning machine,” IEEE Internet of Things Journal, vol. 6, no. 2, pp. 1410–1418, 2019.
[64] Q. Nie, J. Wang, X. Wang, and Y. Liu, “View-invariant human action recognition based on a 3d bio-constrained skeleton model,” IEEE Transactions on Image Processing, vol. 28, no. 8, pp. 3959–3972, 2019.
[65] W. Aly, S. Aly, and S. Almotairi, “User-independent american sign language alphabet recognition based on depth image and pcanet features,” IEEE Access, vol. 7, pp. 123138–123150, 2019.
[66] M. L. Gavrilova, Y. Wang, F. Ahmed, and P. Polash Paul, “Kinect sensor gesture and activity recognition: New applications for consumer cognitive systems,” IEEE Consumer Electronics Magazine, vol. 7, no. 1, pp. 88–94, 2018.
[67] J. Liu, A. Shahroudy, D. Xu, A. C. Kot, and G. Wang, “Skeleton-based action recognition using spatio-temporal lstm network with trust gates,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, no. 12, pp. 3007–3021, 2018.
[68] J. Ba, V. Mnih, and K. Kavukcuoglu, “Multiple object recognition with visual attention,” arXiv preprint arXiv:1412.7755, 2014.
[69] D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” arXiv preprint arXiv:1409.0473, 2014.
[70] S. Song, C. Lan, J. Xing, W. Zeng, and J. Liu, “Spatio-temporal attention-based lstm networks for 3d action recognition and detection,” IEEE Transactions on Image Processing, vol. 27, no. 7, pp. 3459–3471, 2018.
[71] K. Yun, J. Honorio, D. Chattopadhyay, T. L. Berg, and D. Samaras, “Two-person interaction detection using body-pose features and multiple instance learning,” in 2012 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, pp. 28–35, 2012.
[72] A. Shahroudy, J. Liu, T.-T. Ng, and G. Wang, “Ntu rgb+d: A large scale dataset for 3d human activity analysis,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1010–1019, 2016.
[73] Z. Fan, X. Zhao, T. Lin, and H. Su, “Attention-based multiview re-observation fusion network for skeletal action recognition,” IEEE Transactions on Multimedia, vol. 21, no. 2, pp. 363–374, 2019.
[74] I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” arXiv preprint arXiv:1409.3215, 2014.
[75] V. Badrinarayanan, A. Kendall, and R. Cipolla, “Segnet: A deep convolutional encoderdecoder architecture for image segmentation,” IEEE transactions on pattern analysis and machine intelligence, vol. 39, no. 12, pp. 2481–2495, 2017.
[76] D. Halperin, W. Hu, A. Sheth, and D. Wetherall, “Tool release: Gathering 802.11n traces with channel state information,” ACM SIGCOMM CCR, vol. 41, p. 53, Jan. 2011.
[77] Microsoft, “Kinect windows sdk 2.0.” https://www.microsoft.com/enus/ download/details.aspx?id=44561, 2014/10.
[78] S. Yang, X. Yu, and Y. Zhou, “Lstm and gru neural network performance comparison study: Taking yelp review dataset as an example,” in 2020 International Workshop on Electronic Communication and Artificial Intelligence (IWECAI), pp. 98–101, 2020.
[79] K. Cho, B. Van Merriënboer, D. Bahdanau, and Y. Bengio, “On the properties of neural machine translation: Encoder-decoder approaches,” arXiv preprint arXiv:1409.1259, 2014.
[80] M.-T. Luong, H. Pham, and C. D. Manning, “Effective approaches to attention-based neural machine translation,” arXiv preprint arXiv:1508.04025, 2015.
[81] K. Cho, A. Courville, and Y. Bengio, “Describing multimedia content using attentionbased encoder-decoder networks,” IEEE Transactions on Multimedia, vol. 17, no. 11, pp. 1875–1886, 2015.
[82] G. Paolini, A. Peruzzi, A. Mirelman, A. Cereatti, S. Gaukrodger, J. M. Hausdorff, and U. Della Croce, “Validation of a method for real time foot position and orientation tracking with microsoft kinect technology for use in virtual reality and treadmill based gait training programs,” IEEE Transactions on neural systems and rehabilitation engineering, vol. 22, no. 5, pp. 997–1002, 2013.
[83] M. Li, F. Wei, Y. Li, S. Zhang, and G. Xu, “Three-dimensional pose estimation of infants lying supine using data from a kinect sensor with low training cost,” IEEE Sensors Journal, vol. 21, no. 5, pp. 6904–6913, 2020.
[84] L. Chen and R. Ng, “On the marriage of lp-norms and edit distance,” in Proceedings of the Thirtieth international conference on Very large data bases-Volume 30, pp. 792– 803, 2004.