| 研究生: |
蔡雅晴 Tsai, Ya-Ching |
|---|---|
| 論文名稱: |
基於視覺慣性影像穩定之視障者穿戴式輔助導航系統 Visual-Inertial Image Stabilization in Wearable Assistive Navigation Systems for the Visually Impaired |
| 指導教授: |
莊智清
Juang, Jyh-Ching |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 電機工程學系 Department of Electrical Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 118 |
| 中文關鍵詞: | 影片穩定化 、視覺慣性融合 、穿戴式視障輔助導航 、物件偵測 、語意分割 、感知一致性 |
| 外文關鍵詞: | video stabilization, visual-inertial fusion, wearable assistive navigation, object detection, semantic segmentation, perception consistency |
| 相關次數: | 點閱:18 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
隨著人工智慧與電腦視覺技術的發展,穿戴式視覺導航系統近年來已成為輔助科技領域中備受關注的研究方向,其中,針對視障者所設計之輔助導航系統更是重要的應用情境之一。在這類系統中,可透過穿戴於頭部或胸前的相機擷取周圍環境影像,並結合物件偵測與語意分割等人工智慧感知技術,協助視障者獲得即時的環境感知與導航輔助。然而,在實際行走情境中,系統輸出之一致性與安全性仍是重要議題。使用者步伐所造成的高頻影像晃動,會使穿戴式相機所擷取之影像產生不穩定現象,進而影響下游人工智慧感知模型之輸出結果。當物件偵測或語意分割結果在連續影格間產生明顯變化時,將可能降低環境感知資訊的可靠性,並進一步影響後續導航判斷與使用者的行走安全。
為改善上述問題,本研究針對穿戴式視障導航輔助系統中,因行走晃動所造成之影像不穩定問題進行分析,並比較不同影像穩定化方法對人工智慧感知結果之影響。本研究以視覺慣性融合穩定化方法為主要研究方法,結合影像特徵與陀螺儀量測資訊,並以純視覺式及純陀螺儀式穩定化方法作為比較基準。實驗平台採用 Raspberry Pi 4 搭配相機模組與慣性量測單元,建立具時間戳記之同步資料擷取與時間對齊平台,用以蒐集行走場景下之影像與三軸角速度資料。接著,將不同穩定化方法產生之影片輸入至 YOLOv7-tiny 物件偵測模型與 BiSeNet 語意分割模型,以評估影像穩定化對下游人工智慧感知結果穩定性之影響。
不同於傳統影片穩定化研究多以影像平滑度、裁切比例或軌跡誤差作為評估依據,本文進一步從人工智慧感知層級進行分析,採用連續影格間之物件偵測交並比與語意分割遮罩交並比作為評估指標,用以量化不同穩定化方法對感知一致性的改善程度。實驗結果顯示,所提出的方法可在快速相機轉動下提升影像穩定化的穩健性,並改善物件層級感知結果的時間一致性,同時降低具挑戰性運動條件下語意分割輸出的時間波動。上述結果說明,影像穩定化除可改善輸入影像品質外,亦具有提升視障導航輔助系統中人工智慧感知可靠性之潛力。
With the rapid development of artificial intelligence and computer vision technologies, wearable visual navigation systems have recently become an important research topic in the field of assistive technology. Among their applications, wearable assistive navigation systems designed for visually impaired people have received considerable attention. These systems use cameras mounted on the head or chest to capture the surrounding environment and employ AI perception techniques, such as object detection and semantic segmentation, to provide real-time environmental awareness and navigation assistance. However, the consistency and reliability of system outputs remain critical concerns in practical walking scenarios. High-frequency image jitter induced by walking motion can cause instability in the captured images and adversely affect downstream AI perception performance. Significant variations in object detection or semantic segmentation results between consecutive frames may reduce the reliability of environmental perception information and further affect subsequent navigation decisions and user safety.
To address this issue, this study analyzes image instability caused by walking-induced motion in wearable devices for visually impaired people and investigates the effects of different video stabilization methods on downstream AI perception performance. This study primarily focuses on a visual-inertial fusion stabilization method that integrates image features with gyroscope measurements, while vision-only and gyroscope-only stabilization methods are implemented as comparison baselines. The experimental platform consists of a Raspberry Pi 4 equipped with a camera module and an inertial measurement unit (IMU), which is used to collect timestamped video frames and three-axis angular velocity measurements during walking and align them temporally. Videos processed using different stabilization methods are then analyzed using YOLOv7-tiny for object detection and BiSeNet for semantic segmentation to evaluate the influence of video stabilization on downstream AI perception consistency.
Unlike conventional video stabilization studies that primarily evaluate image smoothness, cropping ratio, or trajectory error, this study further analyzes stabilization performance at the AI perception level. The inter-frame Intersection over Union (IoU) of object detection bounding boxes and semantic segmentation masks is employed to quantify the improvement in perception consistency achieved by different stabilization methods. Experimental results demonstrate that the proposed method improves video stabilization robustness under rapid camera rotations. It also improves object-level temporal consistency and reduces temporal fluctuations in semantic segmentation outputs under challenging motion conditions. These findings indicate that image stabilization not only improves input image quality but also enhances the consistency and reliability of downstream AI perception in assistive navigation systems for visually impaired people.
[1] W. Elmannai and K. Elleithy, “Sensor-Based Assistive Devices for Visually-Impaired People: Current Status, Challenges, and Future Directions,” Sensors, vol. 17, no. 3, Art. no. 565, 2017.
[2] R. K. Katzschmann, B. Araki, and D. Rus, “Safe Local Navigation for Visually Impaired Users With a Time-of-Flight and Haptic Feedback Device,” IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol. 26, no. 3, pp. 583-593, 2018.
[3] P. Pfreundschuh, G. Cioffi, C. von Einem, A. Wyss, H. W. van de Venn, C. Cadena, D. Scaramuzza, R. Siegwart, and A. Darvishy, “Sight Guide: A Wearable Assistive Perception and Navigation System for the Vision Assistance Race in the Cybathlon 2024,” arXiv preprint arXiv:2506.02676, 2025.
[4] S. Salman Shah, A. Imran, Saad-Ur-Rehman, A. Arif, K. Khan, M. Arsalan, S. Manzoor, and G. J. Sirewal, “Vision-Based Smart Wearable Assistive Navigation System Using Deep Learning for Visually Impaired People,” Automation, vol. 7, no. 2, Art. no. 41, 2026.
[5] J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You Only Look Once: Unified, Real-Time Object Detection,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 779-788, 2016.
[6] C. Yu, J. Wang, C. Peng, C. Gao, G. Yu, and N. Sang, “BiSeNet: Bilateral Segmentation Network for Real-Time Semantic Segmentation,” in Proceedings of the European Conference on Computer Vision (ECCV), pp. 325-341, 2018.
[7] C. Morimoto and R. Chellappa, “Fast 3D Stabilization and Mosaic Construction,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 660-665, 1997.
[8] Y. Matsushita, E. Ofek, W. Ge, X. Tang, and H.-Y. Shum, “Full-Frame Video Stabilization with Motion Inpainting,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 28, no. 7, pp. 1150-1163, 2006.
[9] F. Liu, M. Gleicher, H. Jin, and A. Agarwala, “Content-Preserving Warps for 3D Video Stabilization,” ACM Transactions on Graphics, vol. 28, no. 3, Art. no. 44, 2009.
[10] F. Liu, M. Gleicher, J. Wang, H. Jin, and A. Agarwala, “Subspace Video Stabilization,” ACM Transactions on Graphics, vol. 30, no. 1, Art. no. 4, 2011.
[11] M. Grundmann, V. Kwatra, and I. Essa, “Auto-Directed Video Stabilization with Robust L1 Optimal Camera Paths,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 225-232, 2011.
[12] S. Liu, P. Tan, L. Yuan, J. Sun, and B. Zeng, “MeshFlow: Minimum Latency Online Video Stabilization,” in Proceedings of the European Conference on Computer Vision (ECCV), pp. 800-815, 2016.
[13] A. Karpenko, D. Jacobs, J. Baek, and M. Levoy, “Digital Video Stabilization and Rolling Shutter Correction Using Gyroscopes,” Stanford Univ., Stanford, CA, USA, Comput. Sci. Tech. Rep. CSTR 2011-03, 2011.
[14] G. Hanning, N. Forsl.w, P.-E. Forss.n, E. Ringaby, D. T.rnqvist, and J. Callmer, “Stabilizing cell phone video using inertial measurement sensors,” in Proceedings of the IEEE International Conference on Computer Vision Workshops (ICCV Workshops), pp. 1–8, 2011.
[15] B. Zhuang, D. Bai, and J. Lee, “5D Video Stabilization Through Sensor Vision Fusion,” in Proceedings of the IEEE International Conference on Image Processing (ICIP), pp. 4340–4344, 2019.
[16] J. Yu, T. Zhang, F. Shi, L. He, and C.-K. Liang, “SensorFlow: Sensor and Image Fused Video Stabilization,” in Proceedings of the Winter Conference on Applications of Computer Vision (WACV), pp. 8443–8452, 2025.
[17] M. Wang, G.-Y. Yang, J.-K. Lin, S.-H. Zhang, A. Shamir, S.-P. Lu, and S.-M. Hu, “Deep Online Video Stabilization with Multi-Grid Warping Transformation Learning,” IEEE Transactions on Image Processing, vol. 28, no. 5, pp. 2283-2292, 2019.
[18] S.-Z. Xu, J. Hu, M. Wang, T.-J. Mu, and S.-M. Hu, “Deep Video Stabilization Using Adversarial Networks,” Computer Graphics Forum, vol. 37, no. 7, pp. 267-276, 2018.
[19] Y.-L. Liu, W.-S. Lai, M.-H. Yang, Y.-Y. Chuang, and J.-B. Huang, “Hybrid Neural Fusion for Full-Frame Video Stabilization,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 2299–2308, 2021.
[20] C. Vincent, T. Kim, and H. Mee., “High Temporal Consistency through Semantic Similarity Propagation in Semi-Supervised Video Semantic Segmentation for Autonomous Flight,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1461–1471, 2025.
[21] S. A. S. Hesham, Y. Liu, G. Sun, H. Ding, J. Yang, E. Konukoglu, X. Geng, and X.Jiang, “Exploiting Temporal State Space Sharing for Video Semantic Segmentation,”in Proceedings of the IEEE/CVF Conference on Computer Vision and PatternRecognition (CVPR), pp. 24211–24221, 2025.
[22] D. L. Hall and J. Llinas, “An Introduction to Multisensor Data Fusion,” Proceedings of the IEEE, vol. 85, no. 1, pp. 6-23, 1997.
[23] B. Khaleghi, A. Khamis, F. O. Karray, and S. N. Razavi, “Multisensor Data Fusion: A Review of the State-of-the-Art,” Information Fusion, vol. 14, no. 1, pp. 28-44, 2013.
[24] C.-Y. Wang, A. Bochkovskiy, and H.-Y. M. Liao, “YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7464-7475, 2023.
[25] S. Liu, L. Yuan, P. Tan, and J. Sun, “Bundled Camera Paths for Video Stabilization,” ACM Transactions on Graphics, vol. 32, no. 4, Art. no. 78, 2013.
[26] Raspberry Pi Foundation, "Raspberry Pi 4 Model B," https://www.raspberrypi.com/products/raspberry-pi-4-model-b/ (accessed Jul. 12, 2026).
[27] Raspberry Pi Foundation, "Camera Module 2," https://www.raspberrypi.com/documentation/accessories/camera.html (accessed Jul. 12, 2026).
[28] STMicroelectronics, "LSM6DS3TR-C iNEMO 6DoF Inertial Measurement Unit (IMU)," https://www.st.com/en/mems-and-sensors/lsm6ds3tr-c.html (accessed Jul. 12, 2026).
[29] J. Shi and C. Tomasi, “Good Features to Track,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 593-600, 1994.
[30] B. D. Lucas and T. Kanade, “An Iterative Image Registration Technique with an Application to Stereo Vision,” in Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), pp. 674-679, 1981.
[31] J.-Y. Bouguet, “Pyramidal Implementation of the Lucas Kanade Feature Tracker: Description of the Algorithm,” Intel Corporation, Microprocessor Research Labs, Tech. Rep., 2001.
[32] M. A. Fischler and R. C. Bolles, “Random Sample Consensus: A Paradigm for Model Fitting with Applications to Image Analysis and Automated Cartography,” Communications of the ACM, vol. 24, no. 6, pp. 381-395, 1981.
[33] D. G. Lowe, “Distinctive Image Features from Scale-Invariant Keypoints,” International Journal of Computer Vision, vol. 60, no. 2, pp. 91-110, 2004.
[34] R. Hartley and A. Zisserman, Multiple View Geometry in Computer Vision, 2nd ed. Cambridge, U.K.: Cambridge Univ. Press, 2004.
[35] D. Scaramuzza and F. Fraundorfer, “Visual Odometry,” IEEE Robotics & Automation Magazine, vol. 18, no. 4, pp. 80-92, 2011.
[36] J. Engel, V. Koltun, and D. Cremers, “Direct Sparse Odometry,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, no. 3, pp. 611-625, 2018.
[37] P. Furgale, J. Rehder, and R. Siegwart, “Unified Temporal and Spatial Calibration for Multi-Sensor Systems,” in Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 1280-1286, 2013.