| 研究生: |
鄧兆偉 Teng, Chao-Wei |
|---|---|
| 論文名稱: |
基於連續時間相對位置注意力之多音束測深儀點雲自動化雜訊過濾方法 Continuous-Time Relative-Position Attention for Automated Outlier Removal in Multibeam Echosounder Point Clouds |
| 指導教授: |
陳奇業
Chen, Chi-Yeh |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 資訊工程學系 Department of Computer Science and Information Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 51 |
| 中文關鍵詞: | 多音束測深儀 、點雲分割 、去噪 、數值高程模型 |
| 外文關鍵詞: | Multibeam Echosounder, Point Cloud Segmentation, Denoising, Digital Elevation Model |
| 相關次數: | 點閱:23 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
水深測量為港埠營運、航道維護、海岸工程與海洋環境監測之基礎,其量測精度與資料完整性直接決定後續決策之可靠程度。多音束測深儀(multibeam echosounder, MBES)可於單一脈衝(ping)沿數百條波束同時取得測深值,描繪高密度、高解析之海床地形。然而受聲速剖面誤差、載具姿態不確定性、海床反射特性差異,以及水中氣泡、魚群與漂浮物等物體之影響,原始點雲時常混入異常測深點(outlier 或 spike)。若未將此類雜訊剔除,將顯著影響最終數值高程模型(digital elevation model, DEM)之精度與可用性。
傳統雜訊清理依靠測量員(hydrographer)逐點檢視並標記而相當耗時。或使用組合不確定度與水深估計演算法(Combined Uncertainty and Bathymetry Estimator, CUBE)等統計濾波方法,但此類方法須隨場域調整參數,於人工結構物與礁岩等複雜海床附近仍易失效,留下需人工介入之區域。為消除此一依賴,本論文將 MBES 雜訊過濾形式化為監督式逐點二元分割,對每一點輸出雜訊點/有效點之判斷,並以建立自動化之端到端模型為目標。
本論文首先驗證場景幾何包含可供雜訊判別的資訊:將每一點至 MBES 條帶(swath)中軸之垂直距離(以主成分分析求得之軸向先驗)作為點-體素骨幹之額外輸入通道,可於三個測區中的兩個提升 AP 與 F1,且兩項指標的三折平均皆隨之提高。依此結果,本論文提出連續時間相對位置注意力(CT-RPA),以局部注意力模組接於 PointNeXt 骨幹之全解析度解碼輸出並以殘差加回。對每一點在內容相似度之外加入空間相對位置偏置與連續時間偏置,使模型不依賴人工特徵工程或離散式分箱流程(ping binning),即可自原始幾何與時間戳學得條帶形狀與跨脈衝關係。於安平亞果、興達、愛河三港之留一港交叉驗證(leave-one-port-out cross-validation)中,深度學習模型之結果整體高於復現之統計式過濾方法(Statistical Filter)。本論文之模型於平均精確率(average precision, AP)與雜訊類別 F1 兩項指標均高於所評估之通用骨幹與分幀式 4D 基線,其中以 AP 作為主要排序指標,F1 則作為操作點表現之互補指標。消融實驗顯示,空間與時間項皆有實質且量級相近之貢獻。第二個較寬之注意力尺度亦具有較高的三折平均 F1。依測線覆蓋次數之子群分析顯示,增益出現於雜訊實際聚集之處,而各港雜訊所聚集之覆蓋層並不相同。於未參與訓練與門檻選擇之第四港(竹圍)上,三折權重之 AP 分別為 0.7668、0.7814 與 0.7507,平均為 0.7663,顯示模型可推廣至未見測區。在固定門檻 0.5 下,其雜訊類別 F1 分別為 0.7362、0.8168 與 0.3925。於生成 DEM 之比較中,本模型於兩種網格尺寸之 RMSE 與最大殘差均小於復現之統計式過濾方法。上述結果顯示 MBES 點雲之時間結構為可端到端學習且具實用價值之訊號,並提供自動化之雜訊過濾方案。
Bathymetric surveying forms the basis of activities such as port operation, navigation-channel maintenance, coastal engineering, and marine environmental monitoring. A multibeam echosounder (MBES) acquires depth measurements along hundreds of beams for every transmitted pulse, or ping, producing a high-density, high-resolution description of the seafloor. However, owing to sound-velocity-profile errors, vessel attitude uncertainty, seabed reflectivity variation, and objects suspended in the water column such as bubbles, fish, and floating debris, the raw point cloud often contains anomalous soundings, namely outliers and spikes. Unless such noisy points are removed, they substantially degrade the accuracy and usability of the resulting digital elevation model (DEM).
Conventional cleaning relies on time-consuming point-by-point inspection and labeling by hydrographers or on statistical filtering methods such as the Combined Uncertainty and Bathymetry Estimator (CUBE). These methods require site-dependent parameter tuning and may perform poorly over complex seabed terrain such as artificial structures and rocky areas, leaving regions that require manual intervention. To reduce this dependence, this thesis formulates MBES noise filtering as supervised per-point binary segmentation and predicts a noise/valid label for every point.
This thesis first tests whether scene geometry provides useful cues for noise classification. Adding the perpendicular distance from each point to its swath centerline, an axis prior obtained by principal component analysis, as an extra input channel of a point-voxel backbone improves AP and F1 on two of the three sites and improves both three-site means. Based on this result, this thesis proposes Continuous-Time Relative-Position Attention (CT-RPA), a local-attention module inserted at the full-resolution decoder output of a PointNeXt backbone and added back as a residual. For each query point, CT-RPA adds spatial relative-position and continuous-time biases to the content similarity. The model uses raw geometry and timestamps without handcrafted input features or discrete ping binning. Under leave-one-port-out cross-validation over Anping Argo, Xingda, and Love River, the evaluated deep learning models have higher aggregate noise-class F1 than the reproduced Statistical Filter. The proposed model has the highest mean AP and noise-class F1 among the evaluated general-purpose backbones and directly adapted frame-based 4D models, with AP used as the primary ranking measure and F1 as a complementary operating-point measure. A matched spatial-only arm indicates that the spatial and temporal terms both make substantial contributions to the module gain, while the wider second attention scale has a higher three-fold mean F1 than the narrower alternatives. In the pass-coverage analysis, larger gains occur in subgroups with higher noise rates, and these subgroups differ among ports. On a fourth port that was not used for training or threshold selection, the three cross-validation weights reach AP values of 0.7668, 0.7814, and 0.7507, indicating generalization to this additional unseen site. At the fixed threshold of 0.5, their noise-class F1 values are 0.7362, 0.8168, and 0.3925. In the DEM comparison, the model has a smaller RMSE and maximum residual than the reproduced Statistical Filter at both grid sizes. These results demonstrate that continuous acquisition time is a useful model input in this dataset, although its effect varies with the acquisition context. The method provides an automatic noise-filtering workflow.
[1] H. Bisquay, X. Freulon, C. de Fouquet, and C. Lajaunie, “Multibeam data cleaning for hydrography using geostatistics,” IEEE Oceanic Engineering Society, OCEANS’98 Conference Proceedings, 1998.
[2] N. Debese, R. Moitié, and N. Seube, “Multibeam echosounder data cleaning through a hierarchic adaptive and robust local surfacing,” Computers & Geosciences, vol. 46, pp. 330–339, 2012.
[3] B. R. Calder and L. A. Mayer, “Automatic processing of high-rate, high-density multibeam echosounder data,” Geochemistry, Geophysics, Geosystems, vol. 4, no. 6, p. 1048, 2003.
[4] T.-Y. Liu, “Semi-Automatic Outlier Identification Applied to Point Clouds from Modernized Bathymetric Surveying using Unmanned Surface Vehicle with Multibeam Echosounder,” Master’s thesis, Department of Geomatics, National Cheng Kung University, Tainan, Taiwan, 2024.
[5] H. Shi, J. Wei, H. Wang, F. Liu, and G. Lin, “Learning Spatial and Temporal Variations for 4D Point Cloud Segmentation,” International Journal of Computer Vision, vol. 132, no. 12, pp. 5603–5617, 2024.
[6] M.-J. Rakotosaona, V. La Barbera, P. Guerrero, N. J. Mitra, and M. Ovsjanikov, “PointCleanNet: Learning to denoise and remove outliers from dense point clouds,” Computer Graphics Forum, vol. 39, no. 1, pp. 185–203, 2020.
[7] T. Ziolkowski, A. Koschmider, and C. W. Devey, “An optimized outlier detection function for multibeam echo-sounder data,” Computers & Geosciences, vol. 186, p. 105 572, 2024.
[8] C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “PointNet: Deep learning on point sets for 3D classification and segmentation,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
[9] C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “PointNet++: Deep hierarchical feature learning on point sets in a metric space,” Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017.
[10] G. Qian, Y. Li, H. Peng, J. Mai, H. Hammoud, M. Elhoseiny, and B. Ghanem, “PointNeXt: Revisiting PointNet++ with improved training and scaling strategies,” Advances in Neural Information Processing Systems (NeurIPS), vol. 35, 2022, pp. 23 192–23 204.
[11] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention Is All You Need,” Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017.
[12] M.-H. Guo, J.-X. Cai, Z.-N. Liu, T.-J. Mu, R. R. Martin, and S.-M. Hu, “PCT: Point Cloud Transformer,” Computational Visual Media, vol. 7, no. 2, pp. 187–199, 2021.
[13] H. Zhao, L. Jiang, J. Jia, P. H. S. Torr, and V. Koltun, “Point Transformer,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
[14] X. Wu, Y. Lao, L. Jiang, X. Liu, and H. Zhao, “Point Transformer V2: Grouped vector attention and partition-based pooling,” Advances in Neural Information Processing Systems (NeurIPS), vol. 35, 2022, pp. 33 330–33 342.
[15] X. Wu, L. Jiang, P.-S. Wang, Z. Liu, X. Liu, Y. Qiao, W. Ouyang, T. He, and H. Zhao, “Point Transformer V3: Simpler, faster, stronger,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024.
[16] C. Choy, J. Gwak, and S. Savarese, “4D spatio-temporal ConvNets: Minkowski convolutional neural networks,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
[17] O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional Networks for Biomedical Image Segmentation,” Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2015.
[18] H. Zhou, X. Zhu, X. Song, Y. Ma, Z. Wang, H. Li, and D. Lin, Cylinder3D: An Effective 3D Framework for Driving-scene LiDAR Semantic Segmentation, arXiv:2008.01550, 2020.
[19] Y. Chen, J. Liu, X. Zhang, X. Qi, and J. Jia, “LargeKernel3D: Scaling up Kernels in 3D Sparse CNNs,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
[20] Z. Liu, H. Tang, Y. Lin, and S. Han, “Point-Voxel CNN for efficient 3D deep learning,” Advances in Neural Information Processing Systems (NeurIPS), vol. 32, 2019.
[21] H. Tang, Z. Liu, S. Zhao, Y. Lin, J. Lin, H. Wang, and S. Han, “Searching Efficient 3D Architectures with Sparse Point-Voxel Convolution,” Proceedings of the European Conference on Computer Vision (ECCV), 2020.
[22] F. Zhang, J. Fang, B. Wah, and P. Torr, “Deep FusionNet for Point Cloud Semantic Segmentation,” Proceedings of the European Conference on Computer Vision (ECCV), 2020.
[23] C. Park, Y. Jeong, M. Cho, and J. Park, “Fast Point Transformer,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
[24] C. Zhang, H. Wan, X. Shen, and Z. Wu, “PVT: Point-Voxel Transformer for Point Cloud Learning,” International Journal of Intelligent Systems, vol. 37, no. 12, pp. 11 985–12 008, 2022.
[25] H. Fan, X. Yu, Y. Ding, Y. Yang, and M. Kankanhalli, “PSTNet: Point Spatio-Temporal Convolution on Point Cloud Sequences,” International Conference on Learning Representations (ICLR), 2021.
[26] H. Fan, X. Yu, Y. Yang, and M. Kankanhalli, “Deep Hierarchical Representation of Point Cloud Videos via Spatio-Temporal Decomposition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 12, pp. 9918–9930, 2022.
[27] H. Fan, Y. Yang, and M. Kankanhalli, “Point 4D Transformer networks for spatio-temporal modeling in point cloud videos,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
[28] H. Fan, Y. Yang, and M. Kankanhalli, “Point Spatio-Temporal Transformer Networks for Point Cloud Video Modeling,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 2, pp. 2181–2192, 2023.
[29] N. Wang, R. Guo, C. Shi, Z. Wang, H. Zhang, H. Lu, Z. Zheng, and X. Chen, “SegNet4D: Efficient Instance-Aware 4D Semantic Segmentation for LiDAR Point Cloud,” IEEE Transactions on Automation Science and Engineering, vol. 22, pp. 15 339–15 350, 2025.
[30] Y. Wei, H. Liu, T. Xie, Q. Ke, and Y. Guo, “Spatial-Temporal Transformer for 3D Point Cloud Sequences,” Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2022.
[31] Y. Zhou, H. Zhu, C. Li, T. Cui, S. Chang, and M. Guo, “TempNet: Online Semantic Segmentation on Large-Scale Point Cloud Series,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 7118–7127.
[32] J. Knights, P. Moghadam, C. Fookes, and S. Sridharan, Point Cloud Segmentation Using Sparse Temporal Local Attention, arXiv:2112.00289, 2021.
[33] S. Luo and W. Hu, “Score-Based Point Cloud Denoising,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
[34] D. Stephens, A. Smith, T. Redfern, A. Talbot, A. Lessnoff, and K. Dempsey, “Using three dimensional convolutional neural networks for denoising echosounder point cloud data,” Applied Computing and Geosciences, vol. 5, p. 100 016, 2020.
[35] L. Ling, Y. Xie, N. Bore, and J. Folkesson, “Score-Based Multibeam Point Cloud Denoising,” 2024 IEEE/OES Autonomous Underwater Vehicles Symposium (AUV), 2024, pp. 1–6.
[36] T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017.
[37] M. Berman, A. Rannen Triki, and M. B. Blaschko, “The Lovász-softmax loss: A tractable surrogate for the optimization of the intersection-over-union measure in neural networks,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
[38] I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” International Conference on Learning Representations (ICLR), 2019.
[39] Maritime Robotics. “The Otter Seadrone by Maritime Robotics,” Accessed: Jul. 15, 2026. [Online]. Available: https://www.maritimerobotics.com/otter
[40] International Hydrographic Organization, IHO Standards for Hydrographic Surveys, ed. 6.1.0, IHO Publication S-44, Monaco, Oct. 2022. [Online]. Available: https://iho.int/uploads/user/pubs/standards/s-44/S-44_Edition_6.1.0.pdf