簡易檢索 / 詳目顯示

研究生: 李承芯
LI, CHENG-HSIN
論文名稱: 應用於無人機影像之地理定位神經輻射域重建框架
Georef-NeRF: Georeferenced Neural Radiance Fields Framework for UAV Imagery
指導教授: 林昭宏
LIN, Chao-Hung
學位類別: 碩士
Master
系所名稱: 工學院 - 測量及空間資訊學系
Department of Geomatics
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 134
中文關鍵詞: 神經輻射域攝影測量地面控制點數值地表模型三維重建
外文關鍵詞: 3D Reconstruction, Neural Radiance Fields, Photogrammetry, Ground Control Points, Orthographic UAV Imagery
相關次數: 點閱:4下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 近年來,神經輻射域(Neural Radiance Fields, NeRF)於三維重建與新視角合成領域展現優秀成果。然而,惟 NeRF 於模型訓練前會將相機位置姿態與場景坐標轉換至正規化訓練坐標系中,以提升模型訓練之穩定性,而此坐標系缺乏與真實地理坐標系統之直接對應關係,因此限制其重建成果於測繪與遙測應用中之定位解釋能力。為此,本研究提出應用於無人機影像之地理參考神經輻射場框架 (Georeferenced Neural Radiance Fields Framework, Georef-NeRF),以建立兼具三維重建與地理參考能力之NeRF 模型,並透過數值地表模型與輸入檢核點之偏誤成果驗證其定位精度。
    本研究以 Nerfstudio 專案中之 Nerfacto 模型為基礎,透過保留相機內外方位參數與地理坐標資訊,建立相機坐標系統、訓練中繼坐標系、訓練坐標系與地理參考坐標系間之轉換機制,使模型雖於正規化訓練坐標系中進行訓練,仍可透過建立之坐標轉換鏈結保留與真實地理坐標系之對應關係,並於成果輸出階段將重建結果轉換回目標地理坐標系。此外,本研究導入具地理座標資訊之地面點位,並依據觀測影像數與再投影誤差篩選可靠之點位作為控制約束,符合條件者之再投影誤差將納入損失函數,作為模型與相機姿態更新之依據,而未納入約束者則作為檢核點,用以評估重建成果之地理定位精度與空間一致性。
    本研究以國立成功大學校園 UAV 空拍影像資料集作為基礎,選取其中70張空拍影像作為實驗資料,並由地理定位成果與影像品質兩方面評估 Georef-NeRF 之表現。實驗影像皆經過$1/8$降採樣處理,解析度為994 imes663,對應之地面解析度由原始約 6.43 公分調整為約 51.44 公分。地理定位評估方面,以攝影測量產製之數值地表模型(Digital Surface Model, DSM)作為參考,並結合 22 筆具地理坐標資訊之地面點位進行誤差分析。由於部分地面點位並非明顯角點或易辨識之特徵點,故人工量測時可能存在約 1--2 像素之點選不確定性,約相當於 0.51--1.03 公尺之地面距離。結果顯示,Georef-NeRF 產製之 DSM 平均高程差異約為0.80 公尺,且差異主要分布於建物邊緣、屋頂結構及都市場景不連續區域,顯示在降採樣影像輸入與人工點選不確定性之限制下,模型之地理定位結果仍受影像解析度、地面點位觀測品質與場景結構複雜度影響。
    影像品質評估方面,相較於未導入地面控制點約束之模型,Georef-NeRF 之 PSNR 由20.71 dB 提升至 22.38 dB,SSIM 由 0.580 提升至 0.667,LPIPS 則由 0.217 降至 0.162,顯示導入地面控制點再投影約束後,模型仍能維持影像重建能力,並提升新視角合成之影像品質。
    綜上所述,本研究初步驗證 Georef-NeRF 在降採樣影像條件下,除可維持一定程度之三維場景重建與新視角合成能力外,亦可透過坐標轉換機制與地面點位約束,建立 NeRF 訓練空間與真實地理坐標系之對應關係,使重建成果具備地理參考應用之可行性。然而,模型之地理定位結果仍受影像解析度、地面點位觀測品質、控制點配置及都市場景幾何複雜度影響。未來研究將進一步改善模型於建物側立面及低紋理區域之重建完整性。由於本研究受限於運算資源與顯示記憶容量,須採降採樣影像進行模型訓練,可能限制細部紋理、建築物邊緣及局部幾何結構之重建表現。因此,後續將於運算資源允許或訓練流程進一步優化之條件下,結合較高解析度之影像資料、量測不確定性較低之地面控制點與檢核點,及更完整之幾何約束機制,以提升 DSM 成果品質、三維場景幾何一致性與實務應用能力。

    Neural Radiance Fields (NeRFs) have demonstrated significant potential for multi-view 3D reconstruction and novel-view synthesis. However, most existing NeRF frameworks perform training and rendering in a normalized coordinate space, which limits their direct applicability to geospatial mapping tasks. To address this limitation, this study proposes a georeferenced NeRF framework, named Georef-NeRF, for generating georeferenced point clouds, digital surface models (DSMs), and orthomosaics from UAV imagery.
    The proposed framework preserves georeferencing information through an explicit coordinate transformation chain that links photogrammetric camera coordinates, normalized NeRF training space, and the target mapping coordinate system. Ground Control Points (GCPs) are further integrated as reprojection constraints during training to enhance the positional consistency of the reconstructed scene. In addition, a camera optimization module and a GCP observation selection mechanism are introduced to improve geometric stability and reduce the influence of unreliable control observations.
    Experiments were conducted using 70 UAV images acquired over the National Cheng Kung University campus, with Metashape results used as the photogrammetric reference. In the DSM-based evaluation, the proposed method achieved a mean difference of 0.80 m, a mean absolute error (MAE) of 2.18 m, and a root mean square error (RMSE) of 3.91 m relative to the reference DSM. The GCP-based evaluation further showed that the reconstructed model maintained object-space positional consistency, with a final geographic error of approximately 0.77 m. In terms of image reconstruction quality, Georef-NeRF achieved a PSNR of 22.38 dB, indicating that the proposed georeferencing constraints can preserve spatial consistency while maintaining reasonable rendering quality.
    Overall, the experimental results demonstrate the feasibility of integrating georeferencing information and GCP constraints into a NeRF-based reconstruction framework for UAV mapping applications. Although local geometric discrepancies remain in complex surface areas such as building edges, vegetation, and regions with sparse or unstable reconstruction, the proposed framework provides a practical foundation for bridging neural radiance field reconstruction and geospatial mapping.

    摘要 i 致謝 viii 目錄 x 表目錄 xiii 圖目錄 xv 第1章 緒論 1 1.1 前言 1 1.2 無人機空拍影像之三維重建與地理資訊應用 2 1.3 研究動機 3 1.4 研究貢獻 4 第2章 文獻回顧 6 2.1 攝影測量(Photogrammetry) 6 2.2 神經輻射場(Neural Radiance Fields, NeRFs) 7 2.2.1 總論 7 2.2.2 資料前處理(Data preprocessing) 8 2.2.3 採樣(Sampling) 9 2.2.4 編碼(Encoding) 9 2.2.5 訓練(Radiance field estimation) 11 2.2.6 體積渲染(Volume Rendering)與最佳化(Optimization) 12 2.3 高斯潑濺(Gaussian Splatting, GS) 12 第3章 研究方法 15 3.1 研究背景 15 3.1.1 總論 15 3.1.2 訓練空間坐標轉換(ray space transformation) 16 3.1.3 採樣(Sampling) 16 3.1.4 編碼(Encoding) 17 3.1.5 訓練(training) 19 3.1.6 體積渲染(Volume Rendering) 20 3.1.7 最佳化(Optimization) 22 3.2 研究架構 23 3.3 資料前處理與坐標轉換 24 3.3.1 資料前處理 25 3.3.2 坐標系統轉換 28 3.3.3 坐標轉換之流程與逆轉換 31 3.4 Georef-NeRF 模型架構 33 3.4.1 模型架構概述 33 3.4.2 Georef-NeRF 之控制點篩選機制與相機優化器 34 3.4.3 Georef-NeRF 之輻射場設計架構 39 3.5 Georef-NeRF 之損失函數 42 3.5.1 Georef-NeRF 之損失函式:Nerfacto 項 42 3.5.2 Georef-NeRF 之損失函式:地面控制點之再投影誤差損失函數項 44 3.5.3 Georef-NeRF 之損失函式:局部相機優化器之姿態約束項 45 3.6 影像品質評估指標 46 第4章 實驗結果與分析 51 4.1 成大校園資料集 51 4.2 模型訓練環境與參數設定 55 4.3 三維重建成果與地理定位成果評估 57 4.3.1 三維重建成果評估 57 4.3.2 地理定位成果評估 65 4.4 影像品質評估 71 第5章 結論與未來展望 75 5.1 結論 75 5.2 未來展望 76 參考文獻 78 附錄 86 A.1 演算法備註 86 A.1.1 高斯潑濺 86 A.1.2 光柵化渲染(Rasterization) 87 A.2 坐標系統轉換實作備註 88 A.2.1 OpenCV 坐標系 88 A.2.2 OpenGL 坐標系 89 A.3 Nerfstudio 專案分析 90 A.4 檢核點之點之記 93 A.4.1 自行量測之點位(P1–P3) 93 A.4.2 總實習乙方總成果報告書中點位 96

    109 級測量系. (2019). 第 107 學年度測量總實習乙方總成果報告書 (測量總實習成果報告書) [提送日期:2019 年 7 月 26 日;指導老師:尤瑞哲、蔡展榮、朱宏杰]. 國立成功大學測量及空間資訊學系.
    湯美華. (2006). 空載光達點雲及地形圖輔助生產真實正射影像之研究 (碩士論文). 國立成功大學. https://hdl.handle.net/11296/85j7vy
    邱式鴻 & 邱俊榮. (2025). 都市區有人機傾斜攝影製作真正射影像精度評估與探討. 國土測繪與空間資訊, 13(1), 109–130.
    陳俊達, 方惠民, 蕭松山, & 康秋桂. (2019). UAV 正射影像應用於未辦地籍整理地區現況測量之研究. 國土測繪與空間資訊, 7(2), 67–85.
    Agisoft LLC. (2023). Agisoft metashape professional.
    Ahmad, J., Gao, P., Delehelle, D., Canio, M., Deshpande, N., Ortiz, J., Caldwell, D. G., & Tefera, Y. T. (2025). From fields to splats: A cross-domain survey of real-time neural scene representations. arXiv preprint arXiv:2509.23555.
    Barron, J. T., Mildenhall, B., Tancik, M., Hedman, P., Martin-Brualla, R., & Srinivasan, P. P. (2021). Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields. Proceedings of the IEEE/CVF international conference on computer vision, 5855–5864.
    Barron, J. T., Mildenhall, B., Verbin, D., Srinivasan, P. P., & Hedman, P. (2022). Mip-nerf 360: Unbounded anti-aliased neural radiance fields. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 5470–5479.
    Barron, J. T., Mildenhall, B., Verbin, D., Srinivasan, P. P., & Hedman, P. (2023). Zip-nerf: Anti-aliased grid-based neural radiance fields. Proceedings of the IEEE/CVF International Conference on Computer Vision, 19697–19705.
    Beer. (1852). Bestimmung der absorption des rothen lichts in farbigen flüssigkeiten. Annalen der Physik, 162(5), 78–88.
    Bergen, J. R., & Adelson, E. H. (1991). The plenoptic function and the elements of early vision. Computational models of visual processing, 1(8), 3.
    Berger, M., Tagliasacchi, A., Seversky, L. M., Alliez, P., Guennebaud, G., Levine, J. A., Sharf, A., & Silva, C. T. (2017). A survey of surface reconstruction from point clouds. Computer graphics forum, 36(1), 301–329.
    Bianco, S., Ciocca, G., & Marelli, D. (2018). Evaluating the performance of structure from motion pipelines. Journal of Imaging, 4(8), 98.
    Chan, E. R., Lin, C. Z., Chan, M. A., Nagano, K., Pan, B., De Mello, S., Gallo, O., Guibas, L. J., Tremblay, J., Khamis, S., et al. (2022). Efficient geometry-aware 3d generative adversarial networks. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 16123–16133.
    Chen, A., Xu, Z., Geiger, A., Yu, J., & Su, H. (2022). Tensorf: Tensorial radiance fields. European conference on computer vision, 333–350.
    Colomina, I., & Molina, P. (2014a). Unmanned aerial systems for photogrammetry and remote sensing: A review. ISPRS Journal of Photogrammetry and Remote Sensing, 92, 79–97. https://doi.org/https://doi.org/10.1016/j.isprsjprs.2014.02.013
    Colomina, I., & Molina, P. (2014b). Unmanned aerial systems for photogrammetry and remote sensing: A review. ISPRS Journal of photogrammetry and remote sensing, 92, 79–97.
    Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., & Fei-Fei, L. (2009). Imagenet: A large-scale hierarchical image database. 2009 IEEE conference on computer vision and pattern recognition, 248–255.
    Erdelj, M., & Natalizio, E. (2016). Uav-assisted disaster management: Applications and open issues. 2016 international conference on computing, networking and communications (ICNC), 1–5.
    Feng, Q., Liu, J., & Gong, J. (2015). Uav remote sensing for urban vegetation mapping using random forest and texture analysis. Remote sensing, 7(1), 1074–1094.
    Fischler, M. A., & Bolles, R. C. (1981). Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography. Communications of the ACM, 24(6), 381–395.
    Fridovich-Keil, S., Yu, A., Tancik, M., Chen, Q., Recht, B., & Kanazawa, A. (2022). Plenoxels: Radiance fields without neural networks. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 5501–5510.
    Furukawa, Y., & Ponce, J. (2009). Accurate, dense, and robust multiview stereopsis. IEEE transactions on pattern analysis and machine intelligence, 32(8), 1362–1376.
    Gao, K., Gao, Y., He, H., Lu, D., Xu, L., & Li, J. (2022). Nerf: Neural radiance field in 3d vision, a comprehensive review. arXiv preprint arXiv:2210.00379.
    Garbin, S. J., Kowalski, M., Johnson, M., Shotton, J., & Valentin, J. (2021). Fastnerf: High-fidelity neural rendering at 200fps. Proceedings of the IEEE/CVF international conference on computer vision, 14346–14355.
    Garland, M., & Heckbert, P. S. (1997). Surface simplification using quadric error metrics. Proceedings of the 24th annual conference on Computer graphics and interactive techniques, 209–216.
    Gortler, S. J., Grzeszczuk, R., Szeliski, R., & Cohen, M. F. (2023). The lumigraph. In Seminal graphics papers: Pushing the boundaries, volume 2 (pp. 453–464).
    Hedman, P., Srinivasan, P. P., Mildenhall, B., Barron, J. T., & Debevec, P. (2021). Baking neural radiance fields for real-time view synthesis. Proceedings of the IEEE/CVF international conference on computer vision, 5875–5884.
    Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. Advances in neural information processing systems, 33, 6840–6851.
    Hu, T., Liu, S., Chen, Y., Shen, T., & Jia, J. (2022). Efficientnerf efficient neural radiance fields. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 12902–12911.
    Hu, W., Wang, Y., Ma, L., Yang, B., Gao, L., Liu, X., & Ma, Y. (2023). Tri-miprf: Tri-mip representation for efficient anti-aliasing neural radiance fields. Proceedings of the IEEE/CVF International Conference on Computer Vision, 19774–19783.
    Iandola, F. N., Han, S., Moskewicz, M. W., Ashraf, K., Dally, W. J., & Keutzer, K. (2016). Squeezenet: Alexnet-level accuracy with 50x fewer parameters and< 0.5 mb model size. arXiv preprint arXiv:1602.07360.
    Kajiya, J. T., & Von Herzen, B. P. (1984). Ray tracing volume densities. ACM SIGGRAPH computer graphics, 18(3), 165–174.
    Kerbl, B., Kopanas, G., Leimkühler, T., Drettakis, G., et al. (2023a). 3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph., 42(4), 139–1.
    Kerbl, B., Kopanas, G., Leimkühler, T., Drettakis, G., et al. (2023b). 3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph., 42(4), 139–1.
    Kitchin, R. (2014). The data revolution: Big data, open data, data infrastructures and their consequences. Sage.
    Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25.
    Levoy, M., & Hanrahan, P. (2023). Light field rendering. In Seminal graphics papers: Pushing the boundaries, volume 2 (pp. 441–452).
    Lin, C.-H., Ma, W.-C., Torralba, A., & Lucey, S. (2021). Barf: Bundle-adjusting neural radiance fields. Proceedings of the IEEE/CVF international conference on computer vision, 5741–5751.
    Lombardi, S., Simon, T., Saragih, J., Schwartz, G., Lehrmann, A., & Sheikh, Y. (2019). Neural volumes: Learning dynamic renderable volumes from images. arXiv preprint arXiv:1906.07751.
    Lowe, D. G. (2004). Distinctive image features from scale-invariant keypoints. International journal of computer vision, 60(2), 91–110.
    Luhmann, T., Robson, S., Kyle, S., & Boehm, J. (2023). Close-range photogrammetry and 3d imaging. Walter de Gruyter GmbH & Co KG.
    Manfreda, S., McCabe, M. F., Miller, P. E., Lucas, R., Pajuelo Madrigal, V., Mallinis, G., Ben Dor, E., Helman, D., Estes, L., Ciraolo, G., et al. (2018). On the use of unmanned aerial systems for environmental monitoring. Remote sensing, 10(4), 641.
    Marcellin, M. W., Gormish, M. J., Bilgin, A., & Boliek, M. P. (2000). An overview of jpeg-2000. Proceedings DCC 2000. Data compression conference, 523–541.
    Martin-Brualla, R., Radwan, N., Sajjadi, M. S., Barron, J. T., Dosovitskiy, A., & Duckworth, D. (2021). Nerf in the wild: Neural radiance fields for unconstrained photo collections. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 7210–7219.
    Max, N. (2002). Optical models for direct volume rendering. IEEE Transactions on Visualization and Computer Graphics, 1(2), 99–108.
    Mikhail, E. M., Bethel, J. S., & McGlone, J. C. (2001). Introduction to modern photogrammetry. John Wiley & Sons.
    Mildenhall, B., Srinivasan, P. P., Ortiz-Cayon, R., Kalantari, N. K., Ramamoorthi, R., Ng, R., & Kar, A. (2019). Local light field fusion: Practical view synthesis with prescriptive sampling guidelines. ACM Transactions on Graphics (ToG), 38(4), 1–14.
    Mildenhall, B., Srinivasan, P. P., Tancik, M., Barron, J. T., Ramamoorthi, R., & Ng, R. (2021). Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1), 99–106.
    Müller, T., Evans, A., Schied, C., & Keller, A. (2022). Instant neural graphics primitives with a multiresolution hash encoding. ACM transactions on graphics (TOG), 41(4), 1–15.
    Nex, F., & Remondino, F. (2014). Uav for 3d mapping applications: A review. Applied geomatics, 6(1), 1–15.
    Nichol, A. Q., & Dhariwal, P. (2021). Improved denoising diffusion probabilistic models. International conference on machine learning, 8162–8171.
    Niemeyer, M., Mescheder, L., Oechsle, M., & Geiger, A. (2020). Differentiable volumetric rendering: Learning implicit 3d representations without 3d supervision. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 3504–3515.
    Pan, J., Wang, M., Li, J., Yuan, S., & Hu, F. (2015). Region change rate-driven seamline determination method. ISPRS Journal of Photogrammetry and Remote Sensing, 105, 141–154.
    Park, J. J., Florence, P., Straub, J., Newcombe, R., & Lovegrove, S. (2019). Deepsdf: Learning continuous signed distance functions for shape representation. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 165–174.
    Peña-Villasenı́n, S., Gil-Docampo, M., & Ortiz-Sanz, J. (2017). 3-d modeling of historic façades using sfm photogrammetry metric documentation of different building types of a historic center. International Journal of Architectural Heritage, 11(6), 871–890.
    Piras, G., Agostinelli, S., & Muzi, F. (2024). Digital twin framework for built environment: A review of key enablers. Energies, 17(2), 436.
    Pix4D SA. (2026). PIX4Dmapper [Photogrammetry software].
    Poole, B., Jain, A., Barron, J. T., & Mildenhall, B. (2022). Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988.
    Rau, J.-Y., Chen, N.-Y., Chen, L.-C., et al. (2002). True orthophoto generation of built-up areas using multi-view images. Photogrammetric Engineering and Remote Sensing, 68(6), 581–588.
    Rebain, D., Jiang, W., Yazdani, S., Li, K., Yi, K. M., & Tagliasacchi, A. (2021). Derf: Decomposed radiance fields. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 14153–14161.
    Reiser, C., Peng, S., Liao, Y., & Geiger, A. (2021). Kilonerf: Speeding up neural radiance fields with thousands of tiny mlps. Proceedings of the IEEE/CVF international conference on computer vision, 14335–14345.
    Reiser, C., Szeliski, R., Verbin, D., Srinivasan, P., Mildenhall, B., Geiger, A., Barron, J., & Hedman, P. (2023). Merf: Memory-efficient radiance fields for real-time view synthesis in unbounded scenes. ACM Transactions on Graphics (ToG), 42(4), 1–12.
    Rho, D., Lee, B., Nam, S., Lee, J. C., Ko, J. H., & Park, E. (2023). Masked wavelet representation for compact neural radiance fields. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 20680–20690.
    Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (2022). High-resolution image synthesis with latent diffusion models. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 10684–10695.
    Schonberger, J. L., & Frahm, J.-M. (2016). Structure-from-motion revisited. Proceedings of the IEEE conference on computer vision and pattern recognition, 4104–4113.
    Schönberger, J. L. (2018). Robust methods for accurate and efficient 3d modeling from unstructured imagery (Doctoral dissertation). ETH Zurich.
    Schops, T., Schonberger, J. L., Galliani, S., Sattler, T., Schindler, K., Pollefeys, M., & Geiger, A. (2017). A multi-view stereo benchmark with high-resolution images and multi-camera videos. Proceedings of the IEEE conference on computer vision and pattern recognition, 3260–3269.
    Shin, S., & Park, J. (2023). Binary radiance fields. Advances in neural information processing systems, 36, 55919–55931.
    Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556.
    Sitzmann, V., Zollhöfer, M., & Wetzstein, G. (2019). Scene representation networks: Continuous 3d-structure-aware neural scene representations. Advances in neural information processing systems, 32.
    Snavely, N., Seitz, S. M., & Szeliski, R. (2006). Photo tourism: Exploring photo collections in 3d. In Acm siggraph 2006 papers (pp. 835–846).
    Sun, C., Sun, M., & Chen, H.-T. (2022). Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 5459–5469.
    Tancik, M., Srinivasan, P., Mildenhall, B., Fridovich-Keil, S., Raghavan, N., Singhal, U., Ramamoorthi, R., Barron, J., & Ng, R. (2020). Fourier features let networks learn high frequency functions in low dimensional domains. Advances in neural information processing systems, 33, 7537–7547.
    Tancik, M., Weber, E., Ng, E., Li, R., Yi, B., Wang, T., Kristoffersen, A., Austin, J., Salahi, K., Ahuja, A., et al. (2023). Nerfstudio: A modular framework for neural radiance field development. ACM SIGGRAPH 2023 conference proceedings, 1–12.
    Turki, H., Zollhöfer, M., Richardt, C., & Ramanan, D. (2023). Pynerf: Pyramidal neural radiance fields. Advances in neural information processing systems, 36, 37670–37681.
    Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30.
    Verbin, D., Hedman, P., Mildenhall, B., Zickler, T., Barron, J. T., & Srinivasan, P. P. (2024). Ref-nerf: Structured view-dependent appearance for neural radiance fields. IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(11), 9426–9437.
    Wang, P., Chen, X., Chen, T., Venugopalan, S., Wang, Z., et al. (2022). Is attention all that nerf needs? arXiv preprint arXiv:2207.13298.
    Wang, Z., Wu, S., Xie, W., Chen, M., & Prisacariu, V. A. (2021). Nerf–: Neural radiance fields without known camera parameters.
    Warburg, F., Weber, E., Tancik, M., Holynski, A., & Kanazawa, A. (2023). Nerfbusters: Removing ghostly artifacts from casually captured nerfs. Proceedings of the IEEE/CVF International Conference on Computer Vision, 18120–18130.
    Weber, E., Holynski, A., Jampani, V., Saxena, S., Snavely, N., Kar, A., & Kanazawa, A. (2024). Nerfiller: Completing scenes via generative 3d inpainting. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 20731–20741.
    Xiao, W., Chierchia, R., Cruz, R. S., Li, X., Ahmedt-Aristizabal, D., Salvado, O., Fookes, C., & Lebrat, L. (2025). Neural radiance fields for the real world: A survey. arXiv preprint arXiv:2501.13104.
    Yao, Y., Luo, Z., Li, S., Fang, T., & Quan, L. (2018). Mvsnet: Depth inference for unstructured multi-view stereo. Proceedings of the European conference on computer vision (ECCV), 767–783.
    Yariv, L., Gu, J., Kasten, Y., & Lipman, Y. (2021). Volume rendering of neural implicit surfaces. Advances in neural information processing systems, 34, 4805–4815.
    Yi, T., Fang, J., Wang, J., Wu, G., Xie, L., Zhang, X., Liu, W., Tian, Q., & Wang, X. (2024). Gaussiandreamer: Fast generation from text to 3d gaussians by bridging 2d and 3d diffusion models. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 6796–6807.
    Yu, A., Li, R., Tancik, M., Li, H., Ng, R., & Kanazawa, A. (2021). Plenoctrees for real-time rendering of neural radiance fields. Proceedings of the IEEE/CVF international conference on computer vision, 5752–5761.
    Zhang, R., Isola, P., Efros, A. A., Shechtman, E., & Wang, O. (2018). The unreasonable effectiveness of deep features as a perceptual metric. Proceedings of the IEEE conference on computer vision and pattern recognition, 586–595.
    Zhang, W., Xing, R., Zeng, Y., Liu, Y.-S., Shi, K., & Han, Z. (2023). Fast learning radiance fields by shooting much fewer rays. IEEE Transactions on Image Processing, 32, 2703–2718.
    ROCMan. (n.d.). NeRF: Representing Scenes as Neural Radience Fields for View Synthesis: 図だくさん解説.

    QR CODE