簡易檢索 / 詳目顯示

研究生: 張聖岳
CHANG, SHENG-YUEH
論文名稱: 以深度神經網路形塑具變數選擇之高階非參數空間自迴歸模型
Flexible Spatial Modeling with Variable Selection: A Deep Neural Network-Based Higher-Order Nonparametric Spatial Autoregressive Framework
指導教授: 李國榮
Lee, Kuo-Jung
學位類別: 碩士
Master
系所名稱: 管理學院 - 統計學系
Department of Statistics
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 58
中文關鍵詞: 變數選擇高階非參數空間自迴歸模型深度神經網路
外文關鍵詞: variable selection, higher-order nonparametric spatial autoregressive model, deep neural network, spatial heterogeneity
相關次數: 點閱:3下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 深度神經網路因具有高度彈性的函數逼近能力,近年來已廣泛應用於複雜資料建模。然而,傳統空間自迴歸模型通常仰賴研究者預先指定空間權重矩陣,而一般深度神經網路亦不易提供變數層級的解釋。基於此,本研究探討具有非參數內生效果之高階非參數空間自迴歸模型中的變數選擇問題,並提出一個整合學習式空間權重機制、非參數內生空間效果與結構化變數選擇的深度神經網路架構。
    在空間結構方面,所提模型透過多個由資料學習的空間權重機制及其對應之非參數內生效果,刻畫可能存在的多重與非線性空間相依關係。在主效應建構方面,將主效應函數分解為線性 skip component 與非線性神經網路分支,並結合 L1 懲罰與階層式限制,使變數在線性係數被壓縮為零時,其於非線性分支中的連結亦同步移除,從而在保留非線性擬合能力的同時進行變數選擇。
    在理論方面,本研究基於固定點定理,推導所提模型之反應值解存在性與唯一性的充分條件,以說明模型在適當條件下的數學可行性。在計算方面,依參數所涉及的最佳化結構,將估計程序區分為一般梯度更新與階層式 proximal 更新,並結合路徑式稀疏化策略與驗證集模型選擇,以同步完成參數估計與重要變數辨識。
    模擬研究與實際資料分析結果顯示,所提方法能有效辨識具有線性與非線性效果的重要變數,並在預測表現上優於比較方法。實證結果亦顯示,學習式空間權重機制與結構化變數選擇可共同提升模型對複雜空間資料的建模能力。整體而言,本研究所提出之方法能兼顧空間相依關係的建模彈性、預測表現與變數層級解釋性,顯示其應用於複雜空間資料分析的可行性。

    This thesis proposes a higher-order nonparametric spatial autoregressive neural network model with variable selection. The proposed framework is motivated by two limitations in spatial data analysis: conventional spatial autoregressive models usually rely on pre-specified spatial weight matrices, while flexible deep neural networks often lack variable-level interpretability.
    The proposed model integrates multiple learnable spatial weight mechanisms, nonparametric endogenous spatial effects, and a structured main-effect component. The spatial component is written as Σ_{m=1}^M W_{space,Φ_m}(X) h_{Γ_m}(Y), allowing the model to learn multiple spatial dependence structures and nonlinear spatial feedback effects from data. The main-effect component is decomposed as f_{θ,Ξ_G}(X) = Xθ + g_{Ξ_G}(X), where the linear skip component supports variable selection and the nonlinear neural network component preserves flexible function approximation. By combining an L1 penalty with a hierarchical constraint, the model can remove unimportant variables from both the linear and nonlinear main-effect branches.
    Theoretical analysis establishes sufficient conditions for the existence and uniqueness of the response solution based on the Banach fixed-point theorem. Computationally, the model is estimated through gradient-based updates and hierarchical proximal updates along a regularization path.
    Simulation studies show that the proposed method can recover important variables and improve predictive performance, especially when a covariate is involved in both spatial weight construction and the true main effect. Empirical analysis using the California Housing Prices data further demonstrates that the proposed method achieves better test performance than the fixed-W comparison method while retaining meaningful variables in the main-effect support.

    中文摘要 i Abstract ii 誌謝 vii 目錄 ix 表目錄 xii 第一章 緒論 1 1-1. 研究背景與動機 1 1-2. 研究問題 3 1-3. 研究目的 3 1-4. 研究貢獻 4 1-5. 論文架構 4 第二章 文獻回顧 6 2-1. 空間自迴歸模型 6 2-2. 高階空間自迴歸模型 6 2-3. 非參數內生空間效果 7 2-4. 深度神經網路與非參數建模 8 2-5. 深度學習下的變數選擇方法 8 2-6. 結構化懲罰與階層式限制 9 2-7. 相關模型比較 10 2-8. 文獻評述 10 第三章 方法 12 3-1. 研究架構概述 12 3-2. 符號定義與資料設定 12 3-3. 高階非參數空間自迴歸模型 13 3-3.1 整體模型形式 13 3-3.2 模型分解與記號 14 3-4. 空間分支:學習式空間機制與非參數內生效果 14 3-4.1 學習式空間權重機制 14 3-4.2 非參數內生空間效果 15 3-5. 主效應分支與變數選擇 16 3-5.1 主效應函數分解 16 3-5.2 階層式結構限制 16 3-6. 理論基礎:固定點解之存在性與唯一性 16 3-6.1 Banach fixed-point theorem 17 3-6.2 本文模型中的固定點映射 17 3-6.3 充分條件 17 3-7. 目標函數與最佳化問題 18 3-7.1 經驗損失函數 (Empirical Loss function) 18 3-7.2 含變數選擇之最佳化問題 19 3-7.3 可行集合非空性 19 3-8. 估計與優化策略 19 3-8.1 Dense warm start 19 3-8.2 Regularization path 20 3-9. 演算法設計 20 3-9.1 Feature-wise hierarchical proximal operator 21 3-10. 超參數與模型選擇 24 3-11. 本章小結 24 第四章 模擬方法與資料分析 25 4-1. 評估目的與整體設計 25 4-2. 評估指標 25 4-3. 模擬研究 26 4-3.1 模擬情境設計 26 4-3.2 比較方法 28 4-3.3 懲罰參數選擇 29 4-3.4 模擬結果 29 4-3.5 模擬結果討論 31 4-4. California Housing Prices 實證分析 32 4-4.1 資料來源與變數說明 32 4-4.2 類別變數處理 32 4-4.3 空間分層與資料切分 33 4-4.4 比較方法與模型選擇 33 4-4.5 實證結果 33 4-4.6 實證結果討論 34 4-5. 綜合討論 35 第五章 結論與未來展望 36 5-1. 主要研究發現與貢獻 36 5-2. 研究限制與未來研究方向 38 5-2.1 空間機制中的變數選擇與多重超參數問題 38 5-2.2 非連續型反應變數與計數資料之延伸 38 5-2.3 時空模型與空間–時間聯合變數選擇 39 5-3. 總結 40 參考文獻 41

    P. A. P. Moran, Notes on continuous stochastic phenomena, Biometrika, 37, 1/2, 17–23, 1950. https://doi.org/10.2307/2332142.
    P. Erdős, A. Rényi, On random graphs I, Publicationes Mathematicae Debrecen, 6, 290–297, 1959.
    P. Erdős, A. Rényi, On the evolution of random graphs, Publications of the Mathematical Institute of the Hungarian Academy of Sciences, 5, 17–61, 1960.
    K. Ord, Estimation methods for models of spatial interaction, Journal of the American Statistical Association, 70, 349, 120–126, 1975. https://doi.org/10.1080/01621459.1975.10480272.
    A. D. Cliff, J. K. Ord, Spatial Processes: Models and Applications, Pion, London, 1981.
    L. Anselin, Spatial Econometrics: Methods and Models, Studies in Operational Regional Science, vol. 4, Kluwer Academic Publishers, Dordrecht, 1988. https://doi.org/10.1007/978-94-015-7799-1.
    L. Anselin, Local indicators of spatial association—LISA, Geographical Analysis, 27, 2, 93–115, 1995. https://doi.org/10.1111/j.1538-4632.1995.tb00338.x.
    N. A. C. Cressie, Statistics for Spatial Data, Revised edition, John Wiley & Sons, New York, 1993.
    H. H. Kelejian, I. R. Prucha, A generalized spatial two-stage least squares procedure for estimating a spatial autoregressive model with autoregressive disturbances, The Journal of Real Estate Finance and Economics, 17, 99–121, 1998. https://doi.org/10.1023/A:1007707430416.
    J. P. LeSage, R. K. Pace, Introduction to Spatial Econometrics, Chapman and Hall/CRC, Boca Raton, 2009. https://doi.org/10.1201/9781420064254.
    J. P. Elhorst, Spatial Econometrics: From Cross-Sectional Data to Spatial Panels, SpringerBriefs in Regional Science, Springer, Berlin, Heidelberg, 2014. https://doi.org/10.1007/978-3-642-40340-8.
    R. Tibshirani, Regression shrinkage and selection via the Lasso, Journal of the Royal Statistical Society: Series B, 58, 1, 267–288, 1996.
    J. Fan, R. Li, Variable selection via nonconcave penalized likelihood and its oracle properties, Journal of the American Statistical Association, 96, 456, 1348–1360, 2001. https://doi.org/10.1198/016214501753382273.
    H. Zou, The adaptive Lasso and its oracle properties, Journal of the American Statistical Association, 101, 476, 1418–1429, 2006. https://doi.org/10.1198/016214506000000735.
    M. Yuan, Y. Lin, Model selection and estimation in regression with grouped variables, Journal of the Royal Statistical Society: Series B, 68, 1, 49–67, 2006.
    T. Hastie, R. Tibshirani, J. Friedman, The Elements of Statistical Learning: Data Mining, Inference, and Prediction, 2nd ed., Springer, New York, 2009. https://doi.org/10.1007/978-0-387-84858-7.
    T. Hastie, R. Tibshirani, M. Wainwright, Statistical Learning with Sparsity: The Lasso and Generalizations, Chapman and Hall/CRC, Boca Raton, 2015. https://doi.org/10.1201/b18401.
    G. Cybenko, Approximation by superpositions of a sigmoidal function, Mathematics of Control, Signals and Systems, 2, 303–314, 1989. https://doi.org/10.1007/BF02551274.
    K. Hornik, M. Stinchcombe, H. White, Multilayer feedforward networks are universal approximators, Neural Networks, 2, 5, 359–366, 1989. https://doi.org/10.1016/0893-6080(89)90020-8.
    C. M. Bishop, Pattern Recognition and Machine Learning, Information Science and Statistics, Springer, New York, 2006.
    I. Goodfellow, Y. Bengio, A. Courville, Deep Learning, MIT Press, Cambridge, MA, 2016. https://www.deeplearningbook.org.
    I. Lemhadri, F. Ruan, L. Abraham, R. Tibshirani, LassoNet: A neural network with feature sparsity, Journal of Machine Learning Research, 22, 127, 1–29, 2021. https://jmlr.org/papers/v22/20-848.html.
    F. Bach, R. Jenatton, J. Mairal, G. Obozinski, Optimization with sparsity-inducing penalties, Foundations and Trends in Machine Learning, 4, 1, 1–106, 2012. https://doi.org/10.1561/2200000015.
    N. Parikh, S. Boyd, Proximal algorithms, Foundations and Trends in Optimization, 1, 3, 127–239, 2014. https://doi.org/10.1561/2400000003.
    S. Boyd, L. Vandenberghe, Convex Optimization, Cambridge University Press, Cambridge, 2004.
    D. P. Kingma, J. Ba, Adam: A method for stochastic optimization, International Conference on Learning Representations, 2015. https://arxiv.org/abs/1412.6980.
    S. Banach, Sur les opérations dans les ensembles abstraits et leur application aux équations intégrales, Fundamenta Mathematicae, 3, 1, 133–181, 1922. https://doi.org/10.4064/fm-3-1-133-181.
    A. Granas, J. Dugundji, Fixed Point Theory, Springer Monographs in Mathematics, Springer, New York, 2003. https://doi.org/10.1007/978-0-387-21593-8.
    H. H. Bauschke, P. L. Combettes, Convex Analysis and Monotone Operator Theory in Hilbert Spaces, Springer, New York, 2011.
    Z. Wang, Y. Song, Deep learning for the spatial additive autoregressive model with nonparametric endogenous effect, Spatial Statistics, 55, 100743, 2023. https://doi.org/10.1016/j.spasta.2023.100743.
    S. Xiao, Y. Song, Z. Wang, Nonparametric spatial autoregressive model using deep neural networks, Spatial Statistics, 57, 100766, 2023. https://doi.org/10.1016/j.spasta.2023.100766.
    Z. Li, Y. Song, L. Jian, Deep learning for higher-order nonparametric spatial autoregressive model, Applied Intelligence, 54, 7570–7580, 2024. https://doi.org/10.1007/s10489-024-05541-8.
    J. Li, Y. Song, L. Jian, Deep neural networks for variable selection of higher-order nonparametric spatial autoregressive model, Statistics and Computing, 34, 194, 2024. https://doi.org/10.1007/s11222-024-10500-x.
    J. Li, Y. Song, Variable selection for nonparametric spatial additive autoregressive model via deep learning, Statistical Papers, 66, 53, 2025. https://doi.org/10.1007/s00362-025-01669-y.
    R. Yang, Y. Song, Variable selection for nonparametric spatial expectile regression using deep neural networks, Networks and Spatial Economics, 25, 847–885, 2025. https://doi.org/10.1007/s11067-025-09685-z.
    T. Xie, R. Cao, J. Du, Variable selection for spatial autoregressive models with a diverging number of parameters, Statistical Papers, 61, 3, 1125–1145, 2020. https://doi.org/10.1007/s00362-018-0984-2.
    T. N. Kipf, M. Welling, Semi-supervised classification with graph convolutional networks, International Conference on Learning Representations, 2017. https://arxiv.org/abs/1609.02907.
    P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Liò, Y. Bengio, Graph attention networks, International Conference on Learning Representations, 2018. https://arxiv.org/abs/1710.10903.
    C. Nugent, California Housing Prices, Kaggle, 2017. Available at: https://www.kaggle.com/datasets/camnugent/california-housing-prices (accessed 29 June 2026).

    QR CODE