簡易檢索 / 詳目顯示

研究生: 何淑雯
Hogo, Wynne
論文名稱: 統計方法與神經網路模型在序位資料分析之比較研究
Comparative Analysis of Statistical and Neural Network Approaches to Ordinal Data
指導教授: 李國榮
Lee, Kuo-Jung
學位類別: 碩士
Master
系所名稱: 管理學院 - 數據科學研究所
Institute of Data Science
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 62
中文關鍵詞: 序位迴歸 、比例勝算模型 、Brant 檢定 、Ordinal ResLogit 、深度學習 、CoGOL 、Integrated Gradients
外文關鍵詞: Ordinal Regression, Proportional Odds Model, Brant Test, Ordinal ResLogit, CoGOL, Deep Learning, Integrated Gradients
相關次數: 點閱:97  下載:0 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 序位分類(Ordinal Classification)是一種監督式學習問題,其目標變數具有自然的排序關係,廣泛應用於房屋品質評估、臨床疾病嚴重程度評估及醫療保險費用預測等領域。比例勝算模型(Proportional Odds Model, POM)是序位迴歸中最常用的方法之一,但其比例勝算假設在具有非線性關係與複雜資料結構的真實資料中常不成立,因此促使研究者發展更具彈性的深度學習方法。
    本研究比較四種序位迴歸模型,包括比例勝算模型(POM)、Ordinal ResLogit、Continuously Generalized Ordinal Logistic(CoGOL),以及本研究提出的 Hybrid ResLogit-CoGOL 模型。研究使用合成資料與真實資料進行實驗。合成資料包含三種情境:前兩種分別在類別不平衡與類別平衡的條件下違反比例勝算假設,第三種則近似滿足比例勝算假設,以進行公平比較。真實資料則採用 Ames Housing 與 Medical Insurance Cost 兩個資料集。模型表現以準確率(Accuracy)、平均絕對誤差(Mean Absolute Error, MAE)、均方誤差(Mean Squared Error, MSE)、Macro F1-score 及 Macro Recall 進行評估,並利用 Integrated Gradients 分析各模型的重要特徵。
    實驗結果顯示,模型表現會受到資料特性的影響。在違反比例勝算假設的合成資料中,Hybrid ResLogit-CoGOL 模型於各項評估指標皆有最佳表現,而 POM 的表現則明顯下降。當比例勝算假設近似成立時,各模型之間的效能差異大幅縮小,顯示深度學習模型的優勢主要來自於比例勝算假設被違反的情況,而非模型架構本身。Brant 檢定結果亦證實,兩個真實資料集皆不符合比例勝算假設。在 Ames Housing 資料集中,Hybrid 模型於三類別分類表現最佳,CoGOL 在四類別分類中表現最佳,而 Ordinal ResLogit 則於五類別分類中獲得最佳結果;在 Medical Insurance Cost 資料集中,Ordinal ResLogit 的整體表現最佳。
    特徵重要性分析結果顯示,相較於 POM,深度學習模型能夠捕捉更多非線性特徵及複雜的特徵交互關係;而 POM 則主要依賴具有較強線性效果的特徵。整體而言,深度學習模型在處理具有複雜資料結構的序位分類問題時,展現出較佳的建模能力與預測效能。

    Ordinal classification is a supervised learning problem where the response categories have a natural order. It is commonly used in applications such as housing quality assessment, clinical severity evaluation, and insurance cost prediction. The Proportional Odds Model (POM) is one of the most widely used methods for ordinal regression, but its proportional odds assumption is often violated in real-world data with nonlinear relationships and complex decision boundaries. This has motivated the development of more flexible deep learning approaches.
    This study compares four ordinal regression models: POM, Ordinal ResLogit, Continuously Generalized Ordinal Logistic (CoGOL), and a Hybrid ResLogit-CoGOL model proposed in this work. The models are evaluated on both synthetic and real-world datasets. Three synthetic scenarios are designed: two that violate the proportional odds assumption under imbalanced and balanced class distributions, and one that approximately satisfies the assumption. Real-world experiments are conducted using the Ames Housing and Medical Insurance Cost datasets. Model performance is evaluated using accuracy, Mean Absolute Error (MAE), Mean Squared Error (MSE), macro F1-score, and macro recall. Integrated Gradients is also used to interpret feature importance.
    The results show that model performance depends on the underlying data characteristics. On the synthetic datasets that violate the proportional odds assumption, the Hybrid ResLogit-CoGOL model achieves the best overall performance, while POM performs substantially worse. However, when the proportional odds assumption is approximately satisfied, the performance differences between models become much smaller, indicating that the advantage of deep learning models mainly appears when the assumption is violated. Brant test results on both real-world datasets further confirm that the proportional odds assumption is violated. In the Ames Housing dataset, the Hybrid model performs best for the three-class problem, CoGOL performs best for the four-class problem, and Ordinal ResLogit achieves the highest performance for the five-class problem. On the Medical Insurance Cost dataset, Ordinal ResLogit produces the best overall results.
    Feature attribution analysis further shows that the deep learning models capture more complex nonlinear relationships than POM. While POM mainly relies on predictors with strong linear effects, the neural network models identify a wider range of important features, demonstrating their ability to model more complicated ordinal decision boundaries.

    1 Introduction 7 1.1 Background and Motivation 7 1.2 Research Questions 8 1.3 Contributions 8 2 Literature Review 9 2.1 Proportional Odds Model (POM) 9 2.2 Brant Test 10 2.3 Generalized Ordinal Logit (GOL) / Partial Proportional Odds 10 2.4 All-thresholds Loss 10 2.5 Consistent Rank Logits (CORAL) 11 2.6 Residual Logit (ResLogit) 11 2.7 Ordinal-ResLogit 12 2.8 Continuously Generalized Ordinal Logit (CoGOL) 12 2.9 Integrated Gradients 12 2.10 Remarks 13 3 Methodology 14 3.1 Overall Research Design 14 3.2 Ordinal Classification Framework 14 3.3 Models Evaluated 15 3.3.1 Proportional-Odds Model (POM) 15 3.3.2 Ordinal ResLogit 15 3.3.3 Deep CoGOL 16 3.3.4 Hybrid ResLogit–CoGOL 17 3.4 Software and Implementation 18 3.5 Loss Functions 19 3.6 Hyperparameter Optimization 20 3.7 Training Configuration 20 3.8 Evaluation Metrics 20 3.9 Feature Attribution via Integrated Gradients 21 3.10 Testing the Proportional Odds Assumption 22 4 Synthetic Simulation 23 4.1 Overview 23 4.2 Data Generation Process 23 4.2.1 Feature Generation 24 4.2.2 Latent Signal Construction 24 4.3 Ordinal Response Construction 24 4.3.1 Scenario A: Imbalanced Class Distribution 25 4.3.2 Scenario B: Balanced Class Distribution 26 4.3.3 Scenario C: Proportional-Odds-Consistent Dataset 27 4.4 Data Splitting 29 5 Real-World Experiment 30 5.1 Dataset Description 30 5.1.1 Ames Housing Dataset 30 5.1.2 Medical Insurance Cost Dataset 30 5.2 Ordinal Response Distribution 30 5.2.1 Ames Housing Dataset 30 5.2.2 Medical Insurance Cost Dataset 31 5.3 Ordinal Response Construction 32 5.3.1 Ames Housing Dataset 32 5.3.2 Medical Insurance Cost Dataset 35 5.4 Feature Selection 37 5.4.1 Ames Housing Dataset 37 5.4.2 Medical Insurance Cost Dataset 37 5.5 Preprocessing 38 6 Results and Discussion 39 6.1 Synthetic Simulation Results 39 6.1.1 Scenario A: Imbalanced Class Distribution 39 6.1.2 POM Degradation Under Complex Data Conditions 40 6.1.3 Scenario B: Balanced Class Distribution 40 6.1.4 Scenario C: Proportional-Odds-Consistent Distribution 41 6.2 Real-World Experiment Results 42 6.2.1 Verification of the Proportional Odds Assumption 42 6.2.2 Ames Housing Dataset: Three-Class Grouping 44 6.2.3 Ames Housing Dataset: Four-Class Grouping 45 6.2.4 Ames Housing Dataset: Five-Class Grouping 46 6.2.5 Medical Insurance Cost Dataset 47 6.3 Feature Attribution via Integrated Gradients 48 6.3.1 Ames Housing Dataset: Three-Class Setting 48 6.3.2 Ames Housing Dataset: Four-Class Setting 49 6.3.3 Ames Housing Dataset: Five-Class Setting 50 6.3.4 Medical Insurance Cost 51 7 Summary 53 7.1 Summary of Findings 53 7.2 Practical Implications 55 7.3 Limitations and Future Work 57 References 59

    [1] Agresti, A. Analysis of Ordinal Categorical Data. Wiley, Hoboken, NJ, 2010.
    [2] Agency for Healthcare Research and Quality (AHRQ). Concentration of Healthcare Expenditures and Selected Characteristics of People with High Expenses, United States Civilian Noninstitutionalized Population, 2018–2022. Medical Expenditure Panel Sur-vey (MEPS), Rockville, MD, Statistical Brief No. 560, 2024.
    [3] Baccianella, S., Esuli, A., & Sebastiani, F. Evaluation measures for ordinal regression. In Proceedings of the 9th IEEE International Conference on Intelligent Systems Design and Applications, pp. 283–287, 2009.
    [4] Bhupathiraju, S. N., & Hu, F. B. Epidemiology of obesity and diabetes and their cardiovascular complications. Circulation Research, 118, 11, 1723–1735, 2016.
    [5] Bin, O. A prediction comparison of housing sales prices by parametric versus semi-parametric regressions. Journal of Housing Economics, 13, 1, 68–84, 2004.
    [6] Brant, R. Assessing proportionality in the proportional odds model for ordinal logistic regression. Biometrics, 46, 4, 1171–1178, 1990.
    [7] Cao, W., Mirjalili, V., & Raschka, S. Rank consistent ordinal regression for neural networks with application to age estimation. Pattern Recognition Letters, 140, 325–331, 2020.
    [8] Clevert, D.-A., Unterthiner, T., & Hochreiter, S. Fast and accurate deep network learning by exponential linear units (ELUs). Proceedings of the 4th International Conference on Learning Representations (ICLR), 2016.
    [9] De Cock, D. Ames, Iowa: Alternative to the Boston housing data as an end of semester regression project. Journal of Statistics Education, 19, 3, 2011.
    [10] Deb, P., & Trivedi, P. K. Demand for medical care by the elderly: A finite mixture approach. Journal of Applied Econometrics, 12, 3, 313–336, 1997.
    [11] Diaz, R., & Marathe, A. Soft labels for ordinal regression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4738–4747, 2019.
    [12] Dormann, C. F., Elith, J., Bacher, S., et al. Collinearity: A review of methods to deal with it and a simulation study evaluating their performance. Ecography, 36, 1, 27–46, 2013.
    [13] Fernandez-Navarro, F., Gutierrez, P. A., Hervas-Martinez, C., & Carbonero-Ruz, M. Ordinal classification using hybrid artificial neural networks with projection and kernel basis functions. Neural Computing and Applications, 33, 10, 4869–4881, 2021.
    [14] Freedman, D. S., Khan, L. K., Serdula, M. K., Dietz, W. H., Srinivasan, S. R., & Berenson, G. S. The relation of childhood BMI to adult adiposity: The Bogalusa Heart Study. Pediatrics, 115, 1, 22–27, 2006.
    [15] Fullerton, A. S. A conceptual framework for ordered logistic regression models. Soci-ological Methods & Research, 38, 2, 306–347, 2009.
    [16] Harris, C. R., Millman, K. J., van der Walt, S. J., et al. Array programming with NumPy. Nature, 585, 357–362, 2020.
    [17] He, H., & Garcia, E. A. Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engineering, 21, 9, 1263–1284, 2009.
    [18] He, K., Zhang, X., Ren, S., & Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770–778, 2016.
    [19] Hunter, J. D. Matplotlib: A 2D graphics environment. Computing in Science & Engineering, 9, 3, 90–95, 2007.
    [20] Johnson, J. M., & Khoshgoftaar, T. M. Survey on deep learning with class imbalance.Journal of Big Data, 6, 1, 27, 2019.
    [21] Kamal, K., & Farooq, B. Ordinal-ResLogit: Interpretable deep residual neural net-works for ordered choices. Journal of Choice Modelling, 50, 100454, 2024.
    [22] Lantz, B. Machine Learning with R: Expert Techniques for Predictive Modeling. Packt Publishing, Birmingham, UK, 2019.
    [23] Li, L., & Lin, H.-T. Ordinal regression by extended binary classification. In Advances in Neural Information Processing Systems, Vol. 19, pp. 865–872, 2007.
    [24] Liu, X. Ordinal regression analysis: Fitting the proportional odds model using Stata, SAS and SPSS. Journal of Modern Applied Statistical Methods, 8, 2, 632–645, 2009.
    [25] Lu, F., Ferraro, F., & Raff, E. Continuously generalized ordinal regression for linear and deep models. In Proceedings of the 2022 SIAM International Conference on Data Mining (SDM), pp. 1038–1046, Society for Industrial and Applied Mathematics (SIAM), 2022.
    [26] Lundberg, S. M., & Lee, S.-I. A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems (NeurIPS), 30, 2017.
    [27] McCullagh, P. Regression models for ordinal data. Journal of the Royal Statistical Society: Series B, 42, 2, 109–142, 1980.
    [28] McKinney, W. Data structures for statistical computing in Python. In Proceedings of the 9th Python in Science Conference, pp. 56–61, 2010.
    [29] Niu, Z., Zhou, M., Wang, L., Gao, X., & Hua, G. Ordinal regression with multiple output CNN for age estimation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4920–4928, 2016.
    [30] Paszke, A., Gross, S., Massa, F., Lerer, A., et al. PyTorch: An imperative style, high-performance deep learning library. Advances in Neural Information Processing Systems, 32, 8024–8035, 2019.
    [31] Pedregosa, F., Varoquaux, G., Gramfort, A., et al. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830, 2011.
    [32] Peterson, B., & Harrell, F. E. Partial proportional odds models for ordinal response variables. Applied Statistics, 39, 2, 205–217, 1990.
    [33] Rennie, J. D. M., & Srebro, N. Loss functions for preference levels: Regression with discrete ordered labels. In Proceedings of the IJCAI Multidisciplinary Workshop on Advances in Preference Handling, pp. 180–186, 2005.
    [34] Ribeiro, M. T., Singh, S., & Guestrin, C. “Why should I trust you?”: Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 1135–1144, 2016.
    [35] Rosen, S. Hedonic prices and implicit markets: Product differentiation in pure com-petition. Journal of Political Economy, 82, 1, 34–55, 1974.
    [36] Sáez, J. A., Luengo, J., Stefanowski, J., & Herrera, F. SMOTE-IPF: Addressing the noisy and borderline examples problem in imbalanced classification by a re-sampling method with filtering. Information Sciences, 291, 184–203, 2016.
    [37] Seabold, S., Perktold, J. Statsmodels: Econometric and statistical modeling with python. Proceedings of the 9th Python in Science Conference, 57–61, 2010.
    [38] Sokolova, M., & Lapalme, G. A systematic analysis of performance measures for classification tasks. Information Processing and Management, 45, 4, 427–437, 2009.
    [39] Sundararajan, M., Taly, A., & Yan, Q. Axiomatic attribution for deep networks. In Proceedings of the 34th International Conference on Machine Learning (ICML), Vol. 70, pp. 3319–3328, 2017.
    [40] Vargas, V. M., Gutierrez, P. A., & Hervas-Martinez, C. Unimodal regularisation based on beta distribution for deep ordinal regression. Pattern Recognition, 122, 108310,2022.
    [41] Williams, R. Generalized ordered logit/partial proportional odds models for ordinal dependent variables. The Stata Journal, 6, 1, 58–82, 2006.
    [42] Wong, M., & Farooq, B. ResLogit: A residual neural network logit model for data-driven choice modelling. Transportation Research Part C: Emerging Technologies,126, 103050, 2021.

    下載圖示
    校外:立即公開
    QR CODE