| 研究生: |
王俊傑 Wang, Chun-Chieh |
|---|---|
| 論文名稱: |
有限混合汙染常態線性混合效應模型之最大概似及貝氏推論 Maximum Likelihood and Bayesian Inference for Finite Mixtures of Contaminated Normal Linear Mixed-effects Models |
| 指導教授: |
王婉倫
Wang, Wan-Lun |
| 學位類別: |
碩士 Master |
| 系所名稱: |
管理學院 - 統計學系 Department of Statistics |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 82 |
| 中文關鍵詞: | AECM 演算法 、分群 、條件共軛先驗 、群組式長期追蹤資料 、MCMC程序 、溫和離群值 |
| 外文關鍵詞: | AECM algorithm, Clustering, Conditional conjugate priors, Grouped longitudinal data, MCMC procedure, Mild outliers |
| 相關次數: | 點閱:81 下載:5 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
有限混合線性混合效應(FM-LME)模型已成為處理具異質性之長期追蹤資料分群的常用工具。然而,FM-LME 模型中傳統的常態假設對於溫和型離群值或受污染數據可能較為敏感。本研究透過考慮群組隨機效應與個體內誤差服從多變量污染常態(MCN)分配,進一步擴展了 FM-LME 模型,此後將其簡稱為有限混合污染常態線性混合效應(FM-CNLME)模型。與多變量常態分配相比,MCN 分配額外包含兩個參數:一個用於控制溫和型離群值的比例,另一個則用於指定污染程度,進而提升參數估計結果的穩健性。在推論模型參數及隱藏資料方面,本研究發展兩套方法:最大概似(ML)估計與貝氏(Bayesian)推論。一個有效率的交替預期條件最大化(AECM)演算法被提供以執行最大概似估計。從完全貝氏觀點,結合吉布斯採樣(Gibbs sampler)與 Metropolis-Hastings(M-H)演算法的馬可夫鏈蒙地卡羅(MCMC)採樣程序被發展以進行後驗推估。數值結果顯示,在處理具有群體結構與污染觀測值的複雜長期追蹤資料時,相較於現有的模型,FM-CNLME 模型能有效識別潛在分群、達成更好的資料擬合優度,並在估計參數和計算配適值上提供更高的精確度。最後,本研究所提出的模型與方法亦應用於實證資料,以證明其在實務上的價值與效率。
Finite mixtures of linear mixed-effects (FM-LME) models have become a commonly used tool for clustering longitudinal data with heterogeneity. However, the conventional assumption of normality in the FM-LME model may be sensitive to mild outliers or contaminated points. This thesis extends the FM-LME model by considering the multivariate contaminated normal (MCN) distributions for component random effects and within-subjects errors, referred to as the finite mixtures of contaminated normal linear mixed-effects (FM-CNLME) model henceforth. Comparing with the multivariate normal distribution, the MCN distribution has two extra parameters: one for controlling the proportion of mild outliers and the other for specifying the degree of contamination, improving the robustness of the estimation results. Regarding the inference of parameters as well as latent data in the model, we develop two approaches: maximum likelihood (ML) estimation and Bayesian inference. For carrying out ML estimation, an efficient alternating expectation conditional maximization (AECM) algorithm is provided. From a fully Bayesian viewpoint, Markov Chain Monte Carlo (MCMC) sampling procedure, which combines the Gibbs sampler and Metropolis-Hastings (M-H) algorithm, is developed to perform posterior inference. Numerical results show that, when dealing with complex longitudinal data that have grouped structures and contaminated observations, the FM-CNLME model can effectively identify hidden groups, achieves better fitness of data, and offer more accurate estimation of parameters as well as fitted responses compared with the existing model. Our proposed model and methods are applied to real-data examples to demonstrate their practical values and efficiency.
Akaike, H. (1998). Information theory and an extension of the maximum likelihood principle. In Selected Papers of Hirotugu Akaike, pages 199–213. Springer.
Azizian, A., Rühlmann, F., Krause, T., Bernhardt, M., Jo, P., König, A., Kleiß, M., Leha, A., Ghadimi, M., and Gaedcke, J. (2020). Ca19-9 for detecting recurrence of pancreatic cancer. Scientific Reports, 10(1):1332.
Azzalini, A. and Capitanio, A. (2003). Distributions generated by perturbation of symmetry with emphasis on a multivariate skew t-distribution. Journal of the Royal Statistical Society: Series B ( Statistical Methodology), 65(2):367–389.
Azzalini, A. and Dalla Valle, A. (1996). The multivariate skew-normal distribution. Biometrika, 83(4):715–726.
Bai, X., Chen, K., and Yao, W. (2016). Mixture of linear mixed models using multivariate t distribution. Journal of Statistical Computation and Simulation, 86(4):771–787.
Brooks, S. P. and Gelman, A. (1998). General methods for monitoring convergence of iterative simulations. Journal of Computational and Graphical Statistics, 7(4):434–455.
Brooks, S. P., Smith, J., Vehtari, A., et al. (2002). Discussion on the paper by Spiegelhalter, Best, Carlin and van der Linde. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 64(4):616–618.
Browne, W. J. and Draper, D. (2006). A comparison of Bayesian and likelihood-based methods for fitting multilevel models. Bayesian Analysis, 1(3):473–514.
Carlin, B. P. and Louis, T. A. (2008). Bayesian Methods for Data Analysis. Texts in Statistical Science. CRC Press, Boca Raton, FL, 3rd edition.
Celeux, G., Martin, O., and Lavergne, C. (2005). Mixture of linear mixed models for clustering gene expression profiles from repeated microarray experiments. Statistical Modelling, 5(3):243–267.
De la Cruz-Mesia, R., Quintana, F. A., and Marshall, G. (2008). Model-based clustering for longitudinal data. Computational Statistics & Data Analysis, 52(3):1441–1457.
Dempster, A. P., Laird, N. M., and Rubin, D. B. (1977). Maximum likelihood from incomplete data via the em algorithm. Journal of the Royal Statistical Society: Series B (Methodological), 39(1):1–22.
Frühwirth-Schnatter, S. (2006). Finite Mixture and Markov Switching Models. Springer, New York.
Frühwirth-Schnatter, S. and Pyne, S. (2010). Bayesian inference for finite mixtures of univariate and multivariate skew-normal and skew-t distributions. Biostatistics, 11(2):317–336.
Gelman, A. and Rubin, D. B. (1992). Inference from iterative simulation using multiple sequences. Statistical Science, 7(4):457–472.
Grün, B. and Leisch, F. (2007). Fitting finite mixtures of generalized linear regressions in r. Computational Statistics & Data Analysis, 51(11):5247–5252.
Hang, J., Wu, L., Zhu, L., Sun, Z., Wang, G., Pan, J., Zheng, S., Xu, K., Du, J., and Jiang, H. (2018). Prediction of overall survival for metastatic pancreatic cancer: Development and validation of a prognostic nomogram with data from open clinical trial and real-world study. Cancer Medicine, 7(7):2974–2984.
Hastings, W. K. (1970). Monte carlo sampling methods using markov chains and their applications. Biometrika, 57(1):97–109.
Kotz, S. and Nadarajah, S. (2004). Multivariate t-Distributions and Their Applications. Cambridge university press, Cambridge.
Lachos, V. H., Castro, L. M., and Dey, D. K. (2013). Bayesian inference in nonlinear mixed-effects models using normal independent distributions. Computational Statistics & Data Analysis, 64:237–252.
Laird, N. M. and Ware, J. H. (1982). Random-effects models for longitudinal data. Biometrics, 38:963–974.
Lin, T.-I. and Wang, W.-L. (2026). Grouped multi-trajectory modeling using finite mixtures of multivariate contaminated normal linear mixed model. Statistical Methods in Medical Research, 35(2):394–412.
Louis, T. A. (1982). Finding the observed information matrix when using the em algorithm. Journal of the Royal Statistical Society: Series B (Methodological), 44(2):226–233.
McLachlan, G. J. and Peel, D. (2000). Finite Mixture Models. Wiley, New York.
Meilijson, I. (1989). A fast improvement to the em algorithm on its own terms. Journal of the Royal Statistical Society: Series B (Methodological), 51(1):127–138.
Meng, X.-L. and Van Dyk, D. (1997). The em algorithm—an old folk-song sung to a fast new tune. Journal of the Royal Statistical Society: Series B (Methodological), 59(3):511–567.
Mirfarah, E., Naderi, M., Lin, T.-I., and Wang, W.-L. (2025). Robust bayesian inference for the censored mixture of experts model using heavy-tailed distributions. Advances in Data Analysis and Classification, 19(4):921–949.
Munoz, A., Carey, V., Schouten, J. P., Segal, M., and Rosner, B. (1992). A parametric family of correlation structures for the analysis of longitudinal data. Biometrics, 48(3):733–742.
Pinheiro, J. C., Liu, C., and Wu, Y. N. (2001). Efficient algorithms for robust estimation in linear mixed-effects models using the multivariate t distribution. Journal of Computational and Graphical Statistics, 10(2):249–276.
Punzo, A. and McNicholas, P. D. (2016). Parsimonious mixtures of multivariate contaminated normal distributions. Biometrical Journal, 58(6):1506–1537.
Schoenberg, I. J. (1938). Metric spaces and positive definite functions. Transactions of the American Mathematical Society, 44(3):522–536.
Schwarz, G. (1978). Estimating the dimension of a model. The Annals of Statistics, 6(2):461–464.
Song, P. X.-K., Zhang, P., and Qu, A. (2007). Maximum likelihood inference in robust linear mixed-effects models using multivariate t distributions. Statistica Sinica, 17:929–943.
Spiegelhalter, D. J., Best, N. G., Carlin, B. P., and van der Linde, A. (2002). Bayesian measures of model complexity and fit. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 64(4):583–639.
Tukey, J. W. (1960). A survey of sampling from contaminated distributions. In Olkin, I., Ghurye, S. G., Hoeffding, W., Madow, W. G., and Mann, H. B., editors, Contributions to Probability and Statistics: Essays in Honor of Harold Hotelling, volume 2 of Stanford studies in mathematics and statistics, pages 448–485. Stanford University Press, Stanford, CA.
Verbeke, G. and Lesaffre, E. (1996). A linear mixed-effects model with heterogeneity in the random-effects population. Journal of the American Statistical Association, 91(433):217–221.
Von Hoff, D. D., Ervin, T., Arena, F. P., Chiorean, E. G., Infante, J., Moore, M., Seay, T., Tjulandin, S. A., Ma, W. W., Saleh, M. N., et al. (2013). Increased survival in pancreatic cancer with nab-paclitaxel plus gemcitabine. New England Journal of Medicine, 369(18):1691–1703.
Wakefield, J. C., Smith, A. F., Racine-Poon, A., and Gelfand, A. E. (1994). Bayesian analysis of linear and non-linear population models by using the gibbs sampler. Journal of the Royal Statistical Society: Series C (Applied Statistics), 43(1):201–221.
Wang, W.-L., Yang, Y.-C., and Lin, T.-I. (2024). Extending finite mixtures of nonlinear mixed-effects models with covariate-dependent mixing weights. Advances in Data Analysis and Classification, 18(2):271–307.
Yang, Y.-C., Lin, T.-I., Castro, L. M., and Wang, W.-L. (2020). Extending finite mixtures of t linear mixed-effects models with concomitant covariates. Computational Statistics & Data Analysis, 148:106961.
Yeh, C., Zhou, M., Sigel, K., Jameson, G., White, R., Safyan, R., Saenger, Y., Hecht, E., Chabot, J., Schreibman, S., et al. (2023). Tumor growth rate informs treatment efficacy in metastatic pancreatic adenocarcinoma: application of a growth and regression model to pivotal trial and real-world data. The Oncologist, 28(2):139–148.