簡易檢索 / 詳目顯示

研究生: 潘皓心
Pang, Hao-Xin
論文名稱: 糖尿病病患服用 Statin 藥物代謝指標之縱向分析
Longitudinal Analysis of Metabolic Parameters of Statin Therapy in Diabetic Patients
指導教授: 馬瀰嘉
Ma, Mi-Chia
學位類別: 碩士
Master
系所名稱: 管理學院 - 統計學系
Department of Statistics
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 148
中文關鍵詞: 第二型糖尿病史他汀電子健康紀錄縱向資料分析深度學習
外文關鍵詞: type 2 diabetes, statin, electronic health records, longitudinal data analysis, deep learning
相關次數: 點閱:5下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 史他汀為成年第二型糖尿病病患心血管風險管理中常用之降脂藥物,但固定史他汀處方狀態與血糖及血脂指標長期變化之關聯,仍須透過真實世界縱向資料審慎評估。本研究採回溯性世代研究設計,使用台灣南部某醫療院所 2020 年 1 月至 2025 年 1 月之去識別化電子健康紀錄。最終納入 1,790 位成年第二型糖尿病病患及 17,252 筆縱向觀測,其中控制組 813 位,史他汀組 977 位。主要反應變數為 HbA1c、Glu-AC、LDL 與 TG。
    統計推論部分以 GEE 評估族群平均關聯,並以 LME-RI+RS 描述病患間基準水準及時間變化速率之異質性。缺失資料以 10 組多重插補處理,固定效果估計依 Rubin 合併法則進行合併。預測分析比較 GEE、LME-RI+RS、LSTM-NARX、TCN 與簡化式 ODE-LSTM,以測試集中每位可評估病患之最後一次實際觀測值為預測目標,並以 RMSE、MAE、MAPE 與 Bias 評估模型表現。此外,本研究建立 S0 至 S4 五種模擬情境,比較資料變異結構改變時之模型預測表現。
    結果顯示,GEE 中四項代謝指標之史他汀處方狀態與追蹤時間交互作用均未達統計顯著。LME-RI+RS 中,Glu-AC、LDL 與 TG 之交互作用未達統計顯著,僅 HbA1c 呈現幅度有限之正向交互作用;由於該結果未於 GEE 中獲得一致支持,應視為模型依賴性之關聯性發現。四項代謝指標之 LME-RI+RS 平均 AIC 均低於 LME-RI,顯示在本研究候選 LME 結構中,納入病患間時間斜率異質性後具有較佳之相對配適。
    在實證預測分析中,三種深度學習模型之主要預測誤差多低於 GEE 與 LME-RI+RS。就整體測試集而言,HbA1c 以 LSTM-NARX、Glu-AC 與 TG 以 TCN、LDL 以簡化式 ODE-LSTM 之 RMSE、MAE 與 MAPE 最低。模擬研究亦顯示,深度學習模型於多數情境下具有較低預測誤差,但未有單一模型於所有情境與代謝指標中均呈現一致優勢。
    本研究結果應解釋為固定史他汀處方狀態與代謝指標縱向變化之統計關聯,不代表史他汀治療之因果效果。GEE 主要用於評估族群平均關聯與時間趨勢;LME-RI+RS 除估計固定效果外,另透過病患層級隨機效果描述病患間基準水準與時間變化速率之異質性;深度學習模型則主要用於比較條件式預測表現。五模型雖使用相同測試病患與預測目標,但其可使用之歷史資訊並不完全相同,因此預測誤差結果應解釋為各模型完整預測流程下之表現,而非模型架構本身之絕對優劣。兩類方法於不等間隔臨床縱向資料分析中具有互補功能。

    Statins are commonly prescribed to adults with type 2 diabetes for lipid management and cardiovascular risk reduction. This retrospective cohort study evaluated the associations between fixed statin prescription status and longitudinal changes in glycated hemoglobin (HbA1c), preprandial blood glucose (Glu-AC), low-density lipoprotein cholesterol (LDL), and triglycerides (TG) using de-identified electronic health records from a medical institution in southern Taiwan between January 2020 and January 2025.
    The primary analysis included 1,790 patients and 17,252 longitudinal observations, including 813 patients in the control group and 977 in the statin group. Generalized estimating equations (GEE) and linear mixed-effects models with random intercepts and random slopes (LME-RI+RS) were used for longitudinal inference. Prediction performance was compared among GEE, LME-RI+RS, LSTM-NARX, temporal convolutional network (TCN), and simplified ODE-LSTM.
    In GEE, no group-by-time interaction was statistically significant for any of the four outcomes. In LME-RI+RS, only the HbA1c interaction was positive and statistically significant, but its magnitude was small, and the corresponding GEE interaction was not significant. This finding was therefore considered model-dependent.
    Deep learning models generally showed lower RMSE, MAE, and MAPE than GEE and LME-RI+RS, but no single model was consistently superior across all outcomes and simulation scenarios.
    These findings represent observational associations rather than causal effects. GEE and LME-RI+RS were used for longitudinal inference, whereas the deep learning models were used to compare conditional prediction performance. These approaches have complementary roles in analyzing irregularly observed clinical longitudinal data.

    摘要 I Extended Abstract II 誌謝 VI 目錄 VII 表目錄 X 圖目錄 XI 第一章 緒論 1 1.1 研究背景 1 1.2 研究動機 3 1.3 研究目的與架構 3 第二章 文獻回顧 6 2.1 縱向資料與不等間隔追蹤資料 6 2.2 廣義估計方程式 7 2.3 線性混合效應模型 8 2.4 深度學習時間序列模型 10 2.4.1 長短期記憶自回歸外生變數模型(LSTM-NARX) 11 2.4.2 時間卷積網路(TCN) 11 2.4.3 簡化式常微分方程長短期記憶模型(簡化式 ODE-LSTM) 12 2.5 主要文獻整合 13 2.6 文獻小結與研究缺口 14 第三章 研究資料與方法 15 3.1 研究設計與研究資料 15 3.1.1 研究對象 16 3.1.2 控制組與史他汀組定義 18 3.2 資料串接與前處理 18 3.2.1 基本資料處理 19 3.2.2 檢驗資料整理與事件層級整併 19 3.2.3 非數值檢驗結果與資料品質檢查 20 3.2.4 追蹤時間建立 20 3.3 研究變數定義 21 3.4 缺失資料處理與多重插補 24 3.4.1 統計推論用多重插補 24 3.4.2 預測分析之缺失資料處理 25 3.4.3 敘述性統計分析 26 3.5 廣義估計方程式 27 3.6 線性混合效應模型 29 3.6.1 隨機截距模型 30 3.6.2 隨機截距與隨機斜率模型 30 3.6.3 模型估計、結構比較與診斷 31 3.6.4 LME 於預測比較之應用 31 3.7 深度學習時間序列模型 32 3.7.1 資料切分與共同輸入設定 32 3.7.2 LSTM-NARX 模型 33 3.7.3 TCN 模型 33 3.7.4 簡化式 ODE-LSTM 模型 34 3.7.5 模型訓練設定 35 3.8 五模型預測比較與評估指標 35 第四章 實證資料分析結果 38 4.1 研究樣本與補值前敘述性統計 38 4.1.1 研究樣本基本特徵 38 4.1.2 主要代謝指標於各追蹤時間區間之敘述性統計與缺失情形 40 4.1.3 主要代謝指標於追蹤期間之分組平均趨勢 47 4.2 縱向統計模型分析結果 48 4.2.1 GEE 模型分析結果 49 4.2.2 LME 模型分析結果 52 4.2.3 GEE 與 LME 結果比較 58 4.3 深度學習模型之預測結果 60 4.3.1 深度學習模型訓練過程 60 4.3.2 深度學習模型之測試集預測表現 62 4.3.3 測試集實際值與預測值之區間平均比較 64 4.4 五模型預測表現整合比較 65 4.4.1 五模型於整體及分組測試集之預測表現 66 4.4.2 各代謝指標之五模型比較 70 4.5 本章小結 71 第五章 模擬研究 73 5.1 模擬研究目的 73 5.2 模擬資料生成機制 74 5.2.1 固定資料骨架之建立 74 5.2.2 反應變數生成模型 75 5.2.3 固定效果參考軌跡 76 5.3 模擬情境設定 78 5.4 模擬流程與模型配適 79 5.4.1 模擬次數與資料切分 79 5.4.2 縱向統計模型配適 80 5.4.3 深度學習模型訓練 80 5.4.4 預測結果輸出與彙整 80 5.5 模擬評估指標 81 5.6 模擬結果 82 5.6.1 基準情境 S0 之模擬結果 83 5.6.2 個體異質性增加情境 S1 之模擬結果 87 5.6.3 個體異質性降低情境 S2 之模擬結果 91 5.6.4 殘差變異增加情境 S3 之模擬結果 95 5.6.5 隨機效果相關方向改變情境 S4 之模擬結果 99 5.6.6 各情境下五模型之整體比較 103 5.7 模擬研究小結 104 第六章 討論與結論 106 6.1 主要研究發現 106 6.2 與既有文獻之比較 109 6.3 方法與臨床實務意涵 111 6.4 研究限制 112 6.5 未來研究方向 114 6.6 結論 115 參考文獻 117 附錄 119 附錄 A 補充敘述性統計結果 119 附錄 B 多重插補診斷結果 122 附錄 C 傳統縱向統計模型補充結果 125 附錄 D 深度學習模型補充設定 128 附錄 E 模擬資料生成、模型設定與殘差診斷 131 E.1 模擬資料骨架與資料生成模型 131 E.2 模擬研究之深度學習模型設定 132 E.3 基準情境 S0 之殘差診斷 132 E.4 模擬平均值之蒙地卡羅標準誤(MCSE) 134

    Akaike, H. (1973). Information theory and an extension of the maximum likelihood principle. In B. N. Petrov & F. Csáki (Eds.), Proceedings of the Second International Symposium on Information Theory (pp. 267–281). Akadémiai Kiadó.
    Alvarez-Jimenez, L., Morales-Palomo, F., Moreno-Cabanas, A., Ortega, J. F., & Mora-Rodriguez, R. (2023). Effects of statin therapy on glycemic control and insulin resistance: A systematic review and meta-analysis. European Journal of Pharmacology, 947, 175672. https://doi.org/10.1016/j.ejphar.2023.175672
    American Diabetes Association Professional Practice Committee. (2025). 10. Cardiovascular disease and risk management: Standards of care in diabetes—2025. Diabetes Care, 48(Supplement_1), S207–S238. https://doi.org/10.2337/dc25-S010
    Bai, S., Kolter, J. Z., & Koltun, V. (2018). An empirical evaluation of generic convolutional and recurrent networks for sequence modeling [Preprint]. arXiv. https://doi.org/10.48550/arXiv.1803.01271
    Brennan, M. B., Huang, E. S., Lobo, J. M., Kang, H., Guihan, M., Basu, A., & Sohn, M. W. (2018). Longitudinal trends and predictors of statin use among patients with diabetes. Journal of Diabetes and Its Complications, 32(1), 27–33. https://doi.org/10.1016/j.jdiacomp.2017.09.014
    Che, Z., Purushotham, S., Cho, K., Sontag, D., & Liu, Y. (2018). Recurrent neural networks for multivariate time series with missing values. Scientific Reports, 8, Article 6085. https://doi.org/10.1038/s41598-018-24271-9
    Chen, R. T. Q., Rubanova, Y., Bettencourt, J., & Duvenaud, D. K. (2018). Neural ordinary differential equations. Advances in Neural Information Processing Systems, 31. https://proceedings.neurips.cc/paper_files/paper/2018/file/69386f6bb1dfed68692a24c8686939b9-Paper.pdf
    Cholesterol Treatment Trialists’ (CTT) Collaborators. (2008). Efficacy of cholesterol-lowering therapy in 18,686 people with diabetes in 14 randomised trials of statins: A meta-analysis. The Lancet, 371(9607), 117–125. https://doi.org/10.1016/S0140-6736(08)60104-X
    DiCiccio, T. J., & Efron, B. (1996). Bootstrap confidence intervals. Statistical Science, 11(3), 189–228. https://doi.org/10.1214/ss/1032280214
    Duan, N. (1983). Smearing estimate: A nonparametric retransformation method. Journal of the American Statistical Association, 78(383), 605–610. https://doi.org/10.2307/2288126
    Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735–1780. https://doi.org/10.1162/neco.1997.9.8.1735
    Hyndman, R. J., & Koehler, A. B. (2006). Another look at measures of forecast accuracy. International Journal of Forecasting, 22(4), 679–688. https://doi.org/10.1016/j.ijforecast.2006.03.001
    Laakso, M., & Fernandes Silva, L. (2023). Statins and risk of type 2 diabetes: Mechanism and clinical implications. Frontiers in Endocrinology, 14, Article 1239335. https://doi.org/10.3389/fendo.2023.1239335
    Laird, N. M., & Ware, J. H. (1982). Random-effects models for longitudinal data. Biometrics, 38(4), 963–974. https://doi.org/10.2307/2529876
    Lechner, M., & Hasani, R. (2020). Learning long-term dependencies in irregularly-sampled time series [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2006.04418
    Lembo, M., Trimarco, V., Pacella, D., Izzo, R., Jankauskas, S. S., Piccinocchi, R., Gallo, P., Bardi, L., Piccinocchi, G., Morisco, C., Cristiano, S., Esposito, G., Giugliano, G., Manzi, M. V., Santulli, G., & Trimarco, B. (2025). A six-year longitudinal study identifies a statin-independent association between low LDL-cholesterol and risk of type 2 diabetes. Cardiovascular Diabetology, 24(1), 429. https://doi.org/10.1186/s12933-025-02964-6
    Liang, K.-Y., & Zeger, S. L. (1986). Longitudinal data analysis using generalized linear models. Biometrika, 73(1), 13–22. https://doi.org/10.2307/2336267
    Lin, T., Horne, B. G., Tino, P., & Giles, C. L. (1996). Learning long-term dependencies in NARX recurrent neural networks. IEEE Transactions on Neural Networks, 7(6), 1329–1338. https://doi.org/10.1109/72.548162
    Rubanova, Y., Chen, R. T. Q., & Duvenaud, D. K. (2019). Latent ordinary differential equations for irregularly-sampled time series. Advances in Neural Information Processing Systems, 32. https://proceedings.neurips.cc/paper/2019/file/42a6845a557bef704ad8ac9cb4461d43-Paper.pdf
    Sattar, N., Preiss, D., Murray, H. M., Welsh, P., Buckley, B. M., de Craen, A. J. M., Seshasai, S. R. K., McMurray, J. J. V., Freeman, D. J., Jukema, J. W., Macfarlane, P. W., Packard, C. J., Stott, D. J., Westendorp, R. G. J., Shepherd, J., Davis, B. R., Pressel, S. L., Marchioli, R., Marfisi, R. M., . . . Ford, I. (2010). Statins and risk of incident diabetes: A collaborative meta-analysis of randomised statin trials. The Lancet, 375(9716), 735–742. https://doi.org/10.1016/S0140-6736(09)61965-6
    Schafer, J. L., & Olsen, M. K. (1998). Multiple imputation for multivariate missing-data problems: A data analyst’s perspective. Multivariate Behavioral Research, 33(4), 545–571. https://doi.org/10.1207/s15327906mbr3304_5
    郝立智、陳鏡方、范麒惠、黃騰慶、張家僖、馬瀰嘉、梁維哲、曾振生、陳昆祥、黃偉輔、楊純宜(2024)。2024 年美國糖尿病學會針對糖尿病合併血脂異常之標準治療建議。臨床醫學月刊,94(5),747–762。https://doi.org/10.6666/ClinMed.202411_94(5).0129

    QR CODE