簡易檢索 / 詳目顯示

研究生: 黃芯鈺
Huang, Xin-Yu
論文名稱: 相對機率密度比估計應用於高維資料下之離群值偵測──基於子空間搜尋方法
Relative Density Ratio Estimation for High-Dimensional Outlier Detection via Subspace Search
指導教授: 蘇佩芳
Su, Pei-Fang
學位類別: 碩士
Master
系所名稱: 管理學院 - 統計學系
Department of Statistics
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 115
中文關鍵詞: 離群值偵測機率密度比估計α相對機率密度比子空間搜尋高維資料
外文關鍵詞: outlier detection, probability density ratio estimation, α-relative density ratio, subspace search, high-dimensional data, dimensionality reduction
相關次數: 點閱:44下載:4
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 在離群值偵測(outlier detection)中,常需比較兩組樣本之間的差異。機率密度比(probability density ratio)是用於描述兩個機率分布間之相對關係的方法之一,也可應用於離群值分布異常偵測問題。然而,現有機率密度比估計方法在高維資料下容易受到不相關變數影響而降低估計效能,且當兩分布差異較大時,機率密度比估計值可能發散,造成數值不穩定。為改善上述問題,本文提出結合子空間搜尋與 α 相對機率密度比(α-relative density ratio)之估計方法,並將估計結果應用於高維資料下之離群值偵測,以提升辨識能力與估計穩定性。
    模擬研究結果顯示,本文方法在不同資料維度與離群樣本比例設定下皆具有良好的辨識能力與估計穩定性,且相較於現有機率密度比估計方法,本文所提方法能在高維資料下維持較佳的偵測效能。在實證分析中,本文將所提之方法應用於身體組成測量資料,結果顯示其能有效辨識不同族群之間的分布差異,驗證本文方法在高維資料離群值偵測問題中的可行性與實務應用。

    Probability density ratio estimation is an effective approach for comparing two probability distributions and has been applied to outlier detection. However, existing methods often suffer from performance degradation in high-dimensional data due to irrelevant variables, and the estimated density ratio may become unstable when the two distributions differ substantially. To address these issues, this study proposes a high-dimensional outlier detection method that combines subspace search with α-relative probability density ratio estimation.
    Simulation results demonstrate that the proposed method achieves stable estimation and favorable detection performance under various data dimensionalities and outlier proportions. The results on body composition measurement data further show that the proposed method effectively identifies distributional differences between different populations, demonstrating its practical applicability to high-dimensional outlier detection.

    摘要 i 誌謝 vi 目錄 vii 表目錄 x 圖目錄 xii 第一章 緒論 1 1.1 研究背景 1 1.2 研究動機資料介紹 1 1.3研究目的 6 第二章 文獻回顧 8 2.1符號定義 8 2.2 機率密度比估計所使用的散度 10 2.3 在原始空間下直接估計的機率密度比估計方法 11 2.3.1 Kullback–Leibler 重要性估計程序(KLIEP) 12 2.3.2 無約束最小平方法重要性擬合法(uLSIF) 12 2.3.3 相對無約束最小平方法重要性擬合(RuLSIF) 13 2.4最小平方異質分布子空間搜尋降維之機率密度比估計(LHSS) 14 第三章 研究方法 18 3.1基於 α-Pearson散度的子空間搜尋與機率密度比估計方法 18 3.1.1問題設定與符號定義 19 3.1.2將原始樣本投影至低維子空間 19 3.1.3 α 相對機率密度比估計與 α-Pearson散度計算 21 3.1.4 子空間搜尋與最終機率密度比估計 23 3.2演算法流程 25 3.3離群值偵測 28 3.4本文方法與LHSS方法之比較 30 第四章 模擬研究 31 4.1 模擬資料生成與實驗設定 31 4.2 模擬實驗結果 33 4.3與其他機率密度比估計方法的比較 36 第五章 實證分析 40 5.1 運動員資料分析 40 5.1.1 分析方法 41 5.1.2 分析結果 43 5.2 病患身體組成資料分析 51 第六章 結論 62 6.1研究貢獻 62 6.2研究限制與未來研究 63 參考文獻 65 附錄A:身體組成測量資料之分布圖 66 附錄B:Kullback–Leibler重要性估計程序(KLIEP) 70 附錄C:無約束最小平方法重要性擬合法(uLSIF) 73 附錄D:相對無約束最小平方法重要性擬合(RuLSIF) 76 附錄E:定理1證明 78 附錄F:定理2證明 80 附錄G:模擬實驗結果 82 附錄H:核均值匹配法(KMM) 92 附錄I:實證分析結果 94

    [1] Boyd, S., & Vandenberghe, L. Convex optimization. Cambridge university press (2004).
    [2] Huang, J., Gretton, A., Borgwardt, K., Schölkopf, B., & Smola, A. Correcting sample selection bias by unlabeled data. Advances in neural information processing systems, 19 (2006).
    [3] Hido, S., Tsuboi, Y., Kashima, H., Sugiyama, M., & Kanamori, T. Statistical outlier detection using direct density ratio estimation. Knowledge and information systems, 26(2), 309-336 (2011).
    [4] Kullback, S., & Leibler, R. A. On information and sufficiency. The annals of mathematical statistics, 22(1), 79-86 (1951).
    [5] Kanamori, T., Hido, S., & Sugiyama, M. A least-squares approach to direct importance estimation. Journal of Machine Learning Research, 10, 1391–1445 (2009).
    [6] Pearson, K. X. On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling. The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science, 50(302), 157-175 (1900).
    [7] Patriksson, M. Nonlinear programming and variational inequality problems: a unified approach (Vol. 23). Springer Science & Business Media (2013).
    [8] Sugiyama, M., Suzuki, T., Nakajima, S., Kashima, H., von Bünau, P., & Kawanabe, M. Direct importance estimation for covariate shift adaptation. Annals of the Institute of Statistical Mathematics, 60, 699–746 (2008).
    [9] Sugiyama, M., Yamada, M., Von Buenau, P., Suzuki, T., Kanamori, T., & Kawanabe, M. Direct density-ratio estimation with dimensionality reduction via least-squares hetero-distributional subspace search. Neural Networks, 24(2), 183-198 (2011).
    [10] Yamada, M., Suzuki, T., Kanamori, T., Hachiya, H., & Sugiyama, M. Relative density-ratio estimation for robust distribution comparison. Neural computation, 25(5), 1324-1370 (2013).

    下載圖示
    校外:立即公開
    QR CODE