| 研究生: |
黃芯鈺 Huang, Xin-Yu |
|---|---|
| 論文名稱: |
相對機率密度比估計應用於高維資料下之離群值偵測──基於子空間搜尋方法 Relative Density Ratio Estimation for High-Dimensional Outlier Detection via Subspace Search |
| 指導教授: |
蘇佩芳
Su, Pei-Fang |
| 學位類別: |
碩士 Master |
| 系所名稱: |
管理學院 - 統計學系 Department of Statistics |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 115 |
| 中文關鍵詞: | 離群值偵測 、機率密度比估計 、α相對機率密度比 、子空間搜尋 、高維資料 |
| 外文關鍵詞: | outlier detection, probability density ratio estimation, α-relative density ratio, subspace search, high-dimensional data, dimensionality reduction |
| 相關次數: | 點閱:44 下載:4 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
在離群值偵測(outlier detection)中,常需比較兩組樣本之間的差異。機率密度比(probability density ratio)是用於描述兩個機率分布間之相對關係的方法之一,也可應用於離群值分布異常偵測問題。然而,現有機率密度比估計方法在高維資料下容易受到不相關變數影響而降低估計效能,且當兩分布差異較大時,機率密度比估計值可能發散,造成數值不穩定。為改善上述問題,本文提出結合子空間搜尋與 α 相對機率密度比(α-relative density ratio)之估計方法,並將估計結果應用於高維資料下之離群值偵測,以提升辨識能力與估計穩定性。
模擬研究結果顯示,本文方法在不同資料維度與離群樣本比例設定下皆具有良好的辨識能力與估計穩定性,且相較於現有機率密度比估計方法,本文所提方法能在高維資料下維持較佳的偵測效能。在實證分析中,本文將所提之方法應用於身體組成測量資料,結果顯示其能有效辨識不同族群之間的分布差異,驗證本文方法在高維資料離群值偵測問題中的可行性與實務應用。
Probability density ratio estimation is an effective approach for comparing two probability distributions and has been applied to outlier detection. However, existing methods often suffer from performance degradation in high-dimensional data due to irrelevant variables, and the estimated density ratio may become unstable when the two distributions differ substantially. To address these issues, this study proposes a high-dimensional outlier detection method that combines subspace search with α-relative probability density ratio estimation.
Simulation results demonstrate that the proposed method achieves stable estimation and favorable detection performance under various data dimensionalities and outlier proportions. The results on body composition measurement data further show that the proposed method effectively identifies distributional differences between different populations, demonstrating its practical applicability to high-dimensional outlier detection.
[1] Boyd, S., & Vandenberghe, L. Convex optimization. Cambridge university press (2004).
[2] Huang, J., Gretton, A., Borgwardt, K., Schölkopf, B., & Smola, A. Correcting sample selection bias by unlabeled data. Advances in neural information processing systems, 19 (2006).
[3] Hido, S., Tsuboi, Y., Kashima, H., Sugiyama, M., & Kanamori, T. Statistical outlier detection using direct density ratio estimation. Knowledge and information systems, 26(2), 309-336 (2011).
[4] Kullback, S., & Leibler, R. A. On information and sufficiency. The annals of mathematical statistics, 22(1), 79-86 (1951).
[5] Kanamori, T., Hido, S., & Sugiyama, M. A least-squares approach to direct importance estimation. Journal of Machine Learning Research, 10, 1391–1445 (2009).
[6] Pearson, K. X. On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling. The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science, 50(302), 157-175 (1900).
[7] Patriksson, M. Nonlinear programming and variational inequality problems: a unified approach (Vol. 23). Springer Science & Business Media (2013).
[8] Sugiyama, M., Suzuki, T., Nakajima, S., Kashima, H., von Bünau, P., & Kawanabe, M. Direct importance estimation for covariate shift adaptation. Annals of the Institute of Statistical Mathematics, 60, 699–746 (2008).
[9] Sugiyama, M., Yamada, M., Von Buenau, P., Suzuki, T., Kanamori, T., & Kawanabe, M. Direct density-ratio estimation with dimensionality reduction via least-squares hetero-distributional subspace search. Neural Networks, 24(2), 183-198 (2011).
[10] Yamada, M., Suzuki, T., Kanamori, T., Hachiya, H., & Sugiyama, M. Relative density-ratio estimation for robust distribution comparison. Neural computation, 25(5), 1324-1370 (2013).