| 研究生: |
鄭富銘 Cheng, Fu-Ming |
|---|---|
| 論文名稱: |
以正則初始化處理交替最小法與次梯度下降法求解最大仿射回歸之離群值問題 Solving max-affine regression by alternating minimization and subgradient descent with regularized initialization to handle outliers |
| 指導教授: |
許瑞麟
Sheu, Ruey-Lin |
| 學位類別: |
碩士 Master |
| 系所名稱: |
理學院 - 數學系應用數學碩博士班 Department of Mathematics |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 74 |
| 中文關鍵詞: | 最大仿射回歸 、交替最小化 、正則化初始化 、應變數訊號截斷 、加權動差初始化 |
| 外文關鍵詞: | max affine regression, alternative minimization, regularized initialization, response clipping, weighted moment initialization |
| 相關次數: | 點閱:77 下載:2 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
最大仿射回歸 (max-affine regression) 是一種非凸迴歸模型,其最佳化表現高度依賴初始化品質。近期的理論研究通常採用基於矩陣的初始化方法,在套用交替最小化(alternating minimization)等局部最佳化方法之前,先估計參數子空間。然而,這類基於矩的譜初始化方法之穩健性質目前仍缺乏充分研究。本論文探討最大仿射回歸初始化方法對幾何結構擾動的敏感性。我們首先研究應變數訊號截斷對矩估計量的影響。數值實驗顯示,對反應變數進行截斷未必能改善最佳化結果,甚至可能扭曲矩矩陣的幾何結構。特別地,我們說明最大仿射回歸中的離群值不應僅由較大的反應值來判定,而應考慮其在矩估計層級上的影響程度,其貢獻可表示為 |y_i||x_i|^2。基於此觀察,我們提出一種矩層級的加權估計量,直接控制個別樣本對矩矩陣的影響。實驗結果顯示,在存在高槓桿離群值(high leverage outliers)的情況下,所提出的加權策略能改善真實參數子空間的估計,而標準的應變數訊號截斷則未能產生明顯效果。研究結果顯示,最大仿射回歸初始化的穩健性,關鍵在於保留矩矩陣中的高槓桿幾何結構,而非僅針對較大的反應值進行處理。
Max-affine regression is a nonconvex regression model whose optimization performance strongly depends on the quality of initialization. Recent theoretical works employ initialization based on moment matrices to recover the parameter subspace before applying local optimization such as alternating minimization. However, the robustness properties of these moment-based spectral methods remain largely unexplored. In this thesis, we investigate the geometric sensitivity of initialization in max affine regression. We first study the effect of response clipping on moment estimators. Numerical experiments show that clipping of the response variable may fail to improve the optimization and can even distort the moment geometry. In particular, we demonstrate that outliers in max affine regression are not characterized by large response values, but rather by large moment-level influence contributions of the form |y_i||x_i|^2. Motivated by this observation, we propose a moment-level weighted estimator that directly controls the influence of individual samples on the moment matrix. Experimental results show that the proposed weighting strategy improves the estimate of true parameter subspace under high-leverage outliers, while standard response clipping remains ineffective. Our findings suggest that robustness in max affine regression initialization is tied to preserving the high-leverage structure of the moment matrix rather than large responses.
[1] Hamed Ahmadi, José R. Martí, and Ali Moshref. Piecewise linear approximation of generators cost functions using max-affine functions. In 2013 IEEE Power & Energy Society General Meeting, pages 1–5. IEEE, 2013.
[2] Theodore W. Anderson. An Introduction to Multivariate Statistical Analysis. John Wiley & Sons, 3 edition, 2003.
[3] Gábor Balázs, András György, and Csaba Szepesvári. Near-optimal max-affine estimators for convex regression. In Proceedings of the Eighteenth International Conference on Artificial Intelligence and Statistics, volume 38 of Proceedings of Machine Learning Research, pages 56–64. PMLR, 2015.
[4] Stephen Boyd and Lieven Vandenberghe. Convex Optimization. Cambridge University Press, 2004.
[5] Leo Breiman. Hinging hyperplanes for regression, classification, and function approximation. IEEE Transactions on Information Theory, 39(3):999–1013, 1993.
[6] Yuxin Chen and Emmanuel J. Candès. Solving random quadratic systems of equations is nearly as easy as solving linear systems. Communications on Pure and Applied Mathematics, 70(5):822–883, 2017.
[7] Koby Crammer and Yoram Singer. On the algorithmic implementation of multiclass kernel-based vector machines. Journal of Machine Learning Research, 2:265–292, 2001.
[8] Amit Daniely, Sivan Sabato, and Shai Shalev-Shwartz. Multiclass learning approaches: A theoretical comparison with implications. In Advances in Neural Information Processing Systems, volume 25, 2012.
[9] Alfred Fischer. Convex-in-the-large representations of piecewise linear functions. Optimization, 50(4):313–321, 2002.
[10] Avishek Ghosh, Ashwin Pananjady, Adityanand Guntuboyina, and Kannan Ramchandran. Max-affine regression: Parameter estimation for gaussian designs. IEEE Transactions on Information Theory, 68(3):1851–1885, 2022.
[11] Gene H. Golub and Charles F. Van Loan. Matrix Computations. Johns Hopkins University Press, 4 edition, 2013.
[12] Ian J. Goodfellow, David Warde-Farley, Mehdi Mirza, Aaron Courville, and Yoshua Bengio. Maxout networks. In Proceedings of the 30th International Conference on Machine Learning, volume 28 of Proceedings of Machine Learning Research, pages 1319–1327. PMLR, 2013.
[13] Lauren A. Hannah and David B. Dunson. Multivariate convex regression with adaptive partitioning. Journal of Machine Learning Research, 14(1):3261–3294, 2013.
[14] Ronald R. Hocking. Methods and Applications of Linear Models: Regression and the Analysis of Variance. Wiley, 3 edition, 2013.
[15] Arthur E. Hoerl and Robert W. Kennard. Ridge regression: Biased estimation for nonorthogonal problems. Technometrics, 12(1):55–67, 1970.
[16] Robert V. Hogg, Joseph W. McKean, and Allen T. Craig. Introduction to Mathematical Statistics. Pearson, 8 edition, 2019.
[17] Roger A. Horn and Charles R. Johnson. Matrix Analysis. Cambridge University Press, 2 edition, 2013.
[18] Harold Hotelling. Analysis of a complex of statistical variables into principal components. Journal of Educational Psychology, 24(6):417–441, 1933.
[19] Ian T. Jolliffe. Principal Component Analysis. Springer, 2 edition, 2002.
[20] Seonho Kim and Kiryung Lee. Max-affine regression via first-order methods. SIAM Journal on Mathematics of Data Science, 6(2):534–552, 2024.
[21] Alessandro Magnani and Stephen P. Boyd. Convex piecewise-linear fitting. Optimization and Engineering, 10(1):1–17, 2009.
[22] Robb J. Muirhead. Aspects of Multivariate Statistical Theory. John Wiley & Sons, 1982.
[23] Vinod Nair and Geoffrey E. Hinton. Rectified linear units improve restricted boltzmann machines. In Proceedings of the 27th International Conference on Machine Learning (ICML), pages 807–814, 2010.
[24] Karl Pearson. On lines and planes of closest fit to systems of points in space. Philosophical Magazine, 2(11):559–572, 1901.
[25] Roger Penrose. A generalized inverse for matrices. Proceedings of the Cambridge Philosophical Society, 51(3):406–413, 1955.
[26] Emilio Seijo and Bodhisattva Sen. Nonparametric least squares estimation of a multivariate convex regression function. The Annals of Statistics, 39(3):1633–1657, 2011.
[27] Charles M. Stein. Estimation of the mean of a multivariate normal distribution. The Annals of Statistics, 9(6):1135–1151, 1981.
[28] Lloyd N. Trefethen and David Bau. Numerical Linear Algebra. SIAM, Philadelphia, PA, 1997.