| 研究生: |
許皓程 Hsu, Hao-Cheng |
|---|---|
| 論文名稱: |
基於深度學習之極值時間序列預測模型 Deep-Learning Approach toward Forecasting of Time Series with Extreme Values |
| 指導教授: |
李昇暾
Li, Sheng-Tun |
| 學位類別: |
碩士 Master |
| 系所名稱: |
管理學院 - 工業與資訊管理學系 Department of Industrial and Information Management |
| 論文出版年: | 2021 |
| 畢業學年度: | 109 |
| 語文別: | 中文 |
| 論文頁數: | 49 |
| 中文關鍵詞: | 極端事件 、深度學習 、注意力機制 、損失函數 |
| 外文關鍵詞: | extreme events, deep learning, attention mechanisms, loss functions |
| 相關次數: | 點閱:246 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
隨著網路世代的快速發展,人工智慧、大數據、機器學習、深度學習等各式各樣的新技術在近年蓬勃的發展,儘管基於上述提到的方法,預測效果已有大幅的改進,但當資料中存在極端事件的情況下,預測的效果普遍不佳。因此,極值時間序列預測分析是一個值得深入研究的主題。極端值在真實的資料中雖然佔比非常小,卻是在資料中實際發生的情況,不能直接忽略它的存在,並且對於時間序料資料的預測效果有顯著的影響。過往的研究方法在預測一般無太大偏差值的資料時成效已經很不錯,但當遇到有極端值存在的情況下,它無法有效的抓到那偏離一般情況的值,例如極端氣候的預測、金融危機的預測、網路攻擊預測等。由於過去的方法在極端值存在的時間序列資料情況下預測效果不好,所以本研究致力於建立一個可以改善傳統模型無法有效預測極端值存在的時間序列資料,期望可以為極端資料預測做出貢獻。本研究將應用極端理論,定義適用於極端情況的損失函數(Extreme Value Loss, EVL),並利用滑動視窗(sliding window)處理時間序列,結合長短期記憶(Long Short-Term Memory, LSTM)和注意力機制(attention mechanism),並改良EVL函數對時間序列資料進行預測,期望可以增加預測準確性。
根據本研究的實驗結果顯示,經過在多數的實驗資料集的測試下,獲得不錯的預測結果,並且同時在預測所使用的時間序列上大多能比先前文獻的方法有更低的平均絕對百分比誤差(Mean Absolute Percentage Error, MAPE),及更高的方向趨勢準確性(Trend Accuracy In Direction, TAD),對時間序列極端事件的判斷也較為準確。
With the rapid development of the Internet generation, various new technologies such as artificial intelligence, big data, machine learning, and deep learning have flourished in recent years. Although the prediction effect has improved substantially based on the above mentioned methods, when extreme events exist in the data, the prediction effect is generally poor. Therefore, extreme value time series prediction analysis is a topic worthy of further study. Although the proportion of extreme values in the real data is very small, they are actual occurrences in the data, and their existence cannot be ignored directly, and have a significant impact on the prediction effect of time series data. The past research methods have been very effective in predicting data without significant deviations in general, but they are not effective in capturing the deviations when extreme values exist, such as extreme climate prediction, financial crisis prediction, and cyber attack prediction. Since the previous methods are not effective in predicting time series data with extreme values, this study aims to build a model that can improve the time series data where traditional models cannot effectively predict the existence of extreme values, in the hope that it can contribute to the prediction of extreme data. In this study, we apply the extreme value theory to define the Extreme Value Loss (EVL) function for the extreme case, and use the sliding window to process the time series, combine the Long Short-Term Memory (LSTM) and the attention mechanism, and improve the EVL function. The sliding window is used to process the time series by combining LSTM and attention mechanism, and the EVL function is improved to predict the time series data, which is expected to increase the accuracy of prediction.
According to the experimental results of this study, we obtained good prediction results on most of the experimental data sets, and most of the time series used in the prediction have lower Mean Absolute Percentage Error (MAPE) and higher Trend Accuracy In Direction (TAD) than the previous methods in the literature, and more accurate judgment of extreme events in time series.
Albeverio, S., Jentsch, V., & Kantz, H. (2006). Extreme events in nature and society: Springer Science & Business Media.
Bahdanau, D., Cho, K., & Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473.
Becker, E. J., Van Den Dool, H., & Peña, M. (2013). Short-term climate extremes: Prediction skill and predictability. Journal of climate, 26(2), 512-531.
Chou, J.-S., & Ngo, N.-T. (2016). Time series analytics using sliding window metaheuristic optimization-based machine learning system for identifying building energy consumption patterns. Applied energy, 177, 751-770.
Chung, J., Gulcehre, C., Cho, K., & Bengio, Y. (2014). Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555.
Clauset, A., Shalizi, C. R., & Newman, M. E. (2009). Power-law distributions in empirical data. SIAM review, 51(4), 661-703.
Clouse, J. A., & Dow, J. (2002). A computational model of banks' optimal reserve management policy. Journal of Economic Dynamics and Control, 26, 1787-1814.
Dadteev, K., Shchukin, B., & Nemeshaev, S. (2020). Using artificial intelligence technologies to predict cash flow. Procedia Computer Science, 169, 264-268.
De Haan, L., & Ferreira, A. (2007). Extreme value theory: an introduction: Springer Science & Business Media.
Ding, D., Zhang, M., Pan, X., Yang, M., & He, X. (2019). Modeling extreme events in time series prediction. Paper presented at the Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining.
Dong, Y., Liu, Y., & Lian, S. (2016). Automatic age estimation based on deep learning algorithm. Neurocomputing, 187, 4-10.
Dovoedo, Y., & Chakraborti, S. (2015). Boxplot-based outlier detection for the location-scale family. Communications in statistics-simulation and computation, 44(6), 1492-1513.
Doya, K. (1993). Bifurcations of recurrent neural networks in gradient descent learning. IEEE Transactions on Neural Networks, 1(75), 218.
Ekinci, Y., Lu, J.-C., & Duman, E. (2015). Optimization of ATM cash replenishment with group-demand forecasts. Expert Systems with Applications, 42(7), 3480-3490.
Fei-Fei, L., Fergus, R., & Perona, P. (2006). One-shot learning of object categories. IEEE transactions on pattern analysis and machine intelligence, 28(4), 594-611.
Fisher, R. A., & Tippett, L. H. C. (1928). Limiting forms of the frequency distribution of the largest or smallest member of a sample. Paper presented at the Mathematical Proceedings of the Cambridge Philosophical Society.
Ghil, M., Yiou, P., Hallegatte, S., Malamud, B., Naveau, P., Soloviev, A., . . . Kossobokov, V. (2010). Extreme Events: Dynamics, Statistics and Prediction.
Gnedenko, B. (1943). Sur la distribution limite du terme maximum d'une serie aleatoire. Annals of mathematics, 423-453.
He, Y., Bai, R., & Dong, J. (2011). Study on the management and control model of cash flow in enterprises. Paper presented at the International Conference on Advances in Education and Management.
Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural computation, 9(8), 1735-1780.
Jamin, A., & Humeau-Heurtier, A. (2020). (Multiscale) Cross-Entropy Methods: A Review. Entropy, 22(1), 45.
Keeling, C. D., & Whorf, T. P. (2004). Atmospheric CO2 concentrations derived from flask air samples at sites in the SIO network. Trends: a compendium of data on Global Change.
Kotz, S., & Nadarajah, S. (2000). Extreme value distributions: theory and applications: World Scientific.
Kumar, S., & Deo, N. (2015). Analysing correlations after the financial crisis of 2008 and multifractality in global financial time series. Pramana, 84(2), 317-325.
Li, K., Liu, L., Zhai, J., Khoshgoftaar, T. M., & Li, T. (2016). The improved grey model based on particle swarm optimization algorithm for time series prediction. Engineering Applications of Artificial Intelligence, 55, 285-291.
Li, S.-T., Kuo, S.-C., Cheng, Y.-C., & Chen, C.-C. (2010). Deterministic vector long-term forecasting for fuzzy time series. Fuzzy Sets and Systems, 161(13), 1852-1870.
Liu, Y., Dong, S., Lu, M., & Wang, J. (2018). LSTM based reserve prediction for bank outlets. Tsinghua Science and Technology, 24(1), 77-85.
Längkvist, M., Karlsson, L., & Loutfi, A. (2014). A review of unsupervised feature learning and deep learning for time-series modeling. Pattern Recognition Letters, 42, 11-24.
Lucas, D., Yver Kwok, C., Cameron-Smith, P., Graven, H., Bergmann, D., Guilderson, T., . . . Keeling, R. (2015). Designing optimal greenhouse gas observing networks that consider performance and cost. Geoscientific Instrumentation, Methods and Data Systems, 4(1), 121-137.
Luong, M.-T., Pham, H., & Manning, C. D. (2015). Effective approaches to attention-based neural machine translation. arXiv preprint arXiv:1508.04025.
Mandić-Rajčević, S., & Colosio, C. (2019). Methods for the Identification of Outliers and Their Influence on Exposure Assessment in Agricultural Pesticide Applicators: A Proposed Approach and Validation Using Biological Monitoring. Toxics, 7(3), 37.
Mikolov, T., & Zweig, G. (2012). Context dependent recurrent neural network language model. Paper presented at the 2012 IEEE Spoken Language Technology Workshop (SLT).
Mohamad, N., Ahmad, N. B. H., & Sulaiman, S. (2018). DATA PRE-PROCESSING: A CASE STUDY IN PREDICTING STUDENt S RETENTION IN MOOC. Journal of Fundamental and Applied Sciences, 9, 598-613.
Moubariki, Z., Beljadid, L., Tirari, M. E. H., Kaicer, M., & Thami, R. O. H. (2019). Enhancing cash management using machine learning. Paper presented at the 2019 1st International Conference on Smart Systems and Data Science (ICSSD).
Nelson, B. K. (1998). Time series analysis using autoregressive integrated moving average (ARIMA) models. Academic emergency medicine, 5(7), 739-744.
Rolski, T., Schmidli, H., Schmidt, V., & Teugels, J. L. (2009). Stochastic processes for insurance and finance (Vol. 505): John Wiley & Sons.
Rusiecki, A. (2019). Trimmed categorical cross-entropy for deep learning with label noise. Electronics Letters, 55(6), 319-320.
Sun, X., Hao, J., & Li, J. (2020). Multi-objective optimization of crude oil-supply portfolio based on interval prediction data. Annals of Operations Research, 1-29.
Torres, J. L., Garcia, A., De Blas, M., & De Francisco, A. (2005). Forecast of hourly average wind speed with ARMA models in Navarre (Spain). Solar energy, 79(1), 65-77.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., . . . Polosukhin, I. (2017). Attention is all you need. Paper presented at the Advances in neural information processing systems.
von Bortkiewicz, L. (1921). Variationsbreite und mittlerer Fehler: Berliner Mathematische Gesellschaft.
Yang, C.-L., Yang, C.-Y., Chen, Z.-X., & Lo, N.-W. (2019). Multivariate time series data transformation for convolutional neural network. Paper presented at the 2019 IEEE/SICE International Symposium on System Integration (SII).
Zhan, Z., Xu, M., & Xu, S. (2015). Predicting cyber attack rates with extreme values. IEEE Transactions on Information Forensics and Security, 10(8), 1666-1677.
Zhang, J.-S., & Xiao, X.-C. (2000). Predicting chaotic time series using recurrent neural network. Chinese Physics Letters, 17(2), 88.