| 研究生: |
許瑋庭 Hsu, Wei-Ting |
|---|---|
| 論文名稱: |
預測自駕車事故傷亡與否:基於小樣本不平衡分類數據 Classification of Autonomous Vehicle Crash Severity: Solving Imbalanced and Small Sample Size Problem |
| 指導教授: |
郭佩棻
Kuo, Pei-Fen |
| 學位類別: |
碩士 Master |
| 系所名稱: |
工學院 - 測量及空間資訊學系 Department of Geomatics |
| 論文出版年: | 2023 |
| 畢業學年度: | 111 |
| 語文別: | 英文 |
| 論文頁數: | 92 |
| 中文關鍵詞: | 自駕車 、小樣本 、不平衡樣本 、興趣點 、肇事嚴重程度 |
| 外文關鍵詞: | autonomous vehicles (AVs), small sample size, imbalance data, point of interest (POI), severity |
| 相關次數: | 點閱:215 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
自駕車(Autonomous Vehicles, AVs)因其便利之特性,近年來發展快速,使用者急速增加。自駕車因有導航與物件偵測等先進技術,普遍認為可降低人為失誤事故、提升道路交通安全、最小化運行成本、提升道路容積與改善交通壅塞。本研究致力於探討其影響自駕車乘客傷亡之關鍵因素。然而,現有自駕車仍不時有傷亡事故發生,一旦自駕車正式上路運行,自動駕駛的不確定因素可能造成災難性的影響。目前,自駕車仍為小規模測試階段,碰撞和傷亡事故資料較少,故其事故集仍為小樣本事件,且資料集具類別分布不均的特性。此外,僅有少數研究定義了影響碰撞嚴重程度的環境因素,例如碰撞地點周圍的設施和道路相關資訊。這些研究在選擇空間特徵時往往忽略了小樣本量和不平衡資料的問題。
本研究使用加州機動車輛管理局 (California Department of Motor Vehicles, CA DMV) 提供之三年(2019-2021)的自動駕駛汽車碰撞報告(AV Level 3)。 該數據集包含 266 份碰撞報告,其中包含 51 份傷亡案件和 215 份無傷亡案件。 除了自駕車碰撞報共中所提供之相關資訊外,碰撞地點周圍環境也是影響自駕車碰撞嚴重程度的重要因素。 因此,本研究從開放街道地圖 (Open Street Map, OSM) 和 DataSF 蒐集多項設施及道路限速資料,建立興趣點 (Point of Interests, POIs) 數據集,以在微觀尺度上精確評估碰撞與周圍環境之間相互作用的影響。
為了解決上述小樣本和數據集不平衡的問題,本研究提出使用兩種方法:隨機過採樣示例(ROSE)和合成少數過採樣技術(SMOTE)。 ROSE 和 SMOTE 皆為常見的過採樣方法。ROSE 方法採用平滑引導在稀缺樣本周圍的特徵空間內生成新樣本;而SMOTE 為基於ROSE的優化方法,結合ROSE的優點與K最近鄰 (K Nearest Neighbors, KNN) 演算法以生成新樣本。SMOTE所產生的新合成樣本由原始樣本與其鄰近樣本組成之線段上的隨機位置所決定。本研究通過採用這些平衡方法,改善不平衡類別的分佈並擴增數據集。
此外,儘管現有研究已經採用多種特徵選擇方法來定義自駕車碰撞事故的影響因子,但仍然缺乏不同分類器之間的性能比較。本研究採用三種不同類型的特徵選擇方法,包含互信息(Mutual Information, MI)、隨機森林(Random Forest, RF)和極限梯度提升(eXtreme Gradient Boosting, XGboost),並比較其特徵選擇的成果。另外,本研究也針對特徵選擇及樣本平衡執行程序的影響進行探討,並設計了兩種不同的實驗程序以進行分析比較。
本研究成果顯示與嚴重 AV 碰撞相關的關鍵變量包含製造商、損壞程度、碰撞類型、運動狀態、衝突方和特定 POI。顯著的 POIs 影響因素包含道路資訊(速限)、交通運輸設施(共享汽車租借站和汽車充電站)、娛樂場所(宗教場所和夜店)、公共場所(公共設施)、教育場所(學校)、醫療設施(牙醫診所)等。實驗結果證明了過採樣方法和不同實驗流程的性能差異,可改善自駕車碰撞資料集小樣本及類別分布不均的問題。本研究成果指出:(1)過採樣方法有助於提升自駕車碰撞傷亡與否的預測能力,其中以SMOTE表現最佳(MODEL 9)。(2)特徵選擇前先進行過採樣的研究程序有助於選擇出更具代表性的解釋變量。(3)優化的演算法不一定為最有效的特徵選擇方法;因此,使用者應依據資料集之特徵選擇合適的分類器。最後,本研究成果有助於促進自駕車正式上路運行,也可以應用於台灣的操作設計領域 (Operation Design Domain, ODD)。本研究選擇之變數可納入未來自駕車場景測試之關鍵參數設計。
Autonomous vehicles (AVs) have gained significant popularity recently due to their convenience. It is believed that the potential advantages of AVs can reduce human error crashes, enhance traffic safety, minimize the operational costs, increased road capacity, and improve road congestion. This study specifically focuses on identifying the influencing factor of AV collision, which typically result in few deaths and injuries. However, the potential ramifications for the future could be catastrophic when AVs start to operate on the road. Currently, AVs are still undergoing small-scale testing, making collisions and casualties rare occurrences. Moreover, only few studies have thoroughly defined the environmental factors, such as surrounding environment and road information, that contribute to the different severity of these collisions. Additionally, these studies often overlook the challenges posed by small sample sizes and imbalanced datasets when selecting spatial features.
The study dataset included three years (2019-2021) of AV collision reports (AV Level 3) from California Department of Motor Vehicles (CA DMV). The dataset consisted of 266 collision reports, comprising 51 injured cases and 215 non-injured cases. In addition to the AV collision-related information, the study also considered surrounding environmental factors as important influencing factors for AV collision severity. Therefore, multiple facilities and road speed limits data were collected from Open Street Map (OSM) and DataSF to create the point of interests (POIs) dataset. This allowed for a precise assessment of the influence of the interaction between the collision and the surrounding environment at a micro scale.
To address the aforementioned issues of small sample size and imbalanced dataset, this study proposes the utilization of two methods: Random Over-Sampling Examples (ROSE) and Synthetic Minority Over-Sampling Technique (SMOTE). ROSE and SMOTE are commonly employed oversampling approaches that aim to tackle these challenges. The ROSE employs bootstrapping method to generate new samples within the feature space surrounding the minority class. On the other hand, SMOTE is an optimized technique that combines the benefits of random over sampling with K Nearest Neighbors (KNN) algorithm to create artificial samples. The new synthetic instance is created in the line segment between the given raw sample and its nearest neighbors in the feature space. By employing these balancing methods, both the distribution of the imbalanced class as well as the overall dataset are improved and expanded.
In addition, despite the utilization of various feature selection approaches to address this issue, there is still a lack of performance comparison among different classifiers. This study aims to fill this gap by conducting three different types of feature selection methods (mutual information, random forest, and XGboost) and comparing their results. Furthermore, the study also investigates the impact of the procedure involving feature selection and balancing. Two distinct experimental procedures were designed to analyze in this topic.
The key variables related to injured AV collisions included manufacturers, damage level, collision type, movement state, conflict parties and specific POIs. The significant POIs encompassed road information (speed limit), transportation (car sharing, and charging stations), entertainment (places of worship, and nightclubs), public place (social facilities), education (schools), medical facility (dentists). Furthermore, the experimental results demonstrated the superior performance of the balancing methods and procedure. The findings in this study demonstrate that: (1) Oversampling method helps to improve the predictive ability and SMOTE performs best (MODEL 9). (2) The experimental procedure that balancing before feature selection can select the most representative influencing factors. (3) The integrated algorithm may not always be the most effective, thereby, user should choose the classifier based on the characteristics of the dataset. Finally, this study contributes the enhancement for AV operation on the public roads and can also be applied to the Operational Design Domain (ODD) in Taiwan. The selected variables can be the key factor to design the AV testing in the future.
[1] Alin, A. (2010). Multicollinearity. Wiley interdisciplinary reviews: computational statistics, 2(3), 370-374.
[2] Arlot, S., & Celisse, A. (2010). A survey of cross-validation procedures for model selection.
[3] Aziz, H.A., Ukkusuri, S.V., Hasan, S. Exploring the determinants of pedestrian–vehicle crash severity in New York City. Accid. Anal. Prev. 2013, 50, 1298–1309.
[4] Banerjee, S. S., Jha, S., Cyriac, J., Kalbarczyk, Z. T., & Iyer, R. K. (2018, June). Hands off the wheel in autonomous vehicles?: A systems perspective on over a million miles of field data. In 2018 48th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN) (pp. 586-597). IEEE.
[5] Banks, V.A.; Stanton, N.A. Keep the driver in control: Automating automobiles of the future. Appl. Ergon. 2016, 53, 389–395
[6] Batista, G. E., Prati, R. C., and Monard, M. C., “A study of the behavior of several methods for balancing machine learning training data,” SIGKD Explorations, vol. 6, no. 1, pp. 20–29, 2004.
[7] Belsley, D. A., “A Guide to using the collinearity diagnostics,” Computer Science in Economics and Management, vol. 4, no. 1, pp. 33–50, 1991.
[8] Blagus, R., & Lusa, L. (2013). SMOTE for high-dimensional class-imbalanced data. BMC bioinformatics, 14, 1-16. https://datascience.stackexchange.com/questions/24189/data-balance-before-or-after-feature-selection-engineering
[9] Boggs, A. M., Wali, B., & Khattak, A. J. (2020). Exploratory analysis of automated vehicle crashes in California: A text analytics & hierarchical Bayesian heterogeneity-based approach. Accident Analysis & Prevention, 135, 105354.
[10] Breiman L. Random forests. Mach Learn. 2001;45(1):5–32.
[11] California Department of Motor Vehicles (CA DMV). Summary of Draft Autonomous Vehicles Deployment Regulations December 16, 2015, available from https://www.dmv.ca.gov/portal/dmv/detail/vr/autonomous/auto
[12] California Department of Motor Vehicles (CA DMV). Article 3.7 –Autonomous Vehicles. Title 13, Division 1, par. 227, Available from https://www.dmv.ca.gov/portal/dmv/detail/vr/autonomous/testing, September 2016.
[13] Chawla NV, Bowyer KW, Hall LO, Kegelmeyer WP (2002) SMOTE: synthetic minority over-sampling technique. J Artif Intell Res 16:321–357
[14] Cerwick, D. M., Gkritza, K., Shaheed, M. S., & Hans, Z. (2014). A comparison of the mixed logit and latent class methods for crash severity analysis. Analytic Methods in Accident Research, 3, 11-27.
[15] Chen, H., Chen, H., Liu, Z., Sun, X., Zhou, R., 2020. Analysis of factors affecting the severity of automated vehicle crashes using xgboost model combining poi data. J. Adv. Transp. 2020.
[16] Chen, P. Built environment factors in explaining the automobile-involved bicycle crash frequencies: A spatial statistic approach. Saf. Sci. 2015, 79, 336–343.
[17] Chen, S., Wang, H., & Meng, Q. (2020). Solving the first‐mile ridesharing problem using autonomous vehicles. Computer‐Aided Civil and Infrastructure Engineering, 35(1), 45-60.
[18] Chen, S., Wang, H., & Meng, Q. (2021). An optimal dynamic lane reversal and traffic control strategy for autonomous vehicles. IEEE Transactions on Intelligent Transportation Systems, 23(4), 3804-3815.
[19] Chen, S., Wang, H., Xiao, L., & Meng, Q. (2022). Random capacity for a single lane with mixed autonomous and human-driven vehicles: Bounds, mean gaps and probability distributions. Transportation research part E: logistics and transportation review, 160, 102650.
[20] Chen, T., & Guestrin, C. (2016, August). Xgboost: A scalable tree boosting system. In Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining (pp. 785-794).
[21] Chen, X., "An Improved Branch and Bound Algorithm for Feature Selection", Pattern Recognition Letters, vol. 24, no. 12, pp. 1925-1933, 2003.
[22] Da L. (2017). Data balance -before or after feature selection/engineering. Data Science. Retrieve from:
[23] Davis, J. and Goadrich, M., "The Relationship between Precision-Recall and ROC Curves", Proc. 23rd Int'l Conf. Machine Learning, pp. 30-38, 2006.
[24] Doan, Q. H., Mai, S. H., Do, Q. T., & Thai, D. K. (2022). A cluster-based data splitting method for small sample and class imbalance problems in impact damage classification. Applied Soft Computing, 120, 108628.
[25] Esposito, C., Landrum, G. A., Schneider, N., Stiefl, N., & Riniker, S. (2021). GHOST: adjusting the decision threshold to handle imbalanced data in machine learning. Journal of Chemical Information and Modeling, 61(6), 2623-2640.
[26] Fagnant, D. J. and Kockelman K., “Preparing a nation for autonomous vehicles: opportunities, barriers and policy recommendations,” Transportation Research Part A: Policy and Practice, vol. 77, pp. 167–181, 2015.
[27] Favarò, F., Eurich, S., & Nader, N. (2018). Autonomous vehicles’ disengagements: Trends, triggers, and regulatory limitations. Accident Analysis & Prevention, 110, 136-148.
[28] Favarò, F. M., Nader, N., Eurich, S. O., Tripp, M., & Varadaraju, N. (2017). Examining accident reports involving autonomous vehicles in California. PLoS one, 12(9), e0184952.
[29] Gao, H., Fam, P. S., Tay, L. T., & Low, H. C. (2020). Three oversampling methods applied in a comparative landslide spatial research in Penang Island, Malaysia. SN Applied Sciences, 2, 1-20.
[30] Ghisloti, R. (2022). Data balance -before or after feature selection/engineering. Data Science. Retrieve from: https://datascience.stackexchange.com/questions/24189/data-balance-before-or-after-feature-selection-engineering
[31] Golub, T.R., Slonim, D.K., Tamayo, P., Huard, C., Gaasenbeek, M., Mesirov, J.P., et al., "Molecular Classification of Cancer: Class Discovery and Class Prediction by Gene Expression Monitoring", Science, vol. 286, pp. 531-537, 1999.
[32] Gourdeau, P., Kanbar, L., Shalish, W., Sant'Anna, G., Kearney, R., & Precup, D. (2015, August). Feature selection and oversampling in analysis of clinical data for extubation readiness in extreme preterm infants. In 2015 37th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC) (pp. 4427-4430). IEEE.
[33] Guillaume Lemaitre, Dayvid Victor, Fernando Nogueira & Christos K. Aridas. Imbalanced-learn: A Python Toolbox to Tackle the Curse of Imbalanced Datasets in Machine Learning. Journal of Machine Learning Research, vol. 18, no. 17, pp. 1-5, 2017
[34] Houseal, L. A., Gaweesh, S. M., Dadvar, S., & Ahmed, M. M. (2022). Causes and effects of autonomous vehicle field test crashes and disengagements using exploratory factor analysis, binary logistic regression, and decision trees. Transportation research record, 2676(8), 571-586.
[35] Hu, Y., Guo, D., Fan, Z., Dong, C., Huang, Q., Xie, S., ... & Xie, Q. (2015). An improved algorithm for imbalanced data and small sample size classification. Journal of Data Analysis and Information Processing, 3(03), 27.
[36] Huelke, D. F., & Compton, C. P. (1995). The effects of seat belts on injury severity of front and rear seat occupants in the same frontal crash. Accident Analysis & Prevention, 27(6), 835-838.
[37] Hulse, J.V., Khoshgoftaar, T.M. and Napolitano, A., "Experimental Perspectives on Learning from Imbalanced Data", Proc. 24th Int'l Conf. Machine Learning, pp. 935-942, 2007.
[38] Imam, T., Kai, M. T., and Kamruzzaman, J., “z-SVM: An SVM for improved classification of imbalanced data,” in Proc. Australas. Joint Conf. Artif. Intell., 2006, pp. 264–273..
[39] Jia, R., Khadka, A., & Kim, I. (2018). Traffic crash analysis with point-of-interest spatial clustering. Accident Analysis & Prevention, 121, 223-230.
[40] Jian, C., Gao, J., and Ao, Y., “A new sampling method for classifying imbalanced data based on support vector machine ensemble,” Neurocomputing, vol. 193, pp. 115–122, 2016.
[41] Johnson, J. M., & Khoshgoftaar, T. M. (2019). Survey on deep learning with class imbalance. Journal of Big Data, 6(1), 1-54.
[42] Kang, Q., Shi, L., Zhou, M. C., Wang, X. S., Wu, Q., and Wei, Z., “A distancebased weighted undersampling scheme for support vector machines and its application to imbalanced classification,” IEEE Trans. Neural Netw. Learn. Syst., vol. 29, no. 9, pp. 4152–4165, Sep. 2018.
[43] Kim, S., Song, T.-J., Rouphail, N.M., Aghdashi, S., Amaro, A., Gonçalves, G.. Exploring the association of rear-end crash propensity and micro-scale driver behavior. Saf. Sci., 89 (2016), pp. 45-54
[44] Koopman, P., & Fratrik, F. (2019). How many operational design domains, objects, and events?. Safeai@ aaai, 4.
[45] Lee, C.; Abdel-Aty, M. Comprehensive analysis of vehicle-pedestrian crashes at intersections in Florida. Accid. Anal. Prev. 2005, 37, 775–786.
[46] Leilabadi, S. H., & Schmidt, S. (2019, October). In-depth analysis of autonomous vehicle collisions in california. In 2019 IEEE Intelligent Transportation Systems Conference (ITSC) (pp. 889-893). IEEE.
[47] Liu, F., & Dai, Y. (2022). Product processing quality classification model for small-sample and imbalanced data environment. Computational Intelligence and Neuroscience, 2022.
[48] Liu, F., Zhao, F., Liu, Z., & Hao, H. (2019). Can autonomous vehicle reduce greenhouse gas emissions? A country-level evaluation. Energy Policy, 132, 462-473.
[49] Liu, T. Y. (2009, August). Easyensemble and feature selection for imbalance data sets. In 2009 international joint conference on bioinformatics, systems biology and intelligent computing (pp. 517-520). IEEE.
[50] Li, Y.; Liu, C.; Ding, L. Impact of pavement conditions on crash severity. Accid. Anal. Prev. 2013, 59, 399–406.
[51] Lin, W. J. and Chen, J. J., “Class-imbalanced classifiers for highdimensional data,” Briefings Bioinf., vol. 14, no. 1, pp. 13–26, 2013.
[52] Ma, D., Sandberg, M., & Jiang, B. (2015). Characterizing the heterogeneity of the OpenStreetMap data and community. ISPRS International Journal of Geo-Information, 4(2), 535-550.
[53] Mahdinia, I.; Mohammadnazar, A.; Arvin, R.; Khattak, A.J. Integration of automated vehicles in mixed traffic: Evaluating changes in performance of following human-driven vehicles. Accid. Anal. Prev. 2021, 152, 106006.
[54] Malik, A. (2020). Sampling before or after feature selection. stack overflow. Retrieve from: https://stackoverflow.com/questions/63375860/sampling-before-or-after-feature-selection
[55] Mathew, J., Pang, C. K., Luo, M., and Leong, W. H., “Classification of imbalanced data by oversampling in kernel space of support vector machines,” IEEE Trans. Neural Netw. Learn. Syst., vol. 29, no. 9, pp. 4065–4076, Sep. 2018.
[56] Menardi, G., & Torelli, N. (2014). Training and assessing classification rules with imbalanced data. Data mining and knowledge discovery, 28(1), 92-122.
[57] Menzel, T., Bagschik, G., Isensee, L., Schomburg, A., & Maurer, M. (2019, June). From functional to logical scenarios: Detailing a keyword-based scenario description for execution in a simulation environment. In 2019 IEEE Intelligent Vehicles Symposium (IV) (pp. 2383-2390). IEEE.
[58] Menzel, T., Bagschik, G., & Maurer, M. (2018, June). Scenarios for development, test and validation of automated vehicles. In 2018 IEEE Intelligent Vehicles Symposium (IV) (pp. 1821-1827). IEEE.
[59] Merat, N.; Jamson, A.H.; Lai, F.; Daly, M.; Carsten, O. Transition to manual: Driver behaviour when resuming control from a highly automated vehicle. Transp. Res. Part F Traffic Psychol. Behav. 2014, 27, 274–282.
[60] Mobilus, S. A. E. (2018). J3016B: Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles-SAE International. Taxonomy and Definitions for Terms Related to On-Road Motor Vehicle Automated Driving Systems.
[61] Mooijman, P., Catal, C., Tekinerdogan, B., Lommen, A., & Blokland, M. (2023). The effects of data balancing approaches: A case study. Applied Soft Computing, 132, 109853.
[62] Naes, T. and Mevik, B. H., “Understanding the collinearity problem in regression and discriminant analysis,” Journal of Chemometrics, vol. 15, no. 4, pp. 413–426, 2001.
[63] NSTC, U. (2020). Ensuring american leadership in automated vehicle technologies: Automated vehicles 4.0. NSTC, USDOT: Washington, DC, USA.
[64] Olsson, U. (1979). Maximum likelihood estimation of the polychoric correlation coefficient. Psychometrika, 44(4), 443-460.
[65] Pan, Y., Chen, S., Li ,T., Niu, S., and Tang, K., “Exploring spatial variation of the bus stop influence zone with multi-source data: a case study in Zhenjiang, China,” Journal of Transport Geography, vol. 76, pp. 166–177, 2019.
[66] Pattern Analysis and Machine Intelligence, vol. 27, no. 8, pp. 1226-1238, Aug. 2005.
[67] Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, Blondel M, Prettenhofer P, Weiss R, Dubourg V, et al. Scikit-learn: machine learning in python. J Mach Learn Res. 2011;12:2825–30.
[68] Peng, H., Long, F. and Ding, C., "Feature Selection Based on Mutual Information: Criteria of Max-Dependency Max-Relevance and Min-Redundancy", IEEE Trans.
[69] Petrere, M., “Pesque-solte,” Ciência Hoje, vol. 53, no. 5, pp. 1189–1232, 2014.
[70] Poon, W. Y., & Lee, S. Y. (1987). Maximum likelihood estimation of multivariate polyserial and polychoric correlation coefficients. Psychometrika, 52(3), 409-430.
[71] Qu, X., Zhu, X., Xiao, X., Wu, H., Guo, B., & Li, D. (2021). Exploring the Influences of Point-of-Interest on Traffic Crashes during Weekdays and Weekends via Multi-Scale Geographically Weighted Regression. ISPRS International Journal of Geo-Information, 10(11), 791
[72] Reddy, S. S., Chao, Y. L., Kotikalapudi, L. P., & Ceesay, E. (2022, June). Accident analysis and severity prediction of road accidents in United States using machine learning algorithms. In 2022 IEEE International IOT, Electronics and Mechatronics Conference (IEMTRONICS) (pp. 1-7). IEEE.
[73] Ren, W., Yu, B., Chen, Y., & Gao, K. (2022). Divergent effects of factors on crash severity under autonomous and conventional driving modes using a hierarchical Bayesian approach. International journal of environmental research and public health, 19(18), 11358.
[74] Retting, R.A.; Kyrychenko, S.Y. Reductions in injury crashes associated with red light camera enforcement in Oxnard, California. Am. J. Public Health 2002, 92, 1822–1825.
[75] SAE. Society of Automotive Engineers. On-Road Automated Vehicle Standards Committee, 2014. Taxonomy and definitions for terms related to on-road motor vehicle automated driving systems.
[76] Shahib, A. Al., Breitling, R. and Gilbert, D., "Feature Selection and the Class Imbalance Problem in Predicting Protein Function from Sequence", Applied Bioinformatics, vol. 4, pp. 195-203, 2005.
[77] Shi, Q., & Zhang, H. (2020). Fault diagnosis of an autonomous vehicle with an improved SVM algorithm subject to unbalanced datasets. IEEE Transactions on Industrial Electronics, 68(7), 6248-6256.
[78] Sauerbier, J., Bock, J., Weber, H., & Eckstein, L. (2019). Definition of scenarios for safety validation of automated driving functions. ATZ worldwide, 121(1), 42-45.
[79] Sarker, I. H. (2021). Machine learning: Algorithms, real-world applications and research directions. SN computer science, 2(3), 160.
[80] Sarker IH, Watters P, Kayes ASM. Effectiveness analysis of machine learning classification models for predicting personalized context-aware smartphone usage. J Big Data. 2019;6(1):1–28.
[81] Schreck, B. (2018). Feature Engineering Vs Feature Selection. Alteryx, Innovation, Engineering. Retrieve from
[82] Sinha, A., Vu, V., Chand, S., Wijayaratna, K., & Dixit, V. (2021). A crash injury model involving autonomous vehicle: Investigating of crash and disengagement reports. Sustainability, 13(14), 7938.
[83] Stilgoe, J., “Machine learning, social learning and the governance of self-driving cars,” Social Studies of Science, vol. 48, no. 1, pp. 25–56, 2018.
[84] Song, Y., Chitturi, M. V., & Noyce, D. A. (2021). Automated vehicle crash sequences: Patterns and potential uses in safety testing. Accident Analysis & Prevention, 153, 106017.
[85] Theofilatos, A., Antoniou, C., & Yannis, G. (2021). Exploring injury severity of children and adolescents involved in traffic crashes in Greece. Journal of traffic and transportation engineering (English edition), 8(4), 596-604.
[86] Veropoulos, K., Cristianini, N., and Campbell, C., “Controlling the sensitivity of support vector machines,” in Proc. 16th Int. Joint Conf. Artif. Intell., 1999, pp. 281–288.
[87] Wali, B., Khattak, A.J., Karnowski, T., The relationship between driving volatility in time to collision and crash-injury severity in a naturalistic driving environment. Anal. Methods Accident Res., 28 (2020), Article 100136
[88] Wang, S., & Li, Z. (2019). Exploring the mechanism of crashes with automated vehicles using statistical modeling approaches. PloS one, 14(3), e0214550.
[89] Wasikowski, M., & Chen, X. W. (2009). Combating the small sample class imbalance problem using feature selection. IEEE Transactions on knowledge and data engineering, 22(10), 1388-1400.
[90] Xu, C.; Ding, Z.; Wang, C.; Li, Z. Statistical analysis of the patterns and characteristics of connected and autonomous vehicle involved crashes. J. Saf. Res. 2019, 71, 41–47.
[91]Yang, J., Qu, Z., & Liu, Z. (2014). Improved feature-selection method considering the imbalance problem in text categorization. The Scientific World Journal, 2014.
[92] Yang, N., Zhao, Y., Chen, J., & Wang, F. (2023). Real-time classification for Φ-OTDR vibration events in the case of small sample size datasets. Optical Fiber Technology, 76, 103217.
[93] Yao, S., Wang, J., Fang, L., & Wu, J. (2018). Identification of vehicle-pedestrian collision hotspots at the micro-level using network kernel density estimation and random forests: A case study in Shanghai, China. Sustainability, 10(12), 4762.
[94] Yu, H., Mu, C., Sun, C., Yang, W., Yang, X., and Xin, Z., “Support vector machine-based optimized decision threshold adjustment strategy for classifying imbalanced data,” Knowl. Syst., vol. 76, no. 1, pp. 67–78, 2015.
[95] Yu, R., & Li, S. (2022). Exploring the associations between driving volatility and autonomous vehicle hazardous scenarios: insights from field operational test data. Accident Analysis & Prevention, 166, 106537.
[96] Zheng, F.; Liu, C.; Liu, X.; Jabari, S.E.; Lu, L. Analyzing the impact of automated vehicles on uncertainty and stability of the mixed traffic flow. Transp. Res. Part C Emerg. Technol. 2020, 112, 203–219.
[97] Zhu, S., & Meng, Q. (2022). What can we learn from autonomous vehicle collision data on crash severity? A cost-sensitive CART approach. Accident Analysis & Prevention, 174, 106769.