| 研究生: |
楊峻豪 Yang, Chun-Hao |
|---|---|
| 論文名稱: |
具自我感知之雲計算任務的細粒度資源管理 Autonomic, Fine-Grained Resource Management in Clouds for Batch and Real-Time Computing |
| 指導教授: |
蕭宏章
Hsiao, Hung-Chang |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 資訊工程學系 Department of Computer Science and Information Engineering |
| 論文出版年: | 2021 |
| 畢業學年度: | 109 |
| 語文別: | 中文 |
| 論文頁數: | 89 |
| 中文關鍵詞: | 分散式計算 、數據計算框架 、雲端資源管理 、細粒度資源管理 |
| 外文關鍵詞: | Distributed Computing, Data Computing Framework, Cloud Resource Management, Fine-Grained Resource Management |
| 相關次數: | 點閱:198 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
隨著大數據蘊涵價值的逐步被釋放,已成為當今火紅的議題之一,對工業領域也具有重要的意義,許多的廠商開始轉型為資料驅動的智慧製造(Data-Driven Smart Manufacturing),目的是讓生產過程中取得的數據轉化為製造智慧,像是讓這些數據能夠即時地回饋在生產作業中,這使得廠商對巨量資料的儲存、管理及分析等技術議題之需求日漸增長,因為這些生產數據量十分的龐大且不斷的快速產生中。
本研究提出了DCFS(Distributed Computing Framework/Service)分散式計算框架,目的是要解決合作廠所提出之對生產作業的數據計算需求,特別是即時高速計算(Real-time Computations)之需求,其也能用於製造產業上其他的大規模數據計算應用場景。在DCFS的設計上,對使用者保有高度的彈性,其可以自行量身打造所要執行之計算工作的任務觸發及計算邏輯,且提供可以透明地直接存取異質儲存體的API存取方式,使用者可以更快速及方便的撰寫其邏輯。對於較常見之計算分析需求,透過DCFS提供的Job Tool能幫助一般使用者或是資料科學家,在無須撰寫額外程式的前提,透過帶入參數的方式,即可執行分散式計算工作。
DCFS提供了兩種細粒度的資源管理機制,可以滿足計算任務的實際資源需求,使提升每個容器的平均記憶體使用率,減少閒置資源造成的浪費,並視計算工作的執行狀況自動化的伸縮容器。實驗結果指出,同時啟用DCFS提出之兩種細粒度資管理機制,能顯著提高整個叢集資源的使用率(實驗指出可提升到近85%),且能加快整個計算工作的執行時間(實驗指出可加快約36%)。
在論文最後,提出了DCFS框架未來的努力目標,像是動態調整容器的CPU資源規格,以及發展成計算平台的形式,讓使用者透過Web UI的方式提交、追蹤及管理計算工作,使這份研究的成果能夠日臻完善。
This paper presents a distributed computing framework (DCFS) for semiconductor manufacturing foundries, considering real-time and batch computing demands. While real-time tasks are generated on-the-fly in a streaming manner, batch jobs are submitted for computing large data sets periodically (or aperiodically). Our framework is not only capable of performing streaming tasks in real-time, but serves batched, big-data jobs. We advocate in DCFS to rely on the state-of-the-art virtualization technology, i.e., containers, to optimize the utilization of compute resources such as processor cycles and memory space. Specifically, we present in DCFS a resource allocation technique that adaptively predicts the resource demands of compute tasks on-the-fly during job computation lifetime. Our prediction technique is unique in that it is transparent by adaptively allocation resources for irregular resource demand patterns without compute job knowledge in advance. On the other hand, DFCS supports failure recovery, guaranteeing compute jobs that can be performed successfully in the presence of failures. We have performed in this study extensive performance measurements for our proposed framework. The performance results reveal that DCFS can significantly improve the resource utilization up-to nearly 85%, while speeding up the job execution time by 36% at most.
[1] Top 20 Big Data Statistics for 2021. [Online]. Available: https://www.sigmacomputing.com/blog/top-20-big-data-statistics/
[2] F. Tao, Q. Qi, A. Liu, and A. Kusiak, "Data-driven Smart Manufacturing," Journal of Manufacturing Systems, vol. 48, pp. 157-169, 2018, doi: 10.1016/j.jmsy.2018.01.006.
[3] 大數據分析如何改變產業鏈?. [Online]. Available: https://www.semi.org/zh/blogs/technology-trends/big-data
[4] P. O’Donovan, K. Leahy, K. Bruton, and D. T. J. O’Sullivan, "An Industrial Big Data Pipeline for Data-driven Analytics Maintenance Applications in Large-scale Smart Manufacturing Facilities," Journal of Big Data, vol. 2, no. 1, p. 25, 2015, doi: 10.1186/s40537-015-0034-z.
[5] Apache Spark. [Online]. Available: https://spark.apache.org
[6] Apache Hadoop - MapReduce Tutorial. [Online]. Available: https://hadoop.apache.org/docs/stable/hadoop-mapreduce-client/hadoop-mapreduce-client-core/MapReduceTutorial.html
[7] Kubernetes. [Online]. Available: https://kubernetes.io
[8] Apache Hadoop YARN. [Online]. Available: https://hadoop.apache.org/docs/current/hadoop-yarn/hadoop-yarn-site/YARN.html
[9] Y. Zhao and G. Wu, "Yadoop: An Elastic Resource Management Solution of YARN," in 2015 IEEE Symposium on Service-Oriented System Engineering, 2015, pp. 276-283, doi: 10.1109/SOSE.2015.20.
[10] Y. Peng, D. Luo, J. Dong, and Z. Wu, "High Concurrent Elastic Resource Allocation in Hadoop YARN," in International Conference on Communications and Networking in China, 2018, pp. 524-534, doi: 10.1007/978-3-319-78130-3_54.
[11] M. Genkin, F. Dehne, M. Pospelova, Y. Chen, and P. Navarro, "Automatic, On-Line Tuning of YARN Container Memory and CPU Parameters," in 2016 IEEE 18th International Conference on High Performance Computing and Communications; IEEE 14th International Conference on Smart City; IEEE 2nd International Conference on Data Science and Systems (HPCC/SmartCity/DSS), 2016, pp. 317-324, doi: 10.1109/HPCC-SmartCity-DSS.2016.0053.
[12] An Introduction to Kubernetes. [Online]. Available: https://www.digitalocean.com/community/tutorials/an-introduction-to-kubernetes
[13] Kubernetes Documentation: What is Kubernetes? [Online]. Available: https://kubernetes.io/docs/concepts/overview/what-is-kubernetes/
[14] B. Burns, B. Grant, D. Oppenheimer, E. Brewer, and J. Wilkes, "Borg, Omega, and Kubernetes: Lessons Learned from Three Container-management Systems Over a Decade," Queue, vol. 14, no. 1, pp. 70–93, 2016, doi: 10.1145/2898442.2898444.
[15] A. Verma, L. Pedrosa, M. Korupolu, D. Oppenheimer, E. Tune, and J. Wilkes, "Large-scale cluster management at Google with Borg," in Proceedings of the Tenth European Conference on Computer Systems, 2015, pp. 1-17, doi: 10.1145/2741948.2741964.
[16] What is Kubernetes? [Online]. Available: https://www.redhat.com/en/topics/containers/what-is-kubernetes
[17] Kubernetes Documentation: Kubernetes Components. [Online]. Available: https://kubernetes.io/docs/concepts/overview/components/
[18] Kubernetes Documentation: kube-apiserver. [Online]. Available: https://kubernetes.io/docs/reference/command-line-tools-reference/kube-apiserver/
[19] Kubernetes core concepts for Azure Kubernetes Service (AKS). [Online]. Available: https://docs.microsoft.com/en-us/azure/aks/concepts-clusters-workloads
[20] Kubernetes Documentation: Kubernetes Scheduler. [Online]. Available: https://kubernetes.io/docs/concepts/scheduling-eviction/kube-scheduler/
[21] Kubernetes Documentation: kube-controller-manager. [Online]. Available: https://kubernetes.io/docs/reference/command-line-tools-reference/kube-controller-manager/
[22] etcd. [Online]. Available: https://etcd.io
[23] What is etcd? [Online]. Available: https://www.ibm.com/cloud/learn/etcd
[24] V. K. Vavilapalli et al., "Apache Hadoop YARN: Yet Another Resource Negotiator," in Proceedings of the 4th annual Symposium on Cloud Computing, 2013, pp. 1-16, doi: 10.1145/2523616.2523633.
[25] What is Apache Hadoop YARN? [Online]. Available: https://searchdatamanagement.techtarget.com/definition/Apache-Hadoop-YARN-Yet-Another-Resource-Negotiator
[26] V. Kalavri and V. Vlassov, "MapReduce: Limitations, Optimizations and Open Issues," in 2013 12th IEEE International Conference on Trust, Security and Privacy in Computing and Communications, 2013, pp. 1031-1038, doi: 10.1109/TrustCom.2013.126.
[27] Cloudera Docs: Understanding YARN architecture. [Online]. Available: https://docs.cloudera.com/runtime/7.2.10/yarn-overview/topics/yarn-apache-yarn.html
[28] What Is Hadoop Yarn Architecture & It’s Components. [Online]. Available: https://www.upgrad.com/blog/what-is-hadoop-yarn-architecture-its-components/
[29] U. Demirbaga et al., "AutoDiagn: An Automated Real-time Diagnosis Framework for Big Data Systems," IEEE Transactions on Computers, 2021, doi: 10.1109/TC.2021.3070639.
[30] A. Alalawi and A. Al-Omary, "A Survey On Cloud-Based Distributed Computing System Frameworks," in 2020 International Conference on Data Analytics for Business and Industry: Way Towards a Sustainable Economy (ICDABI), 2020, pp. 1-6, doi: 10.1109/ICDABI51230.2020.9325662.
[31] How to Create a Business Case for Data Quality Improvement. [Online]. Available: https://www.gartner.com/smarterwithgartner/how-to-create-a-business-case-for-data-quality-improvement
[32] Apache Hadoop. [Online]. Available: https://hadoop.apache.org
[33] [YARN-1197] Support changing resources of an allocated container. [Online]. Available: https://issues.apache.org/jira/browse/YARN-1197
[34] Kubernetes Documentation: Pods. [Online]. Available: https://kubernetes.io/docs/concepts/workloads/pods/
[35] Kubernetes Autoscaler: Vertical Pod Autoscaler. [Online]. Available: https://github.com/kubernetes/autoscaler/tree/master/vertical-pod-autoscaler
[36] Kubernetes Enhancements: In-place Update of Pod Resources (Issue#1287). [Online]. Available: https://github.com/kubernetes/enhancements/tree/master/keps/sig-node/1287-in-place-update-pod-resources
[37] Kubernetes (pull request): In-place Pod Vertical Scaling feature by vinaykul. [Online]. Available: https://github.com/kubernetes/kubernetes/pull/102884
[38] HDFS Architecture. [Online]. Available: https://hadoop.apache.org/docs/current/hadoop-project-dist/hadoop-hdfs/HdfsDesign.html
[39] The Netty project. [Online]. Available: https://netty.io
[40] A. Kuzmanovska, R. H. Mak, and D. Epema, "Dynamically Scheduling a Component-Based Framework in Clusters," 2015, in Job Scheduling Strategies for Parallel Processing, pp. 129-146, doi: 10.1007/978-3-319-15789-4_8.
[41] Using Memory Control in YARN. [Online]. Available: https://hadoop.apache.org/docs/current/hadoop-yarn/hadoop-yarn-site/NodeManagerCGroupsMemory.html
[42] Using CGroups with YARN. [Online]. Available: https://hadoop.apache.org/docs/current/hadoop-yarn/hadoop-yarn-site/NodeManagerCgroups.html