| 研究生: |
鄭凱元 Cheng, Kai-Yuan |
|---|---|
| 論文名稱: |
基於大型語言模型與貝葉斯最佳化的配置調校增強框架:以 RocksDB 為例 A Configuration Tuning Enhancement Framework Based on LLM and Bayesian Optimization: A Case Study of RocksDB |
| 指導教授: |
蕭宏章
Hsiao, Hung-Chang |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 資訊工程學系 Department of Computer Science and Information Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 43 |
| 中文關鍵詞: | 大型語言模型 、貝葉斯最佳化 、RocksDB |
| 外文關鍵詞: | Large Language Model, Bayesian Optimization, RocksDB |
| 相關次數: | 點閱:25 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
RocksDB 作為現代基礎設施中廣泛採用的鍵值存儲系統,擁有超過 100 個可調參數,參數間的複雜交互作用使得手動調校極為困難。現有調校方法各有局限:大型語言模型(Large Language Model, LLM)雖具備豐富的參數語義知識,能在缺乏歷史資料的情況下給出合理初始配置,但缺乏不確定性量化與收斂保證;貝葉斯最佳化(Bayesian Optimization, BO)雖然統計嚴謹,卻在高維度空間下面臨嚴重的冷啟動問題,需要大量評估才能收斂。如何兼取兩者優勢並克服各自缺陷,是本研究的核心問題。
本研究旨在設計一套結合 LLM 語義先驗知識與高斯過程(Gaussian Process, GP)統計篩選的混合調校框架,並應用於 RocksDB 的配置最佳化場景。框架以 GP 代理模型對 LLM 所提出的配置建議進行動態篩選,篩選門檻隨迭代遞增:早期較寬鬆以充分利用 LLM 先驗知識,後期逐步收緊以確保搜尋的嚴謹性。此外,本研究將LLINBO 原有的 UCB 獲取函數替換為期望改善值(Expected Improvement, EI),並推導 EI 下的次線性累積遺憾上界,確保框架的理論收斂保證不受此替換影響。
RocksDB, as a widely adopted key-value storage system in modern infrastructure, possesses over 100 tunable parameters, and the complex interactions among them make manual tuning extremely difficult. Existing tuning approaches each have fundamental limitations: Large Language Models (LLMs), while rich in parameter semantic knowledge and capable of providing reasonable initial configurations without historical data, lack uncertainty quantification and convergence guarantees; Bayesian Optimization (BO), though statistically rigorous, faces severe cold-start problems in high-dimensional spaces, requiring numerous evaluations before converging. How to leverage the strengths of both while overcoming their respective weaknesses is the central problem of this research.
This research designs a hybrid tuning framework that combines LLM semantic prior knowledge with Gaussian Process (GP) statistical filtering for RocksDB configuration optimization. The framework uses a GP surrogate model to dynamically filter LLM configuration suggestions, with a filtering threshold that increases iteratively: relaxed in early stages to fully exploit LLM prior knowledge, then progressively tightened to ensure search rigor. Furthermore, this research replaces the UCB acquisition function in LLINBO with Expected Improvement (EI) and derives the corresponding sub-linear cumulative regret bound, ensuring the theoretical convergence guarantee remains intact under this substitution.
[1] Sami Alabed and Eiko Yoneki. High-dimensional bayesian optimization with multi-task learning for RocksDB. In Proceedings of the 1st Workshop on Machine Learning and Systems (EuroMLSys), pages 111–119, 2021.
[2] Adam D. Bull. Convergence rates of efficient global optimization algorithms. Journal of Machine Learning Research, 12(88):2879–2904, 2011.
[3] Chih-Yu Chang, Milad Azvar, Chinedum Okwudire, and Raed Al Kontar. LLINBO: Trustworthy LLM-in-the-loop bayesian optimization, 2025. arXiv preprint arXiv:2505.14756.
[4] Andy Dent. Getting Started with LevelDB. Packt Publishing Ltd, 2013.
[5] Songyun Duan, Vamsidhar Thummala, and Shivnath Babu. Tuning database configuration parameters with iTuned. Proceedings of the VLDB Endowment, 2(1):1246–1257, 2009.
[6] Facebook. RocksDB tuning guide. https://github.com/facebook/rocksdb/wiki/RocksDB-Tuning-Guide.
[7] Facebook. RocksDB: A persistent key-value store for fast storage environments. https://rocksdb.org, 2021.
[8] Peter I. Frazier. A tutorial on bayesian optimization, 2018. arXiv preprint arXiv:1807.02811.
[9] Google. Google Gemini API Documentation. https://ai.google.dev/gemini-api/docs.
[10] Yichen Jia and Feng Chen. Kill two birds with one stone: Auto-tuning RocksDB for high bandwidth and low latency. In 2020 IEEE 40th International Conference on Distributed Computing Systems (ICDCS), pages 652–664. IEEE, 2020.
[11] Huijun Jin, Won Gi Choi, Jonghwan Choi, Hanseung Sung, and Sanghyun Park. Improvement of RocksDB performance via large-scale parameter analysis and optimization. Journal of Information Processing Systems, 18(3):374–388, 2022.
[12] Donald R. Jones, Matthias Schonlau, and William J. Welch. Efficient global optimization of expensive black-box functions. Journal of Global Optimization, 13(4):455–492, 1998.
[13] Jiale Lao, Yibo Wang, Yufei Li, Jianping Wang, Yunjia Zhang, Zhiyuan Cheng, Wanghu Chen, Mingjie Tang, and Jianguo Wang. GPTuner: A manual-reading database tuning system via GPT-guided bayesian optimization. Proceedings of the VLDB Endowment, 17(8):1939–1952, 2024.
[14] Tennison Liu, Nicolás Astorga, Nabeel Seedat, and Mihaela van der Schaar. Large language models to enhance bayesian optimization. In The Twelfth International Conference on Learning Representations, 2024.
[15] Patrick O'Neil, Edward Cheng, Dieter Gawlick, and Elizabeth O'Neil. The log-structured merge-tree (LSM-tree). Acta Informatica, 33(4):351–385, 1996.
[16] Carl Edward Rasmussen and Christopher K. I. Williams. Gaussian Processes for Machine Learning. MIT Press, Cambridge, MA, 2006.
[17] Jasper Snoek, Hugo Larochelle, and Ryan P. Adams. Practical bayesian optimization of machine learning algorithms. In Advances in Neural Information Processing Systems, volume 25, 2012.
[18] Viraj Thakkar, Madhumitha Sukumar, Jiaxin Dai, Kaushiki Singh, and Zhichao Cao. Can modern LLMs tune and configure LSM-based key-value stores? In Proceedings of the 16th ACM Workshop on Hot Topics in Storage and File Systems, HotStorage '24, pages 116–123, 2024.
[19] Dana Van Aken, Andrew Pavlo, Geoffrey J. Gordon, and Bohan Zhang. Automatic database management system tuning through large-scale machine learning. In Proceedings of the 2017 ACM International Conference on Management of Data (SIGMOD), pages 1009–1024, 2017.
[20] Ting Yao, Yiwen Zhang, Jiguang Wan, Qiu Cui, Liu Tang, Hong Jiang, Changsheng Xie, and Xubin He. MatrixKV: Reducing write stalls and write amplification in LSM-tree based KV stores with a matrix container in NVM. In 2020 USENIX Annual Technical Conference (USENIX ATC 20), pages 17–31, 2020.
[21] Ji Zhang, Yu Liu, Ke Zhou, Guoliang Li, Zhili Xiao, Bin Cheng, Jiashu Xing, Yangtao Wang, Tianheng Cheng, Li Liu, Minwei Ran, and Zekang Li. An end-to-end automatic cloud database tuning system using deep reinforcement learning. In Proceedings of the 2019 ACM International Conference on Management of Data (SIGMOD), pages 415–432, 2019.
[22] Xinyi Zhang, Zhuo Chang, Yang Li, Hong Wu, Jian Tan, Feifei Li, and Bin Cui. Facilitating database tuning with hyper-parameter optimization: A comprehensive experimental evaluation. Proceedings of the VLDB Endowment, 15(9):1808–1821, 2022.
[23] Chenxingyu Zhao, Tapan Chugh, Jaehong Min, Ming Liu, and Arvind Krishnamurthy. Dremel: Adaptive configuration tuning of RocksDB KV-store. Proceedings of the ACM on Measurement and Analysis of Computing Systems, 6(2):1–30, 2022.
[24] Xinyang Zhao, Xuanhe Zhou, and Guoliang Li. Automatic database knob tuning: A survey. IEEE Transactions on Knowledge and Data Engineering, 35(12):12470–12490, 2023.