簡易檢索 / 詳目顯示

研究生: 陳皇隆
Chen, Huang-Lung
論文名稱: 利用隨機森林與 LightGBM 之VVC 畫面內編碼加速
Accelerating VVC Intra Coding via Random Forest and LightGBM Classifiers
指導教授: 戴顯權
Tai, Shen-Chuan
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 電腦與通信工程研究所
Institute of Computer & Communication Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 79
中文關鍵詞: 多功能視訊編碼快速 CU 分割QT-MTT隨機森林LightGBM率失真複雜度權衡
外文關鍵詞: Versatile Video Coding, fast CU partition, QT-MTT, Random Forest, LightGBM, rate-distortion-complexity trade-off
相關次數: 點閱:9下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 本論文針對多功能視訊編碼(Versatile Video Coding, VVC/H.266)編碼端因四元樹加巢狀多型樹(quad-tree plus nested multi-type tree, QT-MTT)分割之窮舉式率失真最佳化(rate-distortion optimization, RDO)搜尋所造成之高運算複雜度問題,探討以輕量化機器學習(machine learning, ML)取代深度學習(deep learning, DL)於大尺寸編碼單元(coding unit, CU)分割決策之可行性與限制。本論文延伸 Tai等人(2025)所提出之拼圖式快速 CU 分割框架(puzzle-based fast CU partitioning,PFCP),維持其整體骨架不變,僅將負責大尺寸 CU(64 × 64、32 × 32)之二元(切/不切)預測器,由原本之輕量卷積神經網路(convolutional neural network,CNN)替換為隨機森林(Random Forest, RF)或 LightGBM(LGBM),其餘較小尺寸維持原有 CNN,構成混合式 ML–CNN 架構。為系統性檢視各替換點之效果,本論文設計涵蓋三種互斥替換策略、兩種機器學習模型、三組不同紋理複雜度之測試序列與四個量化參數(quantization parameter, QP)之 72 組消融實驗,並以 VVC 測試模型(VVC Test Model, VTM)17.2 與原始 DL 框架為雙重比較基準。
    實驗結果顯示,大尺寸 CU 之 CNN 推論時間為整條管線之主要瓶頸,且可由輕量 ML 有效消除:於 64 × 64 區塊,RF 與 LGBM 相對 CNN 之推論加速達約 90 至175 倍,且全程僅以中央處理器(central processing unit, CPU)運行、無需圖形處理器(graphics processing unit, GPU)。預測準確度隨 CU 尺寸增大而提升,於大尺寸(32 × 32、64 × 64)已足以接近 CNN,但於小尺寸明顯不足,此即僅於大尺寸採用 ML、小尺寸保留 CNN 之設計依據。在編碼效能方面,最佳替換組態相對原始DL 框架之 BD-rate(Bjontegaard delta rate)增量僅約 1% 至 8%。若進一步以端到端方式衡量整體效益,最終推薦組態(LGBM both)相對未修改之 VTM 17.2,於單張畫面之編碼時間減少約 49% 至 58%(隨畫面尺寸增大而提高),此節省源自跳過RDO 搜尋,而本論文之貢獻在於在保有該節省的前提下,使大尺寸預測器之運算成本降低兩個數量級且無需 GPU。本論文並明確界定其適用邊界:僅替換 32 × 32 之LGBM 組態於部分序列(尤以 Campfire 與 FourPeople 為甚)出現顯著效能惡化,而同時替換 64 × 64 與 32 × 32 時此異常大幅緩解。整體而言,本論文系統性刻畫以輕量機器學習取代 CNN 於大尺寸 CU 決策之可行性與失效邊界,為延遲敏感、頻寬可容忍之應用開闢新的率失真複雜度操作點。

    This Thesis addresses the high computational complexity of the Versatile Video Coding (VVC/H.266) encoder caused by the exhaustive rate-distortion optimization (RDO) search for quad-tree plus nested multi-type tree (QT-MTT) partitioning, and investigates the feasibility and limits of replacing deep learning (DL) with lightweight machine learning (ML) for the partition decision of large coding units (CUs). This Thesis extends the puzzle-based fast CU partitioning (PFCP) framework proposed by Tai et al. (2025), keeping its overall skeleton unchanged and replacing only the binary (split/non-split) predictors responsible for the large CU sizes (64×64, 32×32)—originally lightweight convolutional neural networks (CNNs)— with Random Forest (RF) or LightGBM (LGBM), while the remaining smaller sizes retain their original CNNs, forming a hybrid ML–CNN architecture. To systematically examine the effect of each substitution point, this Thesis designs 72 ablation experiments spanning three mutually exclusive substitution strategies, two machine-learning models, three test sequences of different texture complexity, and four quantization parameters (QPs), using both the VVC Test Model (VTM) 17.2 and the original DL framework as dual comparison anchors.
    Experimental results show that the CNN inference time on large CUs is the main bottleneck of the entire pipeline and can be effectively eliminated by lightweight ML: on the 64×64 block, the inference speed-up of RF and LGBM relative to the CNN reaches about 90 to 175×, running entirely on a central processing unit (CPU) without a graphics processing unit (GPU). Prediction accuracy rises with CU size, already approaching that of the CNN on the large sizes (32 × 32, 64 × 64) but being clearly insufficient on the small sizes—the basis for adopting ML only on the large CUs while retaining the CNN on the small ones. In terms of coding performance, the Bjontegaard delta rate (BD-rate) increase of the best substitution configurations relative to the original DL framework is only about 1% to 8%. Measured end to end against unmodified VTM 17.2, the recommended configuration (LGBM both) reduces the single-frame encoding time by about 49% to 58%, the saving growing with frame size; this saving originates from skipping the RDO search, and the contribution of this Thesis is that it is preserved while the large-CU predictors become two orders of magnitude cheaper and require no GPU. This Thesis further delineates the applicability boundary: the LGBM configuration that replaces only 32 × 32 suffers a marked performance degradation on some sequences (most notably Campfire and FourPeople), whereas replacing both 64 × 64 and 32×32 substantially mitigates this anomaly. Overall, this Thesis systematically characterizes the feasibility and failure boundary of replacing the CNN with lightweight machine learning for large-CU decisions, opening a new rate-distortion-complexity operating point for latencysensitive, bandwidth-tolerant applications.

    中文摘要 i Abstract iii Acknowledgements v Contents vi List of Tables viii List of Figures x List of Symbols xi 1 Introduction 1 1.1 Overview of Video Compression 1 1.2 The Computational Complexity Problem of VVC 3 1.3 Fast Partition Prediction Based on Deep Learning 4 1.4 Organization of This Thesis 6 2 Related Works 8 2.1 Introduction to H.266/VVC 8 2.1.1 Coding Configurations of VVC 9 2.1.2 Frame Types and the Coding Tree Unit (CTU) 10 2.1.3 Coding Unit (CU) and QT-MTT Partitioning 10 2.1.4 Rate-Distortion Optimization (RDO) 12 2.1.5 Intra Prediction 13 2.2 Learning-Based Fast Partition Methods for VVC 14 2.2.1 Deep-Learning-Based Approaches 14 2.2.2 Random Forest 17 2.2.3 LightGBM 18 3 The Proposed Method 21 3.1 Dataset 21 3.2 VTM Platform and Encoding Configuration 22 3.3 Proposed Hybrid ML–CNN Architecture 24 3.4 Machine Learning Model Training and Hyperparameter Tuning 27 3.4.1 Hyperparameter Tuning of Random Forest 28 3.4.2 Hyperparameter and Threshold Tuning of LightGBM 31 3.4.3 Accuracy Comparison of the Two Models 33 3.5 Block-Size Substitution Strategy Design 34 4 Performance Evaluation 36 4.1 Test Environment 36 4.2 Inference Time Comparison: ML vs. DL 37 4.3 Impact of CU Size on ML Prediction Accuracy 40 4.4 Coding Performance Evaluation and Discussion 42 4.4.1 Campfire (Complex Texture) 44 4.4.2 BasketballDrive (Medium Complexity) 46 4.4.3 FourPeople (Simple Scene) 48 4.4.4 Cross-Sequence Summary 50 4.5 End-to-End Encoding Time 55 5 Conclusions 57 5.1 Conclusion 57 5.2 Future Work 60 References 61

    [1] T. Wiegand, G. J. Sullivan, G. Bjøntegaard, and A. Luthra, “Overview of the H.264/AVC video coding standard,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 13, no. 7, pp. 560–576, 2003.
    [2] G. J. Sullivan, J.-R. Ohm, W.-J. Han, and T. Wiegand, “Overview of the high efficiency video coding (HEVC) standard,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 22, no. 12, pp. 1649–1668, 2012.
    [3] B. Bross, Y.-K. Wang, Y. Ye, S. Liu, J. Chen, G. J. Sullivan, and J.-R. Ohm, “Overview of the versatile video coding (VVC) standard and its applications,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, no. 10, pp. 3736–3764, 2021.
    [4] M. Saldanha, G. Sanchez, C. Marcon, and L. Agostini, “Complexity analysis of VVC intra coding,” in IEEE International Conference on Image Processing (ICIP), 2020, pp. 3119–3123.
    [5] S.-C. Tai, C.-M. Yeh, and W. Huang, “Lightweight CNN-based puzzle algorithm for rapid QT-MTT partitioning in VVC,” Journal of Electronic Imaging, vol. 34, no. 6, p. 063009, 2025.
    [6] T. Li, M. Xu, R. Tang, Y. Chen, and Q. Xing, “DeepQTMT: A deep learning approach for fast QTMT-based CU partition of intra-mode VVC,” IEEE Transactions on Image Processing, vol. 30, pp. 5377–5390, 2021.
    [7] Y. Wang, P. Dai, J. Zhao, and Q. Zhang, “Fast CU partition decision algorithm for VVC intra coding using an MET-CNN,” Electronics, vol. 11, no. 19, p. 3090, 2022.
    [8] A. Feng, K. Liu, D. Liu, L. Li, and F. Wu, “Partition map prediction for fast block partitioning in VVC intra-frame coding,” IEEE Transactions on Image Processing, vol. 32, pp. 2237–2251, 2023.
    [9] J. Zhao, P. Dai, and Q. Zhang, “A complexity reduction method for VVC intra prediction based on statistical analysis and SAE-CNN,” Electronics, vol. 10, no. 24, p. 3112, 2021.
    [10] A. Tissier, W. Hamidouche, S. Belhadj Dit Mdalsi, J. Vanne, F. Galpin, and D. Menard, “Machine learning based efficient QT-MTT partitioning scheme for VVC intra encoders,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 8, pp. 4279–4293, 2023.
    [11] A. Tissier, W. Hamidouche, J. Vanne, F. Galpin, and D. Menard, “CNN oriented complexity reduction of VVC intra encoder,” in IEEE International Conference on Image Processing (ICIP), 2020, pp. 3139–3143.
    [12] S. Zhang, S. Feng, J. Chen, C. Zhou, and F. Yang, “A GCN-based fast CU partition method of intra-mode VVC,” Journal of Visual Communication and Image Representation, vol. 88, p. 103621, 2022.
    [13] Y. Huang, J. Yu, D. Wang, X. Lu, F. Dufaux, H. Guo, and C. Zhu, “Learning-based fast splitting and directional mode decision for VVC intra prediction,” IEEE Transactions on Broadcasting, vol. 70, no. 2, pp. 681–695, 2024.
    [14] B. Abdallah, F. Belghith, M. A. Ben Ayed, and N. Masmoudi, “Fast mode decision for CU coding based on CNN for the VVC standard,” Journal of Electronic Imaging, vol. 31, no. 5, p. 053004, 2022.
    [15] L. Breiman, “Random forests,” Machine Learning, vol. 45, no. 1, pp. 5–32, 2001.
    [16] Q. He, W. Wu, L. Luo, C. Zhu, and H. Guo, “Random forest based fast CU partition for VVC intra coding,” in IEEE International Symposium on Broadband Multimedia Systems and Broadcasting (BMSB), 2021, pp. 1–4.
    [17] J. Zhao, P. Li, and Q. Zhang, “A fast decision algorithm for VVC intra-coding based on texture feature and machine learning,” Computational Intelligence and Neuroscience, vol. 2022, p. 7675749, 2022.
    [18] S. Wen, G. Ding, and D. Ding, “Paired decision trees for fast intra decision in H.266/VVC,” Displays, vol. 80, p. 102545, 2023.
    [19] Q. Zhang, Y. Wang, L. Huang et al., “Fast CU partition decision for H.266/VVC based on the improved DAG-SVM classifier model,” Multimedia Systems, vol. 27, pp. 1–14, 2021.
    [20] L. Chen, B. Cheng, H. Zhu, H. Qin, L. Deng, and L. Luo, “Fast versatile video coding (VVC) intra coding for power-constrained applications,” Electronics, vol. 13, no. 11, p. 2150, 2024.
    [21] J. Li, S. Zhang, and F. Yang, “Random forest accelerated CU partition for inter prediction in H.266/VVC,” in IEEE International Conference on Multimedia and Expo (ICME), 2022, pp. 1–6.
    [22] A. Maraoui, I. Werda, S. Bouaafia, and F. E. Sayadi, “Machine learning-based approaches to reduce HEVC intra coding unit partition decision complexity,” Multimedia Tools and Applications, vol. 81, no. 2, pp. 2777–2802, 2022.
    [23] S. Bouaafia, R. Khemiri, F. E. Sayadi et al., “Fast CU partition-based machine learning approach for reducing HEVC complexity,” Journal of Real-Time Image Processing, vol. 17, pp. 185–196, 2020.
    [24] G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T.-Y. Liu, “LightGBM: A highly efficient gradient boosting decision tree,” in Advances in Neural Information Processing Systems (NeurIPS), 2017, pp. 3146–3154.
    [25] J. H. Friedman, “Greedy function approximation: A gradient boosting machine,” The Annals of Statistics, vol. 29, no. 5, pp. 1189–1232, 2001.
    [26] T. Chen and C. Guestrin, “XGBoost: A scalable tree boosting system,” in Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, 2016, pp. 785–794.
    [27] I. Taabane, D. Menard, A. Mansouri, and A. Ahaitouf, “Machine learning based fast QTMTT partitioning strategy for VVenC encoder in intra coding,” Electronics, vol. 12, no. 6, p. 1338, 2023.
    [28] S. Bakkouri, I. Bakkouri, and A. Elyousfi, “GBM-QTMT: Gradient boosting machinebased fast QTMT partition decision for VVC inter-coding,” Signal, Image and Video Processing, vol. 19, p. 173, 2025.
    [29] M. Xu, T. Li, Z. Wang, X. Deng, R. Yang, and Z. Guan, “Reducing complexity of HEVC: A deep learning approach,” IEEE Transactions on Image Processing, vol. 27, no. 10, pp. 5044–5059, 2018.
    [30] F. Bossen, J. Boyce, K. Sühring, X. Li, and V. Seregin, “JVET common test conditions and software reference configurations for SDR video,” Joint Video Experts Team (JVET) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11, Tech. Rep. JVET-J1010, 2018.
    [31] F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay, “Scikit-learn: Machine learning in Python,” Journal of Machine Learning Research, vol. 12, pp. 2825–2830, 2011.
    [32] G. Bjøntegaard, “Calculation of average PSNR differences between RD-curves,” ITU-T SG16 Q.6 Video Coding Experts Group (VCEG), Austin, TX, USA, Tech. Rep. VCEG-M33, 2001.

    QR CODE