簡易檢索 / 詳目顯示

研究生: 秦沐恩
Chin, Mu-En
論文名稱: 結合語意引導之雙碼本感興趣區域生成式影像壓縮
Semantic-guided ROI-based Generative Image Compression with Dual Codebooks
指導教授: 楊家輝
Yang, Jar-Ferr
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 電腦與通信工程研究所
Institute of Computer & Communication Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 73
中文關鍵詞: 生成式影像壓縮向量量化感興趣區域雙碼本
外文關鍵詞: generative image compression, vector quantization, region of interest, dual codebooks
相關次數: 點閱:88下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 近年來,基於向量量化生成對抗網路的生成式影像壓縮方法,已能在極低位元率下達到良好的感知式重建效果。然而,現有的方法通常對整張影像使用單一共享碼本,難以根據不同區域之語意重要性有效分配表示能力,導致位元率使用效率受限。
    為解決上述問題,本研究提出一個結合語意引導與雙碼本之感興趣區域生程式影像壓縮架構。所提出的方法利用語意遮罩引導潛在特徵於感興趣區域與背景區域之間進行路由,使語意重要區域使用較大的碼本以保留更多細節,而背景區域則使用較小的碼本以降低位元率。此外本研究於重建損失中導入正規化加權機制,以提升感興趣區域之重建品質。
    在熵編碼部分,本研究進一步導入基於Transformer之熵模型進行條件機率估測,以建構量化索引序列之條件機率分佈,並結合算術編碼進行壓縮,以提升整體壓縮效率。實驗結果顯示,所提出之方法能在維持低位元壓縮率的同時,有效保留重要語意區域之重建品質,驗證語意導向表示能力分配餘生程式影像壓縮之可行性。

    Recent advances in generative image compression based on vector quantized generative adversarial networks have demonstrated promising perceptual reconstruction quality at extremely low bitrates. However, existing approaches typically employ a single shared codebook for all image regions, limiting representation allocation according to semantic importance and reducing bitrate utilization efficiency for regions of interest (ROIs).
    To address this issue, this thesis proposes a semantic-guided ROI-based generative image compression framework with dual codebooks. Semantic masks guide latent routing between ROI and background regions, allowing ROI regions to utilize a larger codebook for richer semantic representation while assigning a smaller codebook to background regions for bitrate reduction. A normalized weighted reconstruction loss is further introduced to improve ROI reconstruction quality.
    To improve compression efficiency, a Transformer-based entropy model with two-dimensional positional encoding estimates the conditional probability distribution of quantized index sequences for arithmetic coding. Experimental results demonstrate that the proposed framework maintains low-bitrate compression while effectively preserving ROI reconstruction quality, verifying the feasibility of semantic-aware representation allocation in generative image compression.

    摘要 II Abstract III 誌謝 IV Contents V List of Tables VIII List of Figures IX Chapter 1 Introduction 1 1.1 Research Background 1 1.2 Motivations 2 1.3 Thesis Organization 4 Chapter 2 Related Work 5 2.1 Traditional and Learned Image Compression 5 2.2 VQ-based Generative Image Compression 7 2.3 Semantic-aware and ROI-based Compression 10 2.4 Entropy Modeling for Learned Compression 13 Chapter 3 The Proposed Generative Image Compression Framework 16 3.1 Overview of the Proposed Framework 17 3.2 Semantic-guided ROI Generation 18 3.2.1 Similarity Map Generation 19 3.2.2 ROI Mask Construction 20 3.2.3 Latent Space ROI Mapping 22 3.3 Dual-Codebook VQGAN Architecture 23 3.3.1 Semantic-Guided Codebook Routing 24 3.3.2 Dual-Codebook Quantization 25 3.3.3 Unified Codebook Space Design 26 3.4 Transformer-based Entropy Modeling 28 3.4.1 Conventional Probability Models 29 3.4.2 GPT-like Transformer Entropy Model 30 3.4.3 Spatial Positional Encoding 32 3.5 Loss Functions and Training Strategies 35 3.5.1 Reconstruction Loss 35 3.5.2 Perceptual, Vector Quantization, and Adversarial Loss 36 3.5.3 Overall Training Objectives 37 3.5.4 Transformer Entropy Model Training 38 Chapter 4 Experiment Results 40 4.1 Environment Setup and Datasets 41 4.2 Training Details and Hyperparameter Settings 42 4.3 Evaluation Metrics 43 4.4 Quantitative Results 45 4.4.1 Comparison with Baseline VQGAN 45 4.4.2 Comparison of Entropy Modeling Strategies 47 4.4.3 Analysis of Codebook Utilization 49 4.5 Effect of Spatial Positional Encoding 51 4.6 Qualitative Results 53 Chapter 5 Conclusions 57 Chapter 6 Future Work 58 References 60

    [1] Wallace, G. K. (1991). The JPEG still picture compression standard. Communications of the ACM, 34(4), 30-44.
    [2] Taubman, D. S., Marcellin, M. W., & Rabbani, M. (2002). JPEG2000: Image compression fundamentals, standards and practice. Journal of Electronic Imaging, 11(2), 286-287.
    [3] Sullivan, G. J., Ohm, J. R., Han, W. J., & Wiegand, T. (2012). Overview of the high efficiency video coding (HEVC) standard. IEEE Transactions on circuits and systems for video technology, 22(12), 1649-1668.
    [4] Ahmed, N., Natarajan, T., & Rao, K. R. (1974). Discrete cosine transform. IEEE transactions on Computers, 100(1), 90-93.
    [5] LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. nature, 521(7553), 436-444.
    [6] Esser, P., Rombach, R., & Ommer, B. (2021). Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 12873-12883).
    [7] Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., ... & Bengio, Y. (2014). Generative adversarial nets. Advances in neural information processing systems, 27.
    [8] Zhang, J., Huang, J., Jin, S., & Lu, S. (2024). Vision-language models for vision tasks: A survey. IEEE transactions on pattern analysis and machine intelligence, 46(8), 5625-5644.
    [9] Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., ... & Sutskever, I. (2021, July). Learning transferable visual models from natural language supervision. In International conference on machine learning (pp. 8748-8763). PmLR.
    [10] Li, B., Weinberger, K. Q., Belongie, S., Koltun, V., & Ranftl, R. (2022). Language-driven semantic segmentation. arXiv preprint arXiv:2201.03546.
    [11] Lüddecke, T., & Ecker, A. (2022). Image segmentation using text and image prompts. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 7086-7096).
    [12] Witten, I. H., Neal, R. M., & Cleary, J. G. (1987). Arithmetic coding for data compression. Communications of the ACM, 30(6), 520-540.
    [13] Qian, Y., Lin, M., Sun, X., Tan, Z., & Jin, R. (2022). Entroformer: A transformer-based entropy model for learned image compression. arXiv preprint arXiv:2202.05492.
    [14] Linde, Y., Buzo, A., & Gray, R. (1980). An algorithm for vector quantizer design. IEEE Transactions on communications, 28(1), 84-95.
    [15] Van Den Oord, A., & Vinyals, O. (2017). Neural discrete representation learning. Advances in neural information processing systems, 30.
    [16] Bross, B., Wang, Y. K., Ye, Y., Liu, S., Chen, J., Sullivan, G. J., & Ohm, J. R. (2021). Overview of the versatile video coding (VVC) standard and its applications. IEEE Transactions on Circuits and Systems for Video Technology, 31(10), 3736-3764.
    [17] Ballé, J., Laparra, V., & Simoncelli, E. P. (2016). End-to-end optimized image compression. arXiv preprint arXiv:1611.01704.
    [18] Ballé, J., Minnen, D., Singh, S., Hwang, S. J., & Johnston, N. (2018). Variational image compression with a scale hyperprior. arXiv preprint arXiv:1802.01436.
    [19] Minnen, D., Ballé, J., & Toderici, G. D. (2018). Joint autoregressive and hierarchical priors for learned image compression. Advances in neural information processing systems, 31.
    [20] Samek, W., Montavon, G., Lapuschkin, S., Anders, C. J., & Müller, K. R. (2021). Explaining deep neural networks and beyond: A review of methods and applications. Proceedings of the IEEE, 109(3), 247-278.
    [21] Agustsson, E., Tschannen, M., Mentzer, F., Timofte, R., & Gool, L. V. (2019). Generative adversarial networks for extreme learned image compression. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 221-231).
    [22] Mentzer, F., Toderici, G. D., Tschannen, M., & Agustsson, E. (2020). High-fidelity generative image compression. Advances in neural information processing systems, 33, 11913-11924.
    [23] Kingma, D. P., & Welling, M. (2013). Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114.
    [24] Jia, Z., Li, J., Li, B., Li, H., & Lu, Y. (2024). Generative latent coding for ultra-low bitrate image compression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 26088-26098).
    [25] Mao, Q., Yang, T., Zhang, Y., Wang, Z., Wang, M., Wang, S., ... & Ma, S. (2024, March). Extreme image compression using fine-tuned vqgans. In 2024 Data Compression Conference (DCC) (pp. 203-212). IEEE.
    [26] Xue, N., Mao, Q., Wang, Z., Zhang, Y., & Ma, S. (2024, July). Unifying generation and compression: Ultra-low bitrate image coding via multi-stage transformer. In 2024 IEEE International Conference on Multimedia and Expo (ICME) (pp. 1-6). IEEE.
    [27] Li, A., Li, F., Liu, Y., Cong, R., Zhao, Y., & Bai, H. (2024). Once-for-all: Controllable generative image compression with dynamic granularity adaptation. arXiv preprint arXiv:2406.00758.
    [28] Han, S., & Vasconcelos, N. (2006, October). Image compression using object-based regions of interest. In 2006 International Conference on Image Processing (pp. 3097-3100). IEEE.
    [29] Chang, J., Zhang, J., Li, J., Wang, S., Mao, Q., Jia, C., ... & Gao, W. (2023). Semantic-aware visual decomposition for image coding. International Journal of Computer Vision, 131(9), 2333-2355.
    [30] Chang, J., Mao, Q., Zhao, Z., Wang, S., Wang, S., Zhu, H., & Ma, S. (2019, September). Layered conceptual image compression via deep semantic synthesis. In 2019 IEEE International Conference on Image Processing (ICIP) (pp. 694-698). IEEE.
    [31] Gregor, K., Besse, F., Jimenez Rezende, D., Danihelka, I., & Wierstra, D. (2016). Towards conceptual compression. Advances in neural information processing systems, 29.
    [32] Jin, J., Xia, F., Ding, F., Zhang, X., Liu, M., Zhao, Y., ... & Meng, L. (2025). Customizable ROI-based deep image compression. IEEE Transactions on Circuits and Systems for Video Technology.
    [33] He, D., Zheng, Y., Sun, B., Wang, Y., & Qin, H. (2021). Checkerboard context model for efficient learned image compression. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 14771-14780).
    [34] Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30.
    [35] Zhang, R., Isola, P., Efros, A. A., Shechtman, E., & Wang, O. (2018). The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 586-595).
    [36] Ding, K., Ma, K., Wang, S., & Simoncelli, E. P. (2020). Image quality assessment: Unifying structure and texture similarity. IEEE transactions on pattern analysis and machine intelligence, 44(5), 2567-2581.
    [37] Yu, F., Chen, H., Wang, X., Xian, W., Chen, Y., Liu, F., ... & Darrell, T. (2020). Bdd100k: A diverse driving dataset for heterogeneous multitask learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 2636-2645).
    [38] Warburg, F., Hauberg, S., Lopez-Antequera, M., Gargallo, P., Kuang, Y., & Civera, J. (2020). Mapillary street-level sequences: A dataset for lifelong place recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 2626-2635).

    下載圖示
    校外:立即公開
    QR CODE