| 研究生: |
范瑄育 Fan, Hsuan-Yu |
|---|---|
| 論文名稱: |
應用模擬退火策略於視覺語言任務的提示語搜尋方法 A Simulated Annealing Strategy to Find Optimal Prompts for Vision-Language Tasks |
| 指導教授: |
朱威達
Chu, Wei-Ta |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 資訊工程學系 Department of Computer Science and Information Engineering |
| 論文出版年: | 2025 |
| 畢業學年度: | 113 |
| 語文別: | 英文 |
| 論文頁數: | 45 |
| 中文關鍵詞: | 大型語言模型 、提示學習 、影像分類 、最佳化演算法 |
| 外文關鍵詞: | Large language model, Prompt learning, Image classification, Optimization algorithm |
| 相關次數: | 點閱:107 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
視覺語言模型(Vision-Language Models, VLMs)能有效整合視覺與語言資訊,已廣泛應用於多模態人工智慧領域,成為近年研究關注的核心技術之一。然而在實務應用中,提示(prompt)的設計已成為影響模型表現的關鍵因素。由於 VLM 對提示語句極為敏感,即便只是輕微的措辭變化,也可能導致模型生成截然不同的結果。為此,本研究提出結合大型語言模型(Large Language Model, LLM)之方法,能自動生成多樣且高品質的提示語模板。為進一步提升模板的有效性與多樣性,我們引入模擬退火演算法策略,對生成的提示語進行系統性的篩選與語句改寫。透過此機制,不僅能有效避免現有方法中常見的同質性模板,亦有助於降低模型陷入局部最佳解的風險。實驗結果顯示,該策略可順利整合至既有的提示學習流程,並於十一個影像分類資料集上展現穩定且顯著的效能提升。
Vision-language models (VLMs) provide a new way to connect visual and text information, and how to prompt VLMs becomes the key to getting effective results. However, crafting effective prompt templates remains a non-trivial challenge due to the sensitivity of vision-language models to prompt variations. Even minor changes in prompt phrasing can lead to significantly different outputs. Given the difficulty of manual prompt design and the high sensitivity of VLMs to prompt phrasing, this work proposes to leverage a large language model (LLM) to automatically generate diverse and effective prompt templates. Associated with the designed simulated annealing optimization strategy, the generated templates are selected and revised in a systematic way so that 1) homogeneous prompt templates that often appear in current methods are avoided, and 2) the degree of trapping in local optimum is reduced. Evaluation on eleven image classification datasets demonstrates that the proposed strategy can be seamlessly integrated into existing prompt learning methods and consistently yields performance gains.
[1] Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. Flamingo: a visual language model for few-shot learning. In Proceedings of Annual Conference on Neural Information Processing Systems, 2022.
[2] Luke Bailey, Gustaf Ahdritz, Anat Kleiman, Siddharth Swaroop, Finale Doshi-Velez, and Weiwei Pan. Soft prompting might be a bug, not a feature. In Proceedings of ICML Workshop on Challenges in Deployable Generative AI, 2023.
[3] Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101– mining discriminative components with random forests. In Proceedings of European Conference on Computer Vision, 2014.
[4] Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. Describing textures in the wild. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2014.
[5] Yingjun Du, Wenfang Sun, and Cees G. M. Snoek. IPO: Interpretable prompt optimization for vision-language models. In Proceedings of Annual Conference on Neural Information Processing Systems, 2024.
[6] Li Fei-Fei, R. Fergus, and P. Perona. One-shot learning of object categories. IEEE Transactions on Pattern Analysis and Machine Intelligence, 28(4):594–611, 2006.
[7] Zixian Guo, Ming Liu, Zhilong Ji, Jinfeng Bai, Yiwen Guo, and Wangmeng Zuo. Two optimizers are better than one: LLM catalyst for enhancing gradient-based optimization. CoRR, abs/2405.19732, 2024.
[8] Patrick Helber, Benjamin Bischke, Andreas R. Dengel, and Damian Borth. EuroSAT: A novel dataset and deep learning benchmark for land use and land cover classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 12:2217–2226, 2017.
[9] Matthew Honnibal and Ines Montani. spacy 2: Natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing, 2017.
[10] Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In Proceedings of International Conference on Machine Learning, 2021.
[11] Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz, Salman Khan, and Fahad Shahbaz Khan. MaPLe: Multi-modal prompt learning. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023.
[12] Muhammad Uzair Khattak, Syed Talal Wasim, Muzammal Naseer, Salman Khan, Ming-Hsuan Yang, and Fahad Shahbaz Khan. Self-regulating prompts: Foundational model adaptation without forgetting. In Proceedings of IEEE/CVF International Conference on Computer Vision, 2023.
[13] Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3D object representations for fine-grained categorization. In Proceedings of the IEEE International Conference on Computer Vision Workshops, June 2013.
[14] Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. In Proceedings of Conference on Empirical Methods in Natural Language Processing, 2021.
[15] Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of International Conference on Machine Learning, 2023.
[16] Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of International Conference on Machine Learning, 2022.
[17] Zheng Li, Xiang Li, Xinyi Fu, Xin Zhang, Weiqiang Wang, Shuo Chen, and Jian Yang. PromptKD: Unsupervised prompt distillation for vision-language models. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024.
[18] Zhiqiu Lin, Samuel Yu, Zhiyi Kuang, Deepak Pathak, and Deva Ramanan. Multimodality helps unimodality: Cross-modal few-shot learning with multimodal models. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023.
[19] Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In Proceedings of Advances in Neural Information Processing Systems, 2023.
[20] Shihong Liu, Samuel Yu, Zhiqiu Lin, Deepak Pathak, and Deva Ramanan. Language models as black-box optimizers for vision-language models. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024.
[21] Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, and Lei Zhang. Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection. In Proceedings of European Conference on Computer Vision, 2024.
[22] Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew B. Blaschko, and Andrea Vedaldi. Fine-grained visual classification of aircraft. arXiv preprint arXiv:1306.5151, 2013.
[23] Sachit Menon and Carl Vondrick. Visual classification via description from large language models. In Proceedings of International Conference on Learning Representations, 2022.
[24] Maria-Elena Nilsback and Andrew Zisserman. Automated flower classification over a large number of classes. In Proceedings of the Indian Conference on Computer Vision, Graphics & Image Processing, 2008.
[25] Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman, and C. V. Jawahar. Cats and dogs. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2012.
[26] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of International Conference on Machine Learning, 2021.
[27] Shuvendu Roy and Ali Etemad. Consistency-guided prompt learning for vision-language models. In Proceedings of International Conference on Learning Representations, 2024.
[28] Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet large scale visual recognition challenge. International Journal of Computer Vision, 115:211–252, 2015.
[29] Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. LAION-400M: Open dataset of CLIP-filtered 400 million image-text pairs. In Proceedings of Data Centric AI NeurIPS Workshop, 2021.
[30] Khurram Soomro, Amir Zamir, and Mubarak Shah. UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402, 2012.
[31] Jianfeng Wang, Zhengyuan Yang, Xiaowei Hu, Linjie Li, Kevin Lin, Zhe Gan, Zicheng Liu, Ce Liu, and Lijuan Wang. GIT: A generative image-to-text transformer for vision and language. Transactions on Machine Learning Research, 2022.
[32] Xinyuan Wang, Chenxi Li, Zhen Wang, Fan Bai, Haotian Luo, Jiayou Zhang, Nebojsa Jojic, Eric P. Xing, and Zhiting Hu. PromptAgent: Strategic planning with language models enables expert-level prompt optimization. In Proceedings of International Conference on Learning Representations, 2023.
[33] Mitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li, Simon Kornblith, Rebecca Roelofs, Raphael Gontijo-Lopes, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, and Ludwig Schmidt. Robust fine-tuning of zero-shot models. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
[34] Jianxiong Xiao, Krista A. Ehinger, James Hays, Antonio Torralba, and Aude Oliva. SUN Database: Exploring a large collection of scene categories. International Journal of Computer Vision, 119:3–22, 2014.
[35] Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu. FILIP: Fine-grained interactive language-image pre-training. In Proceedings of International Conference on Learning Representations, 2022.
[36] Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, Ce Liu, Mengchen Liu, Zicheng Liu, Yumao Lu, Yu Shi, Lijuan Wang, Jianfeng Wang, Bin Xiao, Zhen Xiao, Jianwei Yang, Michael Zeng, Luowei Zhou, and Pengchuan Zhang. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432, 2021.
[37] Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer. LiT: Zero-shot transfer with locked-image text tuning. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
[38] Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Conditional prompt learning for vision-language models. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
[39] Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models. International Journal of Computer Vision, 130(9):2337–2348, 2022.
[40] Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. Large language models are human-level prompt engineers. In Proceedings of International Conference on Learning Representations, 2023.