| 研究生: |
林柏均 Lin, Bo-Jiun |
|---|---|
| 論文名稱: |
熵導引自適應視覺搜尋於高效率視覺語言模型推理 SEER: Entropy-Guided Adaptive Visual Search for Efficient Vision-Language Reasoning |
| 指導教授: |
賴槿峰
Lai, Chin-Feng |
| 學位類別: |
碩士 Master |
| 系所名稱: |
工學院 - 工程科學系 Department of Engineering Science |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 80 |
| 中文關鍵詞: | 視覺語言模型 、視覺搜尋 、推論期計算 、預測答案熵 、物件幻覺 |
| 外文關鍵詞: | Vision-Language Model, Visual Search, Test-Time Compute, Predictive Answer Entropy, Object Hallucination |
| 相關次數: | 點閱:63 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
針對視覺語言模型在處理高解析度影像時,面臨「辨識精細細節」與「控制運算成本」難以兩全的困境,本研究提出選擇性熵導引證據推理(Selective Entropy-guided Evidence Reasoning, SEER),一套無需訓練、全程確定性的推論期視覺搜尋框架,該框架核心構想是以模型自身的正規化預測答案熵作為單一內生訊號,統一驅動定位搜尋區域、設定終止時機與判定最終答案。本研究首先實證該熵值與答案正確性高度對齊且能追蹤搜尋進度,彌補了傳統樹搜尋依賴幾何獎勵與固定預算的不足。為將此訊號落實於各項決策,SEER 讓計算量隨題目難度按需分配,聯集聚焦在放大目標時保留物件空間關係,信任邊際隨骨幹能力自適應,最終由熵仲裁在多視角證據間保守選答。
本研究於 V*Bench 與 POPE 資料集上,採用兩個效能不同的 7B 骨幹模型進行評估。實驗結果顯示,相較於 MCTS 樹搜尋,SEER 減少了約七成模型呼叫次數,並將端對端延遲降低五成。在準確度上,SEER 協助較弱的骨幹模型分別提升 5.7 與 2.2 個百分點,並在較強模型上維持基準表現。此外,由於該熵本身即為可靠的信心指標,答案無需額外運算即可附帶信心分數,其可靠度較基準方法提升約 2.5 倍。這些結果顯示,將此一內生訊號落實為確定性的搜尋演算法,較擴大搜尋規模更能提升推論期視覺推理的效益。
This thesis proposes SEER (Selective Entropy-guided Evidence Reasoning), a training-free and deterministic framework that lets a large vision-language model (LVLM) answer fine-grained questions on high-resolution images by searching only where needed. The core idea is to drive all three decisions of inference-time visual search---where to look, when to stop, and which answer to adopt---with a single intrinsic signal: the model's normalized predictive answer entropy. This entropy aligns with answer correctness and tracks search progress, two properties that the driving signals of prior tree search lack. To realize this one signal across every decision, SEER allocates compute by difficulty, crops the union of relevant objects so magnification preserves their spatial relations, adapts its trust margin to backbone capability, and arbitrates conservatively by entropy across views.
On V*Bench and POPE, evaluated against a Monte Carlo tree search baseline with two 7B backbones of differing capability, SEER cuts LVLM calls by about 70% and roughly halves end-to-end latency, while accuracy rises by up to 5.7 points on the weaker backbone and holds on par on the stronger one. Because the driving signal is itself a confidence measure, every answer carries a reliable confidence score at no extra cost, improving risk--coverage quality by a factor of about 2.5. These results indicate that realizing this intrinsic signal as a deterministic search algorithm improves inference-time visual reasoning more than enlarging the search itself.
[1] Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In Advances in Neural Information Processing Systems (NeurIPS), volume 36, 2023.
[2] Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. Qwen2-vl: Enhancing vision-language model's perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024.
[3] Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. Scaling llm test-time compute optimally can be more effective than scaling parameters for reasoning. In International Conference on Learning Representations (ICLR), 2025.
[4] Penghao Wu and Saining Xie. V*: Guided visual search as a core mechanism in multimodal llms. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13084-13094, 2024.
[5] Geng Li, Jinglin Xu, Yunzhen Zhao, and Yuxin Peng. Dyfo: A training-free dynamic focus visual search for enhancing lmms in fine-grained visual understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 19541-19550, 2025.
[6] Cameron B. Browne, Edward Powley, Daniel Whitehouse, Simon M. Lucas, Peter I. Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyridon Samothrakis, and Simon Colton. A survey of monte carlo tree search methods. IEEE Transactions on Computational Intelligence and AI in Games, 4(1):1-43, 2012.
[7] Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating object hallucination in large vision-language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 292-305, 2023.
[8] Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 26296-26306, 2024.
[9] Zhang Li, Biao Yang, Qiang Liu, Zhiyin Ma, Shuo Zhang, Jingxu Yang, Yabo Sun, Yuliang Liu, and Xiang Bai. Monkey: Image resolution and text label are important things for large multi-modal models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 26763-26773, 2024.
[10] Sicong Leng, Hang Zhang, Guanzheng Chen, Xin Li, Shijian Lu, Chunyan Miao, and Lidong Bing. Mitigating object hallucinations in large vision-language models through visual contrastive decoding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13872-13882, 2024.
[11] Qidong Huang, Xiaoyi Dong, Pan Zhang, Bin Wang, Conghui He, Jiaqi Wang, Dahua Lin, Weiming Zhang, and Nenghai Yu. Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13418-13427, 2024.
[12] Haozhan Shen, Kangjia Zhao, Tiancheng Zhao, Ruochen Xu, Zilun Zhang, Mingwei Zhu, and Jianwei Yin. Zoomeye: Enhancing multimodal llms with human-like zooming capabilities through tree-based image exploration. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2025.
[13] Jiarui Zhang, Mahyar Khayatkhoei, Prateek Chhikara, and Filip Ilievski. Mllms know where to look: Training-free perception of small visual details with multimodal llms. In International Conference on Learning Representations (ICLR), 2025.
[14] Claude E. Shannon. A mathematical theory of communication. The Bell System Technical Journal, 27(3):379-423, 1948.
[15] Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221, 2022.
[16] Andrey Malinin and Mark Gales. Uncertainty estimation in autoregressive structured prediction. In International Conference on Learning Representations (ICLR), 2021.
[17] Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning (ICML), pages 1321-1330, 2017.
[18] Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. In International Conference on Learning Representations (ICLR), 2023.
[19] Yarin Gal and Zoubin Ghahramani. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In International Conference on Machine Learning (ICML), pages 1050-1059, 2016.
[20] Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. Simple and scalable predictive uncertainty estimation using deep ensembles. In Advances in Neural Information Processing Systems (NeurIPS), volume 30, 2017.
[21] Sanghwan Kim, Rui Xiao, Stephan Alaniz, Yongqin Xian, and Zeynep Akata. Training-free uncertainty guidance for complex visual tasks with mllms. arXiv preprint arXiv:2510.00705, 2025.
[22] Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. Tree of thoughts: Deliberate problem solving with large language models. In Advances in Neural Information Processing Systems (NeurIPS), volume 36, 2023.
[23] Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu. Reasoning with language model is planning with world model. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8154-8173, 2023.
[24] Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. In International Conference on Learning Representations (ICLR), 2023.
[25] Alex Graves. Adaptive computation time for recurrent neural networks. arXiv preprint arXiv:1603.08983, 2016.
[26] Surat Teerapittayanon, Bradley McDanel, and H. T. Kung. Branchynet: Fast inference via early exiting from deep neural networks. In Proceedings of the 23rd International Conference on Pattern Recognition (ICPR), pages 2464-2469, 2016.
[27] Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin. Deebert: Dynamic early exiting for accelerating bert inference. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 2246-2251, 2020.
[28] Yonatan Geifman and Ran El-Yaniv. Selective classification for deep neural networks. In Advances in Neural Information Processing Systems (NeurIPS), volume 30, 2017.
[29] Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 4015-4026, 2023.
[30] Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, and Lei Zhang. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. In European Conference on Computer Vision (ECCV), pages 38-55, 2024.
[31] Luca Medeiros. Language segment-anything. https://github.com/luca-medeiros/lang-segment-anything, 2023.
[32] Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. spacy: Industrial-strength natural language processing in python. https://spacy.io, 2020.
[33] Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP), pages 611-626, 2023.
[34] Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. GPT3.int8(): 8-bit matrix multiplication for transformers at scale. In Advances in Neural Information Processing Systems (NeurIPS), volume 35, 2022.