簡易檢索 / 詳目顯示

研究生: 林柏均
Lin, Bo-Jiun
論文名稱: 熵導引自適應視覺搜尋於高效率視覺語言模型推理
SEER: Entropy-Guided Adaptive Visual Search for Efficient Vision-Language Reasoning
指導教授: 賴槿峰
Lai, Chin-Feng
學位類別: 碩士
Master
系所名稱: 工學院 - 工程科學系
Department of Engineering Science
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 80
中文關鍵詞: 視覺語言模型視覺搜尋推論期計算預測答案熵物件幻覺
外文關鍵詞: Vision-Language Model, Visual Search, Test-Time Compute, Predictive Answer Entropy, Object Hallucination
相關次數: 點閱:63下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 針對視覺語言模型在處理高解析度影像時,面臨「辨識精細細節」與「控制運算成本」難以兩全的困境,本研究提出選擇性熵導引證據推理(Selective Entropy-guided Evidence Reasoning, SEER),一套無需訓練、全程確定性的推論期視覺搜尋框架,該框架核心構想是以模型自身的正規化預測答案熵作為單一內生訊號,統一驅動定位搜尋區域、設定終止時機與判定最終答案。本研究首先實證該熵值與答案正確性高度對齊且能追蹤搜尋進度,彌補了傳統樹搜尋依賴幾何獎勵與固定預算的不足。為將此訊號落實於各項決策,SEER 讓計算量隨題目難度按需分配,聯集聚焦在放大目標時保留物件空間關係,信任邊際隨骨幹能力自適應,最終由熵仲裁在多視角證據間保守選答。

    本研究於 V*Bench 與 POPE 資料集上,採用兩個效能不同的 7B 骨幹模型進行評估。實驗結果顯示,相較於 MCTS 樹搜尋,SEER 減少了約七成模型呼叫次數,並將端對端延遲降低五成。在準確度上,SEER 協助較弱的骨幹模型分別提升 5.7 與 2.2 個百分點,並在較強模型上維持基準表現。此外,由於該熵本身即為可靠的信心指標,答案無需額外運算即可附帶信心分數,其可靠度較基準方法提升約 2.5 倍。這些結果顯示,將此一內生訊號落實為確定性的搜尋演算法,較擴大搜尋規模更能提升推論期視覺推理的效益。

    This thesis proposes SEER (Selective Entropy-guided Evidence Reasoning), a training-free and deterministic framework that lets a large vision-language model (LVLM) answer fine-grained questions on high-resolution images by searching only where needed. The core idea is to drive all three decisions of inference-time visual search---where to look, when to stop, and which answer to adopt---with a single intrinsic signal: the model's normalized predictive answer entropy. This entropy aligns with answer correctness and tracks search progress, two properties that the driving signals of prior tree search lack. To realize this one signal across every decision, SEER allocates compute by difficulty, crops the union of relevant objects so magnification preserves their spatial relations, adapts its trust margin to backbone capability, and arbitrates conservatively by entropy across views.

    On V*Bench and POPE, evaluated against a Monte Carlo tree search baseline with two 7B backbones of differing capability, SEER cuts LVLM calls by about 70% and roughly halves end-to-end latency, while accuracy rises by up to 5.7 points on the weaker backbone and holds on par on the stronger one. Because the driving signal is itself a confidence measure, every answer carries a reliable confidence score at no extra cost, improving risk--coverage quality by a factor of about 2.5. These results indicate that realizing this intrinsic signal as a deterministic search algorithm improves inference-time visual reasoning more than enlarging the search itself.

    中文摘要 I Abstract II 誌謝 VIII 目錄 IX 表目錄 XII 圖目錄 XIII 符號說明 XIV 第一章 簡介 1 1-1. 研究動機 1 1-1.1 高解析度視覺理解的核心兩難 1 1-1.2 推論期計算與視覺搜尋模式 1 1-1.3 既有樹搜尋方法的結構性侷限 2 1-2. 研究目的 3 1-3. 研究貢獻 4 第二章 相關文獻 6 2-1. 視覺語言模型與高解析度輸入之限制 6 2-2. 推論期視覺搜尋 7 2-3. 不確定性量化與信心校準 8 2-4. 大型語言模型推理中的樹搜尋 10 2-5. 自適應計算與推論期計算擴展 11 2-6. 選擇性預測與拒答機制 11 2-7. 視覺專家:開放詞彙偵測與分割 12 2-8. 研究定位 12 第三章 研究方法 14 3-1. 搜尋驅動訊號:預測答案熵 14 3-1.1 問題設定與符號定義 14 3-1.2 預測答案熵 15 3-1.3 訊號性質與三項決策之參數化 15 3-2. 搜尋與停止機制 17 3-2.1 早停:計算預算的按需配置 18 3-2.2 串級聚焦:語意化的逐層蒐證 18 3-2.3 自適應信任邊際:閾值隨模型自動適配 19 3-3. 選擇策略:以熵為核心的證據仲裁 20 3-4. SEER 演算法 22 3-4.1 系統架構與職責切分 22 3-4.2 失敗處理與跨任務自適應 23 3-4.3 演算法設計 24 第四章 實驗評估:實驗設計、實驗結果與探討 26 4-1. 實驗設計 26 4-1.1 評估基準與資料集 26 4-1.2 比較方法與骨幹模型 27 4-1.3 評估指標 28 4-1.4 實作細節與實驗環境 30 4-2. 熵訊號的有效性驗證 30 4-3. 實驗結果與探討 34 4-3.1 整體效能:Accuracy 與推論成本 35 4-3.2 POPE 逐分割分析與標準指標 38 4-3.3 計算行為分析:自適應分配與等預算比較 40 4-3.4 信心品質與選擇性預測 42 4-3.5 題型層級的增益分析 43 4-3.6 消融與敏感度分析 45 4-3.7 綜合探討 47 第五章 研究結論 50 5-1. 研究總結 50 5-2. 主要研究發現 50 5-3. 研究貢獻 52 第六章 未來工作 54 6-1. 不確定性訊號之推廣:開放式生成與訊號融合 54 6-2. 多專家協作與任務先驗之資料驅動化 54 6-3. 跨任務、跨模態與更深之搜尋結構 55 6-4. 搜尋軌跡之再利用與部署優化 56 參考文獻 57

    [1] Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In Advances in Neural Information Processing Systems (NeurIPS), volume 36, 2023.
    [2] Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. Qwen2-vl: Enhancing vision-language model's perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024.
    [3] Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. Scaling llm test-time compute optimally can be more effective than scaling parameters for reasoning. In International Conference on Learning Representations (ICLR), 2025.
    [4] Penghao Wu and Saining Xie. V*: Guided visual search as a core mechanism in multimodal llms. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13084-13094, 2024.
    [5] Geng Li, Jinglin Xu, Yunzhen Zhao, and Yuxin Peng. Dyfo: A training-free dynamic focus visual search for enhancing lmms in fine-grained visual understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 19541-19550, 2025.
    [6] Cameron B. Browne, Edward Powley, Daniel Whitehouse, Simon M. Lucas, Peter I. Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyridon Samothrakis, and Simon Colton. A survey of monte carlo tree search methods. IEEE Transactions on Computational Intelligence and AI in Games, 4(1):1-43, 2012.
    [7] Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating object hallucination in large vision-language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 292-305, 2023.
    [8] Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 26296-26306, 2024.
    [9] Zhang Li, Biao Yang, Qiang Liu, Zhiyin Ma, Shuo Zhang, Jingxu Yang, Yabo Sun, Yuliang Liu, and Xiang Bai. Monkey: Image resolution and text label are important things for large multi-modal models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 26763-26773, 2024.
    [10] Sicong Leng, Hang Zhang, Guanzheng Chen, Xin Li, Shijian Lu, Chunyan Miao, and Lidong Bing. Mitigating object hallucinations in large vision-language models through visual contrastive decoding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13872-13882, 2024.
    [11] Qidong Huang, Xiaoyi Dong, Pan Zhang, Bin Wang, Conghui He, Jiaqi Wang, Dahua Lin, Weiming Zhang, and Nenghai Yu. Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13418-13427, 2024.
    [12] Haozhan Shen, Kangjia Zhao, Tiancheng Zhao, Ruochen Xu, Zilun Zhang, Mingwei Zhu, and Jianwei Yin. Zoomeye: Enhancing multimodal llms with human-like zooming capabilities through tree-based image exploration. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2025.
    [13] Jiarui Zhang, Mahyar Khayatkhoei, Prateek Chhikara, and Filip Ilievski. Mllms know where to look: Training-free perception of small visual details with multimodal llms. In International Conference on Learning Representations (ICLR), 2025.
    [14] Claude E. Shannon. A mathematical theory of communication. The Bell System Technical Journal, 27(3):379-423, 1948.
    [15] Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221, 2022.
    [16] Andrey Malinin and Mark Gales. Uncertainty estimation in autoregressive structured prediction. In International Conference on Learning Representations (ICLR), 2021.
    [17] Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning (ICML), pages 1321-1330, 2017.
    [18] Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. In International Conference on Learning Representations (ICLR), 2023.
    [19] Yarin Gal and Zoubin Ghahramani. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In International Conference on Machine Learning (ICML), pages 1050-1059, 2016.
    [20] Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. Simple and scalable predictive uncertainty estimation using deep ensembles. In Advances in Neural Information Processing Systems (NeurIPS), volume 30, 2017.
    [21] Sanghwan Kim, Rui Xiao, Stephan Alaniz, Yongqin Xian, and Zeynep Akata. Training-free uncertainty guidance for complex visual tasks with mllms. arXiv preprint arXiv:2510.00705, 2025.
    [22] Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. Tree of thoughts: Deliberate problem solving with large language models. In Advances in Neural Information Processing Systems (NeurIPS), volume 36, 2023.
    [23] Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu. Reasoning with language model is planning with world model. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8154-8173, 2023.
    [24] Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. In International Conference on Learning Representations (ICLR), 2023.
    [25] Alex Graves. Adaptive computation time for recurrent neural networks. arXiv preprint arXiv:1603.08983, 2016.
    [26] Surat Teerapittayanon, Bradley McDanel, and H. T. Kung. Branchynet: Fast inference via early exiting from deep neural networks. In Proceedings of the 23rd International Conference on Pattern Recognition (ICPR), pages 2464-2469, 2016.
    [27] Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin. Deebert: Dynamic early exiting for accelerating bert inference. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 2246-2251, 2020.
    [28] Yonatan Geifman and Ran El-Yaniv. Selective classification for deep neural networks. In Advances in Neural Information Processing Systems (NeurIPS), volume 30, 2017.
    [29] Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 4015-4026, 2023.
    [30] Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, and Lei Zhang. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. In European Conference on Computer Vision (ECCV), pages 38-55, 2024.
    [31] Luca Medeiros. Language segment-anything. https://github.com/luca-medeiros/lang-segment-anything, 2023.
    [32] Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. spacy: Industrial-strength natural language processing in python. https://spacy.io, 2020.
    [33] Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP), pages 611-626, 2023.
    [34] Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. GPT3.int8(): 8-bit matrix multiplication for transformers at scale. In Advances in Neural Information Processing Systems (NeurIPS), volume 35, 2022.

    下載圖示
    校外:立即公開
    QR CODE