| 研究生: |
許宸華 HSU, CHEN-HUA |
|---|---|
| 論文名稱: |
將視覺對比解碼蒸餾至單一模型以提升效能與抑制視覺幻覺 Distilling Visual Contrastive Decoding into a Single Model for Enhanced Inference and Hallucination Mitigation |
| 指導教授: |
賴槿峰
Lai, Chin-Feng |
| 學位類別: |
碩士 Master |
| 系所名稱: |
工學院 - 工程科學系 Department of Engineering Science |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 中文 |
| 論文頁數: | 110 |
| 中文關鍵詞: | 視覺語言模型 、視覺幻覺 、視覺對比解碼 、知識蒸餾 、Smooth-AW |
| 外文關鍵詞: | Vision-Language Model, Visual Hallucination, Visual Contrastive Decoding, Knowledge Distillation, Smooth-AW |
| 相關次數: | 點閱:51 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
本研究針對視覺語言模型於影像問答與描述任務中常見的視覺幻覺問題,提出離線視覺對比解碼蒸餾(VisualContrastiveDecoding Distillation, VCDD)框架。既有視覺對比解碼(VisualContrastive Decoding, VCD)可透過清晰影像與加噪影像之輸出分佈對比,降低模型依賴語言先驗而生成不存在物件或錯誤細節的情形;然而,Online VCD 在每次推論時皆需執行清晰與加噪影像的雙路前向傳播,增加延遲與計算成本,也使其較難部署於即時互動或資源受限場景。
本研究目的如下:
1. 將VCD於推論階段產生的抗幻覺訊號內化至LoRA適配器,使模型於部署時不需執行清晰/加噪影像的雙路前向傳播。
2. 設計兼顧token-level 視覺校正、訓練穩定性與原始能力保存的蒸餾損失,以提升VCD訊號之學習效果。
3. 透過多項量化基準、開放式定性案例與推論效率測試,驗證VCDD在生成可靠性及部署成本上的實際效益。
為達成上述目的,本研究將VCD的token-level對比分佈轉化為蒸餾目標,並內化至LLaVA-1.5-7B 的 LoRA 適配器中。訓練流程採用TeacherForcing與三次前向傳播:先以停用LoRA的凍結基底模型分別取得加噪與清晰影像下的教師分佈,再透過VCD對比、自適應合理性過濾與溫度Softmax產生軟標籤,最後啟用LoRA進行學生模型更新。為提升訓練穩定性,本研究進一步提出Smooth-AW平滑一致性加權損失,利用全局KL均值與per-tokenKL散度之插值調整各答案位置的蒸餾強度,並加入抗遺忘正則化以限制學生模型過度偏離基底模型分佈。
實驗於POPE、MME 與LLaVA-Bench 上評估模型之物件幻覺抑制、綜合感知能力與開放式生成品質。結果顯示,VCDDSmooth-AW在POPE九個物件幻覺設定中的F1均優於原始Baseline,Accuracy 則於八個設定中提升;九個設定的平均Accuracy 由 0.8079 提升至 0.8120,增加 0.41 個百分點,相對提升 0.50%;平均F1則由0.8032 提升至0.8091,增加 0.58 個百分點,相對提升0.73%。部分指標亦超越OnlineVCD,例如AOKVQARandom的Accuracy/F1 達0.8577/0.8512,高於Online VCD 的 0.8550/0.8507。在 MME 中,總分由 Baseline 的 1560.12 提升至1575.89,並於Existence、Position、Landmark、OCR 與 Numerical 等子項目超越Online VCD。LLaVA-Bench 三次訓練 run 的平均分數皆高於 Baseline,Composite平均分數由4.050 提升至 4.219。定性案例亦呈現具體亮點:在電影場景辨識中,VCDD 正確連結至Titanic,而 Baseline 與 Online VCD 均誤判為 Pirates of the Caribbean;在條件式計數案例中,VCDD正確排除已切開水果並回答三個,前兩者則皆回答四個;在藝術風格描述中,VCDD能保留MonaLisa與Renaissance等關鍵視覺語意,同時減少不存在的背景物件。
效率方面,在3,000題POPE短答案評測中,VCDD的總推論時間為7分05秒,與Baseline 相同,Online VCD則需13分27秒;VCDD的吞吐量為7.05q/s,約為Online VCD之1.9倍。雖然VCDD尚未在所有平均指標上全面超越OnlineVCD,但其能以標準單路推論取得接近教師、並在部分量化指標與定性案例中更佳的結果,證明VCD抗幻覺訊號可被離線蒸餾至輕量化適配器,提供兼顧生成可靠性與部署效率的實用方向。
Vision-language models may generate visually unsupported objects, attributes, counts, or scene details. Visual Contrastive Decoding (VCD) mitigates these hallucinations by contrasting clean-image and noisy-image distributions, but Online VCD requires dual-path inference. This thesis proposes Visual Contrastive Decoding Distillation (VCDD), which transfers VCD's correction behavior into a lightweight LoRA adapter. The frozen LLaVA-1.5-7B model produces clean and noisy logits, which are converted into soft targets through VCD, plausibility filtering, and temperature scaling. The student is trained with teacher forcing, Smooth Agreement Weighting (Smooth-AW), and an anti-forgetting regularizer.
Experiments on POPE, MME, and LLaVA-Bench demonstrate both effectiveness and efficiency. VCDD improves F1 over the Baseline in all nine POPE settings and Accuracy in eight. It also surpasses Online VCD on selected metrics: AOKVQA Random reaches 0.8577/0.8512 in Accuracy/F1, compared with 0.8550/0.8507 for Online VCD. The MME total rises from 1560.12 to 1575.89, while Existence, Position, Landmark, OCR, and Numerical exceed Online VCD. The average LLaVA-Bench Composite score increases from 4.050 to 4.219. Qualitative cases show correct extit{Titanic} recognition, conditional counting of three uncut fruits, and fewer unsupported background objects, whereas Baseline and Online VCD fail in these examples. For 3,000 POPE questions, VCDD completes inference in 7:05 versus 13:27 for Online VCD, providing approximately 1.9$ imes$ throughput. Although it does not dominate Online VCD overall, VCDD achieves standard single-path inference with competitive and occasionally superior results.
[1] Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In Advances in Neural Information Processing Systems, volume 36, pages 34892–34916, 2023. DOI: https://doi.org/10.52202/075280-1516.
[2] Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26296–26306, 2024. DOI: https://doi.org/10.1109/CVPR52733.2024.02484.
[3] Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating object hallucination in large vision-language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 292–305, 2023. DOI: https://doi.org/10.18653/v1/2023.emnlp-main.20.
[4] Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, Yunsheng Wu, Rongrong Ji, Caifeng Shan, and Ran He. MME: A comprehensive evaluation benchmark for multimodal large language models. arXiv preprint arXiv:2306.13394, 2023. DOI: https://doi.org/10.48550/arXiv.2306.13394.
[5] Sicong Leng, Hang Zhang, Guanzheng Chen, Xin Li, Shijian Lu, Chunyan Miao, and Lidong Bing. Mitigating object hallucinations in large vision-language models through visual contrastive decoding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13872–13882, 2024. DOI: https://doi.org/10.1109/CVPR52733.2024.01316.
[6] Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015. DOI: https://doi.org/10.48550/arXiv.1503.02531.
[7] Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2022. DOI: https://doi.org/10.48550/arXiv.2106.09685.
[8] Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikołaj Bińkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. Flamingo: A visual language model for few-shot learning. In Advances in Neural Information Processing Systems, volume 35, pages 23716–23736, 2022. DOI: https://doi.org/10.52202/068431-1723.
[9] Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International Conference on Machine Learning, pages 19730–19742, 2023. DOI: https://doi.org/10.48550/arXiv.2301.12597.
[10] Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. InstructBLIP: Towards general-purpose vision-language models with instruction tuning. In Advances in Neural Information Processing Systems, volume 36, pages 49250–49267, 2023. DOI: https://doi.org/10.52202/075280-2142.
[11] Xintong Wang, Jingheng Pan, Liang Ding, and Chris Biemann. Mitigating hallucinations in large vision-language models with instruction contrastive decoding. In Findings of the Association for Computational Linguistics: ACL 2024, pages 15840–15853, 2024. DOI: https://doi.org/10.18653/v1/2024.findings-acl.937.
[12] Avshalom Manevich and Reut Tsarfaty. Mitigating hallucinations in large vision-language models (LVLMs) via language-contrastive decoding (LCD). In Findings of the Association for Computational Linguistics: ACL 2024, pages 6008–6022, 2024. DOI: https://doi.org/10.18653/v1/2024.findings-acl.359.
[13] Laura Fieback, Nishilkumar Balar, Jakob Spiegelberg, and Hanno Gottschalk. Efficient contrastive decoding with probabilistic hallucination detection. arXiv preprint arXiv:2504.12137, 2025. DOI: https://doi.org/10.48550/arXiv.2504.12137.
[14] Junho Kim, Hyunjun Kim, Yeonju Kim, and Yong Man Ro. CODE: Contrasting self-generated description to combat hallucination in large multi-modal models. In Advances in Neural Information Processing Systems, volume 37, pages 133571–133599, 2024. DOI: https://doi.org/10.52202/079017-4246.
[15] Yeji Park, Deokyeong Lee, Junsuk Choe, and Buru Chang. ConVis: Contrastive decoding with hallucination visualization for mitigating hallucinations in multimodal large language models. Proceedings of the AAAI Conference on Artificial Intelligence, 39(6):6434–6442, 2025. DOI: https://doi.org/10.1609/aaai.v39i6.32689.
[16] Zhehan Kan, Ce Zhang, Zihan Liao, Yapeng Tian, Wenming Yang, Junyuan Xiao, Xu Li, Dongmei Jiang, Yaowei Wang, and Qingmin Liao. CATCH: Complementary adaptive token-level contrastive decoding to mitigate hallucinations in LVLMs. arXiv preprint arXiv:2411.12713, 2024. DOI: https://doi.org/10.48550/arXiv.2411.12713.
[17] Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. FitNets: Hints for thin deep nets. arXiv preprint arXiv:1412.6550, 2014. DOI: https://doi.org/10.48550/arXiv.1412.6550.
[18] Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer learning for NLP. In International Conference on Machine Learning, pages 2790–2799, 2019. DOI: https://doi.org/10.48550/arXiv.1902.00751.
[19] Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. QLoRA: Efficient finetuning of quantized LLMs. In Advances in Neural Information Processing Systems, volume 36, pages 10088–10115, 2023. DOI: https://doi.org/10.52202/075280-0441.
[20] Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, volume 33, pages 6840–6851, 2020. DOI: https://doi.org/10.48550/arXiv.2006.11239.
[21] Ronald J. Williams and David Zipser. A learning algorithm for continually running fully recurrent neural networks. Neural Computation, 1(2):270–280, 1989. DOI: https://doi.org/10.1162/neco.1989.1.2.270.
[22] Michael McCloskey and Neal J. Cohen. Catastrophic interference in connectionist networks: The sequential learning problem. In Psychology of Learning and Motivation, volume 24, pages 109–165. Academic Press, 1989. DOI: https://doi.org/10.1016/S0079-7421(08)60536-8.
[23] Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. Microsoft COCO: Common objects in context. In European Conference on Computer Vision, pages 740–755. Springer, 2014. DOI: https://doi.org/10.1007/978-3-319-10602-1_48.
[24] Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. A-OKVQA: A benchmark for visual question answering using world knowledge. In European Conference on Computer Vision, pages 146–162. Springer, 2022. DOI: https://doi.org/10.1007/978-3-031-20074-8_9.
[25] Drew A. Hudson and Christopher D. Manning. GQA: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6700–6709, 2019. DOI: https://doi.org/10.1109/CVPR.2019.00686.