簡易檢索 / 詳目顯示

研究生: 蕭力文
Hsiao, Li-Wun
論文名稱: 超網路生成式 LoRA 適配 PaliGemma 視覺語言模型之純文字條件生成充分性與視覺條件化邊界
Sufficiency of Text-Only Conditioning and the Boundary of Visual Conditioning in Hypernetwork-Generated LoRA Adaptation for the PaliGemma Vision-Language Model
指導教授: 賴槿峰
Lai, Chin-Feng
學位類別: 碩士
Master
系所名稱: 工學院 - 工程科學系
Department of Engineering Science
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 77
中文關鍵詞: 視覺語言模型 、超網路 、低秩適配 、條件式參數生成 、零樣本任務適配
外文關鍵詞: Vision-Language Models, Hypernetworks, Low-Rank Adaptation, Conditional Parameter Generation, Zero-Shot Task Adaptation
相關次數: 點閱:108  下載:1 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 視 覺 語 言 模 型 在 面 對 不 同 下 游 任 務 時, 通 常 仍 需 透 過 Low-Rank Adaptation(LoRA)進行任務專用化,且新增任務往往需要重新蒐集資料與訓練。為降低逐任務訓練與參數管理成本,本研究以 PaliGemma-3B 為基礎,建立可根據任務條件動態生成 LoRA 參數的 Hypernetwork,將文字條件式 LoRA 生成方法延伸至包含視覺編碼器與語言模型的視覺語言模型。
    本研究以二十項視覺語言任務之任務專屬 LoRA 作為監督目標,透過參數重建方式訓練 Hypernetwork。為處理不同模型組件在深度、目標模組與參數維度上的差異,本研究引入組件-模組聯合嵌入、層級位置表示與跨組件輸出機制。此外,亦加入任務代表性影像,並比較向量串接與交叉注意力等視覺融合方式。
    實驗結果顯示,Hypernetwork 能生成具實際下游效用的 LoRA:於二十項訓練任務上,生成 LoRA 之未加權巨集平均為 0.403,高於其直接重建目標逐任務 joint LoRA 的 0.371,並明顯高於未適配基礎模型的 0.093。生成模式消融進一步顯示,主要效益來自語言模型側——僅生成語言模型側參數即可達到 0.401,與同時生成兩側的 0.403 相差 0.002;僅生成視覺編碼器側參數則僅得 0.189,明顯低於對應的逐任務 LoRA 之 0.407,顯示視覺編碼器側參數較難重建。在未見任務上,生成 LoRA 於病理影像 PathVQA 之封閉式問題由 0.095 提升至 0.595、於放射影像 SLAKE 由 0.058提升至 0.674,顯示其具有一定程度的零樣本適配能力。視覺條件化結果則顯示,當文字描述已足以區分任務時,加入影像未能帶來穩定增益(三種條件設定平均皆約0.41);但當所有任務共用同一句通用描述時,純文字條件降至 0.237,加入影像則維持 0.410,回復約 17.3 個百分點。綜合而言,本研究驗證了以 Hypernetwork 動態生成視覺語言模型 LoRA 的可行性,並釐清其跨組件生成、零樣本適配與視覺條件化的能力及限制。

    Vision-language models typically require task specialization through Low-Rank Adaptation (LoRA), and supporting a new task usually means collecting annotated data and running another round of training. To reduce these costs, this study builds a Hypernetwork on PaliGemma-3B that generates LoRA parameters directly from task conditions, extending text-conditioned LoRA generation to a model containing both a vision encoder and a language model.
    The Hypernetwork is trained by parameter reconstruction, using task-specific LoRA parameters from twenty vision-language tasks as supervision targets. To handle differences in model depth, target modules, and parameter dimensions across model components, it introduces component–module joint embeddings, shared layer-wise positional representations, and a variable-dimension output mechanism. Task-representative images are additionally examined as optional visual conditions.
    Across the twenty tasks, the generated LoRA parameters attain an unweighted macro average of 0.403, above their reconstruction target (0.371) and far above the unadapted base model (0.093). The gains come almost entirely from the language-model side: generating only language-model parameters reaches 0.401, whereas generating only vision-encoder parameters reaches 0.189. On unseen tasks, the generated parameters raise closed-form accuracy from 0.095 to 0.595 on PathVQA and from 0.058 to 0.674 on SLAKE. Representative images bring no consistent gain while task descriptions remain distinctive; but when all tasks share one generic description, the text-only condition falls to 0.237 while image conditioning holds at 0.410, indicating that visual conditions mainly supply task identity rather than improved adapter content. Overall, the study demonstrates the feasibility of generating cross-component VLM LoRA parameters from task conditions, and delimits the range over which that ability holds.

    摘要 i 英文延伸摘要 ii 誌謝 vii 目錄 viii 表目錄 x 圖目錄 xi 符號表 xii Chapter 1. 簡介 1 1.1. 研究動機 1 1.2. 研究目標 2 1.3. 本研究的貢獻 3 1.4. 章節提要 5 Chapter 2. 研究背景與相關文獻 6 2.1. 視覺語言模型的任務專用化與 LoRA 6 2.1.1. 視覺語言模型的任務專用化 6 2.1.2. Low-Rank Adaptation 6 2.1.3. 多任務 LoRA 與參數重用方法 7 2.2. Hypernetwork 與條件式參數生成 8 2.2.1. Hypernetwork 基本概念 8 2.2.2. Hypernetwork-based Parameter-Efficient Fine-Tuning 8 2.2.3. Text-to-LoRA 9 2.3. 視覺條件式參數生成 10 2.3.1. 視覺條件資訊的表示與聚合 10 2.3.2. 視覺條件式 Hypernetwork 與動態參數方法 10 2.4. 相關方法比較與研究定位 11 Chapter 3. 研究方法 14 3.1. 架構總覽 14 3.2. 跨組件 LoRA 生成機制 16 3.2.1. 跨模型組件的結構差異 16 3.2.2. 組件-模組聯合嵌入 17 3.2.3. 層級位置表示 17 3.2.4. 不同參數尺寸之輸出與截取 18 3.2.5. LoRA 矩陣之參數共享 18 3.3. 視覺條件化設計 19 3.3.1. 視覺特徵擷取與條件輸入 20 3.3.2. 向量串接 20 3.3.3. 交叉注意力融合 20 3.3.4. 逐圖片交叉注意力融合 21 3.4. 訓練目標 21 3.5. 推論流程 22 3.6. 演算法 23 Chapter 4. 實驗結果與討論 28 4.1. 實驗設計 28 4.1.1. 硬體與軟體規格 28 4.1.2. 資料集分類 28 4.1.3. 資料切分 30 4.1.4. 評估協定 30 4.1.5. 比較基準 31 4.1.6. 訓練與生成之超參數設置 31 4.2. 實驗結果與探討 33 4.2.1. Hypernetwork 生成之 LoRA 與 oracle LoRA 比較 33 4.2.2. 依評估指標型態之穩健性檢查 35 4.2.3. 生成模式與組件貢獻消融 36 4.2.4. 視覺條件化之效益與邊界 38 4.3. 未見任務與外部資料集泛化 45 4.3.1. 外部資料集 45 4.4. 運算與儲存成本分析 47 4.4.1. 參數量與儲存空間 47 4.4.2. 訓練時間與 GPU 記憶體 48 4.4.3. 新任務適配與推論成本 48 4.5. 完整逐任務結果 50 4.6. 限制 50 Chapter 5. 結論與未來展望 55 5.1. 結論 55 5.2. 未來展望 56 參考文獻 57

    [1] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in International Conference on Learning Representations (ICLR), 2022.
    [2] R. Charakorn, E. Cetin, Y. Tang, and R. T. Lange, “Text-to-LoRA: Instant transformer adaption,” in International Conference on Machine Learning (ICML), 2025.
    [3] J. Liang, W. Huang, X. Guo, G. Wan, B. Du, and M. Ye, “ThanoRA: Task heterogeneity-aware multi-task low-rank adaptation,” arXiv preprint arXiv:2505.18640, 2025.
    [4] C. Huang, Q. Liu, B. Y. Lin, T. Pang, C. Du, and M. Lin, “Lorahub: Efficient cross-task generalization via dynamic LoRA composition,” in Conference on Language Modeling (COLM), 2024.
    [5] Y. Gou, Z. Liu, K. Chen, L. Hong, H. Xu, A. Li, D.-Y. Yeung, J. T. Kwok, and Y. Zhang, “Mixture of cluster-conditional LoRA experts for vision-language instruction tuning,” IEEE Transactions on Image Processing, 2026.
    [6] X. Chen, H. Zhang, Y. Qiu, X. Liang, Z. Li, G. Wang, W. Li, T. Mo, H. K.-H. So, and N. Wong, “GuiLoMo: Allocating expert number and rank for lora-moe via bilevel optimization with GuidedSelection vectors,” in Findings of the Association for Computational Linguistics: EMNLP 2025, 2025.
    [7] H. Ivison, A. Bhagia, Y. Wang, H. Hajishirzi, and M. Peters, “HINT: Hypernetwork instruction tuning for efficient zero- and few-shot generalisation,” in Annual Meeting of the Association for Computational Linguistics (ACL), 2023.
    [8] V. Akinwande, M. S. Norouzzadeh, D. Willmott, A. Bair, M. R. Ganesh, and J. Z. Kolter, “HyperCLIP: Adapting vision-language models with hypernetworks,” arXiv preprint arXiv:2412.16777, 2024.
    [9] X. Song, J. Cui, H. Zhang, J. Shi, J. Chen, C. Zhang, and Y.-G. Jiang, “LoRA of change: Learning to generate LoRA for the editing instruction from a single before-after image pair,” arXiv preprint arXiv:2411.19156, 2024.
    [10] H. Jo, H. Choi, M. Cho, and D. Min, “iConFormer: Dynamic parameter-efficient tuning with input-conditioned adaptation,” arXiv preprint arXiv:2409.02838, 2024.
    [11] Y. Wang, C. Xiong, Z. Qin, M. Zhang, K. Xiao, and Z. Li, “Hylovqa: Dynamic hypernetwork-generated low-rank adaptation for continual visual question answering,” in International Joint Conference on Artificial Intelligence (IJCAI), 2026.
    [12] L. Beyer, A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alabdulmohsin, M. Tschannen, E. Bugliarello, T. Unterthiner, D. Keysers, S. Koppula, F. Liu, A. Grycner, A. Gritsenko, N. Houlsby, M. Kumar, K. Rong, J. Eisenschlos, R. Kabra, M. Bauer, M. Bošnjak, X. Chen, M. Minderer, P. Voigtlaender, I. Bica, I. Balazevic, J. Puigcerver, P. Papalampidi, O. Henaff, X. Xiong, R. Soricut, J. Harmsen, and X. Zhai, “PaliGemma: A versatile 3B VLM for transfer,” arXiv preprint arXiv:2407.07726, 2024.
    [13] X. Zhang, Y. Zhang, D. Long, W. Xie, Z. Dai, J. Tang, H. Lin, B. Yang, P. Xie, F. Huang, M. Zhang, W. Li, and M. Zhang, “mGTE: Generalized long-context text representation and reranking models for multilingual text retrieval,” in Conference on Empirical Methods in Natural Language Processing (EMNLP): Industry Track, 2024.
    [14] A. Mishra, S. Shekhar, A. K. Singh, and A. Chakraborty, “OCR-VQA: Visual question answering by reading text in images,” in International Conference on Document Analysis and Recognition (ICDAR), 2019.
    [15] A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach, “Towards VQA models that can read,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
    [16] A. F. Biten, R. Tito, A. Mafla, L. Gomez, M. Rusiñol, E. Valveny, C. V. Jawahar, and D. Karatzas, “Scene text visual question answering,” in IEEE/CVF International Conference on Computer Vision (ICCV), 2019.
    [17] M. Mathew, D. Karatzas, and C. V. Jawahar, “DocVQA: A dataset for VQA on document images,” in IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2021.
    [18] R. Tanaka, K. Nishida, and S. Yoshida, “VisualMRC: Machine reading comprehension on document images,” in AAAI Conference on Artificial Intelligence (AAAI), 2021.
    [19] A. Masry, D. X. Long, J. Q. Tan, S. Joty, and E. Hoque, “ChartQA: A benchmark for question answering about charts with visual and logical reasoning,” in Findings of the Association for Computational Linguistics: ACL, 2022.
    [20] Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the V in VQA matter: Elevating the role of image understanding in visual question answering,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
    [21] D. A. Hudson and C. D. Manning, “GQA: A new dataset for real-world visual reasoning and compositional question answering,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
    [22] D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham, “VizWiz grand challenge: Answering visual questions from blind people,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
    [23] Y. Li, Y. Du, K. Zhou, J. Wang, W. X. Zhao, and J.-R. Wen, “Evaluating object hallucination in large vision-language models,” in Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023.
    [24] A. Suhr, S. Zhou, A. Zhang, I. Zhang, H. Bai, and Y. Artzi, “A corpus for reasoning about natural language grounded in photographs,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), 2019.
    [25] J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. Girshick, “CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
    [26] R. Zellers, Y. Bisk, A. Farhadi, and Y. Choi, “From recognition to cognition: Visual commonsense reasoning,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
    [27] P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao, “MathVista: Evaluating mathematical reasoning of foundation models in visual contexts,” in International Conference on Learning Representations (ICLR), 2024.
    [28] J. Andreas, M. Rohrbach, T. Darrell, and D. Klein, “Neural module networks,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
    [29] K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi, “OK-VQA: A visual question answering benchmark requiring external knowledge,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
    [30] D. Schwenk, A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi, “A-OKVQA: A benchmark for visual question answering using world knowledge,” in European Conference on Computer Vision (ECCV), 2022.
    [31] P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan, “Learn to explain: Multimodal reasoning via thought chains for science question answering,” in Advances in Neural Information Processing Systems (NeurIPS), 2022.
    [32] A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi, “A diagram is worth a dozen images,” in European Conference on Computer Vision (ECCV), 2016.
    [33] J. J. Lau, S. Gayen, A. B. Abacha, and D. Demner-Fushman, “A dataset of clinically generated visual questions and answers about radiology images,” Scientific Data, 2018.

    下載圖示
    校外:立即公開
    QR CODE