簡易檢索 / 詳目顯示

研究生: 李振瑋
Lee, Jen-Wei
論文名稱: SemDINO:結合語意先驗與對抗訓練之強健性潛在空間浮水印框架
SemDINO: A Robust Latent Watermarking Framework Integrating Semantic Priors and Adversarial Training
指導教授: 許志仲
Hsu, Chih-Chung
鄭順林
Jeng, Shuen-Lin
學位類別: 碩士
Master
系所名稱: 管理學院 - 數據科學研究所
Institute of Data Science
論文出版年: 2026
畢業學年度: 114
語文別: 中文
論文頁數: 110
中文關鍵詞: 零樣本影像轉影片浮水印 、時序邏輯平均 、雙尺度頻域載體 、語意先驗 、對抗訓練 、強健性
外文關鍵詞: zero-shot image-to-video watermarking, temporal logit averaging, dual-scale frequency carrier, semantic prior, adversarial training, robustness
相關次數: 點閱:132  下載:2 
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 圖像到影片(Image-to-video, I2V)擴散模型能將任何靜態圖像動畫化為連貫的影片,因此,嵌入在圖像中的來源出處(provenance)資訊,現在必須在由其擁有者完全無法控制的生成器所合成的影片中留存下來。現有的解決方案存在不足之處:承載於精細空間細節中的浮水印在生成式重新合成下會崩潰;影片浮水印技術則假設防禦者就是影片的作者;而近期針對 I2V 所設計的方法,則是針對模擬的生成通道進行訓練,並部署輔助的解碼網路。先前沒有任何方法能夠解決嚴格的零樣本(zero-shot)情境,在此情境中,生成器是一個黑盒子,且嵌入器(embedder)和解碼器(decoder)都從未見過 I2V 通道。
    我們提出 SemDINO,一個穩健的潛在空間浮水印框架(latent watermarking framework),結合凍結的(frozen)DINOv3 語意先驗(semantic prior)與對抗式訓練;其最終形式由實證記錄所塑造,核心立場是:面對此通道,應以解析式的極簡分解取代可學習的複雜度。聯合訓練的潛在生成式嵌入器在冷啟動(cold start)時即陷入死結——我們記錄了這個失敗現象——因此該框架將嵌入器固定為一個純加性、影像域的解析式雙尺度區塊 DFT 載體(analytic dual-scale carrier,訓練與推論皆不含任何 VAE 或擴散骨幹網路),其頻帶依生成式低通通道的存活分析選定;此框架與潛在生成管線的連結在於威脅模型,而非嵌入器的實作機制。對抗式訓練僅保留為不斷升級的失真對抗方(escalating distortion adversary),並以解析式的雙側 PSNR 約束取代保真度對抗方(fidelity adversary);凍結的語意先驗則在嚴格的酬載下限(payload floor)之上閘控嵌入能量。
    本框架最具可遷移性的貢獻在解碼端:一個學習式萃取器(learned extractor)逐一解碼每個生成影格,並在單一次閾值判定(thresholding)之前,透過時序邏輯平均(Temporal Logit Averaging)融合各影格的軟性證據(soft per-frame evidence);此聚合機制在不作任何修改的情況下,同樣能提升 VINE-R 與以通道方式訓練之 LoT-Pass 的表現。載體端的關鍵發現則是振幅預算:完整雙尺度贏得靜態失真的權衡,但在真實 I2V 通道上,其巨集振幅受固定保真度預算壓制而弱於僅巨集變體——我們並以一次下限掃描驗證了提高巨集振幅門檻即可回收此差距。消融實驗(ablation)誠實歸因:酬載下限與序列層級的軟融合(sequence-level soft fusion)是決定性的穩健機制,而顯著性調變(saliency modulation)本身則未貢獻任何可量測的增益(其下限確有支撐作用,並保留為可解釋的控制介面)。SemDINO 以高保真度嵌入 128 位元、具競爭力的靜態穩健性、在轉碼與影格遺失下穩定的零樣本序列級萃取,並能在單張消費級 GPU 上於數小時內完成訓練。

    Image-to-video (I2V) diffusion models can animate any still photograph into a coherent clip, so provenance information embedded in an image must survive a generative re-synthesis performed later by a third-party generator that its owner neither controls nor observes. This thesis addresses the strict zero-shot setting, in which neither the embedder nor the decoder has ever seen the I2V channel. We propose SemDINO, a watermarking framework whose final form is governed by a single stance: for this channel, analytic minimalism should replace learnable complexity. A jointly trained latent generative embedder deadlocks at cold start—a failure we document—so the embedder is fixed to a purely additive, image-domain dual-scale block-DFT carrier containing no VAE or diffusion backbone. A frozen DINOv3 semantic prior gates the embedding energy above a strict payload floor, and all learning capacity is spent on a frequency-disjoint dual-stream blind extractor. At decode time, Temporal Logit Averaging fuses the soft per-frame evidence of an entire generated clip before a single decision threshold. SemDINO embeds 128 bits at 36.87 dB PSNR, attains 98.5% bit accuracy under classical distortions, and decodes 12 of 100 real generated clips above 90% sequence accuracy, while training in under eight hours on a single 24 GB GPU. The aggregation rule transfers unchanged to competing decoders, improving both a state-of-the-art static watermark and a channel-trained I2V baseline on their own per-frame scores.

    中文摘要 i Abstract iii Extended Abstract v 誌謝 viii 目錄 ix 表目錄 xii 圖目錄 xiv 第一章 研究介紹 1 1.1 導論 1 第二章 背景知識 6 2.1 數位影像浮水印基礎 6 2.1.1 系統架構與處理流程 6 2.1.2 評估指標 7 2.1.3 威脅模型 8 2.2 頻域載體與區塊 DFT 嵌入 9 2.2.1 逐區塊 DFT 與中頻帶 9 2.2.2 固定區塊尺寸的低通極限 10 2.3 潛在擴散模型與影像轉影片生成 11 2.3.1 VAE 與潛在擴散 11 2.3.2 影像轉影片擴散與條件畫格 12 2.3.3 零樣本情境 12 2.4 來自自監督視覺 Transformer 的語意顯著性先驗 13 2.4.1 DINO 系列模型與區塊層級顯著性 13 2.4.2 作為能量分配先驗的顯著性 13 2.5 學習式浮水印系統的端對端訓練 14 2.5.1 可微分雜訊層 14 2.5.2 最佳化目標與以 PSNR 為目標的保真度 15 2.5.3 浮水印技術中的對抗式訓練 15 2.5.4 階段訓練策略與冷啟動問題 16 第三章 相關研究 17 3.1 相關研究 17 3.1.1 傳統與基於深度學習之影像浮水印 17 3.1.2 面向生成模型的浮水印技術 18 3.1.3 區域自適應與知覺遮蔽式嵌入 19 3.1.4 影片浮水印與影像轉影片通道 20 3.1.5 浮水印基準測試 21 3.1.6 本論文之定位 22 第四章 方法設計 24 4.1 提出之方法 24 4.1.1 總覽與問題形式化 25 4.1.2 來自凍結 DINOv3 先驗的語意能量閘控 27 4.1.3 雙尺度區塊 DFT 載體 28 4.1.4 與生成先驗路線的關係:本方法為何不需要 VAE 30 4.1.5 雙串流學習式萃取器 32 4.1.6 訓練目標與課程設計 33 4.1.7 透過時序邏輯平均進行零樣本 I2V 解碼 36 第五章 實驗結果與分析 38 5.1 評估 38 5.1.1 實驗設定 38 5.1.2 影像浮水印:保真度與穩健性 40 5.1.3 零樣本 I2V 影片浮水印 47 5.1.4 消融實驗 58 5.1.5 分析與失效模式 66 5.1.6 質化結果 69 第六章 未來展望與結論 79 6.1 結論 79 參考文獻 83

    [1] Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, Varun Jampani, and Robin Rombach. Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023.
    [2] Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. Video generation models as world simulators. OpenAI Technical Report, https://openai.com/research/video-generation-models-as-world-simulators, 2024.
    [3] Shilin Lu, Zihan Zhou, Jiayou Lu, Yuanzhi Zhu, and Adams Wai-Kin Kong. Robust watermarking using generative priors against image editing: From benchmarking to advances. In Proceedings of the International Conference on Learning Representations (ICLR), 2025.
    [4] Kevin Alex Zhang, Lei Xu, Alfredo Cuesta-Infante, and Kalyan Veeramachaneni. Robust invisible video watermarking with attention. arXiv preprint arXiv:1909.01285, 2019.
    [5] Tomáš Souček, Pierre Fernandez, Hady Elsahar, Sylvestre-Alvise Rebuffi, Valeriu Lacatusu, Tuan Tran, Tom Sander, and Alexandre Mourachko. Pixel seal: Adversarial-only training for invisible image and video watermarking, 2025.
    [6] Guanjie Wang, Zehua Ma, Han Fang, and Weiming Zhang. LoT-Pass: Long-term-robust image watermarking for image to video generation, 2025.
    [7] Oriane Siméoni, Huy V. Vo, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Michaël Ramamonjisoa, Francisco Massa, Daniel Haziza, Luca Wehrstedt, Jianyuan Wang, Timothée Darcet, Théo Moutakanni, Leonel Sentana, Claire Roberts, Andrea Vedaldi, Jamie Tolan, John Brandt, Camille Couprie, Julien Mairal, Hervé Jégou, Patrick Labatut, and Piotr Bojanowski. DINOv3, 2025.
    [8] Bang An, Mucong Ding, Tahseen Rabbani, Aakriti Agrawal, Yuancheng Xu, Chenghao Deng, Sicheng Zhu, Abdirisak Mohamed, Yuxin Wen, Tom Goldstein, and Furong Huang. WAVES: Benchmarking the robustness of image watermarks. In Proceedings of the 41st International Conference on Machine Learning (ICML), 2024.
    [9] Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. Image quality assessment: From error visibility to structural similarity. IEEE Transactions on Image Processing, 13(4):600-612, 2004.
    [10] Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 586-595, 2018.
    [11] Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems (NeurIPS), volume 33, pages 6840-6851, 2020.
    [12] Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In Proceedings of the International Conference on Learning Representations (ICLR), 2021.
    [13] Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. Video diffusion models. In Advances in Neural Information Processing Systems (NeurIPS), volume 35, 2022.
    [14] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10684-10695, 2022.
    [15] Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 9650-9660, 2021.
    [16] Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V. Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. DINOv2: Learning robust visual features without supervision. Transactions on Machine Learning Research (TMLR), 2024.
    [17] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations (ICLR), 2021.
    [18] Jiren Zhu, Russell Kaplan, Justin Johnson, and Li Fei-Fei. HiDDeN: Hiding data with deep networks. In Proceedings of the European Conference on Computer Vision (ECCV), pages 657-672, 2018.
    [19] Zhaoyang Jia, Han Fang, and Weiming Zhang. MBRS: Enhancing robustness of DNN-based watermarking by mini-batch of real and simulated JPEG compression. In Proceedings of the 29th ACM International Conference on Multimedia (ACM MM), pages 41-49, 2021.
    [20] Ingemar J. Cox, Joe Kilian, F. Thomson Leighton, and Talal Shamoon. Secure spread spectrum watermarking for multimedia. IEEE Transactions on Image Processing, 6(12):1673-1687, 1997.
    [21] Ali Al-Haj. Combined DWT-DCT digital image watermarking. Journal of Computer Science, 3(9):740-746, 2007.
    [22] K. A. Navas, Mathews Cheriyan Ajay, M. Lekshmi, Tampy S. Archana, and M. Sasikumar. DWT-DCT-SVD based watermarking. In Proceedings of the 3rd International Conference on Communication Systems Software and Middleware (COMSWARE), pages 271-274, 2008.
    [23] Matthew Tancik, Ben Mildenhall, and Ren Ng. StegaStamp: Invisible hyperlinks in physical photographs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2117-2126, 2020.
    [24] Xiaoshuai Wu, Xin Liao, and Bo Ou. SepMark: Deep separable watermarking for unified source tracing and deepfake detection. In Proceedings of the 31st ACM International Conference on Multimedia (ACM MM), 2023.
    [25] Xuanyu Zhang, Runyi Li, Jiwen Yu, Youmin Xu, Weiqi Li, and Jian Zhang. EditGuard: Versatile image watermarking for tamper localization and copyright protection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11964-11974, 2024.
    [26] Pierre Fernandez, Guillaume Couairon, Hervé Jégou, Matthijs Douze, and Teddy Furon. The stable signature: Rooting watermarks in latent diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 22466-22477, 2023.
    [27] Yuxin Wen, John Kirchenbauer, Jonas Geiping, and Tom Goldstein. Tree-ring watermarks: Fingerprints for diffusion images that are invisible and robust. In Advances in Neural Information Processing Systems (NeurIPS), volume 36, 2023.
    [28] Zijin Yang, Kai Zeng, Kejiang Chen, Han Fang, Weiming Zhang, and Nenghai Yu. Gaussian shading: Provable performance-lossless image watermarking for diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024.
    [29] Yehan Sun, Rongrong Ni, Chuangchuang Tan, Huan Liu, Wenhao Ni, Renshuai Tao, and Yao Zhao. RAIN: Redundancy-aware latent injection for quality-preserving image watermarking. In Proceedings of the 40th AAAI Conference on Artificial Intelligence (AAAI), 2026.
    [30] Zhenliang Gan, Chunya Liu, Yichao Tang, Binghao Wang, Shiwen Cui, Weiqiang Wang, and Xinpeng Zhang. GenPTW: Latent image watermarking for provenance tracing and tamper localization. In Proceedings of the 40th AAAI Conference on Artificial Intelligence (AAAI), 2026.
    [31] Runyi Hu, Jie Zhang, Shiqian Zhao, Nils Lukas, Jiwei Li, Qing Guo, Han Qiu, and Tianwei Zhang. Mask image watermarking. In Advances in Neural Information Processing Systems (NeurIPS), 2025.
    [32] Xiyang Luo, Yinxiao Li, Huiwen Chang, Ce Liu, Peyman Milanfar, and Feng Yang. DVMark: A deep multiscale framework for video watermarking, 2021.
    [33] Pierre Fernandez, Hady Elsahar, I. Zeki Yalniz, and Alexandre Mourachko. Video seal: Open and efficient video watermarking, 2024.
    [34] Yuxin Wu and Kaiming He. Group normalization. In Proceedings of the European Conference on Computer Vision (ECCV), pages 3-19, 2018.
    [35] Fisher Yu and Vladlen Koltun. Multi-scale context aggregation by dilated convolutions. In Proceedings of the International Conference on Learning Representations (ICLR), 2016.
    [36] Axel Sauer, Dominik Lorenz, Andreas Blattmann, and Robin Rombach. Adversarial diffusion distillation, 2023.
    [37] Richard Shin and Dawn Song. JPEG-resistant adversarial images. In NIPS Workshop on Machine Learning and Computer Security, 2017.
    [38] Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In Proceedings of the International Conference on Learning Representations (ICLR), 2015.
    [39] Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In Proceedings of the International Conference on Learning Representations (ICLR), 2019.
    [40] Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. Microsoft COCO: Common objects in context. In Proceedings of the European Conference on Computer Vision (ECCV), pages 740-755, 2014.
    [41] Eirikur Agustsson and Radu Timofte. NTIRE 2017 challenge on single image super-resolution: Dataset and study. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 126-135, 2017.
    [42] Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, Tom Duerig, and Vittorio Ferrari. The open images dataset V4: Unified image classification, object detection, and visual relationship detection at scale. International Journal of Computer Vision, 128:1956-1981, 2020.
    [43] Tu Bui, Shruti Agarwal, and John Collomosse. Trust Mark: Universal watermarking for arbitrary resolution images, 2023.
    [44] Runyi Hu, Jie Zhang, Ting Xu, Jiwei Li, and Tianwei Zhang. Robust-wide: Robust watermarking against instruction-driven image editing. In Proceedings of the European Conference on Computer Vision (ECCV), pages 20-37. Springer, 2024.
    [45] Utae Jeong, Sumin In, Hyunju Ryu, Jaewan Choi, Feng Yang, Jongheon Jeong, Seungryong Kim, and Sangpil Kim. WaTeRFlow: Watermark temporal robustness via flow consistency. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 31703-31713, 2026.

    下載圖示
    校外:立即公開
    QR CODE