簡易檢索 / 詳目顯示

研究生: 王駿杰
Wang, Chun-Chieh
論文名稱: 基於軌跡一致性學習之擴散模型
Trajectory-Level Consistency Learning for Diffusion Model
指導教授: 吳宗憲
Wu, Chung-Hsien
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 資訊工程學系
Department of Computer Science and Information Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 74
中文關鍵詞: 擴散語言模型無分類器引導自引導
外文關鍵詞: Diffusion Language Model, Classifier-Free Guidance, Self-Guidance
相關次數: 點閱:49下載:1
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 近年來,擴散模型(Diffusion Models)在生成式人工智慧領域展現出優異的表現,尤其在影像生成方面已成為主流方法。然而,在文字生成任務中,擴散式語言模型(Diffusion Language Models)相較於自回歸模型(Autoregressive Models)仍存在生成品質、收斂效率以及模型表徵利用不足等問題。近期提出的推論階段引導方法,例如 self-guidance 與 auto-guidance,顯示擴散模型本身其實具有潛在的自我修正能力,只需在採樣過程中加入簡單的引導機制,即可顯著提升生成效果。這也暗示目前的訓練方式尚未完全發揮模型的潛在能力。
    為了解決上述問題,本研究提出一種名為 Self-Consistent Guidance(SCG) 的新型訓練框架,將原本只能在推論階段使用的 guidance 機制直接內化至模型訓練過程中。SCG 透過設計一致性導向(consistency-based)的訓練目標,利用模型在不同去噪時間步(denoising timesteps)之間的預測結果,強化其生成過程中的一致性與穩定性。此方法使模型能在訓練期間學習自我引導的行為,而不需修改原始模型架構,也不會增加推論成本。
    實驗結果顯示,SCG 能穩定提升擴散式語言模型的生成品質與語言建模能力,並有效降低 perplexity,同時加速模型收斂速度。此外,SCG 亦具有隱式正則化(implicit regularization)的效果,可降低過擬合並提升模型泛化能力。除了文字生成任務之外,SCG 亦可自然延伸至連續型擴散模型,在影像生成任務中取得更佳的 FID 分數。整體而言,本研究證明了 training-time consistency 對於提升擴散模型效能的重要性,並為縮小擴散模型與自回歸模型之間的性能差距提供了一個有效且具潛力的方向。

    Diffusion-based generative models have recently demonstrated strong potential in a wide range of generation tasks, especially in image synthesis and emerging text generation frameworks. However, compared with autoregressive language models, diffusion language models still face limitations in generation quality, convergence efficiency, and the effective utilization of internal representations. Recent inference-time guidance approaches, such as self-guidance and auto-guidance, reveal that diffusion models possess latent corrective capabilities that can significantly enhance generation performance during sampling. These findings suggest that existing diffusion models may not fully exploit their representational capacity under conventional training objectives.
    To address this limitation, this paper proposes Self-Consistent Guidance (SCG), a novel training framework that incorporates the advantages of guidance mechanisms directly into the training process. Instead of relying on additional inference-time corrections, SCG introduces a consistency-based objective that leverages the model’s own predictions across different denoising timesteps to encourage coherent and stable generation dynamics. By enforcing self-consistency during training, the proposed method enables diffusion models to internalize guidance behavior within model parameters, thereby improving both learning efficiency and generation capability without changing the underlying architecture or increasing inference cost.
    Experimental results demonstrate that SCG consistently improves generation quality and lowers perplexity in diffusion language models while also accelerating convergence during training. In addition, the method acts as an implicit regularizer, reducing overfitting and improving model generalization. Beyond text generation, SCG can be naturally extended to continuous diffusion settings, where it achieves improved FID scores in image generation tasks. Overall, this work highlights training-time consistency as an effective principle for enhancing diffusion-based generative modeling and provides a practical direction for bridging the performance gap between diffusion and autoregressive models.

    摘要 I Abstract III 致謝 V Content VI List of Tables IX List of Figures X Chapter 1 Introduction 1 1.1 Background 1 1.2 Problems 3 1.3 Motivation 6 1.4 Literature Review 7 1.4.1 Fundamental Diffusion Language Models 7 1.4.2 The Diffusion Duality 9 1.4.3 Diffusion Language Model Framework 11 1.4.4 Classifier-Free Guidance and Self-Guidance 13 1.5 Brief Description of Research Methods 14 Chapter 2 Proposed Method 16 2.1 Self-Consistent Guidance 17 2.1.1 Representation-Based Trajectory Formulation 17 2.1.2 Cross-Timestep Consistency Objective 18 2.1.3 Why Cosine Similarity 20 2.2 Adaptive Loss Scheduling 20 2.2.1 Fix Scheduler 22 2.2.2 Linear Scheduler 22 2.2.3 Exponential Scheduler 22 2.2.4 Beta Scheduler 22 2.2.5 Adaptive Loss Scheduler 23 2.3 Relationship to Self-Guidance 25 Chapter 3 Dataset 28 3.1.1 Language Modeling Datasets 28 3.1.2 Image Generation Dataset 29 Chapter 4 Experimental Result 31 4.1 Evaluation Metrics 31 4.1.1 Generation Perplexity 32 4.1.2 Fréchet Inception Distance 32 4.1.3 Evaluation Protocol 33 4.2 Experiment Setting 33 4.2.1 Baseline Models 34 4.2.2 Diffusion Language Model 35 4.2.3 Training Configuration 36 4.2.4 SCG Scheduling Configuration 36 4.2.5 EDM2 Configuration 37 4.3 Results 37 4.3.1 Faster Convergence and Training Efficiency 38 4.3.2 Comparison with Self-Guidance 41 4.3.3 Zero-Shot Generalization 43 4.3.4 Effect of Guidance Scheduling in Discrete Domain 45 4.3.5 Extension to Continuous Diffusion Models 48 4.3.6 Discussion between Discrete and Continuous Models 52 Chapter 5 Conclusion and Future Work 54 5.1 Conclusion 54 5.2 Future Work 56 Reference 58 Appendix 61

    [1] B. M. Tom Brown, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, "Language models are few-shot learners," arXiv preprint arXiv:2005.14165, 2020.
    [2] A. Chowdhery et al., "Palm: Scaling language modeling with pathways," Journal of machine learning research, vol. 24, no. 240, pp. 1–113, 2023.
    [3] J. Ho, A. Jain, and P. Abbeel, "Denoising diffusion probabilistic models," Advances in neural information processing systems, vol. 33, pp. 6840–6851, 2020.
    [4] P. Dhariwal and A. Nichol, "Diffusion models beat gans on image synthesis," Advances in neural information processing systems, vol. 34, pp. 8780–8794, 2021.
    [5] X. Li, J. Thickstun, I. Gulrajani, P. S. Liang, and T. B. Hashimoto, "Diffusion-lm improves controllable text generation," Advances in neural information processing systems, vol. 35, pp. 4328–4343, 2022.
    [6] J. Austin, D. D. Johnson, J. Ho, D. Tarlow, and R. Van Den Berg, "Structured denoising diffusion models in discrete state-spaces," Advances in neural information processing systems, vol. 34, pp. 17981–17993, 2021.
    [7] A. Lou, C. Meng, and S. Ermon, "Discrete diffusion modeling by estimating the ratios of the data distribution," arXiv preprint arXiv:2310.16834, 2023.
    [8] S. S. Sahoo et al., "Simple and effective masked diffusion language models," Advances in Neural Information Processing Systems, vol. 37, pp. 130136–130184, 2024.
    [9] S. Nie et al., "Large language diffusion models," Advances in Neural Information Processing Systems, vol. 38, pp. 50608–50646, 2026.
    [10] T. Bie et al., "Llada2. 0: Scaling up diffusion language models to 100b," arXiv preprint arXiv:2512.15745, 2025.
    [11] S. Gong et al., "Scaling diffusion language models via adaptation from autoregressive models," in International Conference on Learning Representations, 2025, vol. 2025, pp. 5046–5073.
    [12] X. Han, S. Kumar, and Y. Tsvetkov, "Ssd-lm: Semi-autoregressive simplex-based diffusion language model for text generation and modular control," in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 11575–11596.
    [13] R. K. Mahabadi et al., "Tess: Text-to-text self-conditioned simplex diffusion," in Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 2347–2361.
    [14] J. Ho and T. Salimans, "Classifier-free diffusion guidance," arXiv preprint arXiv:2207.12598, 2022.
    [15] T. Li, W. Luo, Z. Chen, L. Ma, and G.-J. Qi, "Self-guidance: Boosting flow and diffusion generation on their own," IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025.
    [16] T. Karras, M. Aittala, T. Kynkäänniemi, J. Lehtinen, T. Aila, and S. Laine, "Guiding a diffusion model with a bad version of itself," Advances in Neural Information Processing Systems, vol. 37, pp. 52996–53021, 2024.
    [17] Y. Schiff et al., "Simple guidance mechanisms for discrete diffusion models," in International Conference on Learning Representations, 2025, vol. 2025, pp. 43776–43821.
    [18] S. S. Sahoo, J. Deschenaux, A. Gokaslan, G. Wang, J. Chiu, and V. Kuleshov, "The diffusion duality," arXiv preprint arXiv:2506.10892, 2025.
    [19] Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, "Score-based generative modeling through stochastic differential equations," arXiv preprint arXiv:2011.13456, 2020.
    [20] E. Hoogeboom, D. Nielsen, P. Jaini, P. Forré, and M. Welling, "Argmax flows and multinomial diffusion: Learning categorical distributions," Advances in neural information processing systems, vol. 34, pp. 12454–12465, 2021.
    [21] M. Xu et al., "Energy-based diffusion language models for text generation," in International Conference on Learning Representations, 2025, vol. 2025, pp. 33769–33789.
    [22] T. Karras, M. Aittala, T. Aila, and S. Laine, "Elucidating the design space of diffusion-based generative models," Advances in neural information processing systems, vol. 35, pp. 26565–26577, 2022.
    [23] Y. Bengio, J. Louradour, R. Collobert, and J. Weston, "Curriculum learning," in Proceedings of the 26th annual international conference on machine learning, 2009, pp. 41–48.
    [24] A. Kendall, Y. Gal, and R. Cipolla, "Multi-task learning using uncertainty to weigh losses for scene geometry and semantics," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7482–7491.
    [25] Z. Chen, V. Badrinarayanan, C.-Y. Lee, and A. Rabinovich, "Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks," in International conference on machine learning, 2018: PMLR, pp. 794–803.
    [26] C. Chelba et al., "One billion word benchmark for measuring progress in statistical language modeling," arXiv preprint arXiv:1312.3005, 2013.
    [27] S. Merity, C. Xiong, J. Bradbury, and R. Socher, "Pointer sentinel mixture models," arXiv preprint arXiv:1609.07843, 2016.
    [28] X. Zhang, J. Zhao, and Y. LeCun, "Character-level convolutional networks for text classification," Advances in neural information processing systems, vol. 28, 2015.
    [29] D. Paperno et al., "The LAMBADA dataset: Word prediction requiring a broad discourse context," in Proceedings of the 54th annual meeting of the association for computational linguistics (volume 1: Long papers), 2016, pp. 1525–1534.
    [30] A. Cohan et al., "A discourse-aware attention model for abstractive summarization of long documents," in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), 2018, pp. 615–621.
    [31] A. Krizhevsky and G. Hinton, "Learning multiple layers of features from tiny images," 2009.
    [32] M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, "Gans trained by a two time-scale update rule converge to a local nash equilibrium," Advances in neural information processing systems, vol. 30, 2017.
    [33] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, "Language models are unsupervised multitask learners," OpenAI blog, vol. 1, no. 8, p. 9, 2019.

    下載圖示
    校外:立即公開
    QR CODE