簡易檢索 / 詳目顯示

研究生: 周子豪
ZHOU, Zi-Hao
論文名稱: Transformer 語言模型中的詞頻驅動猜測: 層級動態分析與詞頻感知訓練介入
Frequency-Driven Guessing in Transformer Language Models: Layer-Wise Dynamics Analysis and Frequency-Aware Training Intervention
指導教授: 謝昀珊
Hsieh, Yun-Shan
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 資訊工程學系
Department of Computer Science and Information Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 49
中文關鍵詞: 變換器語言模型詞元頻率頻率感知訓練深度利用
外文關鍵詞: Transformer Language Models, Token Frequency, Frequency-Aware Training, Depth Utilization
相關次數: 點閱:13下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 近年來,大型語言模型(Large Language Models, LLMs)展現出優異的文字生成能力,過去研究發現,Transformer 語言模型內部運算每一層的預測結果傾向先猜測在訓練資料中常見的詞,再逐步修正為符合輸入上下文語意的結果,此現象被稱為 Guess-then-Refine。然而,模型為何會產生這種偏向高頻詞的猜測,以及是否能減少前期這種冗餘的猜測動作來提早後續的修正過程進行,仍有待實驗分析。為了深入探討這個方向,這篇論文針對GPT-2模型並對於在訓練資料中不同詞的出現頻率以及模型訓練目標損失函數分別從三個角度觀察-分析-干預來進行實驗,主要實驗結果發現:這個模型的猜測行為在訓練初期產生並隨著模型訓練過程逐漸減少,並且透過給予不同詞頻在訓練時的訓練目標不同的權重,能夠有效的減少這個以詞頻驅動的猜測行為,但是無法因此提早最終生成結果的出現。

    In recent years, Large Language Models (LLMs) have demonstrated remarkable text generation capabilities. Previous studies have found that, during the layer-wise computation of Transformer-based language models, intermediate predictions tend to first favor tokens that frequently occur in the training data and then gradually refine these predictions toward results that better fit the semantics of the input context. This phenomenon is referred to as Guess-then-Refine. However, why models exhibit this frequency-driven guessing behavior and whether reducing such redundant early-stage guesses can allow the subsequent refinement process to occur earlier remain open questions that require further experimental investigation.To investigate this phenomenon , this thesis conducts experiments on the GPT-2 model by examining the effects of different token frequencies in the training corpus and the model's training objective loss function from three perspectives: observation, analysis, and intervention The experimental results show that this guessing behavior emerges during the early stages of training and gradually decreases as training progresses. Furthermore, by assigning different weights to the training objectives according to token frequency, the frequency-driven guessing behavior can be effectively reduced. However, this intervention does not effectively accelerate the emergence of the final generation result.

    摘要 i Abstract ii 誌謝 iii Contents iv List of Tables vi List of Figures vii Chapter 1. Introduction 1 1.1. Motivation 1 1.2. Main Contributions 1 1.3. Thesis Organization 2 Chapter 2. Background and Related Work 3 2.1. Hidden State 3 2.2. Probing 5 2.2.1. Logit Lens 5 2.2.2. Tuned Lens 6 2.3. Guess-then-Refine 9 2.4. Inverse Token Frequency Loss 10 2.5. Research Gap 10 Chapter 3. Frequency-Driven Guessing Phase: Observation and Analysis 11 3.1. How did the model’s layer-wise guessing behavior evolve during pretraining? 11 3.2. Cross-Entropy Loss and High-Frequency Tokens 14 Chapter 4. Frequency-Driven Guessing Phase: Intervention 18 4.1. Weakening the High-Frequency Gravity with Weighted-Loss 18 4.1.1. Inverse Log Frequency Loss 19 4.1.2. Inverse Token Frequency Loss 20 4.2. Evaluation Metrics 22 4.2.1. Layer-wise Dynamics Metrics 22 4.2.2. Generative Performance Benchmarks 25 4.3. Setup 29 4.4. Results 31 4.4.1. Layer-wise Dynamics Results 31 4.4.2. Generative Performance Results 32 Chapter 5. Discussion 34 5.1. Ablation on Scaling Factor Lambda with Inverse Token Frequency Loss 34 5.2. FGL Across Different Model Scales 36 Chapter 6. Conclusion 38 References 39

    [1] Nora Belrose, Igor Ostrovsky, Lev McKinney, Zach Furman, Logan Smith, Danny Halawi, Stella Biderman, and Jacob Steinhardt. Eliciting latent edictions from transformers with the tuned lens. arXiv preprint arXiv:2303.08112, 2023.
    [2] Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O'Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. Pythia: A suite for analyzing large language models across training and scaling. In International conference on machine learning, pages 2397–2430. PMLR, 2023.
    [3] Tyler A Chang and Benjamin K Bergen. Word acquisition in neural language models.Transactions of the Association for Computational Linguistics, 10:1–16, 2022.
    [4] Alexander Yom Din, Taelin Karidi, Leshem Choshen, and Mor Geva. Jump to conclusions: Short-cutting transformers with linear transformations. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 9615–9625, 2024.
    [5] Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah.Amathematical framework for transformer
    [6] Wikimedia Foundation. Wikimedia downloads. https://dumps.wikimedia.org, 2023.
    [7] Aaron Gokaslan and Vanya Cohen. Openwebtext corpus. http://Skylion007. github.io/OpenWebTextCorpus, 2019.
    [8] Akshat Gupta, Jay Yeung, Gopala Anumanchipalli, and Anna Ivanova. How do llms use their depth? arXiv preprint arXiv:2510.18871, 2025.
    [9] Felix Hill, Antoine Bordes, Sumit Chopra, and Jason Weston. The goldilocks principle: Reading children’s books with explicit memory representations. arXiv preprint arXiv:1511.02301, 2015.
    [10] Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. What does bert learn about the structure of language? In Proceedings of the 57th annual meeting of the association for computational linguistics, pages 3651–3657, 2019.
    [11] Andrej Karpathy. NanoGPT. https://github.com/karpathy/nanoGPT, 2022.
    [12] Sunwoo Kim, Haneul Yoo, and Alice Oh. On the effect of uncertainty on layer-wise inference dynamics. In Proceedings of the 42nd International Conference on Machine Learning. PMLR, 2025.
    [13] Ryo Nakamura, Katsuhito Sudoh, Koichiro Yoshino, and Satoshi Nakamura. Another diversity-promoting objective function for neural dialogue generation. arXiv preprint arXiv:1811.08100, 2018.
    [14] nostalgebraist. interpreting gpt: the logit lens. Less-Wrong, 2020.
    [15] Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc-Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. The lambada dataset: Word prediction requiring a broad discourse context. In Proceedings of the 54th annual meeting of the association for computational linguistics (volume 1: Long papers), pages 1525–1534, 2016.
    [16] Steven T Piantadosi. Zipf' s word frequency law in natural language: A critical review and future directions. Psychonomic bulletin & review, 21(5):1112–1130, 2014.
    [17] Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
    [18] Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Tran, Yi Tay, and Donald Metzler. Confident adaptive language modeling. Advances in Neural Information Processing Systems, 35:17456–17472, 2022.
    [19] Oscar Skean, Md Rifat Arefin, and Ravid Shwartz-Ziv. Does representation matter? exploring intermediate layers in large language models. In Workshop on Machine Learning and Compression, NeurIPS 2024, 2024.
    [20] Oscar Skean, Md Rifat Arefin, Dan Zhao, Niket Nikul Patel, Jalal Naghiyev, Yann Lecun, and Ravid Shwartz-Ziv. Layer by layer: Uncovering hidden representations in language models. In International Conference on Machine Learning, pages 55854–55875. PMLR, 2025.
    [21] Giulio Starace, Konstantinos Papakostas, Rochelle Choenni, Apostolos Panagiotopou-los, Matteo Rosati, Alina Leidinger, and Ekaterina Shutova. Probing llms for joint encoding of linguistic categories. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 7158–7179, 2023.
    [22] Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin. Deebert: Dynamic early exiting for accelerating bert inference. In Proceedings of the 58th annual meeting of the association for computational linguistics, pages 2246–2251, 2020.

    QR CODE