| 研究生: |
周子豪 ZHOU, Zi-Hao |
|---|---|
| 論文名稱: |
Transformer 語言模型中的詞頻驅動猜測:
層級動態分析與詞頻感知訓練介入 Frequency-Driven Guessing in Transformer Language Models: Layer-Wise Dynamics Analysis and Frequency-Aware Training Intervention |
| 指導教授: |
謝昀珊
Hsieh, Yun-Shan |
| 學位類別: |
碩士 Master |
| 系所名稱: |
電機資訊學院 - 資訊工程學系 Department of Computer Science and Information Engineering |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 49 |
| 中文關鍵詞: | 變換器語言模型 、詞元頻率 、頻率感知訓練 、深度利用 |
| 外文關鍵詞: | Transformer Language Models, Token Frequency, Frequency-Aware Training, Depth Utilization |
| 相關次數: | 點閱:13 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
近年來,大型語言模型(Large Language Models, LLMs)展現出優異的文字生成能力,過去研究發現,Transformer 語言模型內部運算每一層的預測結果傾向先猜測在訓練資料中常見的詞,再逐步修正為符合輸入上下文語意的結果,此現象被稱為 Guess-then-Refine。然而,模型為何會產生這種偏向高頻詞的猜測,以及是否能減少前期這種冗餘的猜測動作來提早後續的修正過程進行,仍有待實驗分析。為了深入探討這個方向,這篇論文針對GPT-2模型並對於在訓練資料中不同詞的出現頻率以及模型訓練目標損失函數分別從三個角度觀察-分析-干預來進行實驗,主要實驗結果發現:這個模型的猜測行為在訓練初期產生並隨著模型訓練過程逐漸減少,並且透過給予不同詞頻在訓練時的訓練目標不同的權重,能夠有效的減少這個以詞頻驅動的猜測行為,但是無法因此提早最終生成結果的出現。
In recent years, Large Language Models (LLMs) have demonstrated remarkable text generation capabilities. Previous studies have found that, during the layer-wise computation of Transformer-based language models, intermediate predictions tend to first favor tokens that frequently occur in the training data and then gradually refine these predictions toward results that better fit the semantics of the input context. This phenomenon is referred to as Guess-then-Refine. However, why models exhibit this frequency-driven guessing behavior and whether reducing such redundant early-stage guesses can allow the subsequent refinement process to occur earlier remain open questions that require further experimental investigation.To investigate this phenomenon , this thesis conducts experiments on the GPT-2 model by examining the effects of different token frequencies in the training corpus and the model's training objective loss function from three perspectives: observation, analysis, and intervention The experimental results show that this guessing behavior emerges during the early stages of training and gradually decreases as training progresses. Furthermore, by assigning different weights to the training objectives according to token frequency, the frequency-driven guessing behavior can be effectively reduced. However, this intervention does not effectively accelerate the emergence of the final generation result.
[1] Nora Belrose, Igor Ostrovsky, Lev McKinney, Zach Furman, Logan Smith, Danny Halawi, Stella Biderman, and Jacob Steinhardt. Eliciting latent edictions from transformers with the tuned lens. arXiv preprint arXiv:2303.08112, 2023.
[2] Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O'Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. Pythia: A suite for analyzing large language models across training and scaling. In International conference on machine learning, pages 2397–2430. PMLR, 2023.
[3] Tyler A Chang and Benjamin K Bergen. Word acquisition in neural language models.Transactions of the Association for Computational Linguistics, 10:1–16, 2022.
[4] Alexander Yom Din, Taelin Karidi, Leshem Choshen, and Mor Geva. Jump to conclusions: Short-cutting transformers with linear transformations. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 9615–9625, 2024.
[5] Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah.Amathematical framework for transformer
[6] Wikimedia Foundation. Wikimedia downloads. https://dumps.wikimedia.org, 2023.
[7] Aaron Gokaslan and Vanya Cohen. Openwebtext corpus. http://Skylion007. github.io/OpenWebTextCorpus, 2019.
[8] Akshat Gupta, Jay Yeung, Gopala Anumanchipalli, and Anna Ivanova. How do llms use their depth? arXiv preprint arXiv:2510.18871, 2025.
[9] Felix Hill, Antoine Bordes, Sumit Chopra, and Jason Weston. The goldilocks principle: Reading children’s books with explicit memory representations. arXiv preprint arXiv:1511.02301, 2015.
[10] Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. What does bert learn about the structure of language? In Proceedings of the 57th annual meeting of the association for computational linguistics, pages 3651–3657, 2019.
[11] Andrej Karpathy. NanoGPT. https://github.com/karpathy/nanoGPT, 2022.
[12] Sunwoo Kim, Haneul Yoo, and Alice Oh. On the effect of uncertainty on layer-wise inference dynamics. In Proceedings of the 42nd International Conference on Machine Learning. PMLR, 2025.
[13] Ryo Nakamura, Katsuhito Sudoh, Koichiro Yoshino, and Satoshi Nakamura. Another diversity-promoting objective function for neural dialogue generation. arXiv preprint arXiv:1811.08100, 2018.
[14] nostalgebraist. interpreting gpt: the logit lens. Less-Wrong, 2020.
[15] Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc-Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. The lambada dataset: Word prediction requiring a broad discourse context. In Proceedings of the 54th annual meeting of the association for computational linguistics (volume 1: Long papers), pages 1525–1534, 2016.
[16] Steven T Piantadosi. Zipf' s word frequency law in natural language: A critical review and future directions. Psychonomic bulletin & review, 21(5):1112–1130, 2014.
[17] Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
[18] Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Tran, Yi Tay, and Donald Metzler. Confident adaptive language modeling. Advances in Neural Information Processing Systems, 35:17456–17472, 2022.
[19] Oscar Skean, Md Rifat Arefin, and Ravid Shwartz-Ziv. Does representation matter? exploring intermediate layers in large language models. In Workshop on Machine Learning and Compression, NeurIPS 2024, 2024.
[20] Oscar Skean, Md Rifat Arefin, Dan Zhao, Niket Nikul Patel, Jalal Naghiyev, Yann Lecun, and Ravid Shwartz-Ziv. Layer by layer: Uncovering hidden representations in language models. In International Conference on Machine Learning, pages 55854–55875. PMLR, 2025.
[21] Giulio Starace, Konstantinos Papakostas, Rochelle Choenni, Apostolos Panagiotopou-los, Matteo Rosati, Alina Leidinger, and Ekaterina Shutova. Probing llms for joint encoding of linguistic categories. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 7158–7179, 2023.
[22] Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin. Deebert: Dynamic early exiting for accelerating bert inference. In Proceedings of the 58th annual meeting of the association for computational linguistics, pages 2246–2251, 2020.