簡易檢索 / 詳目顯示

研究生: 鄭承武
Cheng, Cheng-Wu
論文名稱: 防禦大型語言模型多輪越獄攻擊之策略自適應模擬法庭安全防護機制
CourtGuard: An Adjudication-Based Policy-Adaptive Safeguard Against Multi-Turn LLM Jailbreak Attacks
指導教授: 郭耀煌
Kuo, Yau-Hwang
莊宜勳
Chuang, I-Hsun
學位類別: 碩士
Master
系所名稱: 電機資訊學院 - 資訊工程學系
Department of Computer Science and Information Engineering
論文出版年: 2026
畢業學年度: 114
語文別: 英文
論文頁數: 124
中文關鍵詞: 大型語言模型安全多輪越獄攻擊越獄防禦策略自適應
外文關鍵詞: Large Language Model Safety, Multi-Turn Jailbreak Attacks, Jailbreak Defense, Policy Adaptability
相關次數: 點閱:20下載:0
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 近年來,隨著大型語言模型(LLM)在語意理解、邏輯推理與內容生成能力取得突破性的進展,各大企業正加速將其導入核心業務以推動服務升級。實務上,各企業需制定專屬且具備動態適應能力的安全對齊策略來確保LLM服務的產出符合其企業價值觀與法律規範。然而,攻擊者常採用越獄攻擊(Jailbreak)的方式規避這些策略。其中,新興的多輪越獄攻擊更透過將惡意企圖拆解並隱藏於連續對話的語境中,成功突破了傳統基於分類器的靜態安全防護機制,嚴重威脅系統穩定性。另一方面,現有推理式的安全防護機制雖然能有效識破多輪越獄對話的語境偽裝,卻因其防禦邏輯過度側重於風險阻斷,容易對敏感但合理的正常請求產生過度拒絕,從而大幅削弱了系統的可用性與使用者體驗。
    為此,本論文提出策略自適應模擬法庭安全防護機制(CourtGuard),解決現有LLM服務的安全防禦機制難以適應安全策略迭代,以及在面對多輪越獄攻擊時陷入的高攻擊成功率與高假陽性率的困境。首先,為了有效應對不斷變異的多輪越獄攻擊,CourtGuard採取了推理式安全防護機制,借鑒了法庭審理機制:透過模擬檢察官與律師的角色,針對潛在的攻擊意圖進行正反雙方的舉證與辯論,最終交由法官模型進行裁決,從而有效舒緩了現有機制引發的過度拒絕問題。此外,為了挖掘深藏在多輪對話的攻擊意圖,CourtGuard設立了法庭書記官的角色,透過內容工程技術統整庭審資訊與對話脈絡,引導 LLM 發揮深層推理能力。同時,考量到安全策略可能的變動,CourtGuard設計了司法助理機制,負責引入當時的安全策略與過往判例作為法源依據,確保檢察官、律師與法官能真正依據當前安全策略進行精確的舉證、辯證與裁決。
    最後,本論文採用兩個不同來源的資料集,針對三種最新型的多輪越獄攻擊進行模擬。實驗數據顯示,以抑制攻擊成功率方面,CourtGuard較傳統基於分類器的安全防護機制最多改善36%與43%,而與現行推理式安全機制相比,亦取得最多30%與35%的改善。此外,在良性請求的假陽性率評估中,CourtGuard的誤判率僅為1.6%,遠低於基於分類器的安全防護機制的11.2%與推理式安全機制的16.8%。在策略自適應評估中,CourtGuard於完整策略下維持2.0%的攻擊成功率,當指定策略限制被移除後,CourtGuard在假陽性率評估中最多改善了48%,上述結果說明,CourtGuard不僅能有效抵禦複雜的多輪越獄攻擊並大幅減輕了過度拒絕的問題,也能依據當前生效的策略調整其裁決行為,為LLM服務提供兼具安全性、可用性與策略適應能力的保障。

    In recent years, large language models (LLMs) have achieved remarkable advances in semantic understanding, logical reasoning, and content generation, prompting enterprises to accelerate their adoption in core business operations to enhance service capabilities. In practical deployments, each enterprise must establish customized and dynamically adaptable safety policies to ensure that the outputs of its LLM services comply with organizational values and legal requirements. However, attackers often employ jailbreak attacks to circumvent these policies. In particular, emerging multi-turn jailbreak attacks decompose malicious intent and conceal it within the context of continuous conversations, enabling them to bypass conventional static classifier-based safeguards and posing serious threats to system stability. Meanwhile, although existing reasoning-based safeguards can effectively identify the contextual disguises used in multi-turn jailbreak conversations, their decision logic tends to place excessive emphasis on risk blocking. Consequently, they may over-refuse sensitive yet legitimate requests, substantially reducing system usability and degrading the user experience.
    To address these challenges, this thesis proposes CourtGuard, an adjudication-based, policy-adaptive safeguard designed to overcome the limited adaptability of existing LLM safeguards to evolving safety policies, as well as the high attack success rates and false-positive rates observed under multi-turn jailbreak attacks. First, to effectively address continuously evolving multi-turn jailbreak attacks, CourtGuard adopts a reasoning-based safeguard inspired by courtroom proceedings. It simulates the roles of a prosecutor and an attorney to present and debate opposing analyses of potential attack intent, after which a judge model performs the final adjudication. This adversarial process alleviates the over-refusal problem associated with existing safeguards. In addition, to uncover attack intent hidden across multi-turn conversations, CourtGuard introduces a court clerk role that applies context engineering techniques to organize adjudication information and conversational context, thereby enabling the LLM to perform deeper reasoning. Furthermore, to accommodate changes in safety policies, CourtGuard incorporates a judicial assistant mechanism that introduces the current Deployment Policy and historical adjudication records as the basis for analysis, ensuring that the prosecutor, attorney, and judge conduct their arguments, counterarguments, and adjudication according to the policy currently in effect.
    Finally, this thesis evaluates CourtGuard using two datasets from different sources and three recent multi-turn jailbreak attacks. Experimental results show that CourtGuard reduces attack success rates by up to 36% and 43% compared with conventional classifier-based safeguards on the two datasets, respectively. Compared with existing reasoning-based safeguards, CourtGuard achieves reductions of up to 30% and 35%, respectively. In the false-positive evaluation on benign requests, CourtGuard achieves a false-positive rate of only 1.6%, substantially lower than the 11.2% achieved by classifier-based safeguards and the 16.8% achieved by reasoning-based safeguards. In the policy-adaptability evaluation, CourtGuard maintains an attack success rate of 2.0% under the full policy. When the designated policy restrictions are removed, CourtGuard achieves a false-positive-rate reduction of up to 48% compared with the evaluated safeguards. These results demonstrate that CourtGuard not only effectively mitigates complex multi-turn jailbreak attacks and substantially reduces over-refusal, but also adjusts its adjudication behavior according to the policy currently in effect, providing LLM services with a safeguard that balances safety, usability, and policy adaptability.

    CHAPTER 1 INTRODUCTION 1 1.1 Background 2 1.2 Motivation 7 1.3 Comparison and Contribution 15 1.4 Organization 20 CHAPTER 2 RELATED WORK 21 2.1 Large Language Model Safety Alignment and Guardrails 22 2.2 Jailbreak Attacks against Large Language Models 24 2.3 Existing Jailbreak Defense Methods 29 2.4 Policy-Adaptive Safety Guardrails 35 CHAPTER 3 COURTGUARD: AN ADJUDICATION-BASED POLICY-ADAPTIVE SAFEGUARD AGAINST MULTI-TURN LLM JAILBREAK ATTACKS 37 3.1 System Model and Threat Model 39 3.2 Framework of CourtGuard 43 3.3 Court Clerk Context Manager 46 3.4 Adversarial Trial Inference Module 56 3.5 Judicial Assistant Prompt Generator 70 3.6 Algorithm 79 CHAPTER 4 EXPERIMENTS 80 4.1 Experimental Settings 80 4.2 Performance Analysis 88 4.3 Policy Adaptability Analysis 96 4.4 Ablation Study 99 CHAPTER 5 CONCLUSION 104 CHAPTER 6 FUTURE WORK 107 REFERENCES 108

    [1] T. Brown et al., "Language models are few-shot learners," Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020.
    [2] R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn, "Direct preference optimization: Your language model is secretly a reward model," Advances in neural information processing systems, vol. 36, pp. 53728–53741, 2023.
    [3] L. Ouyang et al., "Training language models to follow instructions with human feedback," Advances in neural information processing systems, vol. 35, pp. 27730–27744, 2022.
    [4] J. Wei et al., "Finetuned language models are zero-shot learners," in International Conference on Learning Representations, 2022.
    [5] A. Bhattacharjee, S. Ghosh, T. Rebedea, and C. Parisien, "Towards inference-time category-wise safety steering for large language models," arXiv preprint arXiv:2410.01174, 2024.
    [6] A. Wei, N. Haghtalab, and J. Steinhardt, "Jailbroken: How does llm safety training fail?," Advances in neural information processing systems, vol. 36, pp. 80079–80110, 2023.
    [7] Q. Ren et al., "Llms know their vulnerabilities: Uncover safety gaps through natural distribution shifts," in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pp. 24763–24785.
    [8] Z. Hao et al., "CHASE: Contextual History for Adaptive and Simple Exploitation in Large Language Model Jailbreaking," in Proceedings of the AAAI Conference on Artificial Intelligence, 2026, vol. 40, no. 1, pp. 345–353.
    [9] M. Russinovich, A. Salem, and R. Eldan, "Great, now write an article about that: The crescendo Multi-Turn LLM jailbreak attack," in 34th USENIX Security Symposium (USENIX Security 25), 2025, pp. 2421–2440.
    [10] D. Yin, H. Qiu, K.-H. Huang, K.-W. Chang, and N. Peng, "Safeworld: Geo-diverse safety alignment," Advances in Neural Information Processing Systems, vol. 37, pp. 128734–128768, 2024.
    [11] T. Mu et al., "Rule based rewards for language model safety," Advances in Neural Information Processing Systems, vol. 37, pp. 108877–108901, 2024.
    [12] C. Zheng et al., "On prompt-driven safeguarding for large language models," in Proceedings of the 41st International Conference on Machine Learning, 2024, vol. 235, pp. 61593–61613.
    [13] P. Röttger, H. Kirk, B. Vidgen, G. Attanasio, F. Bianchi, and D. Hovy, "Xstest: A test suite for identifying exaggerated safety behaviours in large language models," in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 5377–5400.
    [14] C. Shi et al., "Navigating the overkill in large language models," in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 4602–4614.
    [15] J. Cui, W.-L. Chiang, I. Stoica, and C.-J. Hsieh, "Or-bench: An over-refusal benchmark for large language models," in Proceedings of the 42nd International Conference on Machine Learning, 2025, vol. 267, pp. 11515–11542.
    [16] J. Dai et al., "Safe rlhf: Safe reinforcement learning from human feedback," in International Conference on Learning Representations, 2024.
    [17] H. Inan et al., "Llama guard: Llm-based input-output safeguard for human-ai conversations," arXiv preprint arXiv:2312.06674, 2023.
    [18] T. Rebedea, R. Dinu, M. N. Sreedhar, C. Parisien, and J. Cohen, "Nemo guardrails: A toolkit for controllable and safe llm applications with programmable rails," in Proceedings of the 2023 conference on empirical methods in natural language processing: system demonstrations, 2023, pp. 431–445.
    [19] J. Zhang, Y. Zhou, Y. Liu, Z. Li, and S. Hu, "Holistic automated red teaming for large language models through top-down test case generation and multi-turn interaction," in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 13711–13736.
    [20] Z. Weng, X. Jin, J. Jia, and X. Zhang, "Foot-in-the-door: A multi-turn jailbreak for llms," in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, pp. 1939–1950.
    [21] X. Shen, Z. Chen, M. Backes, Y. Shen, and Y. Zhang, "" do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models," in Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, 2024, pp. 1671–1685.
    [22] Y. Zeng, H. Lin, J. Zhang, D. Yang, R. Jia, and W. Shi, "How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms," in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 14322–14350.
    [23] G. Deng et al., "MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots," in Network and Distributed System Security Symposium, 2024, doi: 10.14722/ndss.2024.24188.
    [24] E. Perez et al., "Red teaming language models with language models," in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 3419–3448.
    [25] A. Mehrotra et al., "Tree of attacks: Jailbreaking black-box llms automatically," Advances in Neural Information Processing Systems, vol. 37, pp. 61065–61105, 2024.
    [26] A. Zou, Z. Wang, N. Carlini, M. Nasr, J. Z. Kolter, and M. Fredrikson, "Universal and transferable adversarial attacks on aligned language models," arXiv preprint arXiv:2307.15043, 2023.
    [27] P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong, "Jailbreaking black box large language models in twenty queries," in 2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), 2025, pp. 23–42.
    [28] N. F. Liu et al., "Lost in the middle: How language models use long contexts," Transactions of the association for computational linguistics, vol. 12, pp. 157–173, 2024.
    [29] S. Zhang et al., "JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation," in 34th USENIX Security Symposium (USENIX Security 25), 2025, pp. 8215–8234.
    [30] X. Wang et al., "SelfDefend:LLMs can defend themselves against jailbreaking in a practical manner," in 34th USENIX Security Symposium (USENIX Security 25), 2025, pp. 2441–2460.
    [31] E. Bassani and I. Sanchez, "Guardbench: A large-scale benchmark for guardrail models," in Proceedings of the 2024 conference on empirical methods in natural language processing, 2024, pp. 18393–18409.
    [32] W. Zeng et al., "Shieldgemma: Generative ai content moderation based on gemma," arXiv preprint arXiv:2407.21772, 2024.
    [33] S. Han et al., "Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms," Advances in neural information processing systems, vol. 37, pp. 8093–8131, 2024.
    [34] H. Tong et al., "Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks," arXiv preprint arXiv:2509.22732, 2025.
    [35] M. Liu, I. Baldini, D. Rabinowitz, D. S. Rosenberg, S. Gehrmann, and M. Dredze, "Domain generalizable AI guardrails with augmented policy training," in Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2026, pp. 16452–16469.
    [36] M. Kang et al., "Polyguard: Massive multi-domain safety policy-grounded guardrail dataset," Advances in Neural Information Processing Systems, vol. 38, 2025.
    [37] H. Jiang, Q. Wu, C.-Y. Lin, Y. Yang, and L. Qiu, "Llmlingua: Compressing prompts for accelerated inference of large language models," in Proceedings of the 2023 conference on empirical methods in natural language processing, 2023, pp. 13358–13376.
    [38] H. Jiang et al., "Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression," in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 1658–1677.
    [39] P. Chao et al., "Jailbreakbench: An open robustness benchmark for jailbreaking large language models," Advances in Neural Information Processing Systems, vol. 37, pp. 55005–55029, 2024.
    [40] W. Zhao, X. Ren, J. Hessel, C. Cardie, Y. Choi, and Y. Deng, "Wildchat: 1m chatgpt interaction logs in the wild," in International Conference on Learning Representations, 2024.
    [41] J. Chen, S. Xiao, P. Zhang, K. Luo, D. Lian, and Z. Liu, "M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation," in Findings of the Association for Computational Linguistics: ACL 2024, 2024, pp. 2318–2335, doi: 10.18653/v1/2024.findings-acl.137.
    [42] A. Souly et al., "A strongreject for empty jailbreaks," Advances in Neural Information Processing Systems, vol. 37, pp. 125416–125440, 2024.
    [43] M. Mazeika et al., "Harmbench: A standardized evaluation framework for automated red teaming and robust refusal," in Proceedings of the 41st International Conference on Machine Learning, 2024, vol. 235, pp. 35181–35224.
    [44] L. Zheng et al., "Judging llm-as-a-judge with mt-bench and chatbot arena," Advances in neural information processing systems, vol. 36, pp. 46595–46623, 2023.

    QR CODE