簡易檢索 / 詳目顯示

研究生: 蕭國樹
Hsiao, Kuo-Su
論文名稱: 高效能超純量處理器中指令激發單元的最佳化與取捨
Wakeup Logic Optimizations and Trade-off in High-performance Superscalar Processors
指導教授: 陳中和
Chen, Chung-Ho
學位類別: 博士
Doctor
系所名稱: 電機資訊學院 - 電機工程學系
Department of Electrical Engineering
論文出版年: 2006
畢業學年度: 94
語文別: 英文
論文頁數: 85
中文關鍵詞: 指令激發單元超純量處理器
外文關鍵詞: wakeup logic, superscalar
相關次數: 點閱:201下載:2
分享至:
查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報
  • 在高效能的超純量處理器中,由於複雜的指令激發動作,使得動態指令排程器變的更為複雜且愈加沒有擴展性。為了改進動態指令排程器的能量消耗、激發延遲、硬體成本以及擴展性,在本篇論文中我們提出了三種激發單元的最佳化。首先第一種方法是將來源標籤預先解碼,然後只將來源標籤直接與選中的回應訊號比對,如此可將目的標籤的讀取及多餘的比較從指令激發動作中移除。接著,第二種方法則是利用指令激發動作的區域性來將激發動作限制在小區段裡面,以達到降低負載電容及電路活動量的效果。第三種方法以激發位址來將指令重新排序,在指令佇列窗口中將指令依激發位址排列。在指令激發程序中,指令激發的動作只有在被目的標籤的激發位址選中的區段中才會運作。實驗結果顯示出三種提出的最佳化方案可以有效的節省功率消耗、降低激發延遲。另外從實驗結果也可以看出我們所提出的方案具有很好的擴展性。

    In a high-performance superscalar processor, the dynamic instruction scheduler often comes with poor scalability and high complexity due to the expensive instruction wakeup operation. This thesis presents three optimizations for wakeup logic to improve the power consumption, wakeup latency, area cost, and scalability. First, a wakeup design that pre-decodes the source tag is proposed. This design removes the reads of the destination tag and eliminates the redundant tag matches by matching the source tag directly with only the selected grant line. Next, the second design exploits the wakeup locality that most of the wakeup distances between two dependent instructions are short. By limiting the wakeup operation within a small wakeup range, the load capacitance and circuit activities can be alleviated. Third, a scheduling technique is proposed to schedule instructions into the segmented issue window based on their wakeup addresses. During wakeup process, the wakeup operation is only performed in the segment selected by the wakeup address of the result tag. The experimental results show that the proposed designs save the power consumption, and reduce the wakeup latency compared to the conventional designs. The results also show that the proposed designs have excellent scalability.

    摘要 IV ABSTRACT V ACKNOWLEDGMENTS VI CONTENTS VII LIST OF TABLES IX LIST OF FIGURES X CHAPTER 1 INTRODUCTION 1 1.1 POWER BUDGET FOR DESIGNING PROCESSORS 2 1.2 CRITICAL PATH OF PIPELINE STAGES 5 1.3 INSTRUCTIONS PER CYCLE CONSIDERATION 9 1.4 MAIN CONTRIBUTIONS 10 1.5 ORGANIZATION OF THE DISSERTATION 11 CHAPTER 2 BACKGROUND AND METHODOLOGY 12 2.1 SUPERSCALAR PROCESSOR MODELS: AN OVERVIEW 12 2.2 DYNAMIC SCHEDULING 14 2.3 TWO PRIMARY APPROACHES FOR WAKEUP LOGIC 15 2.3.1 CAM-based Approach 15 2.3.2 Gated-off Design 17 2.3.3 Matrix-based Scheme 18 2.4 RELATED WORK 20 2.5 EXPERIMENTAL METHODOLOGY 23 CHAPTER 3 SELECTIVE MATCH DESIGN 26 3.1 MOTIVATION 26 3.2 REDUCING ENERGY AND LATENCY VIA SELECTIVE MATCH 27 3.3 EXPERIMENTAL EVALUATION AND RESULTS 30 3.3.1 Power Consumption 31 3.3.2 Wakeup Latency 33 3.4 SUMMARY 36 CHAPTER 4 WAKEUP LOCALITY DESIGN 37 4.1 WAKEUP LOCALITY IN INSTRUCTION WAKEUP OPERATION 37 4.2 EXPLORING WAKEUP LOCALITY IN CAM-BASED WAKEUP DESIGN 38 4.3 EXPLORING WAKEUP LOCALITY IN MATRIX-BASED WAKEUP DESIGN 44 4.4 EXPLORING WAKEUP LOCALITY IN STP WAKEUP DESIGN 47 4.5 EXPERIMENTAL EVALUATION AND RESULTS 51 4.5.1 Performance 51 4.5.2 Power Consumption 54 4.5.3 Wakeup Latency 57 4.6 SUMMARY 59 CHAPTER 5 WAKEUP ADDRESS SCHEDULING 61 5.1 MOTIVATION 61 5.2 WAKEUP ADDRESS SCHEDULING 61 5.3 EXPERIMENTAL EVALUATION AND RESULTS 64 5.3.1 Performance 64 5.3.2 Power Consumption 66 5.3.3 Wakeup Latency 67 5.4 SUMMARY 69 CHAPTER 6 DISCUSSION AND TRADE-OFF 71 CHAPTER 7 CONCLUSIONS 78 REFERENCES 80 VITA 84 PUBLICATION LIST 85

    [1] S. Palacharla, N. P. Jouppi, and J. E. Smith, “Quantifying the Complexity of Superscalar Processors,” University of Wisconsin-Madison, Tech. Rep. CS-1328, May 1997.
    [2] K. Wilcox and S. Manne. “Alpha processors: A history of power issues and a look to the future,” in Proc. MICRO, Nov. 1999, Cool Chips Tutorial.
    [3] A. Kumar, “The HP PA8000 RISC CPU,” IEEE Micro, Vol. 17, Apr. 1997, pp. 27-32.
    [4] G. Hinton et al., “The Microarchitecture of the Pentium 4 Processor,” Intel Technology Journal, Feb. 2001.
    [5] K. C. Yeager, “MIPS R10000 Superscalar Microprocessor,” IEEE Micro, vol. 16, Apr. 1996, pp. 28-40.
    [6] R.E. Kessler, “The Alpha 21264 microprocessor,” IEEE Micro, vol. 19, Apr. 1999, pp. 24-36.
    [7] M. Butler and Y. N. Patt, “An Investigation of the Performance of Various Dynamic Scheduling Techniques,” in Proc. MICRO, Dec. 1992, pp. 1-9.
    [8] S. T. Srinivasan and A. R. Lebeck. “Load Latency Tolerance in Dynamically Scheduled Processors,” in Proc. MICRO, Dec. 1998, pp. 148-159.
    [9] L. Gwennap, “Intel’s P6 Uses Decoupled Superscalar Design,” Microprocessor Report, vol. 9, no. 2, Feb. 1995, pp. 1-7.
    [10] S. P. Song, M. Denman, and J. Chang, “The PowerPC 604 RISC Microprocessor,” IEEE Micro, vol. 14, Oct. 1994, pp. 8-17.
    [11] L. Gwennap, “HAL Reveals Multichip SPARC Processor,” Microprocessor Report, vol. 9 no. 3, Mar. 1995, pp. 1-7.
    [12] J. L. Hennessy and D. A. Patterson, Computer Architecture: A Quantitative Approach, 2nd ed., San Francisco, CA: Morgan Kaufmann Publishers, Inc., 1996.
    [13] M. Brown, J. Stark, and Y. Patt. “Select-Free Instruction Scheduling Logic,” in Proc. MICRO, Dec. 2001, pp. 204-213.
    [14] M. Goshima et al., “A High-Speed Dynamic Instruction Scheduling Scheme for Superscalar Processors,” in Proc. MICRO, Dec. 2001, pp. 225-236.
    [15] R. Ho, K. W. Mai, and M. A. Horowitz, “The Future of Wires,” Proceedings of the IEEE, vol. 89, Apr. 2001, pp. 490-504.
    [16] M. S. Hrishikesh, N. P. Jouppi, and K. I. Farkas, “The optimal useful logic depth per pipeline stages is 6-8 FO4,” in Proc. ISCA, May 2002, pp. 14-24.
    [17] D. Folegnani and A. Gonzalez, “Energy-Effective Issue Logic,” in Proc. ISCA, Jul. 2001, pp. 230-239.
    [18] M. A. Ramírez et al., "A Simple Low-Energy Instruction Wakeup Mechanism," in International Symposium on High-Performance Computing (ISHPC), Oct. 2003, pp. 99–112.
    [19] D. Ponomarev, G. Kucuk, and K. Ghose, “Reducing Power Requirements of Instruction Scheduling Through Dynamic Allocation of Multiple Datapath Resources,” in Proc. MICRO, Dec. 2001, pp. 90-101.
    [20] J. Abella and A. González, “Power-Aware Adaptive Issue Queue and Register File,” in Proc. Int. Conf. High-Performance Computing (HiPC), Dec. 2003.
    [21] David H. Albonesi. “Dynamic IPC/Clock Rate Optimization,” in Proc. ISCA, June 1998, pp. 282–292.
    [22] A. Buyuktosunoglu et al., ”A Circuit Level Implementation of an Adaptive Issue Queue for Poweraware microprocessors,” in Proc GLVSLSI, Mar. 2001, pp. 73-83.
    [23] S. Dropsho et al., “Integrating Adaptive On- Chip Storage Structures for Reduced Dynamic Power,” in Proc. 11th Parallel Architectures and Compilation Techniques, Sep. 2002, pp. 141-152.
    [24] D. Ernst and T. M. Austin, “Efficient dynamic scheduling through tag elimination,” in Proc. ISCA, May 2002, pp. 37-46.
    [25] J. J. Sharkey et al., “Instruction packing: reducing power and delay of the dynamic scheduling logic,” in Proc. ISLPED, Aug. 2005, pp. 30-35.
    [26] I. Kim and M. H. Lipasti, “Half-Price Architecture,” in Proc. ISCA, Jun. 2003, pp. 28-38.
    [27] A. Aggarwal, et. al., “Defining Wakeup Width for Efficient Dynamic Scheduling,” in Proc. ICCD, Oct. 2004, pp. 36-41.
    [28] D. Ernst, A. Hamel, and T. Austin, “Cyclone: A Broadcast-Free Dynamic Instruction Scheduler with Selective Replay,” in Proc. ISCA, Jun. 2003, pp. 253-262.
    [29] J. Hu, N. Vijaykrishnan, and M. Irwin, “Exploring Wakeup-Free Instruction Scheduling,” in Proc. HPCA, Feb. 2004, pp. 232-241.
    [30] A.R. Lebeck et al., “A Large, Fast Instruction Window for Tolerating Cache Misses,” in Proc. ISCA, May 2002, pp. 59-70.
    [31] B. Fields, S. Rubin, and R. Bodík, “Focusing Processor Policies via Critical-Path Prediction,” in Proc. ISCA, Jul. 2001, pp. 74-85.
    [32] E. Brekelbaum et al., “Hierarchical Scheduling Windows,” in Proc. MICRO, Nov. 2002, pp. 27-36.
    [33] D. S. Henry, B. C. Kuszmaul, G. H. Loh, and R. Sami, “Circuits for Wide-Window Superscalar Processors,” in Proc. ISCA, Jun. 2000, pp. 236-247.
    [34] K. S. Hsiao and C. H. Chen, "An Efficient Wakeup Design for Energy Reduction in High-Performance Superscalar Processors," in Int. Con. Computing Frontiers (CF), May 2005, pp. 353-360.
    [35] D. V. Ponomarev et al., “Energy-Efficient Issue Queue Design,” in IEEE Transactions on Very Large Scale Integration (VLSI) Systems, vol. 11, Oct. 2003, pp. 789-800.
    [36] M. Huang, J. Renau, and J. Torrellas, “Energy-Efficient Hybrid Wakeup Logic,” in Proc. ISLPED, Aug. 2002, pp. 196-201.
    [37] R. Canal and A. González, “A Low-Complexity Issue Logic,” in Proc. ICS, May 2000, pp. 327-335.
    [38] R. Canal and A. Gonzalez, “Reducing the Complexity of the Issue Logic,” in Proc. ICS, Jun. 2001, pp. 312-320.
    [39] S. Palacharla, N. P. Jouppi, and J. E. Smith, “Complexity-effective superscalar processors,” in Proc. ISCA, Jun. 1997, pp. 206-218.
    [40] P. Michaud and A. Seznec, “Data-flow prescheduling for large instruction windows in out-of-order processors,” in Proc. HPCA, Jan. 2001, pp. 27-36.
    [41] S. E. Raasch, N. L. Binkert, and S. K. Reinhardt, “A Scalable Instruction Queue Design Using Dependence Chains,” in Proc. ISCA, May 2002, pp. 318-329.
    [42] D. Brooks, V. Tiwari, and M. Martonosi, “Wattch: A framework for architectural-level power analysis and optimizations,” in Proc. ISCA, Jun. 2000, pp. 83-94.
    [43] D. Burger and T. M. Austin, “The SimpleScalar tool set, version 2.0,” University of Wisconsin-Madison, Tech. Rep. CS-1342, Jun. 1997.
    [44] C. Lee, M. Potkonjak, and W. Mangione-Smith, “MediaBench: A Tool for Evaluating Multimedia and Communications Systems,” in Proc. MICRO, Dec. 1997, pp. 330-335.
    [45] K. Pagiamtzis and A. Sheikholeslami, “Content-Addressable Memory (CAM) Circuits and Architecture: A Tutorial and Survey,” IEEE Journal of Solid-State Circuits, Vol. 41, No.3, Mar. 2006, pp. 712-727.
    [46] S. R. Kunkel and J. E. Smith, “Optimal Pipelining in Supercomputers,” in Proc. ISCA, Jun. 1986, pp. 404-413.
    [47] N. P. Jouppi and D. W. Wall, “Available Instruction-Level Parallelism for Superscalar and Superpipelined Machines,” in Proceedings of the International Conference on Architectural Support for Programming Languages and Operating Systems, Apr. 1989, pp. 272-282.
    [48] P. K. Dubey and M. J. Flynn, “Optimal Pipelining,” in Journal of Parallel and Distributed Computing, Vol. 8, 1990, pp. 10-19.
    [49] Robert J. Proebsting, “Speed Enhancement Technique for CMOS Circuits,” in United States Patent No. 4,985,643, Jan. 1991.
    [50] T. Chappell, “A 2ns cycle, 4 ns access 512kb CMOS ECL SRAM,” in IEEE International Sold-State Circuits Conference Digest of Technical Papers, Feb. 1991, pp. 50-51.
    [51] D. M. Tullsen et al., “Exploiting Choice: Instruction Fetch and Issue on an Implementable Simultaneous Multithreading Processor,” in Proc. ISCA, May 1996, pp. 191-202.
    [52] J. Keller, “The 21264: A Superscalar Alpha Processor with Out-of-Order Execution,” Presentation at the 9th Annual Microprocessor Forum, Oct. 1996.
    [53] D. Dobberpuhl et al., “A 200-MHz 64-b dual-issue CMOS Microprocessor,” in IEEE Journal of Solid-State Circuits, Vol: 27, 1992, pp. 1555-1557.
    [54] J. L. Henning, “SPEC CPU2000: Measuring CPU performance in the new millennium,” IEEE Computer, Vol: 33, 2000, pp.28-35.
    [55] F. Pollack, “New Microarchitecture Challenges in the Coming Generations of CMOS Process Technologies,” in Proc. MICRO, Nov. 1999, keynote speech.
    [56] S.H. Gunther et al., “Managing the Impact of Increasing Microprocessor Power Consumption,” Intel Technology Journal, Feb. 2001.
    [57] G. Reinman and N. P. Jouppi, “CACTI 2.0: An Integrated Cache Timing and Power Model,” COMPAQ Western Research Lab, Palo Alto, CA, Tech. Rep., Feb. 2000.

    下載圖示
    2007-09-04公開
    QR CODE