| 研究生: |
蕭國樹 Hsiao, Kuo-Su |
|---|---|
| 論文名稱: |
高效能超純量處理器中指令激發單元的最佳化與取捨 Wakeup Logic Optimizations and Trade-off in High-performance Superscalar Processors |
| 指導教授: |
陳中和
Chen, Chung-Ho |
| 學位類別: |
博士 Doctor |
| 系所名稱: |
電機資訊學院 - 電機工程學系 Department of Electrical Engineering |
| 論文出版年: | 2006 |
| 畢業學年度: | 94 |
| 語文別: | 英文 |
| 論文頁數: | 85 |
| 中文關鍵詞: | 指令激發單元 、超純量處理器 |
| 外文關鍵詞: | wakeup logic, superscalar |
| 相關次數: | 點閱:201 下載:2 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
在高效能的超純量處理器中,由於複雜的指令激發動作,使得動態指令排程器變的更為複雜且愈加沒有擴展性。為了改進動態指令排程器的能量消耗、激發延遲、硬體成本以及擴展性,在本篇論文中我們提出了三種激發單元的最佳化。首先第一種方法是將來源標籤預先解碼,然後只將來源標籤直接與選中的回應訊號比對,如此可將目的標籤的讀取及多餘的比較從指令激發動作中移除。接著,第二種方法則是利用指令激發動作的區域性來將激發動作限制在小區段裡面,以達到降低負載電容及電路活動量的效果。第三種方法以激發位址來將指令重新排序,在指令佇列窗口中將指令依激發位址排列。在指令激發程序中,指令激發的動作只有在被目的標籤的激發位址選中的區段中才會運作。實驗結果顯示出三種提出的最佳化方案可以有效的節省功率消耗、降低激發延遲。另外從實驗結果也可以看出我們所提出的方案具有很好的擴展性。
In a high-performance superscalar processor, the dynamic instruction scheduler often comes with poor scalability and high complexity due to the expensive instruction wakeup operation. This thesis presents three optimizations for wakeup logic to improve the power consumption, wakeup latency, area cost, and scalability. First, a wakeup design that pre-decodes the source tag is proposed. This design removes the reads of the destination tag and eliminates the redundant tag matches by matching the source tag directly with only the selected grant line. Next, the second design exploits the wakeup locality that most of the wakeup distances between two dependent instructions are short. By limiting the wakeup operation within a small wakeup range, the load capacitance and circuit activities can be alleviated. Third, a scheduling technique is proposed to schedule instructions into the segmented issue window based on their wakeup addresses. During wakeup process, the wakeup operation is only performed in the segment selected by the wakeup address of the result tag. The experimental results show that the proposed designs save the power consumption, and reduce the wakeup latency compared to the conventional designs. The results also show that the proposed designs have excellent scalability.
[1] S. Palacharla, N. P. Jouppi, and J. E. Smith, “Quantifying the Complexity of Superscalar Processors,” University of Wisconsin-Madison, Tech. Rep. CS-1328, May 1997.
[2] K. Wilcox and S. Manne. “Alpha processors: A history of power issues and a look to the future,” in Proc. MICRO, Nov. 1999, Cool Chips Tutorial.
[3] A. Kumar, “The HP PA8000 RISC CPU,” IEEE Micro, Vol. 17, Apr. 1997, pp. 27-32.
[4] G. Hinton et al., “The Microarchitecture of the Pentium 4 Processor,” Intel Technology Journal, Feb. 2001.
[5] K. C. Yeager, “MIPS R10000 Superscalar Microprocessor,” IEEE Micro, vol. 16, Apr. 1996, pp. 28-40.
[6] R.E. Kessler, “The Alpha 21264 microprocessor,” IEEE Micro, vol. 19, Apr. 1999, pp. 24-36.
[7] M. Butler and Y. N. Patt, “An Investigation of the Performance of Various Dynamic Scheduling Techniques,” in Proc. MICRO, Dec. 1992, pp. 1-9.
[8] S. T. Srinivasan and A. R. Lebeck. “Load Latency Tolerance in Dynamically Scheduled Processors,” in Proc. MICRO, Dec. 1998, pp. 148-159.
[9] L. Gwennap, “Intel’s P6 Uses Decoupled Superscalar Design,” Microprocessor Report, vol. 9, no. 2, Feb. 1995, pp. 1-7.
[10] S. P. Song, M. Denman, and J. Chang, “The PowerPC 604 RISC Microprocessor,” IEEE Micro, vol. 14, Oct. 1994, pp. 8-17.
[11] L. Gwennap, “HAL Reveals Multichip SPARC Processor,” Microprocessor Report, vol. 9 no. 3, Mar. 1995, pp. 1-7.
[12] J. L. Hennessy and D. A. Patterson, Computer Architecture: A Quantitative Approach, 2nd ed., San Francisco, CA: Morgan Kaufmann Publishers, Inc., 1996.
[13] M. Brown, J. Stark, and Y. Patt. “Select-Free Instruction Scheduling Logic,” in Proc. MICRO, Dec. 2001, pp. 204-213.
[14] M. Goshima et al., “A High-Speed Dynamic Instruction Scheduling Scheme for Superscalar Processors,” in Proc. MICRO, Dec. 2001, pp. 225-236.
[15] R. Ho, K. W. Mai, and M. A. Horowitz, “The Future of Wires,” Proceedings of the IEEE, vol. 89, Apr. 2001, pp. 490-504.
[16] M. S. Hrishikesh, N. P. Jouppi, and K. I. Farkas, “The optimal useful logic depth per pipeline stages is 6-8 FO4,” in Proc. ISCA, May 2002, pp. 14-24.
[17] D. Folegnani and A. Gonzalez, “Energy-Effective Issue Logic,” in Proc. ISCA, Jul. 2001, pp. 230-239.
[18] M. A. Ramírez et al., "A Simple Low-Energy Instruction Wakeup Mechanism," in International Symposium on High-Performance Computing (ISHPC), Oct. 2003, pp. 99–112.
[19] D. Ponomarev, G. Kucuk, and K. Ghose, “Reducing Power Requirements of Instruction Scheduling Through Dynamic Allocation of Multiple Datapath Resources,” in Proc. MICRO, Dec. 2001, pp. 90-101.
[20] J. Abella and A. González, “Power-Aware Adaptive Issue Queue and Register File,” in Proc. Int. Conf. High-Performance Computing (HiPC), Dec. 2003.
[21] David H. Albonesi. “Dynamic IPC/Clock Rate Optimization,” in Proc. ISCA, June 1998, pp. 282–292.
[22] A. Buyuktosunoglu et al., ”A Circuit Level Implementation of an Adaptive Issue Queue for Poweraware microprocessors,” in Proc GLVSLSI, Mar. 2001, pp. 73-83.
[23] S. Dropsho et al., “Integrating Adaptive On- Chip Storage Structures for Reduced Dynamic Power,” in Proc. 11th Parallel Architectures and Compilation Techniques, Sep. 2002, pp. 141-152.
[24] D. Ernst and T. M. Austin, “Efficient dynamic scheduling through tag elimination,” in Proc. ISCA, May 2002, pp. 37-46.
[25] J. J. Sharkey et al., “Instruction packing: reducing power and delay of the dynamic scheduling logic,” in Proc. ISLPED, Aug. 2005, pp. 30-35.
[26] I. Kim and M. H. Lipasti, “Half-Price Architecture,” in Proc. ISCA, Jun. 2003, pp. 28-38.
[27] A. Aggarwal, et. al., “Defining Wakeup Width for Efficient Dynamic Scheduling,” in Proc. ICCD, Oct. 2004, pp. 36-41.
[28] D. Ernst, A. Hamel, and T. Austin, “Cyclone: A Broadcast-Free Dynamic Instruction Scheduler with Selective Replay,” in Proc. ISCA, Jun. 2003, pp. 253-262.
[29] J. Hu, N. Vijaykrishnan, and M. Irwin, “Exploring Wakeup-Free Instruction Scheduling,” in Proc. HPCA, Feb. 2004, pp. 232-241.
[30] A.R. Lebeck et al., “A Large, Fast Instruction Window for Tolerating Cache Misses,” in Proc. ISCA, May 2002, pp. 59-70.
[31] B. Fields, S. Rubin, and R. Bodík, “Focusing Processor Policies via Critical-Path Prediction,” in Proc. ISCA, Jul. 2001, pp. 74-85.
[32] E. Brekelbaum et al., “Hierarchical Scheduling Windows,” in Proc. MICRO, Nov. 2002, pp. 27-36.
[33] D. S. Henry, B. C. Kuszmaul, G. H. Loh, and R. Sami, “Circuits for Wide-Window Superscalar Processors,” in Proc. ISCA, Jun. 2000, pp. 236-247.
[34] K. S. Hsiao and C. H. Chen, "An Efficient Wakeup Design for Energy Reduction in High-Performance Superscalar Processors," in Int. Con. Computing Frontiers (CF), May 2005, pp. 353-360.
[35] D. V. Ponomarev et al., “Energy-Efficient Issue Queue Design,” in IEEE Transactions on Very Large Scale Integration (VLSI) Systems, vol. 11, Oct. 2003, pp. 789-800.
[36] M. Huang, J. Renau, and J. Torrellas, “Energy-Efficient Hybrid Wakeup Logic,” in Proc. ISLPED, Aug. 2002, pp. 196-201.
[37] R. Canal and A. González, “A Low-Complexity Issue Logic,” in Proc. ICS, May 2000, pp. 327-335.
[38] R. Canal and A. Gonzalez, “Reducing the Complexity of the Issue Logic,” in Proc. ICS, Jun. 2001, pp. 312-320.
[39] S. Palacharla, N. P. Jouppi, and J. E. Smith, “Complexity-effective superscalar processors,” in Proc. ISCA, Jun. 1997, pp. 206-218.
[40] P. Michaud and A. Seznec, “Data-flow prescheduling for large instruction windows in out-of-order processors,” in Proc. HPCA, Jan. 2001, pp. 27-36.
[41] S. E. Raasch, N. L. Binkert, and S. K. Reinhardt, “A Scalable Instruction Queue Design Using Dependence Chains,” in Proc. ISCA, May 2002, pp. 318-329.
[42] D. Brooks, V. Tiwari, and M. Martonosi, “Wattch: A framework for architectural-level power analysis and optimizations,” in Proc. ISCA, Jun. 2000, pp. 83-94.
[43] D. Burger and T. M. Austin, “The SimpleScalar tool set, version 2.0,” University of Wisconsin-Madison, Tech. Rep. CS-1342, Jun. 1997.
[44] C. Lee, M. Potkonjak, and W. Mangione-Smith, “MediaBench: A Tool for Evaluating Multimedia and Communications Systems,” in Proc. MICRO, Dec. 1997, pp. 330-335.
[45] K. Pagiamtzis and A. Sheikholeslami, “Content-Addressable Memory (CAM) Circuits and Architecture: A Tutorial and Survey,” IEEE Journal of Solid-State Circuits, Vol. 41, No.3, Mar. 2006, pp. 712-727.
[46] S. R. Kunkel and J. E. Smith, “Optimal Pipelining in Supercomputers,” in Proc. ISCA, Jun. 1986, pp. 404-413.
[47] N. P. Jouppi and D. W. Wall, “Available Instruction-Level Parallelism for Superscalar and Superpipelined Machines,” in Proceedings of the International Conference on Architectural Support for Programming Languages and Operating Systems, Apr. 1989, pp. 272-282.
[48] P. K. Dubey and M. J. Flynn, “Optimal Pipelining,” in Journal of Parallel and Distributed Computing, Vol. 8, 1990, pp. 10-19.
[49] Robert J. Proebsting, “Speed Enhancement Technique for CMOS Circuits,” in United States Patent No. 4,985,643, Jan. 1991.
[50] T. Chappell, “A 2ns cycle, 4 ns access 512kb CMOS ECL SRAM,” in IEEE International Sold-State Circuits Conference Digest of Technical Papers, Feb. 1991, pp. 50-51.
[51] D. M. Tullsen et al., “Exploiting Choice: Instruction Fetch and Issue on an Implementable Simultaneous Multithreading Processor,” in Proc. ISCA, May 1996, pp. 191-202.
[52] J. Keller, “The 21264: A Superscalar Alpha Processor with Out-of-Order Execution,” Presentation at the 9th Annual Microprocessor Forum, Oct. 1996.
[53] D. Dobberpuhl et al., “A 200-MHz 64-b dual-issue CMOS Microprocessor,” in IEEE Journal of Solid-State Circuits, Vol: 27, 1992, pp. 1555-1557.
[54] J. L. Henning, “SPEC CPU2000: Measuring CPU performance in the new millennium,” IEEE Computer, Vol: 33, 2000, pp.28-35.
[55] F. Pollack, “New Microarchitecture Challenges in the Coming Generations of CMOS Process Technologies,” in Proc. MICRO, Nov. 1999, keynote speech.
[56] S.H. Gunther et al., “Managing the Impact of Increasing Microprocessor Power Consumption,” Intel Technology Journal, Feb. 2001.
[57] G. Reinman and N. P. Jouppi, “CACTI 2.0: An Integrated Cache Timing and Power Model,” COMPAQ Western Research Lab, Palo Alto, CA, Tech. Rep., Feb. 2000.