Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 29 for “"superscalar"”.

  1. Adding limited reconfigurability to superscalar processors

    … building blocks for scientific supercomputers. Superscalar processors such as these require ever more processing power to run complex simulations, video games or versatile telecoms services. In the case of embedded applications, e.g. for portable devices, both performance and power consumption …

    epfl Repository record for Adding limited reconfigurability to superscalar processors (opens in a new tab)

  2. Superscalar processors via automatic microarchitecture transformations

    Thesis (S.B. and M.Eng.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 2000.

    mit Repository record for Superscalar processors via automatic microarchitecture transformations (opens in a new tab)

  3. Data preload for superscalar and VLIW processors

    Processor design techniques, such as pipelining, superscalar, and VLIW, have dramatically decreased the average number of clock cycles per instruction. As a result, each execution cycle has become more significant to overall system performance. To maximize the effectiveness of each cycle, one must …

    uiuc Repository record for Data preload for superscalar and VLIW processors (opens in a new tab)

  4. Banked microarchitectures for complexity-effective superscalar microprocessors

    High performance superscalar microarchitectures exploit instruction-level parallelism (ILP) to improve processor performance by executing instructions out of program order and by speculating on branch instructions. Monolithic centralized structures with global communications, including issue …

    mit Repository record for Banked microarchitectures for complexity-effective superscalar microprocessors (opens in a new tab)

  5. Chip Multiprocessors With Speculative Multithreading: Design for Performance and Energy Efficiency

    … for SM delivers a speedup of 1.27 over a 3-issue superscalar. The SM CMP is even faster than a 6-issue superscalar at the same frequency, and consumes only 85% of its power. In fact, for the same average power in both chips, the SM CMP is 1.13 times faster than the 6-issue superscalar on average.

    uiuc Repository record for Chip Multiprocessors With Speculative Multithreading: Design for Performance and Energy Efficiency (opens in a new tab)

  6. Speculative Multithreading Architectures

    … However, using direct-execution to model a superscalar processor has been considered an open problem. In the last part of the thesis, we propose a novel direct-execution framework that allows accurate simulation of wide-issue superscalar processors for both a uni- and multi-processor …

    uiuc Repository record for Speculative Multithreading Architectures (opens in a new tab)

  7. Performance evaluation of vector machine architectures

    … it has been claimed that several horizontal superscalar architectures, e.g., VLIW and polycyclic architectures, provide a more balanced performance across a wider range of scientific workloads than do vector machines. The purpose of this research is to study the performance of …

    uiuc Repository record for Performance evaluation of vector machine architectures (opens in a new tab)

  8. Development and evaluation of a simultaneous multithreading processor simulator

    … SMT simulator is derived from SimpleScalar, a superscalar processor simulator widely used in the computer architecture research field. The basic pipeline is expanded to allow multiple threads to be fetched, dispatched, issued, executed, and committed simultaneously. Benchmarks that were …

    njit Repository record for Development and evaluation of a simultaneous multithreading processor simulator (opens in a new tab)

  9. Memory disambiguation to facilitate instruction-level parallelism compilation

    … (ILP) to make effective use of wide-issue superscalar and VLIW processor resources, the compiler must perform aggressive low-level code optimization and scheduling. However, ambiguous memory dependences can significantly limit the compiler's ability to expose ILP. To overcome the problem of …

    uiuc Repository record for Memory disambiguation to facilitate instruction-level parallelism compilation (opens in a new tab)

  10. Enhancing instruction-level parallelism through compiler-controlled speculation

    … programs (1) (2) (3). An effective VLIW or superscalar processor must optimize and schedule instructions across basic block boundaries to achieve higher performance. An effective structure for ILP compilation is the superblock (4). The formation and optimization of superblocks increase ILP …

    uiuc Repository record for Enhancing instruction-level parallelism through compiler-controlled speculation (opens in a new tab)

  11. Efficient spatial and temporal safety for microcontrollers and application-class processors

    … and the first open application of CHERI to a superscalar processor. Tradeoffs in developing the architecture and performant microarchitecture are investigated. The processors are then used as a platform to conduct research in reducing the overheads for achieving temporal safety with CHERI. …

    cambridge Repository record for Efficient spatial and temporal safety for microcontrollers and application-class processors (opens in a new tab)

  12. Compiler-assisted debugging and multiple instruction retry

    … execution can achieve significant speedup for superscalar architectures by utilizing instruction retry to roll back the computation above mispredicted branches. Debugging is another field that can benefit from backward execution. Allowing the user to undo several instructions at specific …

    uiuc Repository record for Compiler-assisted debugging and multiple instruction retry (opens in a new tab)

  13. Processor parallelism considerations and memory latency reduction in shared memory multiprocessors

    … compare the instruction-level parallelism of a superscalar processor and a pipelined processor with the loop-level parallelism of a shared memory multiprocessor. It is shown that the maximum speedup for a loop with a cyclic dependence graph is limited by its critical dependence ratio, …

    uiuc Repository record for Processor parallelism considerations and memory latency reduction in shared memory multiprocessors (opens in a new tab)

  14. AXCIS : rapid processor architectural exploration using canonical instruction segments

    … (IPC). This thesis applies AXCIS to in-order superscalar processors, which are becoming more popular with the emergence of chip multiprocessors (CMP). For 24 SPEC CPU2000 benchmarks and all simulated configurations, AXCIS achieves an average IPC error of 2.6% and is over four orders of …

    mit Repository record for AXCIS : rapid processor architectural exploration using canonical instruction segments (opens in a new tab)

  15. Architectural support for massively-concurrent parallel computing

    … queueing theory, we analyze conventional RISC superscalar processors as a case study of the inadequacy of the class of "ILP" processors. We contrast this to multithreaded processors that exploit both ILP and Thread-Level Parallelism (TLP). As a contribution to parallel programming, we show how …

    concordia Repository record for Architectural support for massively-concurrent parallel computing (opens in a new tab)

  16. Systematic Computer Architecture Prototyping

    … due to multiprogramming. Prototyping of superscalar processors is performed using new statistical sampling techniques in conjunction with two simulation algorithms. The resource usage in an unlimited-resource simulation is used to select processor resource needs, but is only applicable to …

    uiuc Repository record for Systematic Computer Architecture Prototyping (opens in a new tab)

  17. Cluster assignment and instruction scheduling for partitioned register-set machines

    … of their successful designs such as VLIW and superscalar machines use multiple functional units trying to exploit instruction level parallelism in computer programs. As the number of functional units rises, another hardware constraint enters the picture---the number of register-file ports …

    rice Repository record for Cluster assignment and instruction scheduling for partitioned register-set machines (opens in a new tab)

  18. Configurable computer systems can support dataflow computing

    … is the result of a simple, pipelined and superscalar architecture with a very wide data bus and a completely out of order execution of instructions. This creates a program counter less, distributed controlled system design with the realization of intelligent memories. Upon the availability …

    njit Repository record for Configurable computer systems can support dataflow computing (opens in a new tab)

Page 1 of 2