Intel 64 and IA-32 Architectures. Software Developer’s Manual (Collection, 2023) - page 140

 

  Index      Manuals     Intel 64 and IA-32 Architectures. Software Developer’s Manual (Collection, 2023)

 

Search            copyright infringement  

 

   

 

   

 

Content      ..     138      139      140      141     ..

 

 

 

Intel 64 and IA-32 Architectures. Software Developer’s Manual (Collection, 2023) - page 140

 

 

Intel® 64 and IA-32 Architectures
Optimization Reference Manual
Documentation Changes
May 2023
Document Number: 355308-001
Contents
Revision History
4
Preface
5
Nomenclature
5
Summary Tables of Changes
5
Documentation Changes
5
Intel® 64 and IA-32 Architectures Optimization Reference Manual Documentation Changes
3
Revision History
Revision History
Revision
Description
Date
-001
Initial release
May 2023
4
Intel® 64 and IA-32 Architectures Optimization Reference Manual Documentation Changes
Preface
This document is an update to the optimization recommendations contained in the Intel® 64 and IA-32
Architectures Optimization Reference Manual, also known as the Software Optimization Manual. This document
is a compilation of device and documentation errata, specification clarifications and changes. It is intended for
hardware system manufacturers and software developers of applications, operating systems, or tools.
Nomenclature
Documentation Changes include typos, errors, or omissions from the current published specifications. These
will be incorporated in any new release of the specification.
Summary Tables of Changes
The following table indicates documentation changes which apply to the Intel® 64 and IA-32 Architecture
software optimization topics covered by this reference manual.
No.
DOCUMENTATION CHANGES
1
Updates to Chapter 1
2
Updates to Chapter 2
3
Updates to Chapter 3
4
Updates to Chapter 7
5
Updates to Chapter 10
6
Updates to Chapter 11
7
Updates to Chapter 15
8
Updates to Chapter 18
9
Updates to Chapter 20
Documentation Changes
Changes to the Intel® 64 and IA-32 Architectures Optimization Reference Manual volumes follow, and are listed
by chapter. Only chapters with changes are included in this document.
Intel® 64 and IA-32 Architectures Optimization Reference Manual Documentation Changes
5
1.
Updates to Chapter 1
Change bars and violet text show changes to Chapter 1 of the Intel® 64 and IA-32 Architectures Optimization
Reference Manual: Introduction.
------------------------------------------------------------------------------------------
Changes to this chapter:
• Section 1.2:
— References to Intel® Xeon® Scalable Processor Family were updated.
— 13th generation Intel® Core™ processor section was added.
— Updated trademarking as necessary.
Intel® 64 and IA-32 Architectures Optimization Reference Manual Documentation Changes
6
CHAPTER 1
INTRODUCTION
The Intel® 64 and IA-32 Architectures Optimization Reference Manual describes how to optimize soft-
ware to take advantage of the performance characteristics of IA-32 and Intel 64 architecture processors.
The target audience for this manual includes software programmers and compiler writers. This manual
assumes that the reader is familiar with the basics of the IA-32 architecture and has access to the Intel®
64 and IA-32 Architectures Software Developer’s Manual. A detailed understanding of Intel 64 and IA-32
processors is often required. In many cases, knowledge of the underlying microarchitectures is required.
The design guidelines discussed in this manual for developing high-performance software generally
apply to current and future IA-32 and Intel 64 processors. In most cases, coding rules apply to software
running in 64-bit mode of Intel 64 architecture, compatibility mode of Intel 64 architecture, and IA-32
modes (IA-32 modes are supported in IA-32 and Intel 64 architectures). Coding rules specific to 64-bit
modes are noted separately.
NOTE
A public repository is available with open source code samples from select chapters of
this manual. These code samples are released under a 0-Clause BSD license. Intel
provides additional code samples and updates to the repository as the samples are
created and verified.
Public repository: https://github.com/intel/optimization-manual.
Link to license: https://github.com/intel/optimization-manual/blob/master/COPYING.
1.1
TUNING YOUR APPLICATION
Tuning an application for high performance on any Intel 64 or IA-32 processor requires understanding
and basic skills in:
Intel 64 and IA-32 architecture.
C and Assembly language.
Hot-spot regions in the application that impact performance.
Optimization capabilities of the compiler.
Techniques used to evaluate application performance.
The Intel® VTune Performance Analyzer can help you analyze and locate hot-spot regions in your
applications. On the Intel® Core™ i7, Intel® Core™2 Duo, Intel® Core™ Duo, Intel® Core™ Solo,
Pentium® 4, Intel® Xeon®, and Intel® Pentium® M processors, this tool can monitor an application
through a selection of performance monitoring events and analyze the performance event data that is
gathered during code execution.
This manual also describes data that can be gathered using the performance counters through the
processor’s performance monitoring events.
1.2
ABOUT THIS MANUAL
The Intel® Core™ i7 processor and Intel® Xeon® processor 3400, 5500, 7500 series are based on 45 nm
Nehalem microarchitecture. Westmere microarchitecture is a 32 nm version of the Nehalem microarchi-
tecture. Intel® Xeon® processor 5600 series, Intel Xeon processor E7 and various Intel Core i7, i5, i3
processors are based on the Westmere microarchitecture. These processors support Intel 64 architec-
ture.
INTRODUCTION
The Intel® Xeon® processor E5 family, Intel® Xeon® processor E3-1200 family, Intel® Xeon® processor
E7-8800/4800/2800 product families, Intel® Core™ i7-3930K processor, and 2nd generation Intel®
Core™ i7-2xxx, Intel® Core™ i5-2xxx, Intel® Core™ i3-2xxx processor series are based on the Sandy
Bridge microarchitecture and support Intel 64 architecture.
The Intel® Xeon® processor E7-8800/4800/2800 v2 product families, Intel® Xeon® processor E3-1200
v2 product family and 3rd generation Intel® Core™ processors are based on the Ivy Bridge microarchi-
tecture and support Intel 64 architecture.
The Intel® Xeon® processor E5-4600/2600/1600 v2 product families, Intel® Xeon® processor E5-
2400/1400 v2 product families, and Intel® Core™ i7-49xx Processor Extreme Edition are based on the
Ivy Bridge-E microarchitecture and support Intel 64 architecture.
The Intel® Xeon® processor E3-1200 v3 product family and 4th Generation Intel® Core™ processors are
based on the Haswell microarchitecture and support Intel 64 architecture.
The Intel® Xeon® processor E5-2600/1600 v3 product families and the Intel® Core™ i7-59xx Processor
Extreme Edition are based on the Haswell-E microarchitecture and support Intel 64 architecture.
The Intel Atom® processor Z8000 series is based on the Airmont microarchitecture.
The Intel Atom® processor Z3400 series and the Intel Atom® processor Z3500 series are based on the
Silvermont microarchitecture.
The Intel® Core™ M processor family, 5th generation Intel® Core™ processors, Intel® Xeon® processor
D-1500 product family and the Intel® Xeon® processor E5 v4 family are based on the Broadwell microar-
chitecture and support Intel 64 architecture.
The Intel® Xeon® Scalable processor family, Intel® Xeon® processor E3-1500m v5 product family, and
6th generation Intel® Core™ processors are based on the Skylake microarchitecture and support Intel 64
architecture.
The 7th generation Intel® Core™ processors are based on the Kaby Lake microarchitecture and support
Intel 64 architecture.
The Intel Atom® processor C series, the Intel Atom® processor X series, the Intel® Pentium® processor
J series, the Intel® Celeron® processor J series, and the Intel® Celeron® processor N series are based on
the Goldmont microarchitecture.
The Intel® Xeon Phi™ Processor 3200, 5200, 7200 Series is based on the Knights Landing microarchitec-
ture and supports Intel 64 architecture.
The Intel® Pentium® Silver processor series, the Intel® Celeron® processor J series, and the Intel®
Celeron® processor N series are based on the Goldmont Plus microarchitecture.
The 8th generation Intel® Core™ processors, 9th generation Intel® Core™ processors, and Intel® Xeon®
E processors are based on the Coffee Lake microarchitecture and support Intel 64 architecture.
The Intel® Xeon Phi™ Processor 7215, 7285, 7295 Series is based on the Knights Mill microarchitecture
and supports Intel 64 architecture.
The 2nd generation Intel® Xeon® Scalable processor family is based on the Cascade Lake product and
supports Intel 64 architecture.
Some 10th generation Intel® Core™ processors are based on the Ice Lake microarchitecture, and some
are based on the Comet Lake microarchitecture; both support Intel 64 architecture.
Some 11th generation Intel® Core™ processors are based on the Tiger Lake microarchitecture, and
some are based on the Rocket Lake microarchitecture; both support Intel 64 architecture.
Some processors in the 3rd generation Intel® Xeon® Scalable processor family are based on the Cooper
Lake product, and some are based on the Ice Lake microarchitecture; both support Intel 64 architecture.
The 12th generation Intel® Core™ processors are based on the Alder Lake performance hybrid architec-
ture and support Intel 64 architecture.
The 13th generation Intel® Core™ processors are based on the Raptor Lake performance hybrid
architecture and support Intel 64 architecture.
1-2
INTRODUCTION
The 4th generation Intel® Xeon® Scalable processor family is based on the Sapphire Rapids microarchi-
tecture and supports Intel 64 architecture.
The chapters in this manual are summarized as follows:
Chapter 1: Introduction — Defines the purpose and outlines the contents of this manual.
Chapter 2: Intel® 64 and IA-32 Processor Architectures — Describes the microarchitecture of
recent Intel 64 and IA-32 processor families, and other features relevant to software optimization.
Chapter 3: General Optimization Guidelines — Describes general code development and optimi-
zation techniques that apply to all applications designed to take advantage of the common features
of current Intel processors.
Chapter 4: Intel Atom® Processor Architecture — Describes the microarchitecture of recent
Intel Atom processor families, and other features relevant to software optimization.
Chapter 5: Coding for SIMD Architectures — Describes techniques and concepts for using the
SIMD integer and SIMD floating-point instructions provided by the MMX technology, Streaming
SIMD Extensions, Streaming SIMD Extensions 2, Streaming SIMD Extensions 3, SSSE3, and SSE4.1.
Chapter 6: Optimizing for SIMD Integer Applications — Provides optimization suggestions and
common building blocks for applications that use the 128-bit SIMD integer instructions.
Chapter 7: Optimizing for SIMD Floating-point Applications — Provides optimization
suggestions and common building blocks for applications that use the single-precision and double-
precision SIMD floating-point instructions.
Chapter 8: INT8 Deep Learning Inference — Describes INT8 as a data type for Deep learning
Inference on Intel technology. The document covers both AVX-512 implementations and implemen-
tations using the new Intel® DL Boost Instructions.
Chapter 9: Optimizing Cache Usage — Describes how to use the PREFETCH instruction, cache
control management instructions to optimize cache usage, and the deterministic cache parameters.
Chapter 10: Introducing Sub-NUMA Clustering — Describes Sub-NUMA Clustering (SNC), a
mode for improving average latency from last level cache (LLC) to local memory.
Chapter 11: Multicore and Hyper-Threading Technology — Describes guidelines and
techniques for optimizing multithreaded applications to achieve optimal performance scaling. Use
these when targeting multicore processor, processors supporting Hyper-Threading Technology, or
multiprocessor (MP) systems.
Chapter 12: Intel® Optane™ DC Persistent Memory — Provides optimization suggestions for
applications that use Intel® Optane™ DC Persistent Memory.
Chapter 13: 64-Bit Mode Coding Guidelines — This chapter describes a set of additional coding
guidelines for application software written to run in 64-bit mode.
Chapter 14: SSE4.2 and SIMD Programming for Text-Processing/Lexing/Parsing
Describes SIMD techniques of using SSE4.2 along with other instruction extensions to improve
text/string processing and lexing/parsing applications.
Chapter 15: Optimizations for Intel® AVX, FMA, and Intel® AVX2— Provides optimization
suggestions and common building blocks for applications that use Intel® Advanced Vector
Extensions, FMA, and Intel® Advanced Vector Extensions 2 (Intel® AVX2).
Chapter 16: Intel Transactional Synchronization Extensions — Tuning recommendations to
use lock elision techniques with Intel Transactional Synchronization Extensions to optimize multi-
threaded software with contended locks.
Chapter 17: Power Optimization for Mobile Usages — This chapter provides backg11round on
power saving techniques in mobile processors and makes recommendations that developers can
leverage to provide longer battery life.
Chapter 18: Software Optimization for Intel® AVX-512 Instructions— Provides optimization
suggestions and common building blocks for applications that use Intel® Advanced Vector Extensions
512.
1-3
INTRODUCTION
Chapter 19: Intel® Advanced Vector Extensions 512-FP16 Instruction Set for Intel® Xeon®
Processors — Describes the addition of the FP16 ISA for Intel AVX-512 to handle IEEE 754-2019
compliant half-precision floating-point operations.
Chapter 20: Intel® Advanced Matrix Extensions (Intel® AMX) — Describes best practices to
optimally code to the metal on Intel® Xeon® Processors based on Sapphire Rapids SP microarchi-
tecture. It extends the public documentation on Optimizing DL code with DL Boost instructions.
Chapter 21: Cryptography & Finite Field Arithmetic Enhancements — Describes the new
instruction extensions designated for acceleration of cryptography flows and finite field arithmetic.
Chapter 22: Intel® QuickAssist Technology — Describes software development guidelines for
the Intel® QuickAssist Technology (Intel® QAT) API. This API supports both the Cryptographic and
Data Compression services.
Chapter 23: Knights Landing Microarchitecture and Software Optimization — Describes the
microarchitecture of processor families based on the Knights Landing microarchitecture, and
software optimization techniques targeting Intel processors based on the Knights Landing microar-
chitecture.
Appendix A: Application Performance Tools — Introduces tools for analyzing and enhancing
application performance without having to write assembly code.
Appendix B: Using Performance Monitoring Events — Provides information on the Top-Down
Analysis Method and information on how to use performance events specific to the Intel Xeon
processor 5500 series, processors based on Sandy Bridge microarchitecture, and Intel Core Solo and
Intel Core Duo processors.
Appendix C: Intel Architecture Optimization with Large Code Pages — Provides information
on how the performance of runtimes can be improved by using large code pages.
Appendix D: IA-32 Instruction Latency and Throughput — Provides latency and throughput
data for the IA-32 instructions. Instruction timing data specific to recent processor families are
provided.
Appendix E: Earlier Generations of Intel® 64 and IA-32 Processor Architectures
Describes the microarchitecture of earlier generations of Intel 64 and IA-32 processor families, and
other features relevant to software optimization.
Appendix F: Earlier Generations of Intel Atom® Microarchitecture and Software Optimi-
zation — Describes the microarchitecture of earlier generations of processor families based on Intel
Atom microarchitecture, and software optimization techniques targeting Intel Atom microarchi-
tecture.
1.3
RELATED INFORMATION
For more information on the Intel® architecture, techniques, and the processor architecture terminology,
the following are of particular interest:
Intel® 64 and IA-32 Architectures Software Developer’s Manual.
Developing Multi-threaded Applications: A Platform Consistent Approach.
Get Started with Intel® Fortran Compiler Classic and Intel® Fortran Compiler.
Intel® C++ Compiler Classic Developer Guide and Reference.
Intel® Developer Catalog.
Intel® oneAPI Data Analytics Library.
More relevant links include:
AI & Machine Learning: Development tools and resources.
Development Topics & Technologies.
Intel® 64 Architecture Processor Topology Enumeration.
Intel® Distribution of OpenVino™ Toolkit.
1-4
INTRODUCTION
Intel Processor support and information.
Intel® Hyper-Threading Technology (Intel® HT Technology).
Intel® Instruction Set Extensions Technology Support.
Intel® Many Integrated Core Architecture.
Intel® QuickAssist Technology (Intel® QAT).
Intel® SSE4 Programming Reference.
Intel® VTune™ Profiler User Guide.
1-5
INTRODUCTION
1-6
2. Updates to Chapter 2
Change bars and violet text show changes to Chapter 2 of the Intel® 64 and IA-32 Architectures Optimization
Reference Manual: Intel® 64 and IA-32 Processor Architectures.
------------------------------------------------------------------------------------------
Changes to this chapter:
• Corrected branding and style across chapter.
• Section 2.1:
— Updated title of section from Sapphire Rapids Architecture to Sapphire Rapids Microarchitecture.
— Refined technology features associated with the Sapphire Rapids microarchitecture.
— 2.1.1: Changed: Its I/O to the I/O
• Section 2.3:
— Updated to include new performance recommendations.
— Updated Figure 2-1 and 2-3 to include *H in Port 1
— Update Tables 2-1 and 2-2 to include additional footnote regarding *H performance improvements.
Intel® 64 and IA-32 Architectures Optimization Reference Manual
13
CHAPTER 2
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
This chapter gives an overview of features relevant to software optimization for current generations of
Intel® 64 and IA-32 processors1. These features are:
Microarchitectures that enable executing instructions with high throughput at high clock speeds, a
high-speed cache hierarchy, and high-speed system bus.
Intel® Hyper-Threading Technology2 (Intel® HT Technology) support.
Intel 64 architecture on Intel 64 processors.
Single Instruction Multiple Data (SIMD) instruction extensions: MMX™ technology, Streaming SIMD
Extensions (Intel® SSE), Streaming SIMD Extensions 2 (Intel® SSE2), Streaming SIMD Extensions 3
(Intel® SSE3), Supplemental Streaming SIMD Extensions 3 (SSSE3), Intel® SSE4.1, and Intel®
SSE4.2.
Intel® Advanced Vector Extensions (Intel® AVX).
Half-precision floating-point conversion and RDRAND.
Fused Multiply Add Extensions.
Intel® Advanced Vector Extensions 2 (Intel® AVX2).
ADX and RDSEED.
Intel® Advanced Vector Extensions 512 (Intel® AVX-512).
Intel® Thread Director.
2.1
SAPPHIRE RAPIDS MICROARCHITECTURE
Intel processors based on Sapphire Rapids microarchitecture use Golden Cove cores and support the
following additional features:
Intel® Advanced Matrix Extensions (Intel® AMX) (Chapter 20).
Intel® Advanced Vector Extensions 512 (Intel® AVX-512) (Chapter 19).
Intel® Data Streaming Accelerator (Intel® DSA)3.
Intel® In-Memory Analytics Accelerator (Intel® IAA)4.
Intel® Quick Assist Technology (Intel® QAT)(Chapter 22)
2.1.1
4th Generation Intel® Xeon® Scalable Family of Processors
Intel's fourth generation Xeon® Scalable Family of Processors changes from a single-die monolithic
design to multi-die Tiles.
The server products are scalable from dual-socket to eight-socket configurations (Section 3.11).
The I/O is increased with PCI Express 5.0, DDR5 memory, and Compute Express Link 1.1.
1. For previous generations of Intel 64 and IA-32 processors, see Appendix E, “Earlier Generations of Intel® 64 and IA-32
Processor Architectures.” Intel Atom® processors are covered in Chapter 4, “Intel Atom® Processor Architectures.”
2. Intel HT Technology requires a computer system with an Intel processor supporting hyper-threading and an Intel HT
Technology-enabled chipset, BIOS, and operating system. Performance varies depending on the hardware and software
used.
3. Please see the intel® DSA Specification and Intel® DSA User Guide.
4. Please see the Intel® IAA Specification.
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
Packaging includes a multi-die chip with up to 4 tiles. Each tile is a 400mm2 SoC, providing both compute
cores and I/O.
Each tile contains 15 Golden Cove cores (see Section 2.3). Its memory controller provides two channels
of DDR5 with a maximum of eight channels across 4 tiles, and 28 PCIe 5.0 lanes for a maximum of 112
across 4 tiles.
2.2
ALDER LAKE PERFORMANCE HYBRID ARCHITECTURE
The Alder Lake performance hybrid architecture combines two Intel architectures, bringing together the
Golden Cove performant cores and the Gracemont efficient Atom cores onto a single SoC. For details on
the Golden Cove microarchitecture, see Section 2.3, “Golden Cove Microarchitecture.” For details on the
Gracemont microarchitecture, see Section 4.1, “Gracemont Microarchitecture.”
2.2.1
12th Generation Intel® Core™ Processors Supporting Performance Hybrid
Architecture
12th Generation Intel® Core™ processors supporting performance hybrid architecture consist of up to
eight Performance cores (P-cores) and eight Efficient cores (E-cores). These processors also include a
3MB Last Level Cache (LLC) per IDI module, where a module is one P-core or four E-cores. It has
symmetrical ISA and comes in variety of configurations.
P-cores provide single or limited thread performance, while E-cores help provide improved scaling and
multithreaded efficiency. P-cores on these processors can also have Intel Hyper-Threading Technology
enabled. All cores can be active simultaneously when the operating system (OS) decides to schedule on
all processors.
A key OSV requirement for enabling hybrid is symmetric ISA across different core types in a performance
hybrid architecture. In 12th Generation Intel Core processors supporting performance hybrid architec-
ture, ISA is converged to a common baseline between the P-cores and E-cores. In order to maintain
symmetric ISA, the E-cores do not support the following features: Intel AVX-512, Intel AVX-512 FP-16,
and Intel® TSX. The E-cores do support Intel AVX2 and Intel AVX-VNNI.
2.2.2
Hybrid Scheduling
2.2.2.1
Intel® Thread Director
Intel® Thread Director continually monitors software in real-time giving hints to the operating system's
scheduler allowing it to make more intelligent and data-driven decisions on thread scheduling. With Intel
Thread Director, hardware provides runtime feedback to the OS per thread based on various IPC perfor-
mance characteristics, in the form of:
Dynamic performance and energy efficiency capabilities of P-cores and E-cores based on
power/thermal limits.
Idling hints when power and thermal are constrained.
Intel Thread Director is first introduced in desktop and mobile variants of the 12th generation Intel Core
processor based on Alder Lake performance hybrid architecture.
A processor containing both P-cores and E-cores with different performance characteristics creates a
challenge for the operating system’s scheduler. Additionally, different software threads see different
performance ratios between the P-cores and E-cores. For example, the performance ratio between the
P-cores and E-cores for highly vectorized floating-point code is higher than the performance ratio for
scalar integer code. So, when the operating system needs to make an optimal scheduling decision it
needs to be aware of the characteristics of the software threads that are candidates for scheduling. If not
enough P-cores are available and there is a mix of software threads with different characteristics, the
2-2
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
operating system should schedule those threads that benefit most from the P-cores onto those cores and
schedule the others on the E-cores.
Intel Thread Director provides the necessary hint to the operating system about the characteristics of the
software thread executing on each of the logical processors. The hint is dynamic and reflects the recent
characteristics of the thread, i.e., it may change over time based on the dynamic instruction mix of the
thread. The processor also considers microarchitecture factors to define the dynamic software thread
characteristics.
Thread specific hardware support is enumerated via the CPUID instruction and enabled by the operating
system via writing to configuration MSRs. The Intel Thread Director implementation on processors based
on Alder Lake performance hybrid architecture defines four thread classes:
0. Non-vectorized integer or floating-point code.
1. Integer or floating-point vectorized code, excluding Intel® Deep Learning Boost (Intel® DL Boost)
code.
2. Intel DL Boost code.
3. Pause (spin-wait) dominated code.
The dynamic code does not have to be 100% of the class definition. It should be large enough to be
considered belonging to that class. Also, dynamic microarchitectural metrics such as consumed memory
bandwidth or cache bandwidth may move software threads between classes. Example pseudo-code
sequences for the Intel Thread Director classes available on processors based on Alder Lake performance
hybrid architecture are provided in the examples 2-1 through 2-4.
Intel Thread Director also provides a table in system memory, only accessible to the operating system,
that defines the P-core vs. E-core performance ratio per class. This allows the operating system to pick
and choose the right software thread for the right logical processor.
In addition to the performance ratio between P-cores and E-cores, Intel Thread Director provides the
energy efficiency ratio between those cores. The operating system can then use this information when it
prefers energy savings over maximum performance. For example, a backg11round task such as indexing
can be scheduled on the most energy efficient core since its performance is less critical.
Example 2-1. Class 0 Pseudo-code Snippet
while (1)
{
asm(“xor rax, rax;”
“add rax, 5;”
“inc rax;”
);
}
2-3
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
Example 2-2. Class 1 Pseudo-code Snippet
while (1)
{
asm(“vfmaddsub132ps %ymm0, %ymm1, %ymm2;”
“vfmaddsub213ps %ymm0, %ymm1, %ymm3;”
“vfmaddsub231ps %ymm0, %ymm1, %ymm4;”
“vfmaddsub132ps %ymm0, %ymm1, %ymm5;”
“vfmaddsub213ps %ymm0, %ymm1, %ymm6;”
“vfmaddsub231ps %ymm0, %ymm1, %ymm7;”
“vfmaddsub132ps %ymm0, %ymm1, %ymm8;”
“vfmaddsub213ps %ymm0, %ymm1, %ymm9;”
“vfmaddsub231ps %ymm0, %ymm1, %ymm10;”
“vfmaddsub132ps %ymm0, %ymm1, %ymm2;”
“vfmaddsub213ps %ymm0, %ymm1, %ymm3;”
“vfmaddsub231ps %ymm0, %ymm1, %ymm4;”
“vfmaddsub132ps %ymm0, %ymm1, %ymm5;”
“vfmaddsub213ps %ymm0, %ymm1, %ymm6;”
“vfmaddsub231ps %ymm0, %ymm1, %ymm7;”
“vfmaddsub132ps %ymm0, %ymm1, %ymm8;”
“vfmaddsub213ps %ymm0, %ymm1, %ymm9;”
“vfmaddsub231ps %ymm0, %ymm1, %ymm10;”
“vfmaddsub132ps %ymm0, %ymm1, %ymm2;”
“vfmaddsub213ps %ymm0, %ymm1, %ymm3;”
“vfmaddsub231ps %ymm0, %ymm1, %ymm4;”
“vfmaddsub132ps %ymm0, %ymm1, %ymm5;”
“vfmaddsub213ps %ymm0, %ymm1, %ymm6;”
“vfmaddsub231ps %ymm0, %ymm1, %ymm7;”
“vfmaddsub132ps %ymm0, %ymm1, %ymm8;”
“vfmaddsub213ps %ymm0, %ymm1, %ymm9;”
“vfmaddsub231ps %ymm0, %ymm1, %ymm10;”
“vfmaddsub132ps %ymm0, %ymm1, %ymm2;”
“vfmaddsub213ps %ymm0, %ymm1, %ymm3;”
“vfmaddsub231ps %ymm0, %ymm1, %ymm4;”
“vfmaddsub132ps %ymm0, %ymm1, %ymm5;”
“vfmaddsub213ps %ymm0, %ymm1, %ymm6;”
“vfmaddsub231ps %ymm0, %ymm1, %ymm7;”
“vfmaddsub132ps %ymm0, %ymm1, %ymm8;”
“vfmaddsub213ps %ymm0, %ymm1, %ymm9;”
“vfmaddsub231ps %ymm0, %ymm1, %ymm10;”
“vfmaddsub132ps %ymm0, %ymm1, %ymm2;”
);
}
2-4
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
Example 2-3. Class 2 Pseudo-code Snippet
while (1)
{
__asm(
vpdpbusd ymm2, ymm0, ymm1
vpdpbusd ymm3, ymm0, ymm1
vpdpbusd ymm4, ymm0, ymm1
vpdpbusd ymm5, ymm0, ymm1
vpdpbusd ymm6, ymm0, ymm1
vpdpbusd ymm7, ymm0, ymm1
vpdpbusd ymm8, ymm0, ymm1
vpdpbusd ymm9, ymm0, ymm1
vpdpbusd ymm10, ymm0, ymm1
vpdpbusd ymm11, ymm0, ymm1
vpdpbusd ymm12, ymm0, ymm1
vpdpbusd ymm13, ymm0, ymm1
);
}
Example 2-4. Class 3 Pseudo-code Snippet
while (1)
{
asm(“PAUSE;”)
asm(“PAUSE;”)
asm(“PAUSE;”)
asm(“PAUSE;”)
asm(“PAUSE;”)
asm(“PAUSE;”)
asm(“PAUSE;”)
asm(“PAUSE;”)
asm(“PAUSE;”)
asm(“PAUSE;”)
);
}
For more detailed information on this technology, refer to the Intel® 64 and IA-32 Architectures Software
Developer’s Manual.
2.2.2.2
Scheduling with Intel® Hyper-Threading Technology-Enabled on Processors
Supporting x86 Hybrid Architecture
E-cores are designed to provide better performance than a logical P-core with both hardware sibling
hyper-threads busy.
2-5
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
2.2.2.3
Scheduling with a Multi-E-Core Module
E-cores within an idle module help provide better performance than E-cores in a busy module.
2.2.2.4
Scheduling Background Threads on x86 Hybrid Architecture
In most scenarios, backg11round threads can leverage scalability and multithread efficiency of E-cores.
2.2.3
Recommendations for Application Developers
The following are recommendations when using processors supporting performance hybrid architecture:
Stay up to date on updates on operating systems and optimized libraries.
Software needs to avoid setting hard affinities on either threads or processes in order to allow the
operating system to provide the optimal core selection for Intel Hybrid.
Software should replace active spin-waits with lightweight waits ideally using the new
UMWAIT/TPAUSE and older PAUSE instructions which will allow for better hints to the scheduler on
time spinning.
Software can utilize the Windows Power Throttling information using process information and thread
information APIs, to give hints to the scheduler on the Quality of Service (QoS) required for a
particular thread or process to improve both performance and energy efficiency.
Leverage Windows frameworks and media APIs for multimedia application development. Windows
Media Foundation framework is optimized for hybrid architecture and enables media applications to
run efficiently while preventing glitches.
The Windows IrqPolicyMachineDefault policy enables Windows to optimally target interrupts to the
right core, and more so on hybrid architecture.
For additional recommendations and information on performance hybrid architecture, refer to the white
papers on the Performance Hybrid Architecture page.
2.3
GOLDEN COVE MICROARCHITECTURE
The Golden Cove microarchitecture is the successor of Ice Lake microarchitecture. The Golden Cove
microarchitecture introduces the following enhancements:
Wider machine: 56 wide allocation, 1012 execution ports, and 48 wide retirement.
Significant increases in the size of key structures enable deeper OOO execution and expose more
instruction level parallelism.
Greater capabilities per execution port, e.g., 5th integer ALU execution ports with expanded
capability and a new fast floating-point adder.
Intel® Advanced Matrix Extensions (Intel® AMX)1: Built-in integrated Tiled Matrix Multiplication /
Machine Learning Accelerator.
Improved branch prediction.
Improvements for large code footprint workloads, e.g., larger branch prediction structures, enhanced
code prefetcher, and larger instruction TLB.
Wider fetch: legacy decode pipeline fetch bandwidth increase to 32B/cycles, 46 decoders,
increased micro-op cache size, and increased micro-op cache bandwidth.
Maximum load bandwidth increased from 2 loads/cycle to 3 loads/cycle.
Larger 4K Pages DTLB, increase in the number of outstanding Page Miss handlers.
Increased number of outstanding misses (16 FB, 3248 Deeper MLC miss queues).
1. Intel AMX are not available on client parts.
2-6
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
Enhanced data prefetchers for increased memory parallelism.
Mid-level cache size increased to 2MB on server parts; remains 1.25MB on client parts.
2.3.1
Golden Cove Microarchitecture Overview
The basic pipeline functionality of the Golden Cove microarchitecture is depicted in Figure
2-1.
ITLB + 32KB Instruction Cache
BPU
MSROM
Decode
ʅ op Cache
ʅop Queue
Allocate / Rename / Move Elimination / Zero Idiom
Scheduler / Reservation Station
P0
P1
P5
P6
P10
P2
P3
P11
P4
P9
P7
P8
AGU
AGU
AGU
STD
STD
AGU
AGU
ALU
ALU
ALU
ALU
ALU
LEA
LEA
LEA
LEA
LEA
Load Buffer
Store Buffer
INT
Shift
MUL
MULHi
Shift
3x256
2x512
2x256
JMP1
IDIV
JMP2
1x512
*H
LD DTLB
STA DTLB
3x256
2x512
FMA
FMA
FMA512
ALU
ALU
ALU
48KB DOU
VEC
Shift
Shift
AMX
1.25MB Client / 2MB Server MLC
fpDiv
Shuffle
Shuffle
FastADD
FastADD
SOC
Figure 2-1. Processor Core Pipeline Functionality of the Golden Cove Microarchitecture
The Golden Cove front end is depicted in Figure 2-2. The front end is built to feed the wider and deeper
out-of-order core:
Legacy decode pipeline fetch bandwidth increased from 16 to 32 bytes/cycle.
The number of decoders increased from four to six, allowing decode of up to 6 instructions per cycle.
The micro-op cache size increased, and its bandwidth increased to deliver up to 8 micro-ops per
cycle.
Improved branch prediction.
2-7
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
ITLB + 32KB Instruction Cache
BPU
32
bytes
64 bytes
MSROM
Decode
µop Cache
4 uops
6 uops
8 uops
µop Queue
6 uops
Figure 2-2. Processor Front End of the Golden Cove Microarchitecture
Improvements for large code footprint workloads:
Double the size of the instruction TLB: 128256 entries for 4K pages, 1632 entries for 2M/4M
pages.
Bigger branch prediction structures.
Enhanced code prefetcher.
Improved LSD coverage.
The IDQ can hold 144 uops per logical processor in single thread mode, or 72 uops per thread when
SMT is active.
Additional improvements include:
Significant increase in size of key buffer structures to enable deeper OOO execution and expose more
instruction level parallelism.
Wider machine:
— Wider allocation (56 uops per cycle) and retirement (48 uops per cycle) width.
— Increase in number of execution ports (1012).
— Greater capabilities per execution port.
Table 2-1 summarizes the OOO engine's capability to dispatch different types of operations to ports.
2-8
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
Table 2-1. Dispatch Port and Execution Stacks of the Golden Cove Microarchitecture
Port 0
Port 11
Port 2
Port 3
Port 4
Port 52
Port 6
Ports 7, 8
Port 9
Port 10
Port 11
INT ALU
INT
Load
Load
Store
INT ALU
INT ALU
Store
Store
INT ALU
Load
ALU3
Data
Address
Data
LEA
LEA
LEA
LEA
LEA
INT Shift
INT MUL
INT Shift
INT Mul
Hi
Jump1
Jump2
INT Div
FMA
FMA*
FMA**
Vec ALU
Fast
Fast
Adder*
Adder
Vec
Shift
Vec
Vec ALU
ALU*
FP Div
Shuffle
Vec
Shift*
Shuffle*
NOTES:
1. “*” in this table indicates that these features are not available for 512-bit vectors.
2. “**” in this table indicates that these features are not available for 512-bit vectors in Client parts.
3. The Golden Cove microarchitecture implemented performance improvements requiring constraint of the micro-ops which
use *H partial registers (i.e. AH, BH, CH, DH). See Section 3.5.2.3 for more details.
Table 2-2 lists execution units and common representative instructions that rely on these units.
Throughput improvements across the Intel® SSE, Intel AVX, and general-purpose instruction sets are
related to the number of units for the respective operations, and the varieties of instructions that execute
using a particular unit.
Table 2-2. Golden Cove Microarchitecture Execution Units and Representative Instructions1
Execution
# of Unit
Instructions
Unit
add, and, cmp, or, test, xor, movzx, movsx, mov, (v)movdqu, (v)movdqa, (v)movap*,
ALU
52
(v)movup*
SHFT
23
sal, shl, rol, adc, sarx, adcx, adox, etc.
Slow Int
1
mul, imul, bsr, rcl, shld, mulx, pdep, etc.
BM
2
andn, bextr, blsi, blsmsk, bzhi, etc.
2x256-bit
(v)add, (v)cmp. (v)max, (v)min, (v)sub, (v)cvtps2dq, (v)cvtdq2ps, (v)cvtsd2sl, (v)cvtss2sl
1x512-bit
Vec ALU
3x256-bit
(v)pand, (v)por, (v)pxor, (v)movq, (v)movq, (v)movap*, (v)movup*, (v)andp*, (v)orp*,
2x512-bit
(v)paddb/w/d/q, (v)blendv*, (v)blendp*, (v)pblendd
2x256-bit
Vec_Shft
(v)psllv*, (v)psrlv*, vector shift count in imm8
1x512-bit
VEC Add (in
2x256-bit
(v)add*, (v)cmp*, (v)max*, (v)min*, (v)sub*, (v)padds*, (v)paddus*, (v)psign, (v)pabs,
VEC FMA)
1x512-bit
(v)pavgb, (v)pcmpeq*, (v)pmax, (v)cvtps2dq, (v)cvtdq2ps, (v)cvtsd2si, (v)cvtss2si
2-9
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
Table 2-2. Golden Cove Microarchitecture Execution Units and Representative Instructions1 (Contd.)
Execution
# of Unit
Instructions
Unit
VEC Fast
2x256-bit
(v)add*, (v)addsub*, (v)sub*
Add
1x512-bit
Shuffle
2x256-bit
(v)shufp*, vperm*, (v)pack*, (v)unpck*, (v)punpck*, (v)pshuf*, (v)pslldq, (v)alignr,
(v)pmovzx*, vbroadcast*, (v)pslldq, (v)psrldq, (v)pblendw (new cross lane shuffle on
1x512-bit
both ports)
Vec
2x256-bit
(v)mul*, (v)pmul*, (v)pmadd*
Mul/FMA
(1 or
2)x512-bit
SIMD Misc
1
STTNI, (v)pclmulqdq, (v)psadw, vector shift count in xmm
FP Mov
1
(v)movsd/ss, (v)movd gpr
DIVIDE
1
divp*, divs*, vdiv*, sqrt*, vsqrt*, rcp*, vrcp*, rsqrt*, idiv
NOTES:
1. Execution unit mapping to MMX instructions are not covered in this table. See Section 15.16.5 on MMX instruction
throughput remedy.
2. The Golden Cove microarchitecture implemented performance improvements requiring constraint of the micro-ops which
use *H partial registers (i.e. AH, BH, CH, DH). See Section 3.5.2.3 for more details.
3. ibid.
Table 2-3 describes bypass delay in cycles between producer and consumer operations.
Table 2-3. Bypass Delay Between Producer and Consumer Micro-Ops
TO [EU/PORT/Latency]
SHUF/
FROM
Fast
SIMD/0,1/1
FMA/0,1/4
MUL/0,1/4
SIMD/5/1,3
1,5/1,
V2I/0/3
[EU/Port/Latency]
Adder/1,5/3
3
SIMD/0,1/1
0
1
1
1
0
0
0
FMA/0,1/4
1
0
1
0
0
0
0
MUL/0,1/4
1
0
1
0
0
0
0
Fast Adder/0,1/3
1
0
1
-1
0
0
0
SIMD/5/1,3
0
1
1
1
0
0
0
SHUF/1,5/1,3
0
0
1
0
0
0
0
V2I/0/3
0
0
1
0
0
0
0
I2V/5/1
0
1
1
0
0
0
0
The attributes that are relevant to the producer/consumer micro-ops for bypass are a triplet of
abbreviation/one or more port number/latency cycle of the uop. For example:
“SIMD/0,1/1” applies to a 1-cycle vector SIMD uop dispatched to either port 0 or port 1.
“SIMD/5/1,3” applies to either a 1-cycle or 3-cycle non-shuffle uop dispatched to port 5.
“V2I/0/3” applies to a 3-cycle vector-to-integer uop dispatched to port 0.
2-10
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
“I2V/5/1” applies to a 1-cycle integer-to-vector uop dispatched to port 5.
“Fast Adder/1,5/3” applies to either a 3-cycle 256-bit uop dispatched to either port 1 or port 5, or a
512-bit uop dispatched to port 5. This operation supports two cycles back-to-back between a pair of
Fast Adder operations.
A new Fast Adder1 unit is added as 512-bit on port 5 in VEC stack, and as 256-bit on ports 1 and 5. The
Fast Adder performs floating-point ADD/SUB operations in 3 cycles.
Back-to-back ADD/SUB operations that are both executed on the Fast Adder unit perform the operations
in two cycles.
In 128/256-bit, back-to-back ADD/SUB operations executed on the Fast Adder unit perform the
operations in two cycles.
In 512-bit, back-to-back ADD/SUB operations are executed in two cycles if both operations use the
Fast Adder unit on port 5.
The following instructions are executed by the Fast Adder unit:
(V)ADDSUBSS/SD/PS/PD
(V)ADDSS/SD/PS/PD
(V)SUBSS/SD/PS/PD
2.3.1.1
Cache Subsystem and Memory Subsystem
The cache subsystem and memory subsystem changes in the Golden Cove microarchitecture are:
Maximum load bandwidth increased from 2 to 3 loads per cycle. Bandwidth of Intel AVX-512 loads,
Intel AMX loads, and MMX/x87 loads remain at a maximum of 2 loads per cycle.
Simultaneous handling of more loads and stores enabled by enlarged buffers.
Number of entries for 4K pages in the load DTLB increased from 64 to 96.
Page Miss handler can handle up to four D-side page walks in parallel instead of two.
Increased number of outstanding DCU and MLC misses.
Enhanced data prefetchers for increased memory parallelism.
Partial store forwarding allowing forwarding data from store to load also when only part of the load
was covered by the store (in case the load's offset matches the store's offset).
2.3.1.2
Avoiding Destination False Dependency
Some SIMD instructions incur false dependency on the destination operand. The following instructions
are affected:
VFMULCSH, VFMULCPH
VFCMULCSH, VFCMULCPH
VPERMD, VPERMQ, VPERMPS, VPERMPD
VRANGE[SS,PS,SD,PD]
VGETMANTSH, VGETMANTSS, VGETMANTSD
VGETMANTPS, VGETMANTPD (memory versions only)
VPMULLQ
1. The Fast Adder unit is not available on 512-bit vectors in Client parts.
2-11
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
Recommendation: Use dependency breaking zero idioms on the destination register before the
affected instructions to avoid potential slowdown from the false dependency.
Example 2-5. Breaking False Dependency through Zero Idiom
Code with False Dependency Impact
Mitigation: Break False Dependency with Zero Idiom
vaddps zmm3, zmm4, zmm5
vaddps zmm3, zmm4, zmm5
vmovaps [rsi], zmm3
vmovaps [rsi], zmm3
vfmulcph zmm3, zmm2, zmm1
;False dependency on
vpxord zmm3, zmm3, zmm3
;Dependency-breaking
zmm3.
zero idiom.
Will not execute out-of-order
vfmulcph zmm3, zmm2, zmm1
;Execute out-of-order
until vaddps writes zmm3.
without waiting for
vaddps result.
2.4
ICE LAKE CLIENT MICROARCHITECTURE
The Ice Lake client microarchitecture introduces the following new features that allow optimizations of
applications for performance and power consumption:
Targeted vector acceleration.
Crypto acceleration.
Intel® Software Guard Extensions (Intel® SGX) enhancements.
Cache line writeback instruction (CLWB).
2.4.1
Ice Lake Client Microarchitecture Overview
The Ice Lake client microarchitecture builds on the successes of the Skylake client microarchitecture.
The basic pipeline functionality of the Ice Lake Client microarchitecture is depicted in Figure 2-3.
32KB
BPU
Instruction Cache
Legacy Decode
ʅ op Cache
MSROM
Pipeline
ʅop Queue
Allocate / Rename / Move Elimination / Zero Idiom
Scheduler / Reservation Station
P4 + P9
P2
P8
P3
P7
Store Data
Load
STA
Load
STA
Port 0
Port 1
Port 5
Port 6
ALU
ALU
ALU
ULA
LEA
LEA
LEA
LEA
48KB L1 Data Cache
INT
Shift
MUL
MULHi
Shift
JMP1
IDIV
*H
JMP2
*H
512KB L2 Data Cache
FMA
FMA*
SLU
ALU*
ALU
VEC
Shift
Shift*
SOC
fpDIV
Shuffle*
Shuffle
Figure 2-3. Processor Core Pipeline Functionality of the Ice Lake Client Microarchitecture1
2-12
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
NOTES:
1.
“*” in the figure above indicates these features are not available for 512-bit vectors.
2. “INT” represents GPR scalar instructions.
3. “VEC” represents floating-point and integer vector instructions.
4. “MULHi” produces the upper 64 bits of the result of an iMul operation that multiplies two 64-bit registers and places the
result into two 64-bits registers.
5. The “Shuffle” on port 1 is new, and supports only in-lane shuffles that operate within the same 128-bit sub-vector.
6. The “IDIV” unit on port 1 is new, and performs integer divide operations at a reduced latency.
7. The Golden Cove microarchitecture implemented performance improvements requiring constraint of the micro-ops which
use *H partial registers (i.e. AH, BH, CH, DH). See Section 3.5.2.3 for more details.
The Ice Lake client microarchitecture introduces the following new features:
Significant increase in size of key structures enable deeper OOO execution.
Wider machine: 4 5 wide allocation, 8 10 execution ports.
Intel AVX-512 (new for client processors): 512-bit vector operations, 512-bit loads and stores to
memory, and 32 new 512-bit registers.
Greater capabilities per execution port (e.g., SIMD shuffle, LEA), reduced latency Integer Divider.
2×BW for AES-NI peak throughput for existing binaries (microarchitectural).
Rep move string acceleration.
50% increase in size of the L1 data cache.
Reduced effective load latency.
2×L1 store bandwidth: 1 2 stores per cycle.
Enhanced data prefetchers for increased memory parallelism.
Larger 2nd level TLB.
Larger uop cache.
Improved branch predictor.
Large page ITLB size in single thread mode doubled.
Larger L2 cache.
The Ice Lake client microarchitecture supports flexible integration of multiple processor cores with a
shared uncore sub-system consisting of a number of components including a ring interconnect to
multiple slices of L3, processor graphics, integrated memory controller, interconnect fabrics, and more.
2.4.1.1
The Front End
The front end changes in Ice Lake Client microarchitecture include:
Improved branch predictor.
Large page ITLB in single thread mode increased from 8 to 16 entries.
Larger uop cache.
The IDQ can hold 70 uops per logical processor vs. 64 uops per logical processor in previous
generations when two sibling logical processors in the same core are active (2×70 vs. 2×64 per
core). If only one logical processor is active in the core, the IDQ can hold 70 uops vs. 64 uops.
The LSD in the IDQ can detect loops of up to 70 uops per logical processor irrespective single thread
or multi thread operation.
2-13
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
2.4.1.2
The Out of Order and Execution Engines
The Out of Order and execution engines changes in Ice Lake Client microarchitecture include:
A significant increase in size of reorder buffer, load buffer, store buffer, and reservation stations
enable deeper OOO execution and higher cache bandwidth.
Wider machine: 4 5 wide allocation, 8 10 execution ports.
Greater capabilities per execution port (e.g., SIMD shuffle, LEA).
Reduced latency Integer Divider.
A new iDIV unit was added that significantly reduces the latency and improves the of throughput of
integer divide operations.
Table 2-4 summarizes the OOO engine's capability to dispatch different types of operations to ports.
Table 2-4. Dispatch Port and Execution Stacks of the Ice Lake Client Microarchitecture
Port 0
Port 11
Port 2
Port 3
Port 4
Port 5
Port 6
Port 7
Port 8
Port 9
INT ALU
INT ALU
Load
Load
Store
INT ALU
INT ALU
Store
Store
Store
Data
Address
Address
Data
LEA
LEA
LEA
LEA
INT Shift
INT Mul
INT MUL
INT Shift
Hi
Jump1
INT Div
Jump2
FMA
FMA*
Vec ALU
Vec ALU
Vec ALU*
Vec
Shuffle
Vec Shift
Vec
Shift*
FP Div
Vec
Shuffle*
NOTES:
1. “*” in this table indicates these features are not available for 512-bit vectors.
Table 2-5 lists execution units and common representative instructions that rely on these units.
Throughput improvements across the SSE, Intel AVX, and general-purpose instruction sets are related to
the number of units for the respective operations, and the varieties of instructions that execute using a
particular unit.
Table 2-5. Ice Lake Client Microarchitecture Execution Units and Representative Instructions1
Execution
# of
Instructions
Unit
Unit
ALU
4
add, and, cmp, or, test, xor, movzx, movsx, mov, (v)movdqu, (v)movdqa, (v)movap*, (v)movup*
SHFT
2
sal, shl, rol, adc, sarx, adcx, adox, etc.
Slow Int
1
mul, imul, bsr, rcl, shld, mulx, pdep, etc.
BM
2
andn, bextr, blsi, blsmsk, bzhi, etc.
Vec ALU
3
(v)pand, (v)por, (v)pxor, (v)movq, (v)movq, (v)movap*, (v)movup*, (v)andp*, (v)orp*,
(v)paddb/w/d/q, (v)blendv*, (v)blendp*, (v)pblendd
Vec_Shft
2
(v)psllv*, (v)psrlv*, vector shift count in imm8
Vec Add
2
(v)addp*, (v)cmpp*, (v)max*, (v)min*, (v)padds*, (v)paddus*, (v)psign, (v)pabs, (v)pavgb,
(v)pcmpeq*, (v)pmax, (v)cvtps2dq, (v)cvtdq2ps, (v)cvtsd2si, (v)cvtss2si
2-14
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
Table 2-5. Ice Lake Client Microarchitecture Execution Units and Representative Instructions1
Execution
# of
Instructions
Unit
Unit
Shuffle
2
(v)shufp*, vperm*, (v)pack*, (v)unpck*, (v)punpck*, (v)pshuf*, (v)pslldq, (v)alignr, (v)pmovzx*,
vbroadcast*, (v)pslldq, (v)psrldq, (v)pblendw
Vec Mul
2
(v)mul*, (v)pmul*, (v)pmadd*
SIMD Misc
1
STTNI, (v)pclmulqdq, (v)psadw, vector shift count in xmm
FP Mov
1
(v)movsd/ss, (v)movd gpr
DIVIDE
1
divp*, divs*, vdiv*, sqrt*, vsqrt*, rcp*, vrcp*, rsqrt*, idiv
NOTES:
1. Execution unit mapping to MMX instructions are not covered in this table. See Section 15.16.5 on MMX instruction
throughput remedy.
Table 2-6 describes bypass delay in cycles between producer and consumer operations.
Table 2-6. Bypass Delay Between Producer and Consumer Micro-ops
TO [EU/PORT/Latency]
FROM
SIMD/0,1/1
FMA/0,1/4
VIMUL/0,1/4
SIMD/5/1,3
SHUF/5/1,
V2I/0/3
I2V/5/1
[EU/Port/Latency]
3
SIMD/0,1/1
0
1
1
0
0
0
NA
FMA/0,1/4
1
0
1
0
0
0
NA
VIMUL/0,1/4
1
0
1
0
0
0
NA
SIMD/5/1,3
0
1
1
0
0
0
NA
SHUF/5/1,3
0
0
1
0
0
0
NA
V2I/0/3
0
0
1
0
0
0
NA
I2V/5/1
0
1
1
0
0
0
NA
The attributes that are relevant to the producer/consumer micro-ops for bypass are a triplet of abbrevi-
ation/one or more port number/latency cycle of the uop. For example:
“SIMD/0,1/1” applies to 1-cycle vector SIMD uop dispatched to either port 0 or port 1.
“SIMD/5/1,3” applies to either a 1-cycle or 3-cycle non-shuffle uop dispatched to port 5.
“V2I/0/3” applies to a 3-cycle vector-to-integer uop dispatched to port 0.
“I2V/5/1” applies to a 1-cycle integer-to-vector uop to port 5.
2-15
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
2.4.1.3
Cache and Memory Subsystem
The cache hierarchy changes in Ice Lake Client microarchitecture include:
50% increase in size of the L1 data cache.
2×L1 store bandwidth: 3 4 AGUs, 1 2 store data.
Simultaneous handling of more loads and stores enabled by enlarged buffers.
Higher cache bandwidth compared to previous generations.
Larger 2nd level TLB: 1.5K entries 2K entries.
Enhanced data prefetchers for increased memory parallelism.
L2 cache size increased from 256KB to 512KB.
L2 cache associativity increased from 4 ways to 8 ways.
Significant reduction in effective load latency.
Table 2-7. Cache Parameters of the Ice Lake Client Microarchitecture
Capacity /
Line Size
Latency1
Peak Bandwidth
Sustained Bandwidth
Update
Level
Associativity
(bytes)
(cycles)
(bytes/cycles)
(bytes/cycles)
Policy
First Level
48KB/8
64
5
2×64B loads + 1x64B
Same as peak
Writeback
(DCU)
or 2x32B stores
Second
512KB/8
64
13
64
48
Writeback
Level (MLC)
Third Level
Up to 2MB per
64
xx2
32
21
Writeback
(LLC)
core/up to 16 ways
NOTES:
1. Software-visible latency/bandwidth will vary depending on access patterns and other factors.
2. This number depends on core count.
The TLB hierarchy consists of dedicated level one TLB for instruction cache, TLB for L1D, shared L2 TLB
for 4K and 4MB pages and a dedicated L2 TLB for 1GB pages.
Table 2-8. TLB Parameters of the Ice Lake Client Microarchitecture
Per-thread Entries
Level
Page Size
Entries ST
MT Latency
Associativity
Instruction
4KB
128
64
8
Instruction
2MB/4MB
16
8
8
First Level Data (loads)
4KB
64
64 competitively
4
shared
First Level Data (loads)
2MB/4MB
32
32 competitively
4
shared
First Level Data (loads)
1GB
8
8 competitively shared
8
First Level Data (stores)
Shared for all page
16
16 competitively
16
sizes
shared
Second Level
Shared for all page
20481
2048 competitively
16
sizes
shared
NOTES:
1. 4K pages can use all 2048 entries. 2/4MB pages can use 1024 entries (in 8 ways), sharing them with 4K pages. 1GB
pages can use the other 1024 entries (in 8 ways), also sharing them with 4K pages.
2-16
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
Paired Stores
Ice Lake Client microarchitecture includes two store pipelines in the core, with the following features:
Two dedicated AGU for LDs on ports 2 and 3.
Two dedicated AGU for STAs on ports 7 and 8.
Two fully featured STA pipelines.
Two 256-bit wide STD pipelines (AVX-512 store data takes two cycles to write).
Second senior store pipeline to the DCU via store merging.
Ice Lake Client microarchitecture can write two senior stores to the cache in a single cycle if these two
stores can be paired together. That is:
The stores must be to the same cache line.
Both stores are of the same memory type, WB or USWC.
None of the stores cross cache line or page boundary.
In order to maximize performance from the second store port try to:
Align store operations whenever possible.
Place consecutive stores in the same cache line (not necessarily as adjacent instructions).
As seen in Example 2-6, it is important to take into consideration all stores, explicit or not.
Example 2-6. Considering Stores
Stores are Paired Across Loop Iterations
Stores Not Paired Due to Stack Update in Between
Loop:
Loop:
compute reg
call function to compute reg
store [X], reg
store [X], reg
add X, 4
add X, 4
jmp Loop
; stores from different iterations of the
jmp Loop
; stores from different iterations of the
loop can be paired all together because
loop cannot be paired anymore because
they usually would be same line
of the call store to stack
; the call is disturbing pairing
2-17
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
In some cases it is possible to rearrange the code to achieve store pairing. Example 2-7 provides details.
Example 2-7. Rearranging Code to Achieve Store Pairing
Stores to Different Cache Lines - Not Paired
Unrolling May Solve the Problem
Loop:
Loop:
... compute ymm1 …
... compute ymm1 …
vmovaps [x], ymm1
vmovaps [x], ymm1
... compute ymm2 …
... compute new ymm1 …
vmovaps [y], ymm2
vmovaps [x+32], ymm1
add x, 32
... compute ymm2 …
add y, 32
vmovaps [y], ymm2
jmp Loop
; this loop cannot pair any store because
... compute new ymm2 …
of alternating store to different cache
vmovaps [y+32], ymm2
lines [x] and [y]
add x, 64
add y, 64
jmp Loop
; the loop was unrolled 2 times and stores
re-arranged to make sure two stores to
the same cache line are placed one after
another. Now stores to addresses [x] and
[x+32] are to the same cache line and
could be paired together and executed in
same cycle
2.4.1.4
New Instructions
New instructions and architectural changes in Ice Lake Client microarchitecture are listed below. Actual
support may be product dependent.
Crypto acceleration
— SHA NI for acceleration of SHA1 and SHA256 hash algorithms.
— Big-Number Arithmetic (IFMA): VPMADD52 - two new instructions for big number multiplication
for acceleration of RSA vectorized SW and other Crypto algorithms (Public key) performance.
— Galois Field New Instructions (GFNI) for acceleration of various encryption algorithms, error
correction algorithms, and bit matrix multiplications.
— Vector AES and Vector Carry-less Multiply (PCLMULQDQ) instructions to accelerate AES and
AES-GCM.
Security Technologies
— Intel® SGX enhancements to improve usability and applicability: EDMM, multi-package server
support, support for VMM memory oversubscription, performance, larger secure memory.
Sub Page protection for better performance of security VMMs.
Targeted Acceleration
— Vector Bit Manipulation Instructions: VBMI1 (permutes, shifts) and VBMI2 (Expand, Compress,
Shifts)- used for columnar database access, dictionary based decompression, discrete mathe-
matics, and data-mining routines (bit permutation and bit-matrix-multiplication).
— VNNI with support for integer 8 and 16 bits data types- CNN/ML/DL acceleration.
— Bit Algebra (POPCNT, Bit Shuffle).
— Cache line writeback instruction (CLWB) enables fast cache-line update to memory, while
retaining clean copy in cache.
Platform analysis features for more efficient performance software tuning and debug.
— AnyThread removal.
2-18
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
— 2x general counters (up to 8 per-thread).
— Fixed Counter 3 for issue slots.
— New performance metrics for built-in support for Level 1 Top-Down method (% of Issue slots that
are front-end bound, back-end bound, bad speculation, retiring) while leaving the 8 general
purpose counters free for software use.
2.4.1.5
Ice Lake Client Microarchitecture Power Management
Processors based on Ice Lake Client microarchitecture are the first client processors whose cores may
execute at a different frequency from one another. The frequency is selected based on the specific
instruction mix; the type, width and number of vector instructions of the program that executes on each
core, the ratio between active time and idle time of each core, and other considerations such as how
many cores share similar characteristics.
Most of the power management features of Skylake Server Microarchitecture (see Section 2.5) is appli-
cable to Ice Lake Client microarchitecture as well. The main differences are the following:
The typical P0n max frequency difference between Intel® Advanced Vector Extensions (Intel®
AVX-512) and Intel® Advanced Vector Extensions 2 (Intel® AVX2) on Ice Lake Client microarchi-
tecture is much lower than on Skylake Server microarchitecture. Therefore, the negative impact on
overall application performance is much smaller.
All processors based on Ice Lake Client microarchitecture contain a single 512-bit FMA unit, whereas
some of the processors based on Skylake Server microarchitecture contain two such units. Both
processors contain two 256-bit FMA units. The power consumed by Ice Lake Client FMA units is the
same, whereas on Skylake Server the 512-bit units consume twice as much.
Compute heavy workloads, especially those that span multiple Ice Lake client cores, execute at a lower
frequency than P0n, both under Intel AVX-512 and under Intel AVX2 instruction sets, due to power
limitations. In this scenario, Intel AVX-512 architecture, which requires less dynamic instructions to
complete the same task than Intel AVX2 architecture, consumes less power and thus may achieve higher
frequency. The net result may be higher performance due to the shorter path length and a bit higher
frequency.
There are still some cases where coding to the Intel AVX-512 instruction set yields lower performance
than when coding to the Intel AVX2 instruction set. Sometimes it is due to microarchitecture artifacts of
longer vectors, in other cases the natural vectors are just not long enough. Most compilers are still
maturing their Intel AVX-512 support, and it may take them a few more years to generate optimal code.
The general recommendation in the Skylake Server Power Management section (see Section 2.5.3) still
holds. Developers should code to the Intel AVX-512 instruction set and compare the performance to their
Intel AVX2 workload on Ice Lake Client microarchitecture, before making the decision to proceed with a
complete port.
2.5
SKYLAKE SERVER MICROARCHITECTURE
The Intel® Xeon® Processor Scalable Family is based on the Skylake Server microarchitecture. Proces-
sors based on the Skylake microarchitecture can be identified using CPUID’s DisplayFamily_DisplayModel
signature, which can be found in Table 2-1 of CHAPTER 2 of Intel® 64 and IA-32 Architectures Software
Developer’s Manual, Volume 4.
The Skylake Server microarchitecture introduces the following new features1 that allow you to optimize
your application for performance and power consumption.
A new core based on the Skylake Server microarchitecture with process improvements based on the
Kaby Lake microarchitecture.
Intel AVX-512 support.
More cores per socket (max 28 vs. max 22).
1. Some features may not be available on all products.
2-19
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
6 memory channels per socket in Skylake microarchitecture vs. 4 in the Broadwell microarchitecture.
Bigger L2 cache, smaller non inclusive L3 cache.
Intel® Optane™ support.
Intel® Omni-Path Architecture (Intel® OPA).
Sub-NUMA Clustering (SNC) support.
The green stars in Figure 2-4 represent new features in Skylake Server microarchitecture compared to
Skylake microarchitecture for client; a 1MB L2 cache and an additional Intel AVX-512 FMA unit on port 5
which is available on some parts.
Since port 0 and port 1 are 256-bits wide, Intel AVX-512 operations that will be dispatched to port 0 will
execute on both port 0 and port 1; however, other operations such as lea can still execute on port 1 in
parallel. See the red block in Figure 2-8 for the fusion of ports 0 and 1.
Notice that, unlike Skylake microarchitecture for client, the Skylake Server microarchitecture has its
front end loop stream detector (LSD) disabled.
Figure 2-4. Processor Core Pipeline Functionality of the Skylake Server Microarchitecture
2-20
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
2.5.1
Skylake Server Microarchitecture Cache
The Intel Xeon Processor Scalable Family based on Skylake Server microarchitecture has significant
changes in core and uncore architecture to improve performance and scalability of several components
compared with the previous generation of the Intel Xeon processor family based on Broadwell microar-
chitecture.
2.5.1.1
Larger Mid-Level Cache
Skylake Server microarchitecture implements a mid-level (L2) cache of 1 MB capacity with a minimum
load-to-use latency of 14 cycles. The mid-level cache capacity is four times larger than the capacity in
previous Intel Xeon processor family implementations. The line size of the mid-level cache is 64B and it
is 16-way associative. The mid-level cache is private to each core.
Software that has been optimized to place data in mid-level cache may have to be revised to take advan-
tage of the larger mid-level cache available in Skylake Server microarchitecture.
2.5.1.2
Non-Inclusive Last Level Cache
The last level cache (LLC) in Skylake is a non-inclusive, distributed, shared cache. The size of each of the
banks of last level cache has shrunk to 1.375 MB per bank. Because of the non-inclusive nature of the last
level cache, blocks that are present in the mid-level cache of one of the cores may not have a copy resi-
dent in a bank of last level cache. Based on the access pattern, size of the code and data accessed, and
sharing behavior between cores for a cache block, the last level cache may appear as a victim cache of
the mid-level cache and the aggregate cache capacity per core may appear to be a combination of the
private mid-level cache per core and a portion of the last level cache.
2.5.1.3
Skylake Server Microarchitecture Cache Recommendations
A high-level comparison between Skylake Server microarchitecture cache and the previous generation
Broadwell microarchitecture cache is available in the table below.
Table 2-9. Cache Comparison Between Skylake Microarchitecture and Broadwell Microarchitecture
Cache level
Category
Broadwell
Skylake Server
Microarchitecture
Microarchitecture
L1 Data Cache
Size [KB]
32
32
Unit (DCU)
Latency [cycles]
4-6
4-6
Max bandwidth [bytes/cycles]
96
192
Sustained bandwidth [bytes/cycles]
93
133
Associativity [ways]
8
8
L2 Mid-level Cache
Size [KB]
256
1024 (1MB)
(MLC)
Latency [cycles]
12
14
Max bandwidth [bytes/cycles]
32
64
Sustained bandwidth [bytes/cycles]
25
52
Associativity [ways]
8
16
2-21
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
Table 2-9. Cache Comparison Between Skylake Microarchitecture and Broadwell Microarchitecture
L3 Last-level
Size [MB]
Up to 2.5 per core
up to 1.3751 per core
Cache (LLC)
Latency [cycles]
50-60
50-70
Max bandwidth [bytes/cycles]
16
32
Sustained bandwidth [bytes/cycles]
14
15
NOTES:
1. Some Skylake Server parts have some cores disabled and hence have more than 1.375 MB per core of L3 cache.
The figure below shows how Skylake Server microarchitecture shifts the memory balance from
shared-distributed with high latency, to private-local with low latency.
Figure 2-5. Broadwell Microarchitecture and Skylake Server Microarchitecture Cache Structures
The potential performance benefit from the cache changes is high, but software will need to adapt its
memory tiling strategy to be optimal for the new cache sizes.
Recommendation: Rebalance application shared and private data sizes to match the smaller,
non-inclusive L3 cache, and larger L2 cache.
Choice of cache blocking should be based on application bandwidth requirements and changes from one
application to another. Having four times the L2 cache size and twice the L2 cache bandwidth compared
to the previous generation Broadwell microarchitecture enables some applications to block to L2 instead
of L1 and thereby improves performance.
Recommendation: Consider blocking to L2 on Skylake Server microarchitecture if L2 can sustain the
application’s bandwidth requirements.
The change from inclusive last level cache to non-inclusive means that the capacity of mid-level and last
level cache can now be added together. Programs that determine cache capacity per core at run time
should now use a combination of mid-level cache size and last level cache size per core to estimate the
effective cache size per core. Using just the last level cache size per core may result in non-optimal use
of available on-chip cache; see Section 2.5.2 for details.
Recommendation: In case of no data sharing, applications should consider cache capacity per core as
L2 and L3 cache sizes and not only L3 cache size.
2-22
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
2.5.2
Non-Temporal Stores on Skylake Server Microarchitecture
Because of the change in the size of each bank of last level cache on Skylake Server microarchitecture, if
an application, library, or driver only considers the last level cache to determine the size of on-chip
cache-per-core, it may see a reduction with Skylake Server microarchitecture and may use non-temporal
store with smaller blocks of memory writes. Since non-temporal stores evict cache lines back to memory,
this may result in an increase in the number of subsequent cache misses and memory bandwidth
demands on Skylake Server microarchitecture, compared to the previous Intel Xeon processor family.
Also, because of a change in the handling of accesses resulting from non-temporal stores by Skylake
Server microarchitecture, the resources within each core remain busy for a longer duration compared to
similar accesses on the previous Intel Xeon processor family. As a result, if a series of such instructions
are executed, there is a potential that the processor may run out of resources and stall, thus limiting the
memory write bandwidth from each core.
The increase in cache misses due to overuse of non-temporal stores and the limit on the memory write
bandwidth per core for non-temporal stores may result in reduced performance for some applications.
To avoid the performance condition described above with Skylake Server microarchitecture, include
mid-level cache capacity per core in addition to the last level cache per core for applications, libraries, or
drivers that determine the on-chip cache available with each core. Doing so optimizes the available
on-chip cache capacity on Skylake Server microarchitecture as intended, with its non-inclusive last level
cache implementation.
2.5.3
Skylake Server Power Management
This section describes the interaction of Skylake Server's Power Management and its Vector ISA.
Skylake Server microarchitecture dynamically selects the frequency at which each of its cores executes.
The selected frequency depends on the instruction mix; the type, width, and number of vector instruc-
tions that execute over a given period of time. The processor also takes into account the number of cores
that share similar characteristics.
Intel® Xeon® processors based on Broadwell microarchitecture work similarly, but to a lesser extent
since they only support 256-bit vector instructions. Skylake Server microarchitecture supports Intel®
AVX-512 instructions, which can potentially draw more current and more power than Intel® AVX2
instructions.
The processor dynamically adjusts its maximum frequency to higher or lower levels as necessary, there-
fore a program might be limited to different maximum frequencies during its execution.
Table 2-10 includes information about the maximum Intel® Turbo Boost technology core frequency for
each type of instruction executed. The maximum frequency (P0n) is an array of frequencies which
depend on the number of cores within the category. The more cores belonging to a category at any given
time, the lower the maximum frequency.
Table 2-10. Maximum Intel® Turbo Boost Technology Core Frequency Levels
Level
Category
Frequency Level
Max Frequency (P0n)
Instruction Types
0
Intel® AVX2 light
Highest
Max
Scalar, AVX128, SSE, Intel® AVX2 w/o FP
instructions
or INT MUL/FMA
1
Intel® AVX2 heavy
Medium
Max Intel® AVX2
Intel® AVX2 FP + INT MUL/FMA, Intel®
instructions +
AVX-512 without FP or INT MUL/FMA
Intel® AVX-512
light instructions
2
Intel® AVX-512
Lowest
Max Intel® AVX-512
Intel® AVX-512 FP + INT MUL/FMA
heavy instructions
For per SKU max frequency details (reference figure 1-15), refer to the Intel® Xeon® Processor Scalable
Family Technical Resources page.
2-23
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
Figure 2-6 is an example for core frequency range in a given system where each core frequency is deter-
mined independently based on the demand of the workload.
Mixed Workloads
P0n
P0n-AVX2
P0n-AVX-512
P1
AVX512
Cores using Intel®AVX-512
P1-AVX2
AVX2
Cores using Intel® AVX2
P1-AVX-512
Non-AVX
Cores not using Intel®AVX
Cores
SOM00060
Figure 2-6. Mixed Workloads
The following performance monitoring events can be used to determine how many cycles were spent in
each of the three frequency levels.
CORE_POWER.LVL0_TURBO_LICENSE: Core cycles where the core was running in a manner where
the maximum frequency was P0n.
CORE_POWER.LVL1_TURBO_LICENSE: Core cycles where the core was running in a manner where
the maximum frequency was P0n-AVX2.
CORE_POWER.LVL2_TURBO_LICENSE: Core cycles where the core was running in a manner where
the maximum frequency was P0n-AVX-512.
When the core requests a higher license level than its current one, it takes the PCU up to 500
micro-seconds to grant the new license. Until then the core operates at a lower peak capability. During
this time period the PCU evaluates how many cores are executing at the new license level and adjusts
their frequency as necessary, potentially lowering the frequency. Cores that execute at other license
levels are not affected.
A timer of approximately 2ms is applied before going back to a higher frequency level. Any condition that
would have requested a new license resets the timer.
NOTES
A license transition request may occur when executing instructions on a mis-speculated
path.
A large enough mix of Intel AVX-512 light instructions and Intel AVX2 heavy instructions
drives the core to request License 2, despite the fact that they usually map to License 1.
The same is true for Intel AVX2 light instructions and Intel SSE heavy instructions that
may drive the core to License 1 rather than License 0. For example, The Intel® Xeon®
Platinum 8180 processor moves from license 1 to license 2 when executing a mix of 110
Intel AVX-512 light instructions and 20 256-bit heavy instructions over a window of 65
cycles.
2-24
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
Some workloads do not cause the processor to reach its maximum frequency as these workloads are
bound by other factors. For example, the LINPACK benchmark is power limited and does not reach the
processor's maximum frequency. The following graph shows how frequency degrades as vector width
grows, but, despite the frequency drop, performance improves. The data for this graph was collected on
an Intel Xeon Platinum 8180 processor.
LINPACK Performance
3500
3.1
3.5
3259
2.8
3000
3.0
2.5
2500
2.1
2.5
2034
2000
2.0
1500
1178
1.5
760
1000
1.0
500
669
768
791
767
0
SSE4.2
AVX
AVX2
AVX512
GFLOPs
Power (W)
Frequency (GHz)
SOM00061
Figure 2-7. LINPACK Performance
Workloads that execute Intel AVX-512 instructions as a large proportion of their whole instruction count
can gain performance compared to Intel AVX2 instructions, even though they may operate at a lower
frequency. For example, maximum frequency bound Deep Learning workloads that target Intel AVX-512
heavy instructions at a very high percentage can gain 1.3x-1.5x performance improvement vs. the same
workload built to target Intel AVX2 (both operating on Skylake Server microarchitecture).
It is not always easy to predict whether a program's performance will improve from building it to target
Intel AVX-512 instructions. Programs that enjoy high performance gains from the use of xmm or ymm
registers may expect performance improvement by moving to the use of zmm registers. However, some
programs that use zmm registers may not gain as much, or may even lose performance. It is recom-
mended to try multiple build options and measure the performance of the program.
Recommendation: To identify the optimal compiler options to use, build the application with each of the
following set of options and choose the set that provides the best performance.
-xCORE-AVX2 -mtune=skylake-avx512 (Linux* and macOS*)
/QxCORE-AVX2 /tune=skylake-avx512 (Windows*)
-xCORE-AVX512 -qopt-zmm-usage=low (Linux* and macOS*)
/QxCORE-AVX512 /Qopt-zmm-usage:low (Windows*)
-xCORE-AVX512 -qopt-zmm-usage=high (Linux* and macOS*)
/QxCORE-AVX512 /Qopt-zmm-usage:high (Windows*)
See Section 18.26, “CLDEMOTE” for more information about these options.
2-25
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
The GCC Compiler has the option -mprefer-vector-width=none|128|256|512 to control vector width
preference. While -march=skylake-avx512 is designed to provide the best performance for the Skylake
Server microarchitecture some programs can benefit from different vector width preferences. To identify
the optimal compiler options to use, build the application with each of the following set of options and
choose the set that provides the best performance. -mprefer-vector-width=256 is the default for
skylake-avx512.
-march=skylake -mtune=skylake-avx512
-march=skylake-avx512
-march=skylake-avx512 -mprefer-vector-width=512
Clang/LLVM is currently implementing the option -mprefer-vector-width=none|128|256|512, similar
to GCC. To identify the optimal compiler options to use, build the application with each of the following
set of options and choose the set that provides the best performance.
-march=skylake -mtune=skylake-avx512
-march=skylake-avx512 (plus -mprefer-vector-width=256, if available)
-march=skylake-avx512 (plus -mprefer-vector-width=512, if available)
2.6
SKYLAKE CLIENT MICROARCHITECTURE
The Skylake Client microarchitecture builds on the successes of the Haswell and Broadwell microarchitec-
tures. The basic pipeline functionality of the Skylake Client microarchitecture is depicted in Figure 2-8.
32K L1 Instruction
BPU
Cache
MSROM
Decoded Icache
Legacy Decode
(DSB)
Pipeline
4 uops/cycle
6 uops/cycle
5 uops/cycle
Instruction Decode Queue (IDQ,, or micro-op queue)
Allocate/Rename/Retire/MoveElimination/ZeroIdiom
Scheduler
256K L2 Cache
(Unified)
Port 2
Port 0
Port 1
Port 5
Port 6
LD/STA
Int ALU,
Int ALU,
Int ALU,
Int ALU,
Vec FMA,
Fast LEA,
Fast LEA,
Port 3
Int Shft,
Vec MUL,
Vec FMA,
Vec SHUF,
LD/STA
Branch1,
Vec Add,
Vec MUL,
Vec ALU,
Vec ALU,
Vec Add,
CVT
32K L1 Data Cache
Port 4
Vec Shft,
Vec ALU,
STD
Divide,
Vec Shft,
Branch2
Int MUL,
Slow LEA
Port 7
STA
Figure 2-8. CPU Core Pipeline Functionality of the Skylake Client Microarchitecture
The Skylake Client microarchitecture offers the following enhancements:
2-26
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
Larger internal buffers to enable deeper OOO execution and higher cache bandwidth.
Improved front end throughput.
Improved branch predictor.
Improved divider throughput and latency.
Lower power consumption.
Improved SMT performance with Hyper-Threading Technology.
Balanced floating-point ADD, MUL, FMA throughput and latency.
The microarchitecture supports flexible integration of multiple processor cores with a shared uncore
sub-system consisting of a number of components including a ring interconnect to multiple slices of L3
(an off-die L4 is optional), processor graphics, integrated memory controller, interconnect fabrics, etc. A
four-core configuration can be supported similar to the arrangement shown in Appendix E, “Earlier
Generations of Intel® 64 and IA-32 Processor Architectures,” Figure E-2.
2.6.1
The Front End
The front end in the Skylake Client microarchitecture provides the following improvements over previous
generation microarchitectures:
Legacy Decode Pipeline delivery of 5 uops per cycle to the IDQ compared to 4 uops in previous gener-
ations.
The DSB delivers 6 uops per cycle to the IDQ compared to 4 uops in previous generations.
The IDQ can hold 64 uops per logical processor vs. 28 uops per logical processor in previous
generations when two sibling logical processors in the same core are active (2x64 vs. 2x28 per core).
If only one logical processor is active in the core, the IDQ can hold 64 uops (64 vs. 56 uops in ST
operation).
The LSD in the IDQ can detect loops up to 64 uops per logical processor irrespective ST or SMT
operation.
Improved Branch Predictor.
2.6.2
The Out-of-Order Execution Engine
The Out of Order and execution engine changes in Skylake Client microarchitecture include:
Larger buffers enable deeper OOO execution compared to previous generations.
Improved throughput and latency for divide/sqrt and approximate reciprocals.
Identical latency and throughput for all operations running on FMA units.
Longer pause latency enables better power efficiency and better SMT performance resource utili-
zation.
2-27
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
Table 2-11 summarizes the OOO engine’s capability to dispatch different types of operations to various
ports.
Table 2-11. Dispatch Port and Execution Stacks of the Skylake Client Microarchitecture
Port 0
Port 1
Port 2, 3
Port 4
Port 5
Port 6
Port 7
ALU,
ALU,
LD
STD
ALU,
ALU,
STA
Vec ALU
Fast LEA,
STA
Fast LEA,
Shft,
Vec ALU
Vec ALU,
Vec Shft,
Vec Shft,
Vec Shuffle,
Branch1
Vec Add,
Vec Add,
Vec Mul,
Vec Mul,
FMA,
FMA
DIV,
Slow Int
Branch2
Slow LEA
Table 2-12 lists execution units and common representative instructions that rely on these units.
Throughput improvements across the SSE, AVX and general-purpose instruction sets are related to the
number of units for the respective operations, and the varieties of instructions that execute using a
particular unit.
Table 2-12. Skylake Client Microarchitecture Execution Units and Representative Instructions1
Execution
# of
Instructions
Unit
Unit
ALU
4
add, and, cmp, or, test, xor, movzx, movsx, mov, (v)movdqu, (v)movdqa, (v)movap*, (v)movup*
SHFT
2
sal, shl, rol, adc, sarx, adcx, adox, etc.
Slow Int
1
mul, imul, bsr, rcl, shld, mulx, pdep, etc.
BM
2
andn, bextr, blsi, blsmsk, bzhi, etc
Vec ALU
3
(v)pand, (v)por, (v)pxor, (v)movq, (v)movq, (v)movap*, (v)movup*,
(v)andp*, (v)orp*, (v)paddb/w/d/q, (v)blendv*, (v)blendp*, (v)pblendd
Vec_Shft
2
(v)psllv*, (v)psrlv*, vector shift count in imm8
Vec Add
2
(v)addp*, (v)cmpp*, (v)max*, (v)min*, (v)padds*, (v)paddus*, (v)psign, (v)pabs, (v)pavgb,
(v)pcmpeq*, (v)pmax, (v)cvtps2dq, (v)cvtdq2ps, (v)cvtsd2si, (v)cvtss2si
Shuffle
1
(v)shufp*, vperm*, (v)pack*, (v)unpck*, (v)punpck*, (v)pshuf*, (v)pslldq, (v)alignr, (v)pmovzx*,
vbroadcast*, (v)pslldq, (v)psrldq, (v)pblendw
Vec Mul
2
(v)mul*, (v)pmul*, (v)pmadd*,
SIMD Misc
1
STTNI, (v)pclmulqdq, (v)psadw, vector shift count in xmm,
FP Mov
1
(v)movsd/ss, (v)movd gpr,
DIVIDE
1
divp*, divs*, vdiv*, sqrt*, vsqrt*, rcp*, vrcp*, rsqrt*, idiv
NOTES:
1. Execution unit mapping to MMX instructions are not covered in this table. See Section 15.16.5 on MMX instruction
throughput remedy.
2-28
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
A significant portion of the Intel SSE, Intel AVX and general-purpose instructions also have latency
improvements. Appendix C lists the specific details. Software-visible latency exposure of an instruction
sometimes may include additional contributions that depend on the relationship between micro-ops flows
of the producer instruction and the micro-op flows of the ensuing consumer instruction. For example, a
two-uop instruction like VPMULLD may experience two cumulative bypass delays of 1 cycle each from
each of the two micro-ops of VPMULLD.
Table 2-13 describes the bypass delay in cycles between a producer uop and the consumer uop. The
left-most column lists a variety of situations characteristic of the producer micro-op. The top row lists a
variety of situations characteristic of the consumer micro-op.
Table 2-13. Bypass Delay Between Producer and Consumer Micro-ops
SIMD/0,1/1
FMA/0,1/4
VIMUL/0,1/4
SIMD/5/1,3
SHUF/5/1,3
V2I/0/3
I2V/5/1
SIMD/0,1/1
0
1
1
0
0
0
NA
FMA/0,1/4
1
0
1
0
0
0
NA
VIMUL/0,1/4
1
0
1
0
0
0
NA
SIMD/5/1,3
0
1
1
0
0
0
NA
SHUF/5/1,3
0
0
1
0
0
0
NA
V2I/0/3
NA
NA
NA
NA
NA
NA
NA
I2V/5/1
0
0
1
0
0
0
NA
The attributes that are relevant to the producer/consumer micro-ops for bypass are a triplet of abbrevi-
ation/one or more port number/latency cycle of the uop. For example:
“SIMD/0,1/1” applies to 1-cycle vector SIMD uop dispatched to either port 0 or port 1.
“VIMUL/0,1/4” applies to 4-cycle vector integer multiply uop dispatched to either port 0 or port 1.
“SIMD/5/1,3” applies to either 1-cycle or 3-cycle non-shuffle uop dispatched to port 5.
2.6.3
Cache and Memory Subsystem
The cache hierarchy of the Skylake Client microarchitecture has the following enhancements:
Higher Cache bandwidth compared to previous generations.
Simultaneous handling of more loads and stores enabled by enlarged buffers.
Processor can do two page walks in parallel compared to one in Haswell microarchitecture and earlier
generations.
Page split load penalty down from 100 cycles in previous generation to 5 cycles.
L3 write bandwidth increased from 4 cycles per line in previous generation to 2 per line.
Support for the CLFLUSHOPT instruction to flush cache lines and manage memory ordering of flushed
data using SFENCE.
Reduced performance penalty for a software prefetch that specifies a NULL pointer.
L2 associativity changed from 8 ways to 4 ways.
2-29
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
Table 2-14. Cache Parameters of the Skylake Client Microarchitecture
Capacity /
Line Size
Fastest
Peak Bandwidth
Sustained Bandwidth
Update
Level
Associativity
(bytes)
Latency1
(bytes/cyc)
(bytes/cyc)
Policy
First Level Data
32 KB/ 8
64
4 cycle
96 (2x32B Load +
~81
Writeback
1*32B Store)
Instruction
32 KB/8
64
N/A
N/A
N/A
N/A
Second Level
256KB/4
64
12 cycle
64
~29
Writeback
44
Third Level
Up to 2MB
64
32
~18
Writeback
(Shared L3)
per core/Up
to 16 ways
NOTES:
1. Software-visible latency will vary depending on access patterns and other factors.
The TLB hierarchy consists of dedicated level one TLB for instruction cache, TLB for L1D, plus unified TLB
for L2. The partition column of Table 2-15 indicates the resource sharing policy when Hyper-Threading
Technology is active.
Table 2-15. TLB Parameters of the Skylake Client Microarchitecture
Level
Page Size
Entries
Associativity
Partition
Instruction
4KB
128
8 ways
dynamic
Instruction
2MB/4MB
8 per thread
fixed
First Level Data
4KB
64
4
fixed
First Level Data
2MB/4MB
32
4
fixed
First Level Data
1GB
4
4
fixed
Second Level
Shared by 4KB and 2/4MB pages
1536
12
fixed
Second Level
1GB
16
4
fixed
2.6.4
Pause Latency in Skylake Client Microarchitecture
The PAUSE instruction is typically used with software threads executing on two logical processors located
in the same processor core, waiting for a lock to be released. Such short wait loops tend to last between
tens and a few hundreds of cycles, so performance-wise it is better to wait while occupying the CPU than
yielding to the OS. When the wait loop is expected to last for thousands of cycles or more, it is preferable
to yield to the operating system by calling an OS synchronization API function, such as WaitForSingleO-
bject on Windows* OS or futex on Linux.
The PAUSE instruction is intended to:
Temporarily provide the sibling logical processor (ready to make forward progress exiting the spin
loop) with competitively shared hardware resources. The competitively-shared microarchitectural
resources that the sibling logical processor can utilize in the Skylake Client microarchitecture are
listed below.
— Front end slots in the Decode ICache, LSD and IDQ.
— Execution slots in the RS.
Save power consumed by the processor core compared with executing equivalent spin loop
instruction sequence in the following configurations.
— One logical processor is inactive (e.g., entering a C-state).
— Both logical processors in the same core execute the PAUSE instruction.
2-30
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
— HT is disabled (e.g. using BIOS options).
The latency of the PAUSE instruction in prior generation microarchitectures is about 10 cycles, whereas
in Skylake Client microarchitecture it has been extended to as many as 140 cycles.
The increased latency (allowing more effective utilization of competitively-shared microarchitectural
resources to the logical processor ready to make forward progress) has a small positive performance
impact of 1-2% on highly threaded applications. It is expected to have negligible impact on less threaded
applications if forward progress is not blocked executing a fixed number of looped PAUSE instructions.
There's also a small power benefit in 2-core and 4-core systems.
As the PAUSE latency has been increased significantly, workloads that are sensitive to PAUSE latency will
suffer some performance loss.
The following is an example of how to use the PAUSE instruction with a dynamic loop iteration count.
Notice that in the Skylake Client microarchitecture the RDTSC instruction counts at the machine's guar-
anteed P1 frequency independently of the current processor clock (see the INVARIANT TSC property),
and therefore, when running in Intel® Turbo-Boost-enabled mode, the delay will remain constant, but
the number of instructions that could have been executed will change.
Use Poll Delay function in your lock to wait a given amount of guaranteed P1 frequency cycles, specified
in the “clocks” variable.
Example 2-8. Dynamic Pause Loop Example
#include <x86intrin.h>
#include <stdint.h>
/* A useful predicate for dealing with timestamps that may wrap.
Is a before b? Since the timestamps may wrap, this is asking whether it's
shorter to go clockwise from a to b around the clock-face, or anti-clockwise.
Times where going clockwise is less distance than going anti-clockwise
are in the future, others are in the past. e.g. a = MAX-1, b = MAX+1 (=0),
then a > b (true) does not mean a reached b; whereas signed(a) = -2,
signed(b) = 0 captures the actual difference */
static inline bool before(uint64_t a, uint64_t b)
{
return ((int64_t)b - (int64_t)a) > 0;
}
void pollDelay(uint32_t clocks)
{
uint64_t endTime = _rdtsc()+ clocks;
for (; before(_rdtsc(), endTime); )
_mm_pause();
}
For contended spinlocks of the form shown in the baseline example below, we recommend an exponen-
tial back off when the lock is found to be busy, as shown in the improved example, to avoid significant
performance degradation that can be caused by conflicts between threads in the machine. This is more
important as we increase the number of threads in the machine and make changes to the architecture
that might aggravate these conflict conditions. In multi-socket Intel server processors with shared
memory, conflicts across threads take much longer to resolve as the number of threads contending for
the same lock increases. The exponential back off is designed to avoid these conflicts between the
threads thus avoiding the potential performance degradation. Note that in the example below, the
2-31
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
number of PAUSE instructions are increased by a factor of 2 until some MAX_BACKOFF is reached which
is subject to tuning.
Example 2-9. Contended Locks with Increasing Back-off Example
/*******************/
/*Baseline Version */
/*******************/
// atomic {if (lock == free) then change lock state to busy}
while (cmpxchg(lock, free, busy) == fail)
{
while (lock == busy)
{
__asm__ ("pause");
}
}
/*******************/
/*Improved Version */
/*******************/
int mask = 1;
int const max = 64; //MAX_BACKOFF
while (cmpxchg(lock, free, busy) == fail)
{
while (lock == busy)
{
for (int i=mask; i; --i){
__asm__ ("pause");
}
mask = mask < max ? mask<<1 : max;
}
}
2.7
INTEL® HYPER-THREADING TECHNOLOGY (INTEL® HT TECHNOLOGY)
Intel® Hyper-Threading Technology (Intel® HT Technology) enables software to take advantage of
task-level, or thread-level parallelism by providing multiple logical processors within a physical processor
package, or within each processor core in a physical processor package. In its first implementation in the
Intel Xeon processor, Hyper-Threading Technology makes a single physical processor (or a processor
core) appear as two or more logical processors. Intel Xeon Phi processors based on the Knights Landing
microarchitecture support 4 logical processors in each processor core; see Chapter 23 for detailed infor-
mation of Intel HT Technology that is implemented in the Knights Landing microarchitecture.
Most Intel Architecture processor families support Hyper-Threading Technology with two logical proces-
sors in each processor core, or in a physical processor in early implementations. The rest of this section
describes features of the early implementation of Hyper-Threading Technology. Most of the descriptions
also apply to later Hyper-Threading Technology implementations supporting two logical processors. The
microarchitecture sections in this chapter provide additional details to individual microarchitecture and
enhancements to Hyper-Threading Technology.
The two logical processors each have a complete set of architectural registers while sharing one single
physical processor's resources. By maintaining the architecture state of two processors, an Intel HT
2-32
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
Technology capable processor looks like two processors to software, including operating system and
application code.
By sharing resources needed for peak demands between two logical processors, Intel HT Technology is
well suited for multiprocessor systems to provide an additional performance boost in throughput when
compared to traditional MP systems.
Figure 2-9 shows a typical bus-based symmetric multiprocessor (SMP) based on processors supporting
Intel HT Technology. Each logical processor can execute a software thread, allowing a maximum of two
software threads to execute simultaneously on one physical processor. The two software threads execute
simultaneously, meaning that in the same clock cycle an “add” operation from logical processor 0 and
another “add” operation and load from logical processor 1 can be executed simultaneously by the execu-
tion engine.
In the first implementation of Intel HT Technology, the physical execution resources are shared and the
architecture state is duplicated for each logical processor. This minimizes the die area cost of imple-
menting Intel HT Technology while still achieving performance gains for multithreaded applications or
multitasking workloads.
Architectural
Architectural
Architectural
Architectural
State
State
State
State
Execution Engine
Execution Engine
Local APIC
Local APIC
Local APIC
Local APIC
Bus Interface
Bus Interface
System Bus
OM15152
Figure 2-9. Hyper-Threading Technology on an SMP
The performance potential due to HT Technology is due to:
The fact that operating systems and user programs can schedule processes or threads to execute
simultaneously on the logical processors in each physical processor.
The ability to use on-chip execution resources at a higher level than when only a single thread is
consuming the execution resources; higher level of resource utilization can lead to higher system
throughput.
2.7.1
Processor Resources and HT Technology
The majority of microarchitecture resources in a physical processor are shared between the logical
processors. Only a few small data structures were replicated for each logical processor. This section
describes how resources are shared, partitioned or replicated.
2-33
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
2.7.1.1
Replicated Resources
The architectural state is replicated for each logical processor. The architecture state consists of registers
that are used by the operating system and application code to control program behavior and store data
for computations. This state includes the eight general-purpose registers, the control registers, machine
state registers, debug registers, and others. There are a few exceptions, most notably the memory type
range registers (MTRRs) and the performance monitoring resources. For a complete list of the architec-
ture state and exceptions, see the Intel® 64 and IA-32 Architectures Software Developer’s Manual,
Volumes 3A, 3B, 3C, & 3D.
Other resources such as instruction pointers and register renaming tables were replicated to simultane-
ously track execution and state changes of the two logical processors. The return stack predictor is repli-
cated to improve branch prediction of return instructions.
In addition, a few buffers (for example, the 2-entry instruction streaming buffers) were replicated to
reduce complexity.
2.7.1.2
Partitioned Resources
Several buffers are shared by limiting the use of each logical processor to half the entries. These are
referred to as partitioned resources. Reasons for this partitioning include:
Operational fairness.
Permitting the ability to allow operations from one logical processor to bypass operations of the other
logical processor that may have stalled.
For example: a cache miss, a branch misprediction, or instruction dependencies may prevent a logical
processor from making forward progress for some number of cycles. The partitioning prevents the stalled
logical processor from blocking forward progress.
In general, the buffers for staging instructions between major pipe stages are partitioned. These buffers
include µop queues after the execution trace cache, the queues after the register rename stage, the
reorder buffer which stages instructions for retirement, and the load and store buffers.
In the case of load and store buffers, partitioning also provided an easier implementation to maintain
memory ordering for each logical processor and detect memory ordering violations.
2.7.1.3
Shared Resources
Most resources in a physical processor are fully shared to improve the dynamic utilization of the resource,
including caches and all the execution units. Some shared resources which are linearly addressed, like
the DTLB, include a logical processor ID bit to distinguish whether the entry belongs to one logical
processor or the other.
2.7.2
Microarchitecture Pipeline and Intel® HT Technology
This section describes the Intel HT Technology microarchitecture and how instructions from the two
logical processors are handled between the front end and the back end of the pipeline.
Although instructions originating from two programs or two threads execute simultaneously and not
necessarily in program order in the execution core and memory hierarchy, the front end and back end
contain several selection points to select between instructions from the two logical processors. All selec-
tion points alternate between the two logical processors unless one logical processor cannot make use of
a pipeline stage. In this case, the other logical processor has full use of every cycle of the pipeline stage.
Reasons why a logical processor may not use a pipeline stage include cache misses, branch mispredic-
tions, and instruction dependencies.
2-34
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
2.7.3
Execution Core
The core can dispatch up to six µops per cycle, provided the µops are ready to execute. Once the µops
are placed in the queues waiting for execution, there is no distinction between instructions from the two
logical processors. The execution core and memory hierarchy is also oblivious to which instructions
belong to which logical processor.
After execution, instructions are placed in the re-order buffer. The re-order buffer decouples the execu-
tion stage from the retirement stage. The re-order buffer is partitioned such that each uses half the
entries.
2.7.4
Retirement
The retirement logic tracks when instructions from the two logical processors are ready to be retired. It
retires the instruction in program order for each logical processor by alternating between the two logical
processors. If one logical processor is not ready to retire any instructions, then all retirement bandwidth
is dedicated to the other logical processor.
Once stores have retired, the processor needs to write the store data into the level-one data cache.
Selection logic alternates between the two logical processors to commit store data to the cache.
2.8
SIMD TECHNOLOGY
SIMD computations (see Figure 2-10) were introduced to the architecture with MMX technology. MMX
technology allows SIMD computations to be performed on packed byte, word, and doubleword integers.
The integers are contained in a set of eight 64-bit registers called MMX registers (see Figure 2-11).
The Pentium III processor extended the SIMD computation model with the introduction of the Streaming
SIMD Extensions (SSE). SSE allows SIMD computations to be performed on operands that contain four
packed single-precision floating-point data elements. The operands can be in memory or in a set of eight
128-bit XMM registers (see Figure 2-11). SSE also extended SIMD computational capability by adding
additional 64-bit MMX instructions.
Figure 2-10 shows a typical SIMD computation. Two sets of four packed data elements (X1, X2, X3, and
X4, and Y1, Y2, Y3, and Y4) are operated on in parallel, with the same operation being performed on each
corresponding pair of data elements (X1 and Y1, X2 and Y2, X3 and Y3, and X4 and Y4). The results of
the four parallel computations are sorted as a set of four packed data elements.
X4
X3
X2
X1
Y4
Y3
Y2
Y1
OP
OP
OP
OP
X4 op Y4
X3 op Y3
X2 op Y2
X1 op Y1
OM15148
Figure 2-10. Typical SIMD Operations
2-35
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
The Pentium 4 processor further extended the SIMD computation model with the introduction of
Streaming SIMD Extensions 2 (SSE2), Streaming SIMD Extensions 3 (SSE3), and Intel Xeon processor
5100 series introduced Supplemental Streaming SIMD Extensions 3 (SSSE3).
SSE2 works with operands in either memory or in the XMM registers. The technology extends SIMD
computations to process packed double-precision floating-point data elements and 128-bit packed inte-
gers. There are 144 instructions in SSE2 that operate on two packed double-precision floating-point data
elements or on 16 packed byte, 8 packed word, 4 doubleword, and 2 quadword integers.
SSE3 enhances x87, SSE and SSE2 by providing 13 instructions that can accelerate application perfor-
mance in specific areas. These include video processing, complex arithmetics, and thread synchroniza-
tion. SSE3 complements SSE and SSE2 with instructions that process SIMD data asymmetrically,
facilitate horizontal computation, and help avoid loading cache line splits. See Figure 2-11.
SSSE3 provides additional enhancement for SIMD computation with 32 instructions on digital video and
signal processing.
SSE4.1, SSE4.2 and AESNI are additional SIMD extensions that provide acceleration for applications in
media processing, text/lexical processing, and block encryption/decryption.
The SIMD extensions operates the same way in Intel 64 architecture as in IA-32 architecture, with the
following enhancements:
128-bit SIMD instructions referencing XMM register can access 16 XMM registers in 64-bit mode.
Instructions that reference 32-bit general purpose registers can access 16 general purpose registers
in 64-bit mode.
64-bit MMX Registers
128-bit XMM Registers
MM7
XMM7
MM6
XMM6
MM5
XMM5
MM4
XMM4
MM3
XMM3
MM2
XMM2
MM1
XMM1
MM0
XMM0
OM15149
Figure 2-11. SIMD Instruction Register Usage
SIMD improves the performance of 3D graphics, speech recognition, image processing, scientific applica-
tions and applications that have the following characteristics:
Inherently parallel.
Recurring memory access patterns.
Localized recurring operations performed on the data.
Data-independent control flow.
2-36
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
2.9
SUMMARY OF SIMD TECHNOLOGIES AND APPLICATION LEVEL
EXTENSIONS
SIMD floating-point instructions fully support the IEEE Standard 754 for Binary Floating-Point Arithmetic.
They are accessible from all IA-32 execution modes: protected mode, real address mode, and Virtual
8086 mode.
SSE, SSE2, and MMX technologies are architectural extensions. Existing software will continue to run
correctly, without modification on Intel microprocessors that incorporate these technologies. Existing
software will also run correctly in the presence of applications that incorporate SIMD technologies.
SSE and SSE2 instructions also introduced cacheability and memory ordering instructions that can
improve cache usage and application performance.
For more on SSE, SSE2, SSE3 and MMX technologies, see the following chapters in the Intel® 64 and
IA-32 Architectures Software Developer’s Manual, Volume 1:
Chapter 9, “Programming with Intel® MMX Technology.”
Chapter 10, “Programming with Intel® Streaming SIMD Extensions (Intel® SSE).”
Chapter 11, “Programming with Intel® Streaming SIMD Extensions 2 (Intel® SSE2).”
Chapter 12, “Programming with Intel® SSE3, SSSE3, Intel® SSE4, and Intel® AES-NI.”
Chapter 14, “Programming with Intel® AVX, FMA, and Intel® AVX2.”
Chapter 15, “Programming with Intel® AVX-512.”
Chapter 16, “Programming with Intel® Transactional Synchronization Extensions.”
2.9.1
MMX™ Technology
MMX Technology introduced:
64-bit MMX registers.
Support for SIMD operations on packed byte, word, and doubleword integers.
Recommendation: Integer SIMD code written using MMX instructions should consider more efficient
implementations using SSE/Intel AVX instructions.
2.9.2
Streaming SIMD Extensions
Streaming SIMD extensions introduced:
128-bit XMM registers.
128-bit data type with four packed single-precision floating-point operands.
Data prefetch instructions.
Non-temporal store instructions and other cacheability and memory ordering instructions.
Extra 64-bit SIMD integer support.
SSE instructions are useful for 3D geometry, 3D rendering, speech recognition, and video encoding and
decoding.
2.9.3
Streaming SIMD Extensions 2
Streaming SIMD extensions 2 add the following:
128-bit data type with two packed double-precision floating-point operands.
128-bit data types for SIMD integer operation on 16-byte, 8-word, 4-doubleword, or 2-quadword
integers.
Support for SIMD arithmetic on 64-bit integer operands.
2-37
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
Instructions for converting between new and existing data types.
Extended support for data shuffling.
Extended support for cacheability and memory ordering operations.
SSE2 instructions are useful for 3D graphics, video decoding/encoding, and encryption.
2.9.4
Streaming SIMD Extensions 3
Streaming SIMD extensions 3 add the following:
SIMD floating-point instructions for asymmetric and horizontal computation.
A special-purpose 128-bit load instruction to avoid cache line splits.
An x87 FPU instruction to convert to integer independent of the floating-point control word (FCW).
Instructions to support thread synchronization.
SSE3 instructions are useful for scientific, video and multi-threaded applications.
2.9.5
Supplemental Streaming SIMD Extensions 3
The Supplemental Streaming SIMD Extensions 3 introduces 32 new instructions to accelerate eight
types of computations on packed integers. These include:
12 instructions that perform horizontal addition or subtraction operations.
6 instructions that evaluate the absolute values.
2 instructions that perform multiply and add operations and speed up the evaluation of dot products.
2 instructions that accelerate packed-integer multiply operations and produce integer values with
scaling.
2 instructions that perform a byte-wise, in-place shuffle according to the second shuffle control
operand.
6 instructions that negate packed integers in the destination operand if the signs of the corre-
sponding element in the source operand is less than zero.
2 instructions that align data from the composite of two operands.
2.9.6
SSE4.1
SSE4.1 introduces 47 new instructions to accelerate video, imaging and 3D applications. SSE4.1 also
improves compiler vectorization and significantly increase support for packed dword computation. These
include:
Two instructions perform packed dword multiplies.
Two instructions perform floating-point dot products with input/output selects.
One instruction provides a streaming hint for WC loads.
Six instructions simplify packed blending.
Eight instructions expand support for packed integer MIN/MAX.
Four instructions support floating-point round with selectable rounding mode and precision exception
override.
Seven instructions improve data insertion and extractions from XMM registers
Twelve instructions improve packed integer format conversions (sign and zero extensions).
One instruction improves SAD (sum absolute difference) generation for small block sizes.
One instruction aids horizontal searching operations of word integers.
One instruction improves masked comparisons.
2-38
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
One instruction adds qword packed equality comparisons.
One instruction adds dword packing with unsigned saturation.
2.9.7
SSE4.2
SSE4.2 introduces 7 new instructions. These include:
A 128-bit SIMD integer instruction for comparing 64-bit integer data elements.
Four string/text processing instructions providing a rich set of primitives, these primitives can
accelerate:
— Basic and advanced string library functions from strlen, strcmp, to strcspn.
— Delimiter processing, token extraction for lexing of text streams.
— Parser, schema validation including XML processing.
A general-purpose instruction for accelerating cyclic redundancy checksum signature calculations.
A general-purpose instruction for calculating bit count population of integer numbers.
2.9.8
AESNI and PCLMULQDQ
AESNI introduces 7 new instructions, six of them are primitives for accelerating algorithms based on AES
encryption/decryption standard, referred to as AESNI.
The PCLMULQDQ instruction accelerates general-purpose block encryption, which can perform carry-less
multiplication for two binary numbers up to 64-bit wide.
Typically, algorithm based on AES standard involve transformation of block data over multiple iterations
via several primitives. The AES standard supports cipher key of sizes 128, 192, and 256 bits. The respec-
tive cipher key sizes correspond to 10, 12, and 14 rounds of iteration.
AES encryption involves processing 128-bit input data (plain text) through a finite number of iterative
operation, referred to as “AES round”, into a 128-bit encrypted block (ciphertext). Decryption follows the
reverse direction of iterative operation using the “equivalent inverse cipher” instead of the “inverse
cipher”.
The cryptographic processing at each round involves two input data, one is the “state”, the other is the
“round key”. Each round uses a different “round key”. The round keys are derived from the cipher key
using a “key schedule” algorithm. The “key schedule” algorithm is independent of the data processing of
encryption/decryption, and can be carried out independently from the encryption/decryption phase.
The AES extensions provide two primitives to accelerate AES rounds on encryption, two primitives for
AES rounds on decryption using the equivalent inverse cipher, and two instructions to support the AES
key expansion procedure.
2.9.9
Intel® Advanced Vector Extensions (Intel® AVX)
Intel® Advanced Vector Extensions (Intel® AVX) offers comprehensive architectural enhancements over
previous generations of Streaming SIMD Extensions. Intel AVX introduces the following architectural
enhancements:
Support for 256-bit wide vectors and SIMD register set.
256-bit floating-point instruction set enhancement with up to 2X performance gain relative to 128-bit
Streaming SIMD extensions.
Instruction syntax support for generalized three-operand syntax to improve instruction programming
flexibility and efficient encoding of new instruction extensions.
Enhancement of legacy 128-bit SIMD instruction extensions to support three-operand syntax and to
simplify compiler vectorization of high-level language expressions.
2-39
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
Support flexible deployment of 256-bit AVX code, 128-bit AVX code, legacy 128-bit code and scalar
code.
Intel AVX instruction set and 256-bit register state management detail are described in Intel® 64 and
IA-32 Architectures Software Developer’s Manual, Volumes 2A, 2B, 2C, & 2D. Optimization techniques
for Intel AVX are discussed in Chapter 15, “Optimizations for Intel® AVX, Intel® AVX2, and Intel® FMA.”
2.9.10 Half-Precision Floating-Point Conversion (F16C)
VCVTPH2PS and VCVTPS2PH are two instructions supporting half-precision floating-point data type
conversion to and from single-precision floating-point data types. These two instruction extends on the
same programming model as Intel AVX.
2.9.11 RDRAND
The RDRAND instruction retrieves a random number supplied by a cryptographically secure, determin-
istic random bit generator (DBRG). The DBRG is designed to meet NIST SP 800-90A standard.
2.9.12 Fused-Multiply-ADD (FMA) Extensions
FMA extensions enhances Intel AVX with high-throughput, arithmetic capabilities covering fused
multiply-add, fused multiply-subtract, fused multiply add/subtract interleave, signed-reversed multiply
on fused multiply-add and multiply-subtract operations. FMA extensions provide 36 256-bit
floating-point instructions to perform computation on 256-bit vectors and additional 128-bit and scalar
FMA instructions.
2.9.13 Intel® AVX2
Intel AVX2 extends Intel AVX by promoting most of the 128-bit SIMD integer instructions with 256-bit
numeric processing capabilities. AVX2 instructions follow the same programming model as AVX instruc-
tions.
In addition, AVX2 provide enhanced functionalities for broadcast/permute operations on data elements,
vector shift instructions with variable-shift count per data element, and instructions to fetch
non-contiguous data elements from memory.
2.9.14 General-Purpose Bit-Processing Instructions
The fourth generation Intel Core processor family introduces a collection of bit processing instructions
that operate on the general purpose registers. The majority of these instructions uses the VEX-prefix
encoding scheme to provide non-destructive source operand syntax.
There instructions are enumerated by three separate feature flags reported by CPUID. For details, see
Section 5.1 of Intel® 64 and IA-32 Architectures Software Developer’s Manual, Volume 1 and chapters
3, 4 and 5 of the Intel® 64 and IA-32 Architectures Software Developer’s Manual, Volumes 2A, 2B, 2C,
& 2D.
2.9.15 Intel® Transactional Synchronization Extensions
The fourth generation Intel Core processor family introduces Intel® Transactional Synchronization Exten-
sions (Intel® TSX), which aim to improve the performance of lock-protected critical sections of multi-
threaded applications while maintaining the lock-based programming model.
For backg11round and details, see Chapter 16, “Programming with Intel® Transactional Synchronization
Extensions” of Intel® 64 and IA-32 Architectures Software Developer’s Manual, Volume 1.
2-40
INTEL® 64 AND IA-32 PROCESSOR ARCHITECTURES
Software tuning recommendations for using Intel TSX on lock-protected critical sections of multithreaded
applications are described in Chapter 16, “Intel® TSX Recommendations.”
2.9.16 RDSEED
The RDSEED instruction retrieves a random number supplied by a cryptographically secure, enhanced
deterministic random bit generator Enhanced NRBG). The NRBG is designed to meet the NIST SP
800-90B and NIST SP 800-90C standards.
2.9.17 ADCX and ADOX Instructions
The ADCX and ADOX instructions, in conjunction with MULX instruction, enable software to speed up
calculations that require large integer numerics.
2-41
3.
Updates to Chapter 3
Change bars and violet text show changes to Chapter 3of the Intel® 64 and IA-32 Architectures Optimization
Reference Manual: General Optimization Guidelines.
------------------------------------------------------------------------------------------
Changes to this chapter:
• Updated capitalization of headings throughout chapter.
• Updated branding throughout chapter.
• Typo and punctuation corrections as necessary.
• Section 3.4.2.1:
— Changed Micro-fusion to Microfusion in heading to match Macrofusion.
• Section 3.5.2.3:
— Removed outdated technologies references, focusing on information starting from Skylake microarchi-
tecture.
— Provided latest information regarding *H micro-operations.
— Provided a new figure showing location of *H in Port one.
• Section 3.11:
— Updated various items for clarity.
— Section 3.11.2: Added corrected recipes.
— Table 3-8: Changed column orientation to be more legible.
— Table 3-11: Config section now includes workers field.
• Section 3.12: Updated for clarity.
Intel® 64 and IA-32 Architectures Optimization Reference Manual Documentation Changes
13
CHAPTER 3
GENERAL OPTIMIZATION GUIDELINES
This chapter discusses general optimization techniques that can improve the performance of applications
running on Intel® processors. These techniques take advantage of microarchitectural features described
in Chapter 2, “Intel® 64 and IA-32 Processor Architectures.” Optimization guidelines focusing on Intel
multi-core processors, Hyper-Threading Technology, and 64-bit mode applications are discussed in
Chapter 11, “Multicore and Intel® Hyper-Threading Technology (intel® HT),” and Chapter 13, “64-bit
Mode Coding Guidelines.”
Practices that optimize performance focus on three areas:
Tools and techniques for code generation.
Analysis of the performance characteristics of the workload and its interaction with microarchitec-
tural sub-systems.
Tuning code to the target microarchitecture (or families of microarchitecture) to improve perfor-
mance.
Some hints on using tools are summarized first to simplify the first two tasks. The rest of the chapter will
focus on recommendations for code generation or code tuning to the target microarchitectures.
This chapter explains optimization techniques for the Intel® C++ Compiler, the Intel® Fortran Compiler,
and other compilers.
3.1
PERFORMANCE TOOLS
Intel offers several tools to help optimize application performance, including compilers, performance
analysis, and multithreading tools.
3.1.1
Intel® C++ and Fortran Compilers
Intel compilers support multiple operating systems (Windows*, Linux*, Mac OS*, and embedded). The
Intel compilers optimize performance and give application developers access to advanced features,
including:
Flexibility to target 32-bit or 64-bit Intel processors for optimization.
Compatibility with many integrated development environments or third-party compilers.
Automatic optimization features to take advantage of the target processor’s architecture.
Automatic compiler optimization reduces the need to write different code for different processors.
Common compiler features that are supported across Windows, Linux, and Mac OS include:
— General optimization settings.
— Cache-management features.
— Interprocedural optimization (IPO) methods.
— Profile-guided optimization (PGO) methods.
— Multithreading support.
— Floating-point arithmetic precision and consistency support.
— Compiler optimization and vectorization reports.
GENERAL OPTIMIZATION GUIDELINES
3.1.2
General Compiler Recommendations
Generally speaking, a compiler tuned for a target microarchitecture can be expected to match or outper-
form hand-coding. However, if performance problems are noted with the compiled code, some compilers
(like Intel C++ and Fortran compilers) allow the coder to insert intrinsics or inline assembly to exert
control over generated code. If inline assembly is used, the user must verify that the code generated is
high quality and yields good performance.
Default compiler switches are targeted for common cases. An optimization may be made to the compiler
default if it benefits most programs. If the root cause of a performance problem is a poor choice on the
part of the compiler, using different switches or compiling the targeted module with a different compiler
may be the solution. See the “Quick Reference Guide to Optimization with Intel® C++ and Fortran
Compilers” for additional suggestions on compiler Optimization Options, including processor-specific
ones.
3.1.3
VTune Performance Analyzer
VTune uses performance monitoring hardware to collect statistics and coding information about your
application and its interaction with the microarchitecture. This allows software engineers to measure
performance characteristics of the workload for a given microarchitecture. VTune supports all current
and past Intel processor families.
The VTune Performance Analyzer provides two kinds of feedback:
Indication of a performance improvement gained by using a specific coding recommendation or
microarchitectural feature.
Information on whether a change in the program has improved or degraded performance with
respect to a particular metric.
The VTune Performance Analyzer also provides measures for a number of workload characteristics,
including:
Retirement throughput of instruction execution as an indication of the degree of extractable
instruction-level parallelism in the workload.
Data traffic locality as an indication of the stress point of the cache and memory hierarchy.
Data traffic parallelism as an indication of the degree of effectiveness of amortization of data access
latency.
NOTE
Improving performance in one part of the machine does not necessarily bring significant
gains to overall performance. It is possible to degrade overall performance by improving
performance for some particular metric.
Where appropriate, coding recommendations in this chapter include descriptions of the VTune Perfor-
mance Analyzer events that provide measurable data on the performance gain achieved by following the
recommendations. For more on using the VTune analyzer, refer to the application’s online help.
3.2
PROCESSOR PERSPECTIVES
Many coding recommendations work well across current microarchitectures. However, there are situa-
tions where a recommendation may benefit one microarchitecture more than another.
3.2.1
CPUID Dispatch Strategy and Compatible Code Strategy
When optimum performance on all processor generations is desired, applications can take advantage of
the CPUID instruction to identify the processor generation and integrate processor-specific instructions
3-2
GENERAL OPTIMIZATION GUIDELINES
into the source code. The Intel C++ Compiler supports the integration of different versions of the code
for different target processors. The selection of which code to execute at runtime is made based on the
CPU identifiers. Binary code targeted for different processor generations can be generated under the
control of the programmer or by the compiler. Refer to the “Intel® C++ Compiler Classic Developer
Guide and Reference” cpu_dispatch and cpu_specific sections for more information on CPU dispatching
(a.k.a function multi-versioning).
For applications that target multiple generations of microarchitectures, and where minimum binary code
size and single code path is important, a compatible code strategy is the best. Optimizing applications
using techniques developed for the Intel Core microarchitecture combined with Nehalem microarchitec-
ture are likely to improve code efficiency and scalability when running on processors based on current
and future generations of Intel 64 and IA-32 processors.
3.2.2
Transparent Cache-Parameter Strategy
If the CPUID instruction supports function leaf 4, also known as deterministic cache parameter leaf, the
leaf reports cache parameters for each level of the cache hierarchy in a deterministic and
forward-compatible manner across Intel 64 and IA-32 processor families.
For coding techniques that rely on specific parameters of a cache level, using the deterministic cache
parameter allows software to implement techniques in a way that is forward-compatible with future
generations of Intel 64 and IA-32 processors, and cross-compatible with processors equipped with
different cache sizes.
3.2.3
Threading Strategy and Hardware Multithreading Support
Intel 64 and IA-32 processor families offer hardware multithreading support in two forms: multi-core
technology and HT Technology.
To fully harness the performance potential of hardware multithreading in current and future generations
of Intel 64 and IA-32 processors, software must embrace a threaded approach in application design. At
the same time, to address the widest range of installed machines, multithreaded software should be able
to run without failure on a single processor without hardware multithreading support and should achieve
performance on a single logical processor that is comparable to an unthreaded implementation (if such
comparison can be made). This generally requires architecting a multithreaded application to minimize
the overhead of thread synchronization. Additional guidelines on multithreading are discussed in Chapter
11, “Multicore and Intel® Hyper-Threading Technology (intel® HT).”
3.3
CODING RULES, SUGGESTIONS, AND TUNING HINTS
This section includes rules, suggestions, and hints. They are targeted for engineers who are:
Modifying source code to enhance performance (user/source rules).
Writing assemblers or compilers (assembly/compiler rules).
Doing detailed performance tuning (tuning suggestions).
Coding recommendations are ranked in importance using two measures:
Local impact (high, medium, or low) refers to a recommendation’s affect on the performance of a
given instance of code.
Generality (high, medium, or low) measures how often such instances occur across all application
domains. Generality may also be thought of as “frequency.”
These recommendations are approximate. They can vary depending on coding style, application domain,
and other factors.
The purpose of the high, medium, and low (H, M, and L) priorities is to suggest the relative level of
performance gain one can expect if a recommendation is implemented.
3-3
GENERAL OPTIMIZATION GUIDELINES
Because it is not possible to predict the frequency of a particular code instance in applications, priority
hints cannot be directly correlated to application-level performance gain. In cases in which applica-
tion-level performance gain has been observed, we have provided a quantitative characterization of the
gain (for information only). In cases in which the impact has been deemed inapplicable, no priority is
assigned.
3.4
OPTIMIZING THE FRONT END
Optimizing the front end covers two aspects:
Maintaining steady supply of micro-ops to the execution engine — Mispredicted branches can disrupt
streams of micro-ops, or cause the execution engine to waste execution resources on executing
streams of micro-ops in the non-architected code path. Much of the tuning in this respect focuses on
working with the Branch Prediction Unit. Common techniques are covered in Section 3.4.1, “Branch
Prediction Optimization.”
Supplying streams of micro-ops to utilize the execution bandwidth and retirement bandwidth as
much as possible — For Intel Core microarchitecture and Intel Core Duo processor family, this aspect
focuses maintaining high decode throughput. In Sandy Bridge microarchitecture, this aspect focuses
on keeping the hot code running from Decoded ICache. Techniques to maximize decode throughput
for Intel Core microarchitecture are covered in Section 3.4.2, “Fetch and Decode Optimization.”
3.4.1
Branch Prediction Optimization
Branch optimizations have a significant impact on performance. By understanding the flow of branches
and improving their predictability, you can increase the speed of code significantly.
Optimizations that help branch prediction are:
Keep code and data on separate pages. This is very important; see Section 3.6, “Optimizing Memory
Accesses,” for more information.
Eliminate branches whenever possible.
Arrange code to be consistent with the static branch prediction algorithm.
Use the PAUSE instruction in spin-wait loops.
Inline functions and pair up calls and returns.
Unroll as necessary so that repeatedly-executed loops have sixteen or fewer iterations (unless this
causes an excessive code size increase).
Avoid putting multiple conditional branches in the same 8-byte aligned code block (i.e, have their last
bytes' addresses within the same 8-byte aligned code) if the lower 6 bits of their target IPs are the
same. This restriction has been removed in Ice Lake Client and later microarchitectures.
3.4.1.1
Eliminating Branches
Eliminating branches improves performance because:
It reduces the possibility of mispredictions.
It reduces the number of required branch target buffer (BTB) entries. Conditional branches that are
never taken do not consume BTB resources.
There are four principal ways of eliminating branches:
Arrange code to make basic blocks contiguous.
Unroll loops, as discussed in Section 3.4.1.6, “Loop Unrolling.”
Use the CMOV instruction.
Use the SETCC instruction.
3-4
GENERAL OPTIMIZATION GUIDELINES
The following rules apply to branch elimination:
Assembly/Compiler Coding Rule 1. (MH impact, M generality) Arrange code to make basic blocks
contiguous and eliminate unnecessary branches.
Assembly/Compiler Coding Rule 2. (M impact, ML generality) Use the SETCC and CMOV
instructions to eliminate unpredictable conditional branches where possible. Do not do this for
predictable branches. Do not use these instructions to eliminate all unpredictable conditional branches
(because using these instructions will incur execution overhead due to the requirement for executing
both paths of a conditional branch). In addition, converting a conditional branch to SETCC or CMOV
trades off control flow dependence for data dependence and restricts the capability of the out-of-order
engine. When tuning, note that all Intel 64 and IA-32 processors usually have very high branch
prediction rates. Consistently mispredicted branches are generally rare. Use these instructions only if
the increase in computation time is less than the expected cost of a mispredicted branch.
Consider a line of C code that has a condition dependent upon one of the constants:
X = (A < B) CONST1 : CONST2;
This code conditionally compares two values, A and B. If the condition is true, X is set to CONST1; other-
wise it is set to CONST2. An assembly code sequence equivalent to the above C code can contain
branches that are not predictable if there are no correlation in the two values.
Example 3-1 shows the assembly code with unpredictable branches. The unpredictable branches can be
removed with the use of the SETCC instruction. Example 3-2 shows optimized code that has no
branches.
Example 3-1. Assembly Code with an Unpredictable Branch
cmp a, b
; Condition
jbe L30
; Conditional branch
mov ebx const1
; ebx holds X
jmp L31
; Unconditional branch
L30:
mov ebx, const2
L31:
Example 3-2. Code Optimization to Eliminate Branches
xor ebx, ebx
; Clear ebx (X in the C code)
cmp A, B
setge bl
; When ebx = 0 or 1
; OR the complement condition
sub ebx, 1
; ebx=11...11 or 00...00
and ebx, CONST3; CONST3 = CONST1-CONST2
add ebx, CONST2; ebx=CONST1 or CONST2
The optimized code in Example 3-2 sets EBX to zero, then compares A and B. If A is greater than or equal
to B, EBX is set to one. Then EBX is decreased and AND’d with the difference of the constant values. This
sets EBX to either zero or the difference of the values. By adding CONST2 back to EBX, the correct value
is written to EBX. When CONST2 is equal to zero, the last instruction can be deleted.
Another way to remove branches is to use the CMOV and FCMOV instructions. Example 3-3 shows how to
change a TEST and branch instruction sequence using CMOV to eliminate a branch. If the TEST sets the
equal flag, the value in EBX will be moved to EAX. This branch is data-dependent, and is representative
of an unpredictable branch.
3-5

 

 

 

 

 

 

 

Content      ..     138      139      140      141     ..