Intel 64 and IA-32 Architectures. Software Developer’s Manual (Collection, 2023) - page 15

 

  Index      Manuals     Intel 64 and IA-32 Architectures. Software Developer’s Manual (Collection, 2023)

 

Search            copyright infringement  

 

   

 

   

 

Content      ..     13      14      15      16     ..

 

 

 

Intel 64 and IA-32 Architectures. Software Developer’s Manual (Collection, 2023) - page 15

 

 

Intel® 64 and IA-32 Architectures
Software Developer’s Manual
Volume 2A:
Instruction Set Reference, A-L
NOTE: The Intel® 64 and IA-32 Architectures Software Developer's Manual consists of ten volumes:
Basic Architecture, Order Number 253665; Instruction Set Reference, A-L, Order Number 253666;
Instruction Set Reference, M-U, Order Number 253667; Instruction Set Reference, V, Order Number
326018; Instruction Set Reference, W-Z, Order Number 334569; System Programming Guide, Part 1,
Order Number 253668; System Programming Guide, Part
2, Order Number 253669; System
Programming Guide, Part 3, Order Number 326019; System Programming Guide, Part 4, Order Number
332831; Model-Specific Registers, Order Number 335592. Refer to all ten volumes when evaluating
your design needs.
Order Number: 253666-080US
June 2023
CONTENTS
PAGE
CHAPTER 1
ABOUT THIS MANUAL
1.1
INTEL® 64 AND IA-32 PROCESSORS COVERED IN THIS MANUAL
1-1
1.2
OVERVIEW OF VOLUME 2A, 2B, 2C, AND 2D: INSTRUCTION SET REFERENCE
1-4
1.3
NOTATIONAL CONVENTIONS
1-5
1.3.1
Bit and Byte Order
1-5
1.3.2
Reserved Bits and Software Compatibility
1-6
1.3.3
Instruction Operands
1-6
1.3.4
Hexadecimal and Binary Numbers
1-6
1.3.5
Segmented Addressing
1-7
1.3.6
Exceptions
1-7
1.3.7
A New Syntax for CPUID, CR, and MSR Values
1-7
1.4
RELATED LITERATURE
1-8
CHAPTER 2
INSTRUCTION FORMAT
2.1
INSTRUCTION FORMAT FOR PROTECTED MODE, REAL-ADDRESS MODE, AND VIRTUAL-8086 MODE
2-1
2.1.1
Instruction Prefixes
2-1
2.1.2
Opcodes
2-3
2.1.3
ModR/M and SIB Bytes
2-3
2.1.4
Displacement and Immediate Bytes
2-3
2.1.5
Addressing-Mode Encoding of ModR/M and SIB Bytes
2-4
2.2
IA-32E MODE
2-7
2.2.1
REX Prefixes
2-7
2.2.1.1
Encoding
2-8
2.2.1.2
More on REX Prefix Fields
2-8
2.2.1.3
Displacement
2-11
2.2.1.4
Direct Memory-Offset MOVs
2-11
2.2.1.5
Immediates
2-11
2.2.1.6
RIP-Relative Addressing
2-12
2.2.1.7
Default 64-Bit Operand Size
2-12
2.2.2
Additional Encodings for Control and Debug Registers
2-12
2.3
INTEL® ADVANCED VECTOR EXTENSIONS (INTEL® AVX)
2-13
2.3.1
Instruction Format
2-13
2.3.2
VEX and the LOCK prefix
2-13
2.3.3
VEX and the 66H, F2H, and F3H prefixes
2-13
2.3.4
VEX and the REX prefix
2-13
2.3.5
The VEX Prefix
2-14
2.3.5.1
VEX Byte 0, bits[7:0]
2-15
2.3.5.2
VEX Byte 1, bit [7] - ‘R’
2-15
2.3.5.3
3-byte VEX byte 1, bit[6] - ‘X’
2-16
2.3.5.4
3-byte VEX byte 1, bit[5] - ‘B’
2-16
2.3.5.5
3-byte VEX byte 2, bit[7] - ‘W’
2-16
2.3.5.6
2-byte VEX Byte 1, bits[6:3] and 3-byte VEX Byte 2, bits [6:3]- ‘vvvv’ the Source or Dest Register Specifier
2-16
2.3.6
Instruction Operand Encoding and VEX.vvvv, ModR/M
2-17
2.3.6.1
3-byte VEX byte 1, bits[4:0] - “m-mmmm”
2-18
2.3.6.2
2-byte VEX byte 1, bit[2], and 3-byte VEX byte 2, bit [2]- “L”
2-18
2.3.6.3
2-byte VEX byte 1, bits[1:0], and 3-byte VEX byte 2, bits [1:0]- “pp”
2-19
2.3.7
The Opcode Byte
2-19
2.3.8
The ModR/M, SIB, and Displacement Bytes
2-19
2.3.9
The Third Source Operand (Immediate Byte)
2-19
2.3.10
Intel® AVX Instructions and the Upper 128-bits of YMM registers
2-19
2.3.10.1
Vector Length Transition and Programming Considerations
2-19
2.3.11
Intel® AVX Instruction Length
2-20
2.3.12
Vector SIB (VSIB) Memory Addressing
2-20
2.3.12.1
64-bit Mode VSIB Memory Addressing
2-21
2.4
INTEL® ADVANCED MATRIX EXTENSIONS (INTEL® AMX)
2-21
2.5
INTEL® AVX AND INTEL® SSE INSTRUCTION EXCEPTION CLASSIFICATION
2-22
2.5.1
Exceptions Type 1 (Aligned Memory Reference)
2-27
2.5.2
Exceptions Type 2 (>=16 Byte Memory Reference, Unaligned)
2-28
Vol. 2A iii
CONTENTS
PAGE
2.5.3
Exceptions Type 3 (<16 Byte Memory Argument)
2-29
2.5.4
Exceptions Type 4 (>=16 Byte Mem Arg, No Alignment, No Floating-point Exceptions)
2-30
2.5.5
Exceptions Type 5 (<16 Byte Mem Arg and No FP Exceptions)
2-31
2.5.6
Exceptions Type 6 (VEX-Encoded Instructions without Legacy SSE Analogues)
2-32
2.5.7
Exceptions Type 7 (No FP Exceptions, No Memory Arg)
2-33
2.5.8
Exceptions Type 8 (AVX and No Memory Argument)
2-33
2.5.9
Exceptions Type 11 (VEX-only, Mem Arg, No AC, Floating-point Exceptions)
2-34
2.5.10
Exceptions Type 12 (VEX-only, VSIB Mem Arg, No AC, No Floating-point Exceptions)
2-35
2.6
VEX ENCODING SUPPORT FOR GPR INSTRUCTIONS
2-35
2.6.1
Exceptions Type 13 (VEX-Encoded GPR Instructions)
2-36
2.7
INTEL® AVX-512 ENCODING
2-36
2.7.1
Instruction Format and EVEX
2-37
2.7.2
Register Specifier Encoding and EVEX
2-39
2.7.3
Opmask Register Encoding
2-40
2.7.4
Masking Support in EVEX
2-40
2.7.5
Compressed Displacement (disp8*N) Support in EVEX
2-41
2.7.6
EVEX Encoding of Broadcast/Rounding/SAE Support
2-42
2.7.7
Embedded Broadcast Support in EVEX
2-42
2.7.8
Static Rounding Support in EVEX
2-42
2.7.9
SAE Support in EVEX
2-42
2.7.10
Vector Length Orthogonality
2-42
2.7.11
#UD Equations for EVEX
2-43
2.7.11.1
State Dependent #UD
2-43
2.7.11.2
Opcode Independent #UD
2-43
2.7.11.3
Opcode Dependent #UD
2-44
2.7.12
Device Not Available
2-45
2.7.13
Scalar Instructions
2-45
2.8
EXCEPTION CLASSIFICATIONS OF EVEX-ENCODED INSTRUCTIONS
2-45
2.8.1
Exceptions Type E1 and E1NF of EVEX-Encoded Instructions
2-49
2.8.2
Exceptions Type E2 of EVEX-Encoded Instructions
2-51
2.8.3
Exceptions Type E3 and E3NF of EVEX-Encoded Instructions
2-52
2.8.4
Exceptions Type E4 and E4NF of EVEX-Encoded Instructions
2-54
2.8.5
Exceptions Type E5 and E5NF
2-56
2.8.6
Exceptions Type E6 and E6NF
2-58
2.8.7
Exceptions Type E7NM
2-60
2.8.8
Exceptions Type E9 and E9NF
2-61
2.8.9
Exceptions Type E10 and E10NF
2-63
2.8.10
Exceptions Type E11 (EVEX-only, Mem Arg, No AC, Floating-point Exceptions)
2-65
2.8.11
Exceptions Type E12 and E12NP (VSIB Mem Arg, No AC, No Floating-point Exceptions)
2-66
2.9
EXCEPTION CLASSIFICATIONS OF OPMASK INSTRUCTIONS, TYPE K20 AND TYPE K21
2-68
2.9.1
Exceptions Type K20
2-68
2.9.2
Exceptions Type K21
2-69
2.10
INTEL® AMX INSTRUCTION EXCEPTION CLASSES
2-70
CHAPTER 3
INSTRUCTION SET REFERENCE, A-L
3.1
INTERPRETING THE INSTRUCTION REFERENCE PAGES
3-1
3.1.1
Instruction Format
3-1
3.1.1.1
Opcode Column in the Instruction Summary Table (Instructions without VEX Prefix)
3-2
3.1.1.2
Opcode Column in the Instruction Summary Table (Instructions with VEX prefix)
3-3
3.1.1.3
Instruction Column in the Opcode Summary Table
3-5
3.1.1.4
Operand Encoding Column in the Instruction Summary Table
3-8
3.1.1.5
64/32-bit Mode Column in the Instruction Summary Table
3-8
3.1.1.6
CPUID Support Column in the Instruction Summary Table
3-9
3.1.1.7
Description Column in the Instruction Summary Table
3-9
3.1.1.8
Description Section
3-9
3.1.1.9
Operation Section
3-9
3.1.1.10
Intel® C/C++ Compiler Intrinsics Equivalents Section
3-12
3.1.1.11
Flags Affected Section
3-14
3.1.1.12
FPU Flags Affected Section
3-14
3.1.1.13
Protected Mode Exceptions Section
3-14
3.1.1.14
Real-Address Mode Exceptions Section
3-15
3.1.1.15
Virtual-8086 Mode Exceptions Section
3-15
3.1.1.16
Floating-Point Exceptions Section
3-16
iv Vol. 2A
CONTENTS
PAGE
3.1.1.17
SIMD Floating-Point Exceptions Section
3-16
3.1.1.18
Compatibility Mode Exceptions Section
3-16
3.1.1.19
64-Bit Mode Exceptions Section
3-16
3.2
INTEL® AMX CONSIDERATIONS
3-16
3.2.1
Implementation Parameters
3-17
3.2.2
Helper Functions
3-17
3.3
INSTRUCTIONS (A-L)
3-18
AAA-ASCII Adjust After Addition
3-19
AAD-ASCII Adjust AX Before Division
3-21
AAM-ASCII Adjust AX After Multiply
3-23
AAS-ASCII Adjust AL After Subtraction
3-25
ADC-Add With Carry
3-27
ADCX-Unsigned Integer Addition of Two Operands With Carry Flag
3-30
ADD-Add
3-32
ADDPD-Add Packed Double Precision Floating-Point Values
3-34
ADDPS-Add Packed Single Precision Floating-Point Values
3-37
ADDSD-Add Scalar Double Precision Floating-Point Values
3-40
ADDSS-Add Scalar Single Precision Floating-Point Values
3-42
ADDSUBPD-Packed Double Precision Floating-Point Add/Subtract
3-44
ADDSUBPS-Packed Single Precision Floating-Point Add/Subtract
3-46
ADOX - Unsigned Integer Addition of Two Operands With Overflow Flag
3-49
AESDEC-Perform One Round of an AES Decryption Flow
3-51
AESDEC128KL-Perform Ten Rounds of AES Decryption Flow With Key Locker Using 128-Bit Key
3-53
AESDEC256KL-Perform 14 Rounds of AES Decryption Flow With Key Locker Using 256-Bit Key
3-55
AESDECLAST-Perform Last Round of an AES Decryption Flow
3-57
AESDECWIDE128KL-Perform Ten Rounds of AES Decryption Flow With Key Locker on 8 Blocks Using 128-Bit Key 3-59
AESDECWIDE256KL-Perform 14 Rounds of AES Decryption Flow With Key Locker on 8 Blocks Using 256-Bit Key . 3-61
AESENC-Perform One Round of an AES Encryption Flow
3-63
AESENC128KL-Perform Ten Rounds of AES Encryption Flow With Key Locker Using 128-Bit Key
3-65
AESENC256KL-Perform 14 Rounds of AES Encryption Flow With Key Locker Using 256-Bit Key
3-67
AESENCLAST-Perform Last Round of an AES Encryption Flow
3-69
AESENCWIDE128KL-Perform Ten Rounds of AES Encryption Flow With Key Locker on 8 Blocks Using 128-Bit Key 3-71
AESENCWIDE256KL-Perform 14 Rounds of AES Encryption Flow With Key Locker on 8 Blocks Using 256-Bit Key . 3-73
AESIMC-Perform the AES InvMixColumn Transformation
3-75
AESKEYGENASSIST-AES Round Key Generation Assist
3-76
AND-Logical AND
3-78
ANDN-Logical AND NOT
3-80
ANDPD-Bitwise Logical AND of Packed Double Precision Floating-Point Values
3-81
ANDPS-Bitwise Logical AND of Packed Single Precision Floating-Point Values
3-84
ANDNPD-Bitwise Logical AND NOT of Packed Double Precision Floating-Point Values
3-87
ANDNPS-Bitwise Logical AND NOT of Packed Single Precision Floating-Point Values
3-90
ARPL-Adjust RPL Field of Segment Selector
3-93
BEXTR-Bit Field Extract
3-95
BLENDPD-Blend Packed Double Precision Floating-Point Values
3-96
BLENDPS-Blend Packed Single Precision Floating-Point Values
3-98
BLENDVPD-Variable Blend Packed Double Precision Floating-Point Values
3-100
BLENDVPS-Variable Blend Packed Single Precision Floating-Point Values
3-102
BLSI-Extract Lowest Set Isolated Bit
3-105
BLSMSK-Get Mask Up to Lowest Set Bit
3-106
BLSR-Reset Lowest Set Bit
3-107
BNDCL-Check Lower Bound
3-108
BNDCU/BNDCN-Check Upper Bound
3-110
BNDLDX-Load Extended Bounds Using Address Translation
3-112
BNDMK-Make Bounds
3-115
BNDMOV-Move Bounds
3-117
BNDSTX-Store Extended Bounds Using Address Translation
3-120
BOUND-Check Array Index Against Bounds
3-123
BSF-Bit Scan Forward
3-125
BSR-Bit Scan Reverse
3-127
Vol. 2A v
CONTENTS
PAGE
BSWAP-Byte Swap
3-129
BT-Bit Test
3-130
BTC-Bit Test and Complement
3-132
BTR-Bit Test and Reset
3-134
BTS-Bit Test and Set
3-136
BZHI-Zero High Bits Starting with Specified Bit Position
3-138
CALL-Call Procedure
3-139
CBW/CWDE/CDQE-Convert Byte to Word/Convert Word to Doubleword/Convert Doubleword to Quadword
3-156
CLAC-Clear AC Flag in EFLAGS Register
3-157
CLC-Clear Carry Flag
3-158
CLD-Clear Direction Flag
3-159
CLDEMOTE-Cache Line Demote
3-160
CLFLUSH-Flush Cache Line
3-162
CLFLUSHOPT-Flush Cache Line Optimized
3-164
CLI-Clear Interrupt Flag
3-166
CLRSSBSY-Clear Busy Flag in a Supervisor Shadow Stack Token
3-168
CLTS-Clear Task-Switched Flag in CR0
3-170
CLUI-Clear User Interrupt Flag
3-171
CLWB-Cache Line Write Back
3-172
CMC-Complement Carry Flag
3-174
CMOVcc-Conditional Move
3-175
CMP-Compare Two Operands
3-179
CMPPD-Compare Packed Double Precision Floating-Point Values
3-181
CMPPS-Compare Packed Single Precision Floating-Point Values
3-188
CMPS/CMPSB/CMPSW/CMPSD/CMPSQ-Compare String Operands
3-195
CMPSD-Compare Scalar Double Precision Floating-Point Value
3-199
CMPSS-Compare Scalar Single Precision Floating-Point Value
3-203
CMPXCHG-Compare and Exchange
3-208
CMPXCHG8B/CMPXCHG16B-Compare and Exchange Bytes
3-210
COMISD-Compare Scalar Ordered Double Precision Floating-Point Values and Set EFLAGS
3-213
COMISS-Compare Scalar Ordered Single Precision Floating-Point Values and Set EFLAGS
3-215
CPUID-CPU Identification
3-217
CRC32-Accumulate CRC32 Value
3-262
CVTDQ2PD-Convert Packed Doubleword Integers to Packed Double Precision Floating-Point Values
3-265
CVTDQ2PS-Convert Packed Doubleword Integers to Packed Single Precision Floating-Point Values
3-269
CVTPD2DQ-Convert Packed Double Precision Floating-Point Values to Packed Doubleword Integers
3-272
CVTPD2PI-Convert Packed Double Precision Floating-Point Values to Packed Dword Integers
3-276
CVTPD2PS-Convert Packed Double Precision Floating-Point Values to Packed Single Precision Floating-Point
Values
3-277
CVTPI2PD-Convert Packed Dword Integers to Packed Double Precision Floating-Point Values
3-281
CVTPI2PS-Convert Packed Dword Integers to Packed Single Precision Floating-Point Values
3-282
CVTPS2DQ-Convert Packed Single Precision Floating-Point Values to Packed Signed Doubleword Integer Values . 3-283
CVTPS2PD-Convert Packed Single Precision Floating-Point Values to Packed Double Precision Floating-Point
Values
3-286
CVTPS2PI-Convert Packed Single Precision Floating-Point Values to Packed Dword Integers
3-289
CVTSD2SI-Convert Scalar Double Precision Floating-Point Value to Doubleword Integer
3-290
CVTSD2SS-Convert Scalar Double Precision Floating-Point Value to Scalar Single Precision Floating-Point Value . . 3-292
CVTSI2SD-Convert Doubleword Integer to Scalar Double Precision Floating-Point Value
3-294
CVTSI2SS-Convert Doubleword Integer to Scalar Single Precision Floating-Point Value
3-296
CVTSS2SD-Convert Scalar Single Precision Floating-Point Value to Scalar Double Precision Floating-Point Value . . 3-298
CVTSS2SI-Convert Scalar Single Precision Floating-Point Value to Doubleword Integer
3-300
CVTTPD2DQ-Convert with Truncation Packed Double Precision Floating-Point Values to Packed Doubleword
Integers
3-302
CVTTPD2PI-Convert With Truncation Packed Double Precision Floating-Point Values to Packed Dword Integers . . 3-306
CVTTPS2DQ-Convert With Truncation Packed Single Precision Floating-Point Values to Packed Signed Doubleword
Integer Values
3-307
CVTTPS2PI-Convert With Truncation Packed Single Precision Floating-Point Values to Packed Dword Integers . . . 3-310
CVTTSD2SI-Convert With Truncation Scalar Double Precision Floating-Point Value to Signed Integer
3-311
CVTTSS2SI-Convert With Truncation Scalar Single Precision Floating-Point Value to Integer
3-313
vi Vol. 2A
CONTENTS
PAGE
CWD/CDQ/CQO-Convert Word to Doubleword/Convert Doubleword to Quadword
3-315
DAA-Decimal Adjust AL After Addition
3-316
DAS-Decimal Adjust AL After Subtraction
3-318
DEC-Decrement by 1
3-320
DIV-Unsigned Divide
3-322
DIVPD-Divide Packed Double Precision Floating-Point Values
3-325
DIVPS-Divide Packed Single Precision Floating-Point Values
3-328
DIVSD-Divide Scalar Double Precision Floating-Point Value
3-331
DIVSS-Divide Scalar Single Precision Floating-Point Values
3-333
DPPD-Dot Product of Packed Double Precision Floating-Point Values
3-335
DPPS-Dot Product of Packed Single Precision Floating-Point Values
3-337
EMMS-Empty MMX Technology State
3-340
ENCODEKEY128-Encode 128-Bit Key With Key Locker
3-341
ENCODEKEY256-Encode 256-Bit Key With Key Locker
3-343
ENDBR32-Terminate an Indirect Branch in 32-bit and Compatibility Mode
3-345
ENDBR64-Terminate an Indirect Branch in 64-bit Mode
3-346
ENTER-Make Stack Frame for Procedure Parameters
3-347
ENQCMD-Enqueue Command
3-350
ENQCMDS-Enqueue Command Supervisor
3-353
EXTRACTPS-Extract Packed Floating-Point Values
3-356
F2XM1-Compute 2x-1
3-358
FABS-Absolute Value
3-360
FADD/FADDP/FIADD-Add
3-361
FBLD-Load Binary Coded Decimal
3-364
FBSTP-Store BCD Integer and Pop
3-366
FCHS-Change Sign
3-368
FCLEX/FNCLEX-Clear Exceptions
3-370
FCMOVcc-Floating-Point Conditional Move
3-372
FCOM/FCOMP/FCOMPP-Compare Floating-Point Values
3-374
FCOMI/FCOMIP/ FUCOMI/FUCOMIP-Compare Floating-Point Values and Set EFLAGS
3-377
FCOS-Cosine
3-380
FDECSTP-Decrement Stack-Top Pointer
3-382
FDIV/FDIVP/FIDIV-Divide
3-383
FDIVR/FDIVRP/FIDIVR-Reverse Divide
3-386
FFREE-Free Floating-Point Register
3-389
FICOM/FICOMP-Compare Integer
3-390
FILD-Load Integer
3-392
FINCSTP-Increment Stack-Top Pointer
3-394
FINIT/FNINIT-Initialize Floating-Point Unit
3-395
FIST/FISTP-Store Integer
3-397
FISTTP-Store Integer With Truncation
3-400
FLD-Load Floating-Point Value
3-402
FLD1/FLDL2T/FLDL2E/FLDPI/FLDLG2/FLDLN2/FLDZ-Load Constant
3-404
FLDCW-Load x87 FPU Control Word
3-406
FLDENV-Load x87 FPU Environment
3-408
FMUL/FMULP/FIMUL-Multiply
3-410
FNOP-No Operation
3-413
FPATAN-Partial Arctangent
3-414
FPREM-Partial Remainder
3-416
FPREM1-Partial Remainder
3-418
FPTAN-Partial Tangent
3-420
FRNDINT-Round to Integer
3-422
FRSTOR-Restore x87 FPU State
3-423
FSAVE/FNSAVE-Store x87 FPU State
3-425
FSCALE-Scale
3-428
FSIN-Sine
3-430
FSINCOS-Sine and Cosine
3-432
FSQRT-Square Root
3-434
FST/FSTP-Store Floating-Point Value
3-436
Vol. 2A vii
CONTENTS
PAGE
FSTCW/FNSTCW-Store x87 FPU Control Word
3-438
FSTENV/FNSTENV-Store x87 FPU Environment
3-440
FSTSW/FNSTSW-Store x87 FPU Status Word
3-442
FSUB/FSUBP/FISUB-Subtract
3-444
FSUBR/FSUBRP/FISUBR-Reverse Subtract
3-447
FTST-TEST
3-450
FUCOM/FUCOMP/FUCOMPP-Unordered Compare Floating-Point Values
3-452
FXAM-Examine Floating-Point
3-454
FXCH-Exchange Register Contents
3-456
FXRSTOR-Restore x87 FPU, MMX, XMM, and MXCSR State
3-458
FXSAVE-Save x87 FPU, MMX Technology, and SSE State
3-461
FXTRACT-Extract Exponent and Significand
3-469
FYL2X-Compute y * log2x
3-471
FYL2XP1-Compute y * log2(x +1)
3-473
GF2P8AFFINEINVQB-Galois Field Affine Transformation Inverse
3-475
GF2P8AFFINEQB-Galois Field Affine Transformation
3-478
GF2P8MULB-Galois Field Multiply Bytes
3-480
HADDPD-Packed Double Precision Floating-Point Horizontal Add
3-482
HADDPS-Packed Single Precision Floating-Point Horizontal Add
3-485
HLT-Halt
3-488
HRESET-History Reset
3-489
HSUBPD-Packed Double Precision Floating-Point Horizontal Subtract
3-491
HSUBPS-Packed Single Precision Floating-Point Horizontal Subtract
3-494
IDIV-Signed Divide
3-497
IMUL-Signed Multiply
3-500
IN-Input From Port
3-504
INC-Increment by 1
3-506
INCSSPD/INCSSPQ-Increment Shadow Stack Pointer
3-508
INS/INSB/INSW/INSD-Input from Port to String
3-510
INSERTPS-Insert Scalar Single Precision Floating-Point Value
3-513
INT n/INTO/INT3/INT1-Call to Interrupt Procedure
3-516
INVD-Invalidate Internal Caches
3-531
INVLPG-Invalidate TLB Entries
3-533
INVPCID-Invalidate Process-Context Identifier
3-535
IRET/IRETD/IRETQ-Interrupt Return
3-538
Jcc-Jump if Condition Is Met
3-547
JMP-Jump
3-552
KADDW/KADDB/KADDQ/KADDD-ADD Two Masks
3-561
KANDW/KANDB/KANDQ/KANDD-Bitwise Logical AND Masks
3-563
KANDNW/KANDNB/KANDNQ/KANDND-Bitwise Logical AND NOT Masks
3-564
KMOVW/KMOVB/KMOVQ/KMOVD-Move From and to Mask Registers
3-565
KNOTW/KNOTB/KNOTQ/KNOTD-NOT Mask Register
3-567
KORW/KORB/KORQ/KORD-Bitwise Logical OR Masks
3-568
KORTESTW/KORTESTB/KORTESTQ/KORTESTD-OR Masks and Set Flags
3-569
KSHIFTLW/KSHIFTLB/KSHIFTLQ/KSHIFTLD-Shift Left Mask Registers
3-571
KSHIFTRW/KSHIFTRB/KSHIFTRQ/KSHIFTRD-Shift Right Mask Registers
3-573
KTESTW/KTESTB/KTESTQ/KTESTD-Packed Bit Test Masks and Set Flags
3-575
KUNPCKBW/KUNPCKWD/KUNPCKDQ-Unpack for Mask Registers
3-577
KXNORW/KXNORB/KXNORQ/KXNORD-Bitwise Logical XNOR Masks
3-578
KXORW/KXORB/KXORQ/KXORD-Bitwise Logical XOR Masks
3-579
LAHF-Load Status Flags Into AH Register
3-580
LAR-Load Access Rights Byte
3-581
LDDQU-Load Unaligned Integer 128 Bits
3-584
LDMXCSR-Load MXCSR Register
3-586
LDS/LES/LFS/LGS/LSS-Load Far Pointer
3-587
LDTILECFG-Load Tile Configuration
3-591
LEA-Load Effective Address
3-594
LEAVE-High Level Procedure Exit
3-596
LFENCE-Load Fence
3-598
viii Vol. 2A
CONTENTS
PAGE
LGDT/LIDT-Load Global/Interrupt Descriptor Table Register
3-599
LLDT-Load Local Descriptor Table Register
3-602
LMSW-Load Machine Status Word
3-604
LOADIWKEY-Load Internal Wrapping Key With Key Locker
3-606
LOCK-Assert LOCK# Signal Prefix
3-609
LODS/LODSB/LODSW/LODSD/LODSQ-Load String
3-611
LOOP/LOOPcc-Loop According to ECX Counter
3-614
LSL-Load Segment Limit
3-617
LTR-Load Task Register
3-620
LZCNT-Count the Number of Leading Zero Bits
3-622
CHAPTER 4
INSTRUCTION SET REFERENCE, M-U
4.1
IMM8 CONTROL BYTE OPERATION FOR PCMPESTRI / PCMPESTRM / PCMPISTRI / PCMPISTRM
4-1
4.1.1
General Description
4-1
4.1.2
Source Data Format
4-2
4.1.3
Aggregation Operation
4-2
4.1.4
Polarity
4-3
4.1.5
Output Selection
4-4
4.1.6
Valid/Invalid Override of Comparisons
4-4
4.1.7
Summary of Im8 Control byte
4-5
4.1.8
Diagram Comparison and Aggregation Process
4-6
4.2
COMMON TRANSFORMATION AND PRIMITIVE FUNCTIONS FOR SHA1XXX AND SHA256XXX
4-6
4.3
INSTRUCTIONS (M-U)
4-7
MASKMOVDQU-Store Selected Bytes of Double Quadword
4-8
MASKMOVQ-Store Selected Bytes of Quadword
4-10
MAXPD-Maximum of Packed Double Precision Floating-Point Values
4-12
MAXPS-Maximum of Packed Single Precision Floating-Point Values
4-15
MAXSD-Return Maximum Scalar Double Precision Floating-Point Value
4-18
MAXSS-Return Maximum Scalar Single Precision Floating-Point Value
4-20
MFENCE-Memory Fence
4-22
MINPD-Minimum of Packed Double Precision Floating-Point Values
4-23
MINPS-Minimum of Packed Single Precision Floating-Point Values
4-26
MINSD-Return Minimum Scalar Double Precision Floating-Point Value
4-29
MINSS-Return Minimum Scalar Single Precision Floating-Point Value
4-31
MONITOR-Set Up Monitor Address
4-33
MOV-Move
4-35
MOV-Move to/from Control Registers
4-39
MOV-Move to/from Debug Registers
4-42
MOVAPD-Move Aligned Packed Double Precision Floating-Point Values
4-44
MOVAPS-Move Aligned Packed Single Precision Floating-Point Values
4-48
MOVBE-Move Data After Swapping Bytes
4-52
MOVD/MOVQ-Move Doubleword/Move Quadword
4-55
MOVDDUP-Replicate Double Precision Floating-Point Values
4-58
MOVDIRI-Move Doubleword as Direct Store
4-61
MOVDIR64B-Move 64 Bytes as Direct Store
4-63
MOVDQA,VMOVDQA32/64-Move Aligned Packed Integer Values
4-65
MOVDQU,VMOVDQU8/16/32/64-Move Unaligned Packed Integer Values
4-70
MOVDQ2Q-Move Quadword from XMM to MMX Technology Register
4-77
MOVHLPS-Move Packed Single Precision Floating-Point Values High to Low
4-78
MOVHPD-Move High Packed Double Precision Floating-Point Value
4-80
MOVHPS-Move High Packed Single Precision Floating-Point Values
4-82
MOVLHPS-Move Packed Single Precision Floating-Point Values Low to High
4-84
MOVLPD-Move Low Packed Double Precision Floating-Point Value
4-86
MOVLPS-Move Low Packed Single Precision Floating-Point Values
4-88
MOVMSKPD-Extract Packed Double Precision Floating-Point Sign Mask
4-90
MOVMSKPS-Extract Packed Single Precision Floating-Point Sign Mask
4-92
MOVNTDQA-Load Double Quadword Non-Temporal Aligned Hint
4-94
MOVNTDQ-Store Packed Integers Using Non-Temporal Hint
4-96
MOVNTI-Store Doubleword Using Non-Temporal Hint
4-98
Vol. 2A ix
CONTENTS
PAGE
MOVNTPD-Store Packed Double Precision Floating-Point Values Using Non-Temporal Hint
4-100
MOVNTPS-Store Packed Single Precision Floating-Point Values Using Non-Temporal Hint
4-102
MOVNTQ-Store of Quadword Using Non-Temporal Hint
4-104
MOVQ-Move Quadword
4-105
MOVQ2DQ-Move Quadword from MMX Technology to XMM Register
4-108
MOVS/MOVSB/MOVSW/MOVSD/MOVSQ-Move Data From String to String
4-110
MOVSD-Move or Merge Scalar Double Precision Floating-Point Value
4-114
MOVSHDUP-Replicate Single Precision Floating-Point Values
4-117
MOVSLDUP-Replicate Single Precision Floating-Point Values
4-120
MOVSS-Move or Merge Scalar Single Precision Floating-Point Value
4-123
MOVSX/MOVSXD-Move With Sign-Extension
4-126
MOVUPD-Move Unaligned Packed Double Precision Floating-Point Values
4-128
MOVUPS-Move Unaligned Packed Single Precision Floating-Point Values
4-132
MOVZX-Move With Zero-Extend
4-136
MPSADBW-Compute Multiple Packed Sums of Absolute Difference
4-138
MUL-Unsigned Multiply
4-146
MULPD-Multiply Packed Double Precision Floating-Point Values
4-148
MULPS-Multiply Packed Single Precision Floating-Point Values
4-151
MULSD-Multiply Scalar Double Precision Floating-Point Value
4-154
MULSS-Multiply Scalar Single Precision Floating-Point Values
4-156
MULX-Unsigned Multiply Without Affecting Flags
4-158
MWAIT-Monitor Wait
4-160
NEG-Two's Complement Negation
4-163
NOP-No Operation
4-165
NOT-One's Complement Negation
4-166
OR-Logical Inclusive OR
4-168
ORPD-Bitwise Logical OR of Packed Double Precision Floating-Point Values
4-170
ORPS-Bitwise Logical OR of Packed Single Precision Floating-Point Values
4-173
OUT-Output to Port
4-176
OUTS/OUTSB/OUTSW/OUTSD-Output String to Port
4-178
PABSB/PABSW/PABSD/PABSQ-Packed Absolute Value
4-182
PACKSSWB/PACKSSDW-Pack With Signed Saturation
4-188
PACKUSDW-Pack With Unsigned Saturation
4-196
PACKUSWB-Pack With Unsigned Saturation
4-201
PADDB/PADDW/PADDD/PADDQ-Add Packed Integers
4-206
PADDSB/PADDSW-Add Packed Signed Integers with Signed Saturation
4-213
PADDUSB/PADDUSW-Add Packed Unsigned Integers With Unsigned Saturation
4-217
PALIGNR-Packed Align Right
4-221
PAND-Logical AND
4-224
PANDN-Logical AND NOT
4-227
PAUSE-Spin Loop Hint
4-230
PAVGB/PAVGW-Average Packed Integers
4-231
PBLENDVB-Variable Blend Packed Bytes
4-235
PBLENDW-Blend Packed Words
4-239
PCLMULQDQ-Carry-Less Multiplication Quadword
4-242
PCMPEQB/PCMPEQW/PCMPEQD- Compare Packed Data for Equal
4-245
PCMPEQQ-Compare Packed Qword Data for Equal
4-251
PCMPESTRI-Packed Compare Explicit Length Strings, Return Index
4-254
PCMPESTRM-Packed Compare Explicit Length Strings, Return Mask
4-256
PCMPGTB/PCMPGTW/PCMPGTD-Compare Packed Signed Integers for Greater Than
4-258
PCMPGTQ-Compare Packed Data for Greater Than
4-264
PCMPISTRI-Packed Compare Implicit Length Strings, Return Index
4-267
PCMPISTRM-Packed Compare Implicit Length Strings, Return Mask
4-269
PCONFIG-Platform Configuration
4-271
PDEP-Parallel Bits Deposit
4-277
PEXT-Parallel Bits Extract
4-279
PEXTRB/PEXTRD/PEXTRQ-Extract Byte/Dword/Qword
4-281
PEXTRW-Extract Word
4-284
PHADDW/PHADDD-Packed Horizontal Add
4-287
x Vol. 2A
CONTENTS
PAGE
PHADDSW-Packed Horizontal Add and Saturate
4-291
PHMINPOSUW-Packed Horizontal Word Minimum
4-293
PHSUBW/PHSUBD-Packed Horizontal Subtract
4-295
PHSUBSW-Packed Horizontal Subtract and Saturate
4-298
PINSRB/PINSRD/PINSRQ-Insert Byte/Dword/Qword
4-300
PINSRW-Insert Word
4-303
PMADDUBSW-Multiply and Add Packed Signed and Unsigned Bytes
4-305
PMADDWD-Multiply and Add Packed Integers
4-308
PMAXSB/PMAXSW/PMAXSD/PMAXSQ-Maximum of Packed Signed Integers
4-311
PMAXUB/PMAXUW-Maximum of Packed Unsigned Integers
4-318
PMAXUD/PMAXUQ-Maximum of Packed Unsigned Integers
4-323
PMINSB/PMINSW-Minimum of Packed Signed Integers
4-327
PMINSD/PMINSQ-Minimum of Packed Signed Integers
4-332
PMINUB/PMINUW-Minimum of Packed Unsigned Integers
4-336
PMINUD/PMINUQ-Minimum of Packed Unsigned Integers
4-341
PMOVMSKB-Move Byte Mask
4-345
PMOVSX-Packed Move With Sign Extend
4-347
PMOVZX-Packed Move With Zero Extend
4-356
PMULDQ-Multiply Packed Doubleword Integers
4-366
PMULHRSW-Packed Multiply High With Round and Scale
4-369
PMULHUW-Multiply Packed Unsigned Integers and Store High Result
4-373
PMULHW-Multiply Packed Signed Integers and Store High Result
4-377
PMULLD/PMULLQ-Multiply Packed Integers and Store Low Result
4-381
PMULLW-Multiply Packed Signed Integers and Store Low Result
4-385
PMULUDQ-Multiply Packed Unsigned Doubleword Integers
4-389
POP-Pop a Value From the Stack
4-392
POPA/POPAD-Pop All General-Purpose Registers
4-397
POPCNT-Return the Count of Number of Bits Set to 1
4-399
POPF/POPFD/POPFQ-Pop Stack Into EFLAGS Register
4-401
POR-Bitwise Logical OR
4-405
PREFETCHh-Prefetch Data Into Caches
4-408
PREFETCHW-Prefetch Data Into Caches in Anticipation of a Write
4-410
PSADBW-Compute Sum of Absolute Differences
4-412
PSHUFB-Packed Shuffle Bytes
4-416
PSHUFD-Shuffle Packed Doublewords
4-420
PSHUFHW-Shuffle Packed High Words
4-424
PSHUFLW-Shuffle Packed Low Words
4-427
PSHUFW-Shuffle Packed Words
4-430
PSIGNB/PSIGNW/PSIGND-Packed SIGN
4-431
PSLLDQ-Shift Double Quadword Left Logical
4-435
PSLLW/PSLLD/PSLLQ-Shift Packed Data Left Logical
4-437
PSRAW/PSRAD/PSRAQ-Shift Packed Data Right Arithmetic
4-449
PSRLDQ-Shift Double Quadword Right Logical
4-459
PSRLW/PSRLD/PSRLQ-Shift Packed Data Right Logical
4-461
PSUBB/PSUBW/PSUBD-Subtract Packed Integers
4-473
PSUBQ-Subtract Packed Quadword Integers
4-481
PSUBSB/PSUBSW-Subtract Packed Signed Integers With Signed Saturation
4-484
PSUBUSB/PSUBUSW-Subtract Packed Unsigned Integers With Unsigned Saturation
4-488
PTEST-Logical Compare
4-492
PTWRITE-Write Data to a Processor Trace Packet
4-494
PUNPCKHBW/PUNPCKHWD/PUNPCKHDQ/PUNPCKHQDQ- Unpack High Data
4-496
PUNPCKLBW/PUNPCKLWD/PUNPCKLDQ/PUNPCKLQDQ-Unpack Low Data
4-506
PUSH-Push Word, Doubleword, or Quadword Onto the Stack
4-516
PUSHA/PUSHAD-Push All General-Purpose Registers
4-519
PUSHF/PUSHFD/PUSHFQ-Push EFLAGS Register Onto the Stack
4-521
PXOR-Logical Exclusive OR
4-523
RCL/RCR/ROL/ROR-Rotate
4-526
RCPPS-Compute Reciprocals of Packed Single Precision Floating-Point Values
4-531
RCPSS-Compute Reciprocal of Scalar Single Precision Floating-Point Values
4-533
Vol. 2A xi
CONTENTS
PAGE
RDFSBASE/RDGSBASE-Read FS/GS Segment Base
4-535
RDMSR-Read From Model Specific Register
4-537
RDPID-Read Processor ID
4-539
RDPKRU-Read Protection Key Rights for User Pages
4-540
RDPMC-Read Performance-Monitoring Counters
4-542
RDRAND-Read Random Number
4-544
RDSEED-Read Random SEED
4-546
RDSSPD/RDSSPQ-Read Shadow Stack Pointer
4-548
RDTSC-Read Time-Stamp Counter
4-550
RDTSCP-Read Time-Stamp Counter and Processor ID
4-552
REP/REPE/REPZ/REPNE/REPNZ-Repeat String Operation Prefix
4-554
RET-Return From Procedure
4-558
RORX - Rotate Right Logical Without Affecting Flags
4-571
ROUNDPD-Round Packed Double Precision Floating-Point Values
4-572
ROUNDPS-Round Packed Single Precision Floating-Point Values
4-575
ROUNDSD-Round Scalar Double Precision Floating-Point Values
4-577
ROUNDSS-Round Scalar Single Precision Floating-Point Values
4-579
RSM-Resume From System Management Mode
4-581
RSQRTPS-Compute Reciprocals of Square Roots of Packed Single Precision Floating-Point Values
4-583
RSQRTSS-Compute Reciprocal of Square Root of Scalar Single Precision Floating-Point Value
4-585
RSTORSSP-Restore Saved Shadow Stack Pointer
4-587
SAHF-Store AH Into Flags
4-590
SAL/SAR/SHL/SHR-Shift
4-592
SARX/SHLX/SHRX-Shift Without Affecting Flags
4-597
SAVEPREVSSP-Save Previous Shadow Stack Pointer
4-599
SBB-Integer Subtraction With Borrow
4-601
SCAS/SCASB/SCASW/SCASD-Scan String
4-604
SENDUIPI-Send User Interprocessor Interrupt
4-608
SERIALIZE-Serialize Instruction Execution
4-610
SETcc-Set Byte on Condition
4-611
SETSSBSY-Mark Shadow Stack Busy
4-614
SFENCE-Store Fence
4-616
SGDT-Store Global Descriptor Table Register
4-617
SHA1RNDS4-Perform Four Rounds of SHA1 Operation
4-619
SHA1NEXTE-Calculate SHA1 State Variable E After Four Rounds
4-621
SHA1MSG1-Perform an Intermediate Calculation for the Next Four SHA1 Message Dwords
4-622
SHA1MSG2-Perform a Final Calculation for the Next Four SHA1 Message Dwords
4-623
SHA256RNDS2-Perform Two Rounds of SHA256 Operation
4-624
SHA256MSG1-Perform an Intermediate Calculation for the Next Four SHA256 Message Dwords
4-626
SHA256MSG2-Perform a Final Calculation for the Next Four SHA256 Message Dwords
4-627
SHLD-Double Precision Shift Left
4-628
SHRD-Double Precision Shift Right
4-631
SHUFPD-Packed Interleave Shuffle of Pairs of Double Precision Floating-Point Values
4-634
SHUFPS-Packed Interleave Shuffle of Quadruplets of Single Precision Floating-Point Values
4-639
SIDT-Store Interrupt Descriptor Table Register
4-643
SLDT-Store Local Descriptor Table Register
4-645
SMSW-Store Machine Status Word
4-647
SQRTPD-Square Root of Double Precision Floating-Point Values
4-649
SQRTPS-Square Root of Single Precision Floating-Point Values
4-652
SQRTSD-Compute Square Root of Scalar Double Precision Floating-Point Value
4-655
SQRTSS-Compute Square Root of Scalar Single Precision Value
4-657
STAC-Set AC Flag in EFLAGS Register
4-659
STC-Set Carry Flag
4-660
STD-Set Direction Flag
4-661
STI-Set Interrupt Flag
4-662
STMXCSR-Store MXCSR Register State
4-664
STOS/STOSB/STOSW/STOSD/STOSQ-Store String
4-665
STR-Store Task Register
4-668
STTILECFG-Store Tile Configuration
4-670
xii Vol. 2A
CONTENTS
PAGE
STUI-Set User Interrupt Flag
4-672
SUB-Subtract
4-673
SUBPD-Subtract Packed Double Precision Floating-Point Values
4-675
SUBPS-Subtract Packed Single Precision Floating-Point Values
4-678
SUBSD-Subtract Scalar Double Precision Floating-Point Value
4-681
SUBSS-Subtract Scalar Single Precision Floating-Point Value
4-683
SWAPGS-Swap GS Base Register
4-685
SYSCALL-Fast System Call
4-687
SYSENTER-Fast System Call
4-690
SYSEXIT-Fast Return from Fast System Call
4-693
SYSRET-Return From Fast System Call
4-696
TDPBF16PS-Dot Product of BF16 Tiles Accumulated into Packed Single Precision Tile
4-699
TDPBSSD/TDPBSUD/TDPBUSD/TDPBUUD-Dot Product of Signed/Unsigned Bytes with Dword Accumulation
4-701
TEST-Logical Compare
4-703
TESTUI-Determine User Interrupt Flag
4-705
TILELOADD/TILELOADDT1-Load Tile
4-706
TILERELEASE-Release Tile
4-708
TILESTORED-Store Tile
4-709
TILEZERO-Zero Tile
4-710
TPAUSE-Timed PAUSE
4-711
TZCNT-Count the Number of Trailing Zero Bits
4-713
UCOMISD-Unordered Compare Scalar Double Precision Floating-Point Values and Set EFLAGS
4-715
UCOMISS-Unordered Compare Scalar Single Precision Floating-Point Values and Set EFLAGS
4-717
UD-Undefined Instruction
4-719
UIRET-User-Interrupt Return
4-720
UMONITOR-User Level Set Up Monitor Address
4-722
UMWAIT-User Level Monitor Wait
4-724
UNPCKHPD-Unpack and Interleave High Packed Double Precision Floating-Point Values
4-726
UNPCKHPS-Unpack and Interleave High Packed Single Precision Floating-Point Values
4-730
UNPCKLPD-Unpack and Interleave Low Packed Double Precision Floating-Point Values
4-734
UNPCKLPS-Unpack and Interleave Low Packed Single Precision Floating-Point Values
4-738
CHAPTER 5
INSTRUCTION SET REFERENCE, V
5.1
TERNARY BIT VECTOR LOGIC TABLE
5-1
5.2
INSTRUCTIONS (V)
5-4
VADDPH-Add Packed FP16 Values
5-5
VADDSH-Add Scalar FP16 Values
5-7
VALIGND/VALIGNQ-Align Doubleword/Quadword Vectors
5-8
VBLENDMPD/VBLENDMPS-Blend Float64/Float32 Vectors Using an OpMask Control
5-11
VBROADCAST-Load with Broadcast Floating-Point Data
5-13
VCMPPH-Compare Packed FP16 Values
5-21
VCMPSH-Compare Scalar FP16 Values
5-23
VCOMISH-Compare Scalar Ordered FP16 Values and Set EFLAGS
5-25
VCOMPRESSPD-Store Sparse Packed Double Precision Floating-Point Values Into Dense Memory
5-27
VCOMPRESSPS-Store Sparse Packed Single Precision Floating-Point Values Into Dense Memory
5-29
VCVTDQ2PH-Convert Packed Signed Doubleword Integers to Packed FP16 Values
5-31
VCVTNE2PS2BF16-Convert Two Packed Single Data to One Packed BF16 Data
5-33
VCVTNEPS2BF16-Convert Packed Single Data to Packed BF16 Data
5-35
VCVTPD2PH-Convert Packed Double Precision FP Values to Packed FP16 Values
5-37
VCVTPD2QQ-Convert Packed Double Precision Floating-Point Values to Packed Quadword Integers
5-39
VCVTPD2UDQ-Convert Packed Double Precision Floating-Point Values to Packed Unsigned Doubleword Integers. . 5-41
VCVTPD2UQQ-Convert Packed Double Precision Floating-Point Values to Packed Unsigned Quadword Integers . . . 5-43
VCVTPH2DQ-Convert Packed FP16 Values to Signed Doubleword Integers
5-45
VCVTPH2PD-Convert Packed FP16 Values to FP64 Values
5-47
VCVTPH2PS/VCVTPH2PSX-Convert Packed FP16 Values to Single Precision Floating-Point Values
5-49
VCVTPH2QQ-Convert Packed FP16 Values to Signed Quadword Integer Values
5-53
VCVTPH2UDQ-Convert Packed FP16 Values to Unsigned Doubleword Integers
5-55
VCVTPH2UQQ-Convert Packed FP16 Values to Unsigned Quadword Integers
5-57
Vol. 2A xiii
CONTENTS
PAGE
VCVTPH2UW-Convert Packed FP16 Values to Unsigned Word Integers
5-59
VCVTPH2W-Convert Packed FP16 Values to Signed Word Integers
5-61
VCVTPS2PH-Convert Single-Precision FP Value to 16-bit FP Value
5-63
VCVTPS2PHX-Convert Packed Single Precision Floating-Point Values to Packed FP16 Values
5-67
VCVTPS2QQ-Convert Packed Single Precision Floating-Point Values to Packed Signed Quadword Integer Values . . 5-69
VCVTPS2UDQ-Convert Packed Single Precision Floating-Point Values to Packed Unsigned Doubleword Integer
Values
5-71
VCVTPS2UQQ-Convert Packed Single Precision Floating-Point Values to Packed Unsigned Quadword Integer
Values
5-74
VCVTQQ2PD-Convert Packed Quadword Integers to Packed Double Precision Floating-Point Values
5-76
VCVTQQ2PH-Convert Packed Signed Quadword Integers to Packed FP16 Values
5-78
VCVTQQ2PS-Convert Packed Quadword Integers to Packed Single Precision Floating-Point Values
5-80
VCVTSD2SH-Convert Low FP64 Value to an FP16 Value
5-82
VCVTSD2USI-Convert Scalar Double Precision Floating-Point Value to Unsigned Doubleword Integer
5-83
VCVTSH2SD-Convert Low FP16 Value to an FP64 Value
5-85
VCVTSH2SI-Convert Low FP16 Value to Signed Integer
5-86
VCVTSH2SS-Convert Low FP16 Value to FP32 Value
5-87
VCVTSH2USI-Convert Low FP16 Value to Unsigned Integer
5-88
VCVTSI2SH-Convert a Signed Doubleword/Quadword Integer to an FP16 Value
5-89
VCVTSS2SH-Convert Low FP32 Value to an FP16 Value
5-91
VCVTSS2USI-Convert Scalar Single Precision Floating-Point Value to Unsigned Doubleword Integer
5-92
VCVTTPD2QQ-Convert With Truncation Packed Double Precision Floating-Point Values to Packed Quadword
Integers
5-94
VCVTTPD2UDQ-Convert With Truncation Packed Double Precision Floating-Point Values to Packed Unsigned
Doubleword Integers
5-96
VCVTTPD2UQQ-Convert With Truncation Packed Double Precision Floating-Point Values to Packed Unsigned
Quadword Integers
5-98
VCVTTPH2DQ-Convert with Truncation Packed FP16 Values to Signed Doubleword Integers
5-100
VCVTTPH2QQ-Convert with Truncation Packed FP16 Values to Signed Quadword Integers
5-102
VCVTTPH2UDQ-Convert with Truncation Packed FP16 Values to Unsigned Doubleword Integers
5-104
VCVTTPH2UQQ-Convert with Truncation Packed FP16 Values to Unsigned Quadword Integers
5-106
VCVTTPH2UW-Convert Packed FP16 Values to Unsigned Word Integers
5-108
VCVTTPH2W-Convert Packed FP16 Values to Signed Word Integers
5-110
VCVTTPS2QQ-Convert With Truncation Packed Single Precision Floating-Point Values to Packed Signed Quadword
Integer Values
5-112
VCVTTPS2UDQ-Convert With Truncation Packed Single Precision Floating-Point Values to Packed Unsigned
Doubleword Integer Values
5-114
VCVTTPS2UQQ-Convert With Truncation Packed Single Precision Floating-Point Values to Packed Unsigned
Quadword Integer Values
5-116
VCVTTSD2USI-Convert With Truncation Scalar Double Precision Floating-Point Value to Unsigned Integer
5-118
VCVTTSH2SI-Convert with Truncation Low FP16 Value to a Signed Integer
5-119
VCVTTSH2USI-Convert with Truncation Low FP16 Value to an Unsigned Integer
5-120
VCVTTSS2USI-Convert With Truncation Scalar Single Precision Floating-Point Value to Unsigned Integer
5-121
VCVTUDQ2PD-Convert Packed Unsigned Doubleword Integers to Packed Double Precision Floating-Point Values. 5-122
VCVTUDQ2PH-Convert Packed Unsigned Doubleword Integers to Packed FP16 Values
5-124
VCVTUDQ2PS-Convert Packed Unsigned Doubleword Integers to Packed Single Precision Floating-Point Values . . 5-126
VCVTUQQ2PD-Convert Packed Unsigned Quadword Integers to Packed Double Precision Floating-Point Values . . 5-128
VCVTUQQ2PH-Convert Packed Unsigned Quadword Integers to Packed FP16 Values
5-130
VCVTUSI2SH-Convert Unsigned Doubleword Integer to an FP16 Value
5-132
VCVTUQQ2PS-Convert Packed Unsigned Quadword Integers to Packed Single Precision Floating-Point Values . . . 5-134
VCVTUSI2SD-Convert Unsigned Integer to Scalar Double Precision Floating-Point Value
5-136
VCVTUSI2SS-Convert Unsigned Integer to Scalar Single Precision Floating-Point Value
5-138
VCVTUW2PH-Convert Packed Unsigned Word Integers to FP16 Values
5-140
VCVTW2PH-Convert Packed Signed Word Integers to FP16 Values
5-142
VDBPSADBW-Double Block Packed Sum-Absolute-Differences (SAD) on Unsigned Bytes
5-144
VDIVPH-Divide Packed FP16 Values
5-147
VDIVSH-Divide Scalar FP16 Values
5-149
VDPBF16PS-Dot Product of BF16 Pairs Accumulated Into Packed Single Precision
5-150
VEXPANDPD-Load Sparse Packed Double Precision Floating-Point Values From Dense Memory
5-152
xiv Vol. 2A
CONTENTS
PAGE
VEXPANDPS-Load Sparse Packed Single Precision Floating-Point Values From Dense Memory
5-154
VERR/VERW-Verify a Segment for Reading or Writing
5-156
VEXTRACTF128/VEXTRACTF32x4/VEXTRACTF64x2/VEXTRACTF32x8/VEXTRACTF64x4- Extract Packed
Floating-Point Values
5-158
VEXTRACTI128/VEXTRACTI32x4/VEXTRACTI64x2/VEXTRACTI32x8/VEXTRACTI64x4-Extract Packed Integer
Values
5-164
VFCMADDCPH/VFMADDCPH-Complex Multiply and Accumulate FP16 Values
5-170
VFCMADDCSH/VFMADDCSH-Complex Multiply and Accumulate Scalar FP16 Values
5-173
VFCMULCPH/VFMULCPH-Complex Multiply FP16 Values
5-175
VFCMULCSH/VFMULCSH-Complex Multiply Scalar FP16 Values
5-178
VFIXUPIMMPD-Fix Up Special Packed Float64 Values
5-180
VFIXUPIMMPS-Fix Up Special Packed Float32 Values
5-184
VFIXUPIMMSD-Fix Up Special Scalar Float64 Value
5-188
VFIXUPIMMSS-Fix Up Special Scalar Float32 Value
5-191
VFMADD132PD/VFMADD213PD/VFMADD231PD-Fused Multiply-Add of Packed Double Precision Floating-Point
Values
5-194
VF[,N]MADD[132,213,231]PH-Fused Multiply-Add of Packed FP16 Values
5-201
VFMADD132PS/VFMADD213PS/VFMADD231PS-Fused Multiply-Add of Packed Single Precision Floating-Point
Values
5-207
VFMADD132SD/VFMADD213SD/VFMADD231SD-Fused Multiply-Add of Scalar Double Precision Floating-Point
Values
5-213
VF[,N]MADD[132,213,231]SH-Fused Multiply-Add of Scalar FP16 Values
5-216
VFMADD132SS/VFMADD213SS/VFMADD231SS-Fused Multiply-Add of Scalar Single Precision Floating-Point
Values
5-219
VFMADDSUB132PD/VFMADDSUB213PD/VFMADDSUB231PD-Fused Multiply-Alternating Add/Subtract of Packed
Double Precision Floating-Point Values
5-222
VFMADDSUB132PH/VFMADDSUB213PH/VFMADDSUB231PH-Fused Multiply-Alternating Add/Subtract of Packed
FP16 Values
5-229
VFMADDSUB132PS/VFMADDSUB213PS/VFMADDSUB231PS-Fused Multiply-Alternating Add/Subtract of Packed
Single Precision Floating-Point Values
5-234
VFMSUB132PD/VFMSUB213PD/VFMSUB231PD-Fused Multiply-Subtract of Packed Double Precision Floating-Point
Values
5-241
VF[,N]MSUB[132,213,231]PH-Fused Multiply-Subtract of Packed FP16 Values
5-247
VFMSUB132PS/VFMSUB213PS/VFMSUB231PS-Fused Multiply-Subtract of Packed Single Precision Floating-Point
Values
5-253
VFMSUB132SD/VFMSUB213SD/VFMSUB231SD-Fused Multiply-Subtract of Scalar Double Precision Floating-Point
Values
5-259
VF[,N]MSUB[132,213,231]SH-Fused Multiply-Subtract of Scalar FP16 Values
5-262
VFMSUB132SS/VFMSUB213SS/VFMSUB231SS-Fused Multiply-Subtract of Scalar Single Precision Floating-Point
Values
5-265
VFMSUBADD132PD/VFMSUBADD213PD/VFMSUBADD231PD-Fused Multiply-Alternating Subtract/Add of Packed
Double Precision Floating-Point Values
5-268
VFMSUBADD132PH/VFMSUBADD213PH/VFMSUBADD231PH-Fused Multiply-Alternating Subtract/Add of Packed
FP16 Values
5-275
VFMSUBADD132PS/VFMSUBADD213PS/VFMSUBADD231PS-Fused Multiply-Alternating Subtract/Add of Packed
Single Precision Floating-Point Values
5-280
VFNMADD132PD/VFNMADD213PD/VFNMADD231PD-Fused Negative Multiply-Add of Packed Double Precision
Floating-Point Values
5-287
VFNMADD132PS/VFNMADD213PS/VFNMADD231PS-Fused Negative Multiply-Add of Packed Single Precision
Floating-Point Values
5-294
VFNMADD132SD/VFNMADD213SD/VFNMADD231SD-Fused Negative Multiply-Add of Scalar Double Precision
Floating-Point Values
5-300
VFNMADD132SS/VFNMADD213SS/VFNMADD231SS-Fused Negative Multiply-Add of Scalar Single Precision
Floating-Point Values
5-303
VFNMSUB132PD/VFNMSUB213PD/VFNMSUB231PD-Fused Negative Multiply-Subtract of Packed Double Precision
Floating-Point Values
5-306
VFNMSUB132PS/VFNMSUB213PS/VFNMSUB231PS-Fused Negative Multiply-Subtract of Packed Single Precision
Floating-Point Values
5-313
VFNMSUB132SD/VFNMSUB213SD/VFNMSUB231SD-Fused Negative Multiply-Subtract of Scalar Double Precision
Vol. 2A xv
CONTENTS
PAGE
Floating-Point Values
5-319
VFNMSUB132SS/VFNMSUB213SS/VFNMSUB231SS-Fused Negative Multiply-Subtract of Scalar Single Precision
Floating-Point Values
5-322
VFPCLASSPD-Tests Types of Packed Float64 Values
5-325
VFPCLASSPH-Test Types of Packed FP16 Values
5-328
VFPCLASSPS-Tests Types of Packed Float32 Values
5-331
VFPCLASSSD-Tests Type of a Scalar Float64 Value
5-333
VFPCLASSSH-Test Types of Scalar FP16 Values
5-335
VFPCLASSSS-Tests Type of a Scalar Float32 Value
5-336
VGATHERDPD/VGATHERQPD-Gather Packed Double Precision Floating-Point Values Using Signed Dword/Qword
Indices
5-338
VGATHERDPS/VGATHERQPS-Gather Packed Single Precision Floating-Point Values Using Signed Dword/Qword
Indices
5-342
VGATHERDPS/VGATHERDPD-Gather Packed Single, Packed Double with Signed Dword Indices
5-346
VGATHERQPS/VGATHERQPD-Gather Packed Single, Packed Double with Signed Qword Indices
5-349
VGETEXPPD-Convert Exponents of Packed Double Precision Floating-Point Values to Double Precision Floating-Point
Values
5-352
VGETEXPPH-Convert Exponents of Packed FP16 Values to FP16 Values
5-355
VGETEXPPS-Convert Exponents of Packed Single Precision Floating-Point Values to Single Precision Floating-Point
Values
5-358
VGETEXPSD-Convert Exponents of Scalar Double Precision Floating-Point Value to Double Precision Floating-Point
Value
5-362
VGETEXPSH-Convert Exponents of Scalar FP16 Values to FP16 Values
5-364
VGETEXPSS-Convert Exponents of Scalar Single Precision Floating-Point Value to Single Precision Floating-Point
Value
5-366
VGETMANTPD-Extract Float64 Vector of Normalized Mantissas From Float64 Vector
5-368
VGETMANTPH-Extract FP16 Vector of Normalized Mantissas from FP16 Vector
5-372
VGETMANTPS-Extract Float32 Vector of Normalized Mantissas From Float32 Vector
5-376
VGETMANTSD-Extract Float64 of Normalized Mantissas From Float64 Scalar
5-379
VGETMANTSH-Extract FP16 of Normalized Mantissa from FP16 Scalar
5-381
VGETMANTSS-Extract Float32 Vector of Normalized Mantissa From Float32 Vector
5-383
VINSERTF128/VINSERTF32x4/VINSERTF64x2/VINSERTF32x8/VINSERTF64x4-Insert Packed Floating-Point
Values
5-385
VINSERTI128/VINSERTI32x4/VINSERTI64x2/VINSERTI32x8/VINSERTI64x4-Insert Packed Integer Values
5-389
VMASKMOV-Conditional SIMD Packed Loads and Stores
5-393
VMAXPH-Return Maximum of Packed FP16 Values
5-396
VMAXSH-Return Maximum of Scalar FP16 Values
5-398
VMINPH-Return Minimum of Packed FP16 Values
5-400
VMINSH-Return Minimum Scalar FP16 Value
5-402
VMOVSH-Move Scalar FP16 Value
5-404
VMOVW-Move Word
5-406
VMULPH-Multiply Packed FP16 Values
5-407
VMULSH-Multiply Scalar FP16 Values
5-409
VP2INTERSECTD/VP2INTERSECTQ-Compute Intersection Between DWORDS/QUADWORDS to a Pair of Mask
Registers
5-410
VPBLENDD-Blend Packed Dwords
5-412
VPBLENDMB/VPBLENDMW-Blend Byte/Word Vectors Using an Opmask Control
5-414
VPBLENDMD/VPBLENDMQ-Blend Int32/Int64 Vectors Using an OpMask Control
5-416
VPBROADCASTB/W/D/Q-Load With Broadcast Integer Data From General Purpose Register
5-419
VPBROADCAST-Load Integer and Broadcast
5-422
VPBROADCASTM-Broadcast Mask to Vector Register
5-431
VPCMPB/VPCMPUB-Compare Packed Byte Values Into Mask
5-433
VPCMPD/VPCMPUD-Compare Packed Integer Values Into Mask
5-436
VPCMPQ/VPCMPUQ-Compare Packed Integer Values Into Mask
5-439
VPCMPW/VPCMPUW-Compare Packed Word Values Into Mask
5-442
VPCOMPRESSB/VCOMPRESSW-Store Sparse Packed Byte/Word Integer Values Into Dense Memory/Register
5-445
VPCOMPRESSD-Store Sparse Packed Doubleword Integer Values Into Dense Memory/Register
5-448
VPCOMPRESSQ-Store Sparse Packed Quadword Integer Values Into Dense Memory/Register
5-450
VPCONFLICTD/Q-Detect Conflicts Within a Vector of Packed Dword/Qword Values Into Dense Memory/ Register . 5-452
xvi Vol. 2A
CONTENTS
PAGE
VPDPBUSD-Multiply and Add Unsigned and Signed Bytes
5-455
VPDPBUSDS-Multiply and Add Unsigned and Signed Bytes With Saturation
5-457
VPDPWSSD-Multiply and Add Signed Word Integers
5-459
VPDPWSSDS-Multiply and Add Signed Word Integers With Saturation
5-461
VPERM2F128-Permute Floating-Point Values
5-463
VPERM2I128-Permute Integer Values
5-465
VPERMB-Permute Packed Bytes Elements
5-467
VPERMD/VPERMW-Permute Packed Doubleword/Word Elements
5-469
VPERMI2B-Full Permute of Bytes From Two Tables Overwriting the Index
5-472
VPERMI2W/D/Q/PS/PD-Full Permute From Two Tables Overwriting the Index
5-474
VPERMILPD-Permute In-Lane of Pairs of Double Precision Floating-Point Values
5-480
VPERMILPS-Permute In-Lane of Quadruples of Single Precision Floating-Point Values
5-485
VPERMPD-Permute Double Precision Floating-Point Elements
5-490
VPERMPS-Permute Single Precision Floating-Point Elements
5-493
VPERMQ-Qwords Element Permutation
5-496
VPERMT2B-Full Permute of Bytes From Two Tables Overwriting a Table
5-499
VPERMT2W/D/Q/PS/PD-Full Permute From Two Tables Overwriting One Table
5-501
VPEXPANDB/VPEXPANDW-Expand Byte/Word Values
5-506
VPEXPANDD-Load Sparse Packed Doubleword Integer Values From Dense Memory/Register
5-509
VPEXPANDQ-Load Sparse Packed Quadword Integer Values From Dense Memory/Register
5-511
VPGATHERDD/VPGATHERQD-Gather Packed Dword Values Using Signed Dword/Qword Indices
5-513
VPGATHERDD/VPGATHERDQ-Gather Packed Dword, Packed Qword With Signed Dword Indices
5-517
VPGATHERDQ/VPGATHERQQ-Gather Packed Qword Values Using Signed Dword/Qword Indices
5-520
VPGATHERQD/VPGATHERQQ-Gather Packed Dword, Packed Qword with Signed Qword Indices
5-524
VPLZCNTD/Q-Count the Number of Leading Zero Bits for Packed Dword, Packed Qword Values
5-527
VPMADD52HUQ-Packed Multiply of Unsigned 52-Bit Unsigned Integers and Add High 52-Bit Products to 64-Bit
Accumulators
5-530
VPMADD52LUQ-Packed Multiply of Unsigned 52-Bit Integers and Add the Low 52-Bit Products to Qword
Accumulators
5-532
VPMASKMOV-Conditional SIMD Integer Packed Loads and Stores
5-534
VPMOVB2M/VPMOVW2M/VPMOVD2M/VPMOVQ2M-Convert a Vector Register to a Mask
5-537
VPMOVDB/VPMOVSDB/VPMOVUSDB-Down Convert DWord to Byte
5-540
VPMOVDW/VPMOVSDW/VPMOVUSDW-Down Convert DWord to Word
5-544
VPMOVM2B/VPMOVM2W/VPMOVM2D/VPMOVM2Q-Convert a Mask Register to a Vector Register
5-548
VPMOVQB/VPMOVSQB/VPMOVUSQB-Down Convert QWord to Byte
5-551
VPMOVQD/VPMOVSQD/VPMOVUSQD-Down Convert QWord to DWord
5-555
VPMOVQW/VPMOVSQW/VPMOVUSQW-Down Convert QWord to Word
5-559
VPMOVWB/VPMOVSWB/VPMOVUSWB-Down Convert Word to Byte
5-563
VPMULTISHIFTQB-Select Packed Unaligned Bytes From Quadword Sources
5-567
VPOPCNT-Return the Count of Number of Bits Set to 1 in BYTE/WORD/DWORD/QWORD
5-569
VPROLD/VPROLVD/VPROLQ/VPROLVQ-Bit Rotate Left
5-572
VPRORD/VPRORVD/VPRORQ/VPRORVQ-Bit Rotate Right
5-576
VPSCATTERDD/VPSCATTERDQ/VPSCATTERQD/VPSCATTERQQ-Scatter Packed Dword, Packed Qword with Signed
Dword, Signed Qword Indices
5-580
VPSHLD-Concatenate and Shift Packed Data Left Logical
5-584
VPSHLDV-Concatenate and Variable Shift Packed Data Left Logical
5-587
VPSHRD-Concatenate and Shift Packed Data Right Logical
5-590
VPSHRDV-Concatenate and Variable Shift Packed Data Right Logical
5-593
VPSHUFBITQMB-Shuffle Bits From Quadword Elements Using Byte Indexes Into Mask
5-596
VPSLLVW/VPSLLVD/VPSLLVQ-Variable Bit Shift Left Logical
5-597
VPSRAVW/VPSRAVD/VPSRAVQ-Variable Bit Shift Right Arithmetic
5-602
VPSRLVW/VPSRLVD/VPSRLVQ-Variable Bit Shift Right Logical
5-607
VPTERNLOGD/VPTERNLOGQ-Bitwise Ternary Logic
5-612
VPTESTMB/VPTESTMW/VPTESTMD/VPTESTMQ-Logical AND and Set Mask
5-615
VPTESTNMB/W/D/Q-Logical NAND and Set
5-618
VRANGEPD-Range Restriction Calculation for Packed Pairs of Float64 Values
5-621
VRANGEPS-Range Restriction Calculation for Packed Pairs of Float32 Values
5-625
VRANGESD-Range Restriction Calculation From a Pair of Scalar Float64 Values
5-628
VRANGESS-Range Restriction Calculation From a Pair of Scalar Float32 Values
5-631
Vol. 2A xvii
CONTENTS
PAGE
VRCP14PD-Compute Approximate Reciprocals of Packed Float64 Values
5-634
VRCP14SD-Compute Approximate Reciprocal of Scalar Float64 Value
5-636
VRCP14PS-Compute Approximate Reciprocals of Packed Float32 Values
5-638
VRCP14SS-Compute Approximate Reciprocal of Scalar Float32 Value
5-640
VRCPPH-Compute Reciprocals of Packed FP16 Values
5-642
VRCPSH-Compute Reciprocal of Scalar FP16 Value
5-644
VREDUCEPD-Perform Reduction Transformation on Packed Float64 Values
5-645
VREDUCEPH-Perform Reduction Transformation on Packed FP16 Values
5-648
VREDUCEPS-Perform Reduction Transformation on Packed Float32 Values
5-651
VREDUCESD-Perform a Reduction Transformation on a Scalar Float64 Value
5-653
VREDUCESH-Perform Reduction Transformation on Scalar FP16 Value
5-655
VREDUCESS-Perform a Reduction Transformation on a Scalar Float32 Value
5-657
VRNDSCALEPD-Round Packed Float64 Values to Include a Given Number of Fraction Bits
5-659
VRNDSCALEPH-Round Packed FP16 Values to Include a Given Number of Fraction Bits
5-662
VRNDSCALEPS-Round Packed Float32 Values to Include a Given Number of Fraction Bits
5-665
VRNDSCALESD-Round Scalar Float64 Value to Include a Given Number of Fraction Bits
5-668
VRNDSCALESH-Round Scalar FP16 Value to Include a Given Number of Fraction Bits
5-670
VRNDSCALESS-Round Scalar Float32 Value to Include a Given Number of Fraction Bits
5-672
VRSQRT14PD-Compute Approximate Reciprocals of Square Roots of Packed Float64 Values
5-674
VRSQRT14SD-Compute Approximate Reciprocal of Square Root of Scalar Float64 Value
5-676
VRSQRT14PS-Compute Approximate Reciprocals of Square Roots of Packed Float32 Values
5-678
VRSQRT14SS-Compute Approximate Reciprocal of Square Root of Scalar Float32 Value
5-680
VRSQRTPH-Compute Reciprocals of Square Roots of Packed FP16 Values
5-682
VRSQRTSH-Compute Approximate Reciprocal of Square Root of Scalar FP16 Value
5-684
VSCALEFPD-Scale Packed Float64 Values With Float64 Values
5-685
VSCALEFPH-Scale Packed FP16 Values with FP16 Values
5-688
VSCALEFPS-Scale Packed Float32 Values With Float32 Values
5-690
VSCALEFSD-Scale Scalar Float64 Values With Float64 Values
5-693
VSCALEFSH-Scale Scalar FP16 Values with FP16 Values
5-695
VSCALEFSS-Scale Scalar Float32 Value With Float32 Value
5-697
VSCATTERDPS/VSCATTERDPD/VSCATTERQPS/VSCATTERQPD-Scatter Packed Single, Packed Double with Signed
Dword and Qword Indices
5-699
VSHUFF32x4/VSHUFF64x2/VSHUFI32x4/VSHUFI64x2-Shuffle Packed Values at 128-Bit Granularity
5-703
VSQRTPH-Compute Square Root of Packed FP16 Values
5-708
VSQRTSH-Compute Square Root of Scalar FP16 Value
5-710
VSUBPH-Subtract Packed FP16 Values
5-711
VSUBSH-Subtract Scalar FP16 Value
5-713
VTESTPD/VTESTPS-Packed Bit Test
5-714
VUCOMISH-Unordered Compare Scalar FP16 Values and Set EFLAGS
5-717
VZEROALL-Zero XMM, YMM, and ZMM Registers
5-718
VZEROUPPER-Zero Upper Bits of YMM and ZMM Registers
5-719
CHAPTER 6
INSTRUCTION SET REFERENCE, W-Z
INSTRUCTIONS (W-Z)
6.1
6-1
WAIT/FWAIT-Wait
6-2
WBINVD-Write Back and Invalidate Cache
6-3
WBNOINVD-Write Back and Do Not Invalidate Cache
6-5
WRFSBASE/WRGSBASE-Write FS/GS Segment Base
6-7
WRMSR-Write to Model Specific Register
6-9
WRPKRU-Write Data to User Page Key Register
6-11
WRSSD/WRSSQ-Write to Shadow Stack
6-13
WRUSSD/WRUSSQ-Write to User Shadow Stack
6-15
XABORT-Transactional Abort
6-17
XACQUIRE/XRELEASE-Hardware Lock Elision Prefix Hints
6-19
XADD-Exchange and Add
6-23
XBEGIN-Transactional Begin
6-25
XCHG-Exchange Register/Memory With Register
6-28
XEND-Transactional End
6-30
xviii Vol. 2A
CONTENTS
PAGE
XGETBV-Get Value of Extended Control Register
6-32
XLAT/XLATB-Table Look-up Translation
6-34
XOR-Logical Exclusive OR
6-36
XORPD-Bitwise Logical XOR of Packed Double Precision Floating-Point Values
6-38
XORPS-Bitwise Logical XOR of Packed Single Precision Floating-Point Values
6-41
XRESLDTRK-Resume Tracking Load Addresses
6-44
XRSTOR-Restore Processor Extended States
6-45
XRSTORS-Restore Processor Extended States Supervisor
6-50
XSAVE-Save Processor Extended States
6-54
XSAVEC-Save Processor Extended States With Compaction
6-57
XSAVEOPT-Save Processor Extended States Optimized
6-60
XSAVES-Save Processor Extended States Supervisor
6-63
XSETBV-Set Extended Control Register
6-66
XSUSLDTRK-Suspend Tracking Load Addresses
6-68
XTEST-Test if in Transactional Execution
6-69
CHAPTER 7
SAFER MODE EXTENSIONS REFERENCE
7.1
OVERVIEW
7-1
7.2
SMX FUNCTIONALITY
7-1
7.2.1
Detecting and Enabling SMX
7-1
7.2.2
SMX Instruction Summary
7-2
7.2.2.1
GETSEC[CAPABILITIES]
7-3
7.2.2.2
GETSEC[ENTERACCS]
7-3
7.2.2.3
GETSEC[EXITAC]
7-3
7.2.2.4
GETSEC[SENTER]
7-4
7.2.2.5
GETSEC[SEXIT]
7-4
7.2.2.6
GETSEC[PARAMETERS]
7-4
7.2.2.7
GETSEC[SMCTRL]
7-4
7.2.2.8
GETSEC[WAKEUP]
7-4
7.2.3
Measured Environment and SMX
7-5
7.3
GETSEC LEAF FUNCTIONS
7-5
GETSEC[CAPABILITIES]-Report the SMX Capabilities
7-7
GETSEC[ENTERACCS]-Execute Authenticated Chipset Code
7-10
GETSEC[EXITAC]-Exit Authenticated Code Execution Mode
7-18
GETSEC[SENTER]-Enter a Measured Environment
7-21
GETSEC[SEXIT]-Exit Measured Environment
7-30
GETSEC[PARAMETERS]-Report the SMX Parameters
7-33
GETSEC[SMCTRL]-SMX Mode Control
7-37
GETSEC[WAKEUP]-Wake Up Sleeping Processors in Measured Environment
7-40
CHAPTER 8
INSTRUCTION SET REFERENCE UNIQUE TO INTEL® XEON PHI™ PROCESSORS
PREFETCHWT1-Prefetch Vector Data Into Caches With Intent to Write and T1 Hint
8-2
V4FMADDPS/V4FNMADDPS-Packed Single Precision Floating-Point Fused Multiply-Add (4-Iterations)
8-4
V4FMADDSS/V4FNMADDSS-Scalar Single Precision Floating-Point Fused Multiply-Add (4-Iterations)
8-6
VEXP2PD-Approximation to the Exponential 2^x of Packed Double Precision Floating-Point Values With Less
Than 2^-23 Relative Error
8-8
VEXP2PS-Approximation to the Exponential 2^x of Packed Single Precision Floating-Point Values With Less
Than 2^-23 Relative Error
8-10
VGATHERPF0DPS/VGATHERPF0QPS/VGATHERPF0DPD/VGATHERPF0QPD-Sparse Prefetch Packed SP/DP Data
Values With Signed Dword, Signed Qword Indices Using T0 Hint
8-12
VGATHERPF1DPS/VGATHERPF1QPS/VGATHERPF1DPD/VGATHERPF1QPD-Sparse Prefetch Packed SP/DP Data
Values With Signed Dword, Signed Qword Indices Using T1 Hint
8-14
VP4DPWSSDS-Dot Product of Signed Words With Dword Accumulation and Saturation (4-Iterations)
8-16
VP4DPWSSD-Dot Product of Signed Words With Dword Accumulation (4-Iterations)
8-18
VRCP28PD-Approximation to the Reciprocal of Packed Double Precision Floating-Point Values With Less Than
2^-28 Relative Error
8-20
VRCP28SD-Approximation to the Reciprocal of Scalar Double Precision Floating-Point Value With Less Than 2^-28
Relative Error
8-22
Vol. 2A xix
CONTENTS
PAGE
VRCP28PS-Approximation to the Reciprocal of Packed Single Precision Floating-Point Values With Less Than 2^-28
Relative Error
8-24
VRCP28SS-Approximation to the Reciprocal of Scalar Single Precision Floating-Point Value With Less Than 2^-28
Relative Error
8-26
VRSQRT28PD-Approximation to the Reciprocal Square Root of Packed Double Precision Floating-Point Values With
Less Than 2^-28 Relative Error
8-28
VRSQRT28SD-Approximation to the Reciprocal Square Root of Scalar Double Precision Floating-Point Value With Less
Than 2^-28 Relative Error
8-30
VRSQRT28PS-Approximation to the Reciprocal Square Root of Packed Single Precision Floating-Point Values With
Less Than 2^-28 Relative Error
8-32
VRSQRT28SS-Approximation to the Reciprocal Square Root of Scalar Single Precision Floating-Point Value With Less
Than 2^-28 Relative Error
8-34
VSCATTERPF0DPS/VSCATTERPF0QPS/VSCATTERPF0DPD/VSCATTERPF0QPD-Sparse Prefetch Packed SP/DP Data
Values with Signed Dword, Signed Qword Indices Using T0 Hint With Intent to Write
8-36
VSCATTERPF1DPS/VSCATTERPF1QPS/VSCATTERPF1DPD/VSCATTERPF1QPD-Sparse Prefetch Packed SP/DP Data
Values With Signed Dword, Signed Qword Indices Using T1 Hint With Intent to Write
8-38
APPENDIX A
OPCODE MAP
A.1
USING OPCODE TABLES
A-1
A.2
KEY TO ABBREVIATIONS
A-1
A.2.1
Codes for Addressing Method
A-1
A.2.2
Codes for Operand Type
A-2
A.2.3
Register Codes
A-3
A.2.4
Opcode Look-up Examples for One, Two, and Three-Byte Opcodes
A-3
A.2.4.1
One-Byte Opcode Instructions
A-3
A.2.4.2
Two-Byte Opcode Instructions
A-4
A.2.4.3
Three-Byte Opcode Instructions
A-5
A.2.4.4
VEX Prefix Instructions
A-5
A.2.5
Superscripts Utilized in Opcode Tables
A-6
A.3
ONE, TWO, AND THREE-BYTE OPCODE MAPS
A-6
A.4
OPCODE EXTENSIONS FOR ONE-BYTE AND TWO-BYTE OPCODES
A-17
A.4.1
Opcode Look-up Examples Using Opcode Extensions
A-17
A.4.2
Opcode Extension Tables
A-17
A.5
ESCAPE OPCODE INSTRUCTIONS
A-20
A.5.1
Opcode Look-up Examples for Escape Instruction Opcodes
A-20
A.5.2
Escape Opcode Instruction Tables
A-20
A.5.2.1
Escape Opcodes with D8 as First Byte
A-20
A.5.2.2
Escape Opcodes with D9 as First Byte
A-21
A.5.2.3
Escape Opcodes with DA as First Byte
A-22
A.5.2.4
Escape Opcodes with DB as First Byte
A-23
A.5.2.5
Escape Opcodes with DC as First Byte
A-24
A.5.2.6
Escape Opcodes with DD as First Byte
A-25
A.5.2.7
Escape Opcodes with DE as First Byte
A-26
A.5.2.8
Escape Opcodes with DF As First Byte
A-27
APPENDIX B
INSTRUCTION FORMATS AND ENCODINGS
B.1
MACHINE INSTRUCTION FORMAT
B-1
B.1.1
Legacy Prefixes
B-1
B.1.2
REX Prefixes
B-2
B.1.3
Opcode Fields
B-2
B.1.4
Special Fields
B-2
B.1.4.1
Reg Field (reg) for Non-64-Bit Modes
B-3
B.1.4.2
Reg Field (reg) for 64-Bit Mode
B-4
B.1.4.3
Encoding of Operand Size (w) Bit
B-4
B.1.4.4
Sign-Extend (s) Bit
B-5
B.1.4.5
Segment Register (sreg) Field
B-5
B.1.4.6
Special-Purpose Register (eee) Field
B-5
B.1.4.7
Condition Test (tttn) Field
B-6
B.1.4.8
Direction (d) Bit
B-6
B.1.5
Other Notes
B-6
B.2
GENERAL-PURPOSE INSTRUCTION FORMATS AND ENCODINGS FOR NON-64-BIT MODES
B-7
xx Vol. 2A
CONTENTS
PAGE
B.2.1
General Purpose Instruction Formats and Encodings for 64-Bit Mode
B-18
B.3
PENTIUM® PROCESSOR FAMILY INSTRUCTION FORMATS AND ENCODINGS
B-37
B.4
64-BIT MODE INSTRUCTION ENCODINGS FOR SIMD INSTRUCTION EXTENSIONS
B-37
B.5
MMX INSTRUCTION FORMATS AND ENCODINGS
B-38
B.5.1
Granularity Field (gg)
B-38
B.5.2
MMX Technology and General-Purpose Register Fields (mmxreg and reg)
B-38
B.5.3
MMX Instruction Formats and Encodings Table
B-38
B.6
PROCESSOR EXTENDED STATE INSTRUCTION FORMATS AND ENCODINGS
B-41
B.7
P6 FAMILY INSTRUCTION FORMATS AND ENCODINGS
B-41
B.8
SSE INSTRUCTION FORMATS AND ENCODINGS
B-42
B.9
SSE2 INSTRUCTION FORMATS AND ENCODINGS
B-47
B.9.1
Granularity Field (gg)
B-47
B.10
SSE3 FORMATS AND ENCODINGS TABLE
B-57
B.11
SSSE3 FORMATS AND ENCODING TABLE
B-58
B.12
AESNI AND PCLMULQDQ INSTRUCTION FORMATS AND ENCODINGS
B-60
B.13
SPECIAL ENCODINGS FOR 64-BIT MODE
B-61
B.14
SSE4.1 FORMATS AND ENCODING TABLE
B-64
B.15
SSE4.2 FORMATS AND ENCODING TABLE
B-69
B.16
AVX FORMATS AND ENCODING TABLE
B-70
B.17
FLOATING-POINT INSTRUCTION FORMATS AND ENCODINGS
B-108
B.18
VMX INSTRUCTIONS
B-112
B.19
SMX INSTRUCTIONS
B-113
APPENDIX C
INTEL® C/C++ COMPILER INTRINSICS AND FUNCTIONAL EQUIVALENTS
C.1
SIMPLE INTRINSICS
C-2
C.2
COMPOSITE INTRINSICS
C-14
Vol. 2A xxi
CONTENTS
PAGE
FIGURES
Figure 1-1.
Bit and Byte Order
1-5
Figure 1-2.
Syntax for CPUID, CR, and MSR Data Presentation
1-8
Figure 2-1.
Intel 64 and IA-32 Architectures Instruction Format
2-1
Figure 2-2.
Table Interpretation of ModR/M Byte (C8H)
2-4
Figure 2-3.
Prefix Ordering in 64-bit Mode
2-8
Figure 2-4.
Memory Addressing Without an SIB Byte; REX.X Not Used
2-9
Figure 2-5.
Register-Register Addressing (No Memory Operand); REX.X Not Used
2-9
Figure 2-6.
Memory Addressing With a SIB Byte
2-10
Figure 2-7.
Register Operand Coded in Opcode Byte; REX.X & REX.R Not Used
2-10
Figure 2-8.
Instruction Encoding Format with VEX Prefix
2-13
Figure 2-9.
VEX bit fields
2-15
Figure 2-10.
Intel® AVX-512 Instruction Format and the EVEX Prefix
2-37
Figure 2-11.
Bit Field Layout of the EVEX Prefix
2-37
Figure 3-1.
Bit Offset for BIT[RAX, 21]
3-11
Figure 3-2.
Memory Bit Indexing
3-12
Figure 3-3.
ADDSUBPD-Packed Double Precision Floating-Point Add/Subtract
3-45
Figure 3-4.
ADDSUBPS-Packed Single Precision Floating-Point Add/Subtract
3-47
Figure 3-5.
Memory Layout of BNDMOV to/from Memory
3-117
Figure 3-6.
Version Information Returned by CPUID in EAX
3-240
Figure 3-7.
Feature Information Returned in the ECX Register
3-242
Figure 3-8.
Feature Information Returned in the EDX Register
3-244
Figure 3-9.
Determination of Support for the Processor Brand String
3-254
Figure 3-10.
Algorithm for Extracting Processor Frequency
3-255
Figure 3-11.
CVTDQ2PD (VEX.256 encoded version)
3-266
Figure 3-12.
VCVTPD2DQ (VEX.256 encoded version)
3-273
Figure 3-13.
VCVTPD2PS (VEX.256 encoded version)
3-278
Figure 3-14.
CVTPS2PD (VEX.256 encoded version)
3-287
Figure 3-15.
VCVTTPD2DQ (VEX.256 encoded version)
3-303
Figure 3-16.
64-Byte Data Written to Enqueue Registers
3-350
Figure 3-17.
HADDPD-Packed Double Precision Floating-Point Horizontal Add
3-482
Figure 3-18.
VHADDPD Operation
3-483
Figure 3-19.
HADDPS-Packed Single Precision Floating-Point Horizontal Add
3-486
Figure 3-20.
VHADDPS Operation
3-486
Figure 3-21.
HSUBPD-Packed Double Precision Floating-Point Horizontal Subtract
3-491
Figure 3-22.
VHSUBPD operation
3-492
Figure 3-23.
HSUBPS-Packed Single Precision Floating-Point Horizontal Subtract
3-495
Figure 3-24.
VHSUBPS Operation
3-495
Figure 3-25.
INVPCID Descriptor
3-535
Figure 4-1.
Operation of PCMPSTRx and PCMPESTRx
4-6
Figure 4-2.
VMOVDDUP Operation
4-59
Figure 4-3.
MOVSHDUP Operation
4-117
Figure 4-4.
MOVSLDUP Operation
4-120
Figure 4-5.
256-bit VMPSADBW Operation
4-139
Figure 4-6.
Operation of the PACKSSDW Instruction Using 64-Bit Operands
4-189
Figure 4-7.
256-bit VPALIGN Instruction Operation
4-222
Figure 4-8.
PDEP Example
4-277
Figure 4-9.
PEXT Example
4-279
Figure 4-10.
256-bit VPHADDD Instruction Operation
4-288
Figure 4-11.
PMADDWD Execution Model Using 64-bit Operands
4-309
Figure 4-12.
PMULHUW and PMULHW Instruction Operation Using 64-bit Operands
4-374
Figure 4-13.
PMULLU Instruction Operation Using 64-bit Operands
4-386
Figure 4-14.
PSADBW Instruction Operation Using 64-bit Operands
4-413
Figure 4-15.
PSHUFB with 64-Bit Operands
4-418
Figure 4-16.
256-bit VPSHUFD Instruction Operation
4-421
Figure 4-17.
PSLLW, PSLLD, and PSLLQ Instruction Operation Using 64-bit Operand
4-439
Figure 4-18.
PSRAW and PSRAD Instruction Operation Using a 64-bit Operand
4-451
Figure 4-19.
PSRLW, PSRLD, and PSRLQ Instruction Operation Using 64-bit Operand
4-463
Figure 4-20.
PUNPCKHBW Instruction Operation Using 64-bit Operands
4-498
Figure 4-21.
256-bit VPUNPCKHDQ Instruction Operation
4-498
Figure 4-22.
PUNPCKLBW Instruction Operation Using 64-bit Operands
4-508
Figure 4-23.
256-bit VPUNPCKLDQ Instruction Operation
4-508
Figure 4-24.
Bit Control Fields of Immediate Byte for ROUNDxx Instruction
4-573
xxii Vol. 2A
CONTENTS
PAGE
Figure 4-25.
256-bit VSHUFPD Operation of Four Pairs of Double Precision Floating-Point Values
4-635
Figure 4-26.
256-bit VSHUFPS Operation of Selection from Input Quadruplet and Pair-wise Interleaved Result
4-640
Figure 4-27.
VUNPCKHPS Operation
4-731
Figure 4-28.
VUNPCKLPS Operation
4-739
Figure 5-1.
VBROADCASTSS Operation (VEX.256 encoded version)
5-15
Figure 5-2.
VBROADCASTSS Operation (VEX.128-bit version)
5-15
Figure 5-3.
VBROADCASTSD Operation (VEX.256-bit version)
5-15
Figure 5-4.
VBROADCASTF128 Operation (VEX.256-bit version)
5-15
Figure 5-5.
VBROADCASTF64X4 Operation (512-bit version with writemask all 1s)
5-16
Figure 5-6.
VCVTPH2PS (128-bit Version)
5-50
Figure 5-7.
VCVTPS2PH (128-bit Version)
5-63
Figure 5-8.
64-bit Super Block of SAD Operation in VDBPSADBW
5-145
Figure 5-9.
VFIXUPIMMPD Immediate Control Description
5-182
Figure 5-10.
VFIXUPIMMPS Immediate Control Description
5-186
Figure 5-11.
VFIXUPIMMSD Immediate Control Description
5-190
Figure 5-12.
VFIXUPIMMSS Immediate Control Description
5-193
Figure 5-13.
Imm8 Byte Specifier of Special Case Floating-Point Values for VFPCLASSPD/SD/PS/SS
5-325
Figure 5-14.
VGETEXPPS Functionality On Normal Input values
5-359
Figure 5-15.
Imm8 Controls for VGETMANTPD/SD/PS/SS
5-368
Figure 5-16.
VPBROADCASTD Operation (VEX.256 encoded version)
5-424
Figure 5-17.
VPBROADCASTD Operation (128-bit version)
5-424
Figure 5-18.
VPBROADCASTQ Operation (256-bit version)
5-424
Figure 5-19.
VBROADCASTI128 Operation (256-bit version)
5-425
Figure 5-20.
VBROADCASTI256 Operation (512-bit version)
5-425
Figure 5-21.
VPERM2F128 Operation
5-463
Figure 5-22.
VPERM2I128 Operation
5-465
Figure 5-23.
VPERMILPD Operation
5-481
Figure 5-24.
VPERMILPD Shuffle Control
5-481
Figure 5-25.
VPERMILPS Operation
5-486
Figure 5-26.
VPERMILPS Shuffle Control
5-486
Figure 5-27.
Imm8 Controls for VRANGEPD/SD/PS/SS
5-621
Figure 5-28.
Imm8 Controls for VREDUCEPD/SD/PS/SS
5-645
Figure 5-29.
Imm8 Controls for VRNDSCALEPD/SD/PS/SS
5-660
Figure 8-1.
Register Source-Block Dot Product of Two Signed Word Operands With Doubleword Accumulation
8-18
Figure A-1.
ModR/M Byte nnn Field (Bits 5, 4, and 3)
A-17
Figure B-1.
General Machine Instruction Format
B-1
Figure B-2.
Hybrid Notation of VEX-Encoded Key Instruction Bytes
B-70
Vol. 2A xxiii
CONTENTS
PAGE
TABLES
Table 2-1.
16-Bit Addressing Forms with the ModR/M Byte
2-5
Table 2-2.
32-Bit Addressing Forms with the ModR/M Byte
2-6
Table 2-3.
32-Bit Addressing Forms with the SIB Byte
2-7
Table 2-4.
REX Prefix Fields [BITS: 0100WRXB]
2-9
Table 2-6.
Direct Memory Offset Form of MOV
2-11
Table 2-5.
Special Cases of REX Encodings
2-11
Table 2-7.
RIP-Relative Addressing
2-12
Table 2-8.
VEX.vvvv to register name mapping
2-17
Table 2-9.
Instructions with a VEX.vvvv destination
2-17
Table 2-10.
VEX.m-mmmm interpretation
2-18
Table 2-11.
VEX.L interpretation
2-19
Table 2-12.
VEX.pp interpretation
2-19
Table 2-13.
32-Bit VSIB Addressing Forms of the SIB Byte
2-21
Table 2-14.
Exception Class Description
2-23
Table 2-15.
Instructions in each Exception Class
2-24
Table 2-16.
#UD Exception and VEX.W=1 Encoding
2-25
Table 2-17.
#UD Exception and VEX.L Field Encoding
2-26
Table 2-18.
Type 1 Class Exception Conditions
2-27
Table 2-19.
Type 2 Class Exception Conditions
2-28
Table 2-20.
Type 3 Class Exception Conditions
2-29
Table 2-21.
Type 4 Class Exception Conditions
2-30
Table 2-22.
Type 5 Class Exception Conditions
2-31
Table 2-23.
Type 6 Class Exception Conditions
2-32
Table 2-24.
Type 7 Class Exception Conditions
2-33
Table 2-25.
Type 8 Class Exception Conditions
2-33
Table 2-26.
Type 11 Class Exception Conditions
2-34
Table 2-27.
Type 12 Class Exception Conditions
2-35
Table 2-28.
VEX-Encoded GPR Instructions
2-36
Table 2-29.
Type 13 Class Exception Conditions
2-36
Table 2-30.
EVEX Prefix Bit Field Functional Grouping
2-38
Table 2-31.
32-Register Support in 64-bit Mode Using EVEX with Embedded REX Bits
2-39
Table 2-32.
EVEX Encoding Register Specifiers in 32-bit Mode
2-39
Table 2-33.
Opmask Register Specifier Encoding
2-40
Table 2-34.
Compressed Displacement (DISP8*N) Affected by Embedded Broadcast
2-41
Table 2-35.
EVEX DISP8*N for Instructions Not Affected by Embedded Broadcast
2-41
Table 2-36.
EVEX Embedded Broadcast/Rounding/SAE and Vector Length on Vector Instructions
2-43
Table 2-37.
OS XSAVE Enabling Requirements of Instruction Categories
2-43
Table 2-38.
Opcode Independent, State Dependent EVEX Bit Fields
2-43
Table 2-39.
#UD Conditions of Operand-Encoding EVEX Prefix Bit Fields
2-44
Table 2-40.
#UD Conditions of Opmask Related Encoding Field
2-44
Table 2-41.
#UD Conditions Dependent on EVEX.b Context
2-45
Table 2-42.
EVEX-Encoded Instruction Exception Class Summary
2-45
Table 2-43.
EVEX Instructions in Each Exception Class
2-46
Table 2-44.
Type E1 Class Exception Conditions
2-49
Table 2-45.
Type E1NF Class Exception Conditions
2-50
Table 2-46.
Type E2 Class Exception Conditions
2-51
Table 2-47.
Type E3 Class Exception Conditions
2-52
Table 2-48.
Type E3NF Class Exception Conditions
2-53
Table 2-49.
Type E4 Class Exception Conditions
2-54
Table 2-50.
Type E4NF Class Exception Conditions
2-55
Table 2-51.
Type E5 Class Exception Conditions
2-56
Table 2-52.
Type E5NF Class Exception Conditions
2-57
Table 2-53.
Type E6 Class Exception Conditions
2-58
Table 2-54.
Type E6NF Class Exception Conditions
2-59
Table 2-55.
Type E7NM Class Exception Conditions
2-60
Table 2-56.
Type E9 Class Exception Conditions
2-61
Table 2-57.
Type E9NF Class Exception Conditions
2-62
xxiv Vol. 2A
CONTENTS
PAGE
Table 2-58.
Type E10 Class Exception Conditions
2-63
Table 2-59.
Type E10NF Class Exception Conditions
2-64
Table 2-60.
Type E11 Class Exception Conditions
2-65
Table 2-61.
Type E12 Class Exception Conditions
2-66
Table 2-62.
Type E12NP Class Exception Conditions
2-67
Table 2-63.
TYPE K20 Exception Definition (VEX-Encoded OpMask Instructions w/o Memory Arg)
2-68
Table 2-64.
TYPE K21 Exception Definition (VEX-Encoded OpMask Instructions Addressing Memory)
2-69
Table 2-65.
Intel® AMX Exception Classes
2-70
Table 3-1.
Register Codes Associated With +rb, +rw, +rd, +ro
3-2
Table 3-2.
Range of Bit Positions Specified by Bit Offset Operands
3-12
Table 3-3.
Standard and Non-Standard Data Types
3-14
Table 3-4.
Intel 64 and IA-32 General Exceptions
3-15
Table 3-5.
x87 FPU Floating-Point Exceptions
3-16
Table 3-6.
SIMD Floating-Point Exceptions
3-16
Table 3-7.
Decision Table for CLI Results
3-166
Table 3-1.
Comparison Predicate for CMPPD and CMPPS Instructions
3-182
Table 3-2.
Pseudo-Op and CMPPD Implementation
3-183
Table 3-3.
Pseudo-Op and VCMPPD Implementation
3-184
Table 3-4.
Pseudo-Op and CMPPS Implementation
3-189
Table 3-5.
Pseudo-Op and VCMPPS Implementation
3-190
Table 3-6.
Pseudo-Op and CMPSD Implementation
3-200
Table 3-7.
Pseudo-Op and VCMPSD Implementation
3-200
Table 3-8.
Pseudo-Op and CMPSS Implementation
3-204
Table 3-9.
Pseudo-Op and VCMPSS Implementation
3-204
Table 3-8.
Information Returned by CPUID Instruction
3-218
Table 3-9.
Processor Type Field
3-240
Table 3-10.
Feature Information Returned in the ECX Register
3-242
Table 3-11.
More on Feature Information Returned in the EDX Register
3-245
Table 3-12.
Encoding of CPUID Leaf 2 Descriptors
3-247
Table 3-13.
Processor Brand String Returned with Pentium 4 Processor
3-254
Table 3-14.
Mapping of Brand Indices; and Intel 64 and IA-32 Processor Brand Strings
3-256
Table 3-15.
DIV Action
3-322
Table 3-16.
Results Obtained from F2XM1
3-358
Table 3-17.
Results Obtained from FABS
3-360
Table 3-18.
FADD/FADDP/FIADD Results
3-362
Table 3-19.
FBSTP Results
3-366
Table 3-20.
FCHS Results
3-368
Table 3-21.
FCOM/FCOMP/FCOMPP Results
3-374
Table 3-22.
FCOMI/FCOMIP/ FUCOMI/FUCOMIP Results
3-377
Table 3-23.
FCOS Results
3-380
Table 3-24.
FDIV/FDIVP/FIDIV Results
3-384
Table 3-25.
FDIVR/FDIVRP/FIDIVR Results
3-387
Table 3-26.
FICOM/FICOMP Results
3-390
Table 3-27.
FIST/FISTP Results
3-397
Table 3-28.
FISTTP Results
3-400
Table 3-29.
FMUL/FMULP/FIMUL Results
3-411
Table 3-30.
FPATAN Results
3-414
Table 3-31.
FPREM Results
3-416
Table 3-32.
FPREM1 Results
3-418
Table 3-33.
FPTAN Results
3-420
Table 3-34.
FSCALE Results
3-428
Table 3-35.
FSIN Results
3-430
Table 3-36.
FSINCOS Results
3-432
Table 3-37.
FSQRT Results
3-434
Table 3-38.
FSUB/FSUBP/FISUB Results
3-445
Table 3-39.
FSUBR/FSUBRP/FISUBR Results
3-448
Table 3-40.
FTST Results
3-450
Table 3-41.
FUCOM/FUCOMP/FUCOMPP Results
3-452
Table 3-42.
FXAM Results
3-454
Vol. 2A xxv
CONTENTS
PAGE
Table 3-43.
Non-64-Bit-Mode Layout of FXSAVE and FXRSTOR Memory Region
3-461
Table 3-44.
Field Definitions
3-462
Table 3-45.
Recreating FSAVE Format
3-464
Table 3-46.
Layout of the 64-Bit Mode FXSAVE64 Map (Requires REX.W = 1)
3-465
Table 3-47.
Layout of the 64-Bit Mode FXSAVE Map (REX.W = 0)
3-466
Table 3-48.
FYL2X Results
3-471
Table 3-49.
FYL2XP1 Results
3-473
Table 3-50.
Inverse Byte Listings
3-476
Table 3-51.
IDIV Results
3-497
Table 3-52.
Decision Table
3-517
Table 3-53.
Segment and Gate Types
3-582
Table 3-10.
Memory Area Layout
3-591
Table 3-54.
Non-64-bit Mode LEA Operation with Address and Operand Size Attributes
3-594
Table 3-55.
64-bit Mode LEA Operation with Address and Operand Size Attributes
3-594
Table 3-56.
Segment and Gate Descriptor Types
3-618
Table 4-1.
Source Data Format
4-2
Table 4-2.
Aggregation Operation
4-2
Table 4-3.
Aggregation Operation
4-3
Table 4-4.
Polarity
4-3
Table 4-5.
Output Selection
4-4
Table 4-6.
Output Selection
4-4
Table 4-7.
Comparison Result for Each Element Pair BoolRes[i.j]
4-4
Table 4-8.
Summary of Imm8 Control Byte
4-5
Table 4-9.
MUL Results
4-146
Table 4-10.
MWAIT Extension Register (ECX)
4-161
Table 4-11.
MWAIT Hints Register (EAX)
4-161
Table 4-12.
Recommended Multi-Byte Sequence of NOP Instruction
4-165
Table 4-13.
PCLMULQDQ Quadword Selection of Immediate Byte
4-243
Table 4-14.
Pseudo-Op and PCLMULQDQ Implementation
4-243
Table 4-15.
MKTME_KEY_PROGRAM_STRUCT Format
4-271
Table 4-16.
Effect of POPF/POPFD on the EFLAGS Register
4-402
Table 4-17.
Repeat Prefixes
4-555
Table 4-18.
Rounding Modes and Encoding of Rounding Control (RC) Field
4-573
Table 4-19.
Decision Table for STI Results
4-662
Table 4-20.
TPAUSE Input Register Bit Definitions
4-711
Table 4-21.
UMWAIT Input Register Bit Definitions
4-724
Table 5-1.
Lower 8 columns of the 16x16 Map of VPTERNLOG Boolean Logic Operations
5-2
Table 5-2.
Upper 8 columns of the 16x16 Map of VPTERNLOG Boolean Logic Operations
5-3
Table 5-3.
Immediate Byte Encoding for 16-bit Floating-Point Conversion Instructions
5-64
Table 5-1.
NaN Propagation Priorities
5-150
Table 5-2.
VF[,N]MADD[132,213,231]PH Notation for Operands
5-202
Table 5-3.
VF[,N]MADD[132,213,231]SH Notation for Operands
5-216
Table 5-4.
VFMADDSUB[132,213,231]PH Notation for Odd and Even Elements
5-230
Table 5-5.
VF[,N]MSUB[132,213,231]PH Notation for Operands
5-248
Table 5-6.
VF[,N]MSUB[132,213,231]SH Notation for Operands
5-262
Table 5-7.
VFMSUBADD[132,213,231]PH Notation for Odd and Even Elements
5-276
Table 5-4.
Classifier Operations for VFPCLASSPD/SD/PS/SS
5-325
Table 5-8.
Classifier Operations for VFPCLASSPH/VFPCLASSSH
5-328
Table 5-5.
VGETEXPPD/SD Special Cases
5-352
Table 5-6.
VGETEXPPH/VGETEXPSH Special Cases
5-355
Table 5-7.
VGETEXPPS/SS Special Cases
5-358
Table 5-8.
GetMant() Special Float Values Behavior
5-369
Table 5-9.
imm8 Controls for VGETMANTPH/VGETMANTSH
5-372
Table 5-10.
GetMant() Special Float Values Behavior
5-373
Table 5-11.
Pseudo-Op and VPCMP* Implementation
5-434
Table 5-12.
Examples of VPTERNLOGD/Q Imm8 Boolean Function and Input Index Values
5-613
Table 5-13.
Signaling of Comparison Operation of One or More NaN Input Values and Effect of Imm8[3:2]
5-622
Table 5-14.
Comparison Result for Opposite-Signed Zero Cases for MIN, MIN_ABS, and MAX, MAX_ABS
5-622
Table 5-15.
Comparison Result of Equal-Magnitude Input Cases for MIN_ABS and MAX_ABS, (|a| = |b|, a>0, b<0)
5-622
xxvi Vol. 2A
CONTENTS
PAGE
Table 5-16.
VRCP14PD/VRCP14SD Special Cases
5-634
Table 5-17.
VRCP14PS/VRCP14SS Special Cases
5-638
Table 5-18.
VRCPPH/VRCPSH Special Cases
5-642
Table 5-19.
VREDUCEPD/SD/PS/SS Special Cases
5-646
Table 5-20.
VREDUCEPH/VREDUCESH Special Cases
5-649
Table 5-21.
VRNDSCALEPD/SD/PS/SS Special Cases
5-660
Table 5-22.
Imm8 Controls for VRNDSCALEPH/VRNDSCALESH
5-663
Table 5-23.
VRNDSCALEPH/VRNDSCALESH Special Cases
5-663
Table 5-24.
VRSQRT14PD Special Cases
5-675
Table 5-25.
VRSQRT14SD Special Cases
5-677
Table 5-26.
VRSQRT14PS Special Cases
5-679
Table 5-27.
VRSQRT14SS Special Cases
5-681
Table 5-28.
VRSQRTPH/VRSQRTSH Special Cases
5-682
Table 5-29.
VSCALEFPD/SD/PS/SS Special Cases
5-685
Table 5-30.
Additional VSCALEFPD/SD Special Cases
5-686
Table 5-31.
VSCALEFPH/VSCALEFSH Special Cases
5-688
Table 5-32.
Additional VSCALEFPH/VSCALEFSH Special Cases
5-688
Table 5-33.
Additional VSCALEFPS/SS Special Cases
5-690
Table 7-1.
Layout of IA32_FEATURE_CONTROL
7-2
Table 7-2.
GETSEC Leaf Functions
7-3
Table 7-3.
GETSEC Capability Result Encoding (EBX = 0)
7-7
Table 7-4.
Register State Initialization After GETSEC[ENTERACCS]
7-12
Table 7-5.
IA32_MISC_ENABLE MSR Initialization by ENTERACCS and SENTER
7-13
Table 7-6.
Register State Initialization After GETSEC[SENTER] and GETSEC[WAKEUP]
7-24
Table 7-7.
SMX Reporting Parameters Format
7-33
Table 7-8.
TXT Feature Extensions Flags
7-34
Table 7-9.
External Memory Types Using Parameter 3
7-35
Table 7-10.
Default Parameter Values
7-35
Table 7-11.
Supported Actions for GETSEC[SMCTRL(0)]
7-37
Table 7-12.
RLP MVMM JOIN Data Structure
7-40
Table 8-1.
Special Values Behavior
8-9
Table 8-2.
Special Values Behavior
8-11
Table 8-3.
VRCP28PD Special Cases
8-21
Table 8-4.
VRCP28SD Special Cases
8-23
Table 8-5.
VRCP28PS Special Cases
8-25
Table 8-6.
VRCP28SS Special Cases
8-27
Table 8-7.
VRSQRT28PD Special Cases
8-29
Table 8-8.
VRSQRT28SD Special Cases
8-31
Table 8-9.
VRSQRT28PS Special Cases
8-33
Table 8-10.
VRSQRT28SS Special Cases
8-35
Table A-1.
Superscripts Utilized in Opcode Tables
A-6
Table A-2.
One-byte Opcode Map: (00H - F7H) *
A-7
Table A-3.
Two-byte Opcode Map: 00H - 77H (First Byte is 0FH) *
A-9
Table A-4.
Three-byte Opcode Map: 00H - F7H (First Two Bytes are 0F 38H) *
A-13
Table A-5.
Three-byte Opcode Map: 00H - F7H (First two bytes are 0F 3AH) *
A-15
Table A-6.
Opcode Extensions for One- and Two-byte Opcodes by Group Number *
A-18
Table A-7.
D8 Opcode Map When ModR/M Byte is Within 00H to BFH *
A-20
Table A-8.
D8 Opcode Map When ModR/M Byte is Outside 00H to BFH *
A-21
Table A-9.
D9 Opcode Map When ModR/M Byte is Within 00H to BFH *
A-21
Table A-10.
D9 Opcode Map When ModR/M Byte is Outside 00H to BFH *
A-22
Table A-11.
DA Opcode Map When ModR/M Byte is Within 00H to BFH *
A-22
Table A-12.
DA Opcode Map When ModR/M Byte is Outside 00H to BFH *
A-23
Table A-13.
DB Opcode Map When ModR/M Byte is Within 00H to BFH *
A-23
Table A-14.
DB Opcode Map When ModR/M Byte is Outside 00H to BFH *
A-24
Table A-15.
DC Opcode Map When ModR/M Byte is Within 00H to BFH *
A-24
Table A-16.
DC Opcode Map When ModR/M Byte is Outside 00H to BFH *
A-25
Table A-17.
DD Opcode Map When ModR/M Byte is Within 00H to BFH *
A-25
Table A-18.
DD Opcode Map When ModR/M Byte is Outside 00H to BFH *
A-26
Table A-19.
DE Opcode Map When ModR/M Byte is Within 00H to BFH *
A-26
Vol. 2A xxvii
CONTENTS
PAGE
Table A-20.
DE Opcode Map When ModR/M Byte is Outside 00H to BFH *
A-27
Table A-21.
DF Opcode Map When ModR/M Byte is Within 00H to BFH *
A-27
Table A-22.
DF Opcode Map When ModR/M Byte is Outside 00H to BFH *
A-28
Table B-1.
Special Fields Within Instruction Encodings
B-2
Table B-2.
Encoding of reg Field When w Field is Not Present in Instruction
B-3
Table B-3.
Encoding of reg Field When w Field is Present in Instruction
B-3
Table B-4.
Encoding of reg Field When w Field is Not Present in Instruction
B-4
Table B-5.
Encoding of reg Field When w Field is Present in Instruction
B-4
Table B-6.
Encoding of Operand Size (w) Bit
B-4
Table B-7.
Encoding of Sign-Extend (s) Bit
B-5
Table B-8.
Encoding of the Segment Register (sreg) Field
B-5
Table B-9.
Encoding of Special-Purpose Register (eee) Field
B-5
Table B-10.
Encoding of Conditional Test (tttn) Field
B-6
Table B-11.
Encoding of Operation Direction (d) Bit
B-6
Table B-13.
General Purpose Instruction Formats and Encodings for Non-64-Bit Modes
B-7
Table B-12.
Notes on Instruction Encoding
B-7
Table B-14.
Special Symbols
B-18
Table B-15.
General Purpose Instruction Formats and Encodings for 64-Bit Mode
B-18
Table B-16.
Pentium® Processor Family Instruction Formats and Encodings, Non-64-Bit Modes
B-37
Table B-17.
Pentium® Processor Family Instruction Formats and Encodings, 64-Bit Mode
B-37
Table B-18.
Encoding of Granularity of Data Field (gg)
B-38
Table B-19.
MMX Instruction Formats and Encodings
B-38
Table B-20.
Formats and Encodings of XSAVE/XRSTOR/XGETBV/XSETBV Instructions
B-41
Table B-21.
Formats and Encodings of P6 Family Instructions
B-41
Table B-22.
Formats and Encodings of SSE Floating-Point Instructions
B-42
Table B-23.
Formats and Encodings of SSE Integer Instructions
B-46
Table B-25.
Encoding of Granularity of Data Field (gg)
B-47
Table B-24.
Format and Encoding of SSE Cacheability & Memory Ordering Instructions
B-47
Table B-26.
Formats and Encodings of SSE2 Floating-Point Instructions
B-48
Table B-27.
Formats and Encodings of SSE2 Integer Instructions
B-52
Table B-28.
Format and Encoding of SSE2 Cacheability Instructions
B-56
Table B-29.
Formats and Encodings of SSE3 Floating-Point Instructions
B-57
Table B-30.
Formats and Encodings for SSE3 Event Management Instructions
B-57
Table B-31.
Formats and Encodings for SSE3 Integer and Move Instructions
B-57
Table B-32.
Formats and Encodings for SSSE3 Instructions
B-58
Table B-33.
Formats and Encodings of AESNI and PCLMULQDQ Instructions
B-61
Table B-34.
Special Case Instructions Promoted Using REX.W
B-61
Table B-35.
Encodings of SSE4.1 instructions
B-64
Table B-36.
Encodings of SSE4.2 instructions
B-69
Table B-37.
Encodings of AVX instructions
B-71
Table B-38.
General Floating-Point Instruction Formats
B-108
Table B-39.
Floating-Point Instruction Formats and Encodings
B-108
Table B-40.
Encodings for VMX Instructions
B-112
Table B-41.
Encodings for SMX Instructions
B-113
Table C-1.
Simple Intrinsics
C-2
Table C-2.
Composite Intrinsics
C-14
xxviii Vol. 2A
CHAPTER 1
ABOUT THIS MANUAL
The Intel® 64 and IA-32 Architectures Software Developer’s Manual, Volumes 2A, 2B, 2C, & 2D: Instruction Set
Reference (order numbers 253666, 253667, 326018, and 334569), is part of a set that describes the architecture
and programming environment of all Intel 64 and IA-32 architecture processors. Other volumes in this set are:
The Intel® 64 and IA-32 Architectures Software Developer’s Manual, Volume 1: Basic Architecture (Order
Number 253665).
The Intel® 64 and IA-32 Architectures Software Developer’s Manual, Volumes 3A, 3B, 3C, & 3D: System
Programming Guide (order numbers 253668, 253669, 326019, and 332831).
The Intel® 64 and IA-32 Architectures Software Developer’s Manual, Volume 4: Model-Specific Registers
(order number 335592).
The Intel® 64 and IA-32 Architectures Software Developer’s Manual, Volume 1, describes the basic architecture
and programming environment of Intel 64 and IA-32 processors. The Intel® 64 and IA-32 Architectures Software
Developer’s Manual, Volumes 2A, 2B, 2C, & 2D, describes the instruction set of the processor and the opcode struc-
ture. These volumes apply to application programmers and to programmers who write operating systems or exec-
utives. The Intel® 64 and IA-32 Architectures Software Developer’s Manual, Volumes 3A, 3B, 3C, & 3D, describes
the operating-system support environment of Intel 64 and IA-32 processors. These volumes target operating-
system and BIOS designers. In addition, the Intel® 64 and IA-32 Architectures Software Developer’s Manual,
Volume 3B, addresses the programming environment for classes of software that host operating systems. The
Intel® 64 and IA-32 Architectures Software Developer’s Manual, Volume 4, describes the model-specific registers
of Intel 64 and IA-32 processors.
1.1
INTEL® 64 AND IA-32 PROCESSORS COVERED IN THIS MANUAL
This manual set includes information pertaining primarily to the most recent Intel 64 and IA-32 processors, which
include:
Pentium® processors
P6 family processors
Pentium® 4 processors
Pentium® M processors
Intel® Xeon® processors
Pentium® D processors
Pentium® processor Extreme Editions
64-bit Intel® Xeon® processors
Intel® Core™ Duo processor
Intel® Core™ Solo processor
Dual-Core Intel® Xeon® processor LV
Intel® Core™2 Duo processor
Intel® Core™2 Quad processor Q6000 series
Intel® Xeon® processor 3000, 3200 series
Intel® Xeon® processor 5000 series
Intel® Xeon® processor 5100, 5300 series
Intel® Core™2 Extreme processor X7000 and X6800 series
Intel® Core™2 Extreme processor QX6000 series
Intel® Xeon® processor 7100 series
Vol. 2A
1-1
ABOUT THIS MANUAL
Intel® Pentium® Dual-Core processor
Intel® Xeon® processor 7200, 7300 series
Intel® Xeon® processor 5200, 5400, 7400 series
Intel® Core™2 Extreme processor QX9000 and X9000 series
Intel® Core™2 Quad processor Q9000 series
Intel® Core™2 Duo processor E8000, T9000 series
Intel Atom® processor family
Intel Atom® processors 200, 300, D400, D500, D2000, N200, N400, N2000, E2000, Z500, Z600, Z2000,
C1000 series are built from 45 nm and 32 nm processes
Intel® Core™ i7 processor
Intel® Core™ i5 processor
Intel® Xeon® processor E7-8800/4800/2800 product families
Intel® Core™ i7-3930K processor
2nd generation Intel® Core™ i7-2xxx, Intel® Core™ i5-2xxx, Intel® Core™ i3-2xxx processor series
Intel® Xeon® processor E3-1200 product family
Intel® Xeon® processor E5-2400/1400 product family
Intel® Xeon® processor E5-4600/2600/1600 product family
3rd generation Intel® Core™ processors
Intel® Xeon® processor E3-1200 v2 product family
Intel® Xeon® processor E5-2400/1400 v2 product families
Intel® Xeon® processor E5-4600/2600/1600 v2 product families
Intel® Xeon® processor E7-8800/4800/2800 v2 product families
4th generation Intel® Core™ processors
The Intel® Core™ M processor family
Intel® Core™ i7-59xx Processor Extreme Edition
Intel® Core™ i7-49xx Processor Extreme Edition
Intel® Xeon® processor E3-1200 v3 product family
Intel® Xeon® processor E5-2600/1600 v3 product families
5th generation Intel® Core™ processors
Intel® Xeon® processor D-1500 product family
Intel® Xeon® processor E5 v4 family
Intel Atom® processor X7-Z8000 and X5-Z8000 series
Intel Atom® processor Z3400 series
Intel Atom® processor Z3500 series
6th generation Intel® Core™ processors
Intel® Xeon® processor E3-1500m v5 product family
7th generation Intel® Core™ processors
Intel® Xeon Phi™ Processor 3200, 5200, 7200 Series
Intel® Xeon® Scalable Processor Family
8th generation Intel® Core™ processors
Intel® Xeon Phi™ Processor 7215, 7285, 7295 Series
Intel® Xeon® E processors
9th generation Intel® Core™ processors
2nd generation Intel® Xeon® Scalable Processor Family
1-2
Vol. 2A
ABOUT THIS MANUAL
10th generation Intel® Core™ processors
11th generation Intel® Core™ processors
3rd generation Intel® Xeon® Scalable Processor Family
12th generation Intel® Core™ processors
13th generation Intel® Core™ processors
4th generation Intel® Xeon® Scalable Processor Family
P6 family processors are IA-32 processors based on the P6 family microarchitecture. This includes the Pentium®
Pro, Pentium® II, Pentium® III, and Pentium® III Xeon® processors.
The Pentium® 4, Pentium® D, and Pentium® processor Extreme Editions are based on the Intel NetBurst® microar-
chitecture. Most early Intel® Xeon® processors are based on the Intel NetBurst® microarchitecture. Intel Xeon
processor 5000, 7100 series are based on the Intel NetBurst® microarchitecture.
The Intel® Core™ Duo, Intel® Core™ Solo and dual-core Intel® Xeon® processor LV are based on an improved
Pentium® M processor microarchitecture.
The Intel® Xeon® processor 3000, 3200, 5100, 5300, 7200, and 7300 series, Intel® Pentium® dual-core, Intel®
Core™2 Duo, Intel® Core™2 Quad, and Intel® Core™2 Extreme processors are based on Intel® Core™ microarchi-
tecture.
The Intel® Xeon® processor 5200, 5400, 7400 series, Intel® Core™2 Quad processor Q9000 series, and Intel®
Core™2 Extreme processors QX9000, X9000 series, Intel® Core™2 processor E8000 series are based on Enhanced
Intel® Core™ microarchitecture.
The Intel Atom® processors 200, 300, D400, D500, D2000, N200, N400, N2000, E2000, Z500, Z600, Z2000,
C1000 series are based on the Intel Atom® microarchitecture and supports Intel 64 architecture.
P6 family, Pentium® M, Intel® Core™ Solo, Intel® Core™ Duo processors, dual-core Intel® Xeon® processor LV,
and early generations of Pentium 4 and Intel Xeon processors support IA-32 architecture. The Intel® AtomTM
processor Z5xx series support IA-32 architecture.
The Intel® Xeon® processor 3000, 3200, 5000, 5100, 5200, 5300, 5400, 7100, 7200, 7300, 7400 series, Intel®
Core™2 Duo, Intel® Core™2 Extreme, Intel® Core™2 Quad processors, Pentium® D processors, Pentium® Dual-
Core processor, newer generations of Pentium 4 and Intel Xeon processor family support Intel® 64 architecture.
The Intel® Core™ i7 processor and Intel® Xeon® processor 3400, 5500, 7500 series are based on 45 nm Nehalem
microarchitecture. Westmere microarchitecture is a 32 nm version of the Nehalem microarchitecture. Intel®
Xeon® processor 5600 series, Intel Xeon processor E7 and various Intel Core i7, i5, i3 processors are based on the
Westmere microarchitecture. These processors support Intel 64 architecture.
The Intel® Xeon® processor E5 family, Intel® Xeon® processor E3-1200 family, Intel® Xeon® processor E7-
8800/4800/2800 product families, Intel® Core™ i7-3930K processor, and 2nd generation Intel® Core™ i7-2xxx,
Intel® CoreTM i5-2xxx, Intel® Core™ i3-2xxx processor series are based on the Sandy Bridge microarchitecture and
support Intel 64 architecture.
The Intel® Xeon® processor E7-8800/4800/2800 v2 product families, Intel® Xeon® processor E3-1200 v2 product
family and 3rd generation Intel® Core™ processors are based on the Ivy Bridge microarchitecture and support
Intel 64 architecture.
The Intel® Xeon® processor E5-4600/2600/1600 v2 product families, Intel® Xeon® processor E5-2400/1400 v2
product families and Intel® Core™ i7-49xx Processor Extreme Edition are based on the Ivy Bridge-E microarchitec-
ture and support Intel 64 architecture.
The Intel® Xeon® processor E3-1200 v3 product family and 4th Generation Intel® Core™ processors are based on
the Haswell microarchitecture and support Intel 64 architecture.
The Intel® Xeon® processor E5-2600/1600 v3 product families and the Intel® Core™ i7-59xx Processor Extreme
Edition are based on the Haswell-E microarchitecture and support Intel 64 architecture.
The Intel Atom® processor Z8000 series is based on the Airmont microarchitecture.
The Intel Atom® processor Z3400 series and the Intel Atom® processor Z3500 series are based on the Silvermont
microarchitecture.
Vol. 2A
1-3
ABOUT THIS MANUAL
The Intel® Core™ M processor family, 5th generation Intel® Core™ processors, Intel® Xeon® processor D-1500
product family and the Intel® Xeon® processor E5 v4 family are based on the Broadwell microarchitecture and
support Intel 64 architecture.
The Intel® Xeon® Scalable Processor Family, Intel® Xeon® processor E3-1500m v5 product family and 6th gener-
ation Intel® Core™ processors are based on the Skylake microarchitecture and support Intel 64 architecture.
The 7th generation Intel® Core™ processors are based on the Kaby Lake microarchitecture and support Intel 64
architecture.
The Intel Atom® processor C series, the Intel Atom® processor X series, the Intel® Pentium® processor J series,
the Intel® Celeron® processor J series, and the Intel® Celeron® processor N series are based on the Goldmont
microarchitecture.
The Intel® Xeon Phi™ Processor 3200, 5200, 7200 Series is based on the Knights Landing microarchitecture and
supports Intel 64 architecture.
The Intel® Pentium® Silver processor series, the Intel® Celeron® processor J series, and the Intel® Celeron®
processor N series are based on the Goldmont Plus microarchitecture.
The 8th generation Intel® Core™ processors, 9th generation Intel® Core™ processors, and Intel® Xeon® E proces-
sors are based on the Coffee Lake microarchitecture and support Intel 64 architecture.
The Intel® Xeon Phi™ Processor 7215, 7285, 7295 Series is based on the Knights Mill microarchitecture and
supports Intel 64 architecture.
The 2nd generation Intel® Xeon® Scalable Processor Family is based on the Cascade Lake product and supports
Intel 64 architecture.
Some 10th generation Intel® Core™ processors are based on the Ice Lake microarchitecture, and some are based
on the Comet Lake microarchitecture; both support Intel 64 architecture.
Some 11th generation Intel® Core™ processors are based on the Tiger Lake microarchitecture, and some are
based on the Rocket Lake microarchitecture; both support Intel 64 architecture.
Some 3rd generation Intel® Xeon® Scalable Processor Family processors are based on the Cooper Lake product,
and some are based on the Ice Lake microarchitecture; both support Intel 64 architecture.
The 12th generation Intel® Core™ processors are based on the Alder Lake performance hybrid architecture and
support Intel 64 architecture.
The 13th generation Intel® Core™ processors are based on the Raptor Lake performance hybrid architecture and
support Intel 64 architecture.
The 4th generation Intel® Xeon® Scalable Processor Family is based on Sapphire Rapids microarchitecture and
supports Intel 64 architecture.
IA-32 architecture is the instruction set architecture and programming environment for Intel's 32-bit microproces-
sors. Intel® 64 architecture is the instruction set architecture and programming environment which is the superset
of Intel’s 32-bit and 64-bit architectures. It is compatible with the IA-32 architecture.
1.2
OVERVIEW OF VOLUME 2A, 2B, 2C, AND 2D: INSTRUCTION SET
REFERENCE
A description of Intel® 64 and IA-32 Architectures Software Developer’s Manual, Volumes 2A, 2B, 2C, & 2D,
content follows:
Chapter 1 - About This Manual. Gives an overview of all ten volumes of the Intel® 64 and IA-32 Architectures
Software Developer’s Manual. It also describes the notational conventions in these manuals and lists related Intel®
manuals and documentation of interest to programmers and hardware designers.
Chapter 2 - Instruction Format. Describes the machine-level instruction format used for all IA-32 instructions
and gives the allowable encodings of prefixes, the operand-identifier byte (ModR/M byte), the addressing-mode
specifier byte (SIB byte), and the displacement and immediate bytes.
Chapter 3 - Instruction Set Reference, A-L. Describes Intel 64 and IA-32 instructions in detail, including an
algorithmic description of operations, the effect on flags, the effect of operand- and address-size attributes, and
1-4
Vol. 2A
ABOUT THIS MANUAL
the exceptions that may be generated. The instructions are arranged in alphabetical order. General-purpose, x87
FPU, Intel MMX™ technology, SSE/SSE2/SSE3/SSSE3/SSE4 extensions, and system instructions are included.
Chapter 4 - Instruction Set Reference, M-U. Continues the description of Intel 64 and IA-32 instructions
started in Chapter 3. It starts Intel® 64 and IA-32 Architectures Software Developer’s Manual, Volume 2B.
Chapter 5 - Instruction Set Reference, V. Continues the description of Intel 64 and IA-32 instructions started
in chapters 3 and 4. This chapter starts Intel® 64 and IA-32 Architectures Software Developer’s Manual, Volume
2C.
Chapter 6 - Instruction Set Reference, W-Z. Continues the description of Intel 64 and IA-32 instructions
started in chapters 3, 4, and 5. It provides the balance of the alphabetized list of instructions and starts Intel® 64
and IA-32 Architectures Software Developer’s Manual, Volume 2D.
Chapter 7 - Safer Mode Extensions Reference. Describes the safer mode extensions (SMX). SMX is intended
for a system executive to support launching a measured environment in a platform where the identity of the soft-
ware controlling the platform hardware can be measured for the purpose of making trust decisions.
Chapter 8- Instruction Set Reference Unique to Intel® Xeon Phi™ Processors. Describes the instruction
set that is unique to Intel® Xeon Phi™ processors based on the Knights Landing and Knights Mill microarchitec-
tures. The set is not supported in any other Intel processors.
Appendix A - Opcode Map. Gives an opcode map for the IA-32 instruction set.
Appendix B - Instruction Formats and Encodings. Gives the binary encoding of each form of each IA-32
instruction.
Appendix C - Intel® C/C++ Compiler Intrinsics and Functional Equivalents. Lists the Intel® C/C++ compiler
intrinsics and their assembly code equivalents for each of the IA-32 MMX and SSE/SSE2/SSE3 instructions.
1.3
NOTATIONAL CONVENTIONS
This manual uses specific notation for data-structure formats, for symbolic representation of instructions, and for
hexadecimal and binary numbers. A review of this notation makes the manual easier to read.
1.3.1
Bit and Byte Order
In illustrations of data structures in memory, smaller addresses appear toward the bottom of the figure; addresses
increase toward the top. Bit positions are numbered from right to left. The numerical value of a set bit is equal to
two raised to the power of the bit position. IA-32 processors are “little endian” machines; this means the bytes of
a word are numbered starting from the least significant byte. Figure 1-1 illustrates these conventions.
Highest
Data Structure
Address
31
24 23
16 15
8 7
0
Bit offset
28
24
20
16
12
8
4
Lowest
Byte 3
Byte 2
Byte 1
Byte 0
0
Address
Byte Offset
Figure 1-1. Bit and Byte Order
Vol. 2A
1-5
ABOUT THIS MANUAL
1.3.2
Reserved Bits and Software Compatibility
In many register and memory layout descriptions, certain bits are marked as reserved. When bits are marked as
reserved, it is essential for compatibility with future processors that software treat these bits as having a future,
though unknown, effect. The behavior of reserved bits should be regarded as not only undefined, but unpredict-
able. Software should follow these guidelines in dealing with reserved bits:
Do not depend on the states of any reserved bits when testing the values of registers which contain such bits.
Mask out the reserved bits before testing.
Do not depend on the states of any reserved bits when storing to memory or to a register.
Do not depend on the ability to retain information written into any reserved bits.
When loading a register, always load the reserved bits with the values indicated in the documentation, if any, or
reload them with values previously read from the same register.
NOTE
Avoid any software dependence upon the state of reserved bits in IA-32 registers. Depending upon
the values of reserved register bits will make software dependent upon the unspecified manner in
which the processor handles these bits. Programs that depend upon reserved values risk incompat-
ibility with future processors.
1.3.3
Instruction Operands
When instructions are represented symbolically, a subset of the IA-32 assembly language is used. In this subset,
an instruction has the following format:
label: mnemonic argument1, argument2, argument3
where:
A label is an identifier which is followed by a colon.
A mnemonic is a reserved name for a class of instruction opcodes which have the same function.
The operands argument1, argument2, and argument3 are optional. There may be from zero to three operands,
depending on the opcode. When present, they take the form of either literals or identifiers for data items.
Operand identifiers are either reserved names of registers or are assumed to be assigned to data items
declared in another part of the program (which may not be shown in the example).
When two operands are present in an arithmetic or logical instruction, the right operand is the source and the left
operand is the destination.
For example:
LOADREG: MOV EAX, SUBTOTAL
In this example, LOADREG is a label, MOV is the mnemonic identifier of an opcode, EAX is the destination operand,
and SUBTOTAL is the source operand. Some assembly languages put the source and destination in reverse order.
1.3.4
Hexadecimal and Binary Numbers
Base 16 (hexadecimal) numbers are represented by a string of hexadecimal digits followed by the character H (for
example, F82EH). A hexadecimal digit is a character from the following set: 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, A, B, C, D,
E, and F.
Base 2 (binary) numbers are represented by a string of 1s and 0s, sometimes followed by the character B (for
example, 1010B). The “B” designation is only used in situations where confusion as to the type of number might
arise.
1-6
Vol. 2A
ABOUT THIS MANUAL
1.3.5
Segmented Addressing
The processor uses byte addressing. This means memory is organized and accessed as a sequence of bytes.
Whether one or more bytes are being accessed, a byte address is used to locate the byte or bytes in memory. The
range of memory that can be addressed is called an address space.
The processor also supports segmented addressing. This is a form of addressing where a program may have many
independent address spaces, called segments. For example, a program can keep its code (instructions) and stack
in separate segments. Code addresses would always refer to the code space, and stack addresses would always
refer to the stack space. The following notation is used to specify a byte address within a segment:
Segment-register:Byte-address
For example, the following segment address identifies the byte at address FF79H in the segment pointed by the DS
register:
DS:FF79H
The following segment address identifies an instruction address in the code segment. The CS register points to the
code segment and the EIP register contains the address of the instruction.
CS:EIP
1.3.6
Exceptions
An exception is an event that typically occurs when an instruction causes an error. For example, an attempt to
divide by zero generates an exception. However, some exceptions, such as breakpoints, occur under other condi-
tions. Some types of exceptions may provide error codes. An error code reports additional information about the
error. An example of the notation used to show an exception and error code is shown below:
#PF(fault code)
This example refers to a page-fault exception under conditions where an error code naming a type of fault is
reported. Under some conditions, exceptions which produce error codes may not be able to report an accurate
code. In this case, the error code is zero, as shown below for a general-protection exception:
#GP(0)
1.3.7
A New Syntax for CPUID, CR, and MSR Values
Obtain feature flags, status, and system information by using the CPUID instruction, by checking control register
bits, and by reading model-specific registers. We are moving toward a new syntax to represent this information.
See Figure 1-2.
Vol. 2A
1-7
ABOUT THIS MANUAL
CPUID Input and Output
CPUID.01H:EDX.SSE[bit 25] = 1
Input value for EAX register
Output register and feature flag or field
name with bit position(s)
Value (or range) of output
Control Register Values
CR4.OSFXSR[bit 9] = 1
Example CR name
Feature flag or field name
with bit position(s)
Value (or range) of output
Model-Specific Register Values
IA32_MISC_ENABLE.ENABLEFOPCODE[bit 2] = 1
Example MSR name
Feature flag or field name with bit position(s)
Value (or range) of output
SDM29002
Figure 1-2. Syntax for CPUID, CR, and MSR Data Presentation
1.4
RELATED LITERATURE
Literature related to Intel 64 and IA-32 processors is listed and viewable on-line at:
See also:
The latest security information on Intel® products:
Software developer resources, guidance, and insights for security advisories:
The data sheet for a particular Intel 64 or IA-32 processor
The specification update for a particular Intel 64 or IA-32 processor
Intel® C++ Compiler documentation and online help:
1-8
Vol. 2A
ABOUT THIS MANUAL
Intel® Fortran Compiler documentation and online help:
Intel® Software Development Tools:
Intel® 64 and IA-32 Architectures Software Developer’s Manual (in one, four or ten volumes):
Intel® 64 and IA-32 Architectures Optimization Reference Manual:
Intel® Trusted Execution Technology Measured Launched Environment Programming Guide:
Intel® Software Guard Extensions (Intel® SGX) Information
Developing Multi-threaded Applications: A Platform Consistent Approach:
tions.pdf
Using Spin-Loops on Intel® Pentium® 4 Processor and Intel® Xeon® Processor:
Performance Monitoring Unit Sharing Guide
Literature related to select features in future Intel processors are available at:
Intel® Architecture Instruction Set Extensions Programming Reference
More relevant links are:
Intel® Developer Zone:
Developer centers:
Processor support general link:
Intel® Hyper-Threading Technology (Intel® HT Technology):
Vol. 2A
1-9
ABOUT THIS MANUAL
1-10
Vol. 2A
CHAPTER 2
INSTRUCTION FORMAT
This chapter describes the instruction format for all Intel 64 and IA-32 processors. The instruction format for
protected mode, real-address mode and virtual-8086 mode is described in Section 2.1. Increments provided for IA-
32e mode and its sub-modes are described in Section 2.2.
2.1
INSTRUCTION FORMAT FOR PROTECTED MODE, REAL-ADDRESS MODE,
AND VIRTUAL-8086 MODE
The Intel 64 and IA-32 architectures instruction encodings are subsets of the format shown in Figure 2-1. Instruc-
tions consist of optional instruction prefixes (in any order), primary opcode bytes (up to three bytes), an
addressing-form specifier (if required) consisting of the ModR/M byte and sometimes the SIB (Scale-Index-Base)
byte, a displacement (if required), and an immediate data field (if required).
Instruction
Opcode
ModR/M
SIB
Displacement
Immediate
Prefixes
Prefixes of
1-, 2-, or 3-byte
1 byte
1 byte
Address
Immediate
1 byte each
opcode
(if required)
(if required)
displacement
data of
(optional)1, 2
of 1, 2, or 4
1, 2, or 4
bytes or none3
bytes or none3
7
6 5
3
2
0
7
6 5
3
2
0
Reg/
Mod
R/M
Scale
Index
Base
Opcode
1. The REX prefix is optional, but if used must be immediately before the opcode; see Section
2.2.1, “REX Prefixes” for additional information.
2. For VEX encoding information, see Section 2.3, “Intel® Advanced Vector Extensions (Intel®
AVX)”.
3. Some rare instructions can take an 8B immediate or 8B displacement.
Figure 2-1. Intel 64 and IA-32 Architectures Instruction Format
2.1.1
Instruction Prefixes
Instruction prefixes are divided into four groups, each with a set of allowable prefix codes. For each instruction, it
is only useful to include up to one prefix code from each of the four groups (Groups 1, 2, 3, 4). Groups 1 through 4
may be placed in any order relative to each other.
Group 1
- Lock and repeat prefixes:
LOCK prefix is encoded using F0H.
REPNE/REPNZ prefix is encoded using F2H. Repeat-Not-Zero prefix applies only to string and
input/output instructions. (F2H is also used as a mandatory prefix for some instructions.)
REP or REPE/REPZ is encoded using F3H. The repeat prefix applies only to string and input/output
instructions. (F3H is also used as a mandatory prefix for some instructions.)
Vol. 2A
2-1
INSTRUCTION FORMAT
- BND prefix is encoded using F2H if the following conditions are true:
CPUID.(EAX=07H, ECX=0):EBX.MPX[bit 14] is set.
BNDCFGU.EN and/or IA32_BNDCFGS.EN is set.
When the F2 prefix precedes a near CALL, a near RET, a near JMP, a short Jcc, or a near Jcc instruction
(see Appendix E, “Intel® Memory Protection Extensions,” of the Intel® 64 and IA-32 Architectures
Software Developer’s Manual, Volume 1).
Group 2
- Segment override prefixes:
2EH-CS segment override (use with any branch instruction is reserved).
36H-SS segment override prefix (use with any branch instruction is reserved).
3EH-DS segment override prefix (use with any branch instruction is reserved).
26H-ES segment override prefix (use with any branch instruction is reserved).
64H-FS segment override prefix (use with any branch instruction is reserved).
65H-GS segment override prefix (use with any branch instruction is reserved).
— Branch hints1:
2EH-Branch not taken (used only with Jcc instructions).
3EH-Branch taken (used only with Jcc instructions).
Group 3
Operand-size override prefix is encoded using 66H (66H is also used as a mandatory prefix for some
instructions).
Group 4
67H-Address-size override prefix.
The LOCK prefix (F0H) forces an operation that ensures exclusive use of shared memory in a multiprocessor envi-
ronment. See “LOCK-Assert LOCK# Signal Prefix” in Chapter 3, “Instruction Set Reference, A-L,” for a description
of this prefix.
Repeat prefixes (F2H, F3H) cause an instruction to be repeated for each element of a string. Use these prefixes
only with string and I/O instructions (MOVS, CMPS, SCAS, LODS, STOS, INS, and OUTS). Use of repeat prefixes
and/or undefined opcodes with other Intel 64 or IA-32 instructions is reserved; such use may cause unpredictable
behavior.
Some instructions may use F2H,F3H as a mandatory prefix to express distinct functionality.
Branch hint prefixes (2EH, 3EH) allow a program to give a hint to the processor about the most likely code path for
a branch. Use these prefixes only with conditional branch instructions (Jcc). Other use of branch hint prefixes
and/or other undefined opcodes with Intel 64 or IA-32 instructions is reserved; such use may cause unpredictable
behavior.
The operand-size override prefix allows a program to switch between 16- and 32-bit operand sizes. Either size can
be the default; use of the prefix selects the non-default size.
Some SSE2/SSE3/SSSE3/SSE4 instructions and instructions using a three-byte sequence of primary opcode bytes
may use 66H as a mandatory prefix to express distinct functionality.
Other use of the 66H prefix is reserved; such use may cause unpredictable behavior.
The address-size override prefix (67H) allows programs to switch between 16- and 32-bit addressing. Either size
can be the default; the prefix selects the non-default size. Using this prefix and/or other undefined opcodes when
operands for the instruction do not reside in memory is reserved; such use may cause unpredictable behavior.
1. Some earlier microarchitectures used these as branch hints, but recent generations have not and they are reserved for future hint
usage.
2-2
Vol. 2A
INSTRUCTION FORMAT
2.1.2
Opcodes
A primary opcode can be 1, 2, or 3 bytes in length. An additional 3-bit opcode field is sometimes encoded in the
ModR/M byte. Smaller fields can be defined within the primary opcode. Such fields define the direction of opera-
tion, size of displacements, register encoding, condition codes, or sign extension. Encoding fields used by an
opcode vary depending on the class of operation.
Two-byte opcode formats for general-purpose and SIMD instructions consist of one of the following:
An escape opcode byte 0FH as the primary opcode and a second opcode byte.
A mandatory prefix (66H, F2H, or F3H), an escape opcode byte, and a second opcode byte (same as previous
bullet).
For example, CVTDQ2PD consists of the following sequence: F3 0F E6. The first byte is a mandatory prefix (it is not
considered as a repeat prefix).
Three-byte opcode formats for general-purpose and SIMD instructions consist of one of the following:
An escape opcode byte 0FH as the primary opcode, plus two additional opcode bytes.
A mandatory prefix (66H, F2H, or F3H), an escape opcode byte, plus two additional opcode bytes (same as
previous bullet).
For example, PHADDW for XMM registers consists of the following sequence: 66 0F 38 01. The first byte is the
mandatory prefix.
Valid opcode expressions are defined in Appendix A and Appendix B.
2.1.3
ModR/M and SIB Bytes
Many instructions that refer to an operand in memory have an addressing-form specifier byte (called the ModR/M
byte) following the primary opcode. The ModR/M byte contains three fields of information:
The mod field combines with the r/m field to form 32 possible values: eight registers and 24 addressing modes.
The reg/opcode field specifies either a register number or three more bits of opcode information. The purpose
of the reg/opcode field is specified in the primary opcode.
The r/m field can specify a register as an operand or it can be combined with the mod field to encode an
addressing mode. Sometimes, certain combinations of the mod field and the r/m field are used to express
opcode information for some instructions.
Certain encodings of the ModR/M byte require a second addressing byte (the SIB byte). The base-plus-index and
scale-plus-index forms of 32-bit addressing require the SIB byte. The SIB byte includes the following fields:
The scale field specifies the scale factor.
The index field specifies the register number of the index register.
The base field specifies the register number of the base register.
See Section 2.1.5 for the encodings of the ModR/M and SIB bytes.
2.1.4
Displacement and Immediate Bytes
Some addressing forms include a displacement immediately following the ModR/M byte (or the SIB byte if one is
present). If a displacement is required, it can be 1, 2, or 4 bytes.
If an instruction specifies an immediate operand, the operand always follows any displacement bytes. An imme-
diate operand can be 1, 2 or 4 bytes.
Vol. 2A
2-3
INSTRUCTION FORMAT
2.1.5
Addressing-Mode Encoding of ModR/M and SIB Bytes
The values and corresponding addressing forms of the ModR/M and SIB bytes are shown in Table 2-1 through Table
2-3: 16-bit addressing forms specified by the ModR/M byte are in Table 2-1 and 32-bit addressing forms are in
Table 2-2. Table 2-3 shows 32-bit addressing forms specified by the SIB byte. In cases where the reg/opcode field
in the ModR/M byte represents an extended opcode, valid encodings are shown in Appendix B.
In Table 2-1 and Table 2-2, the Effective Address column lists 32 effective addresses that can be assigned to the
first operand of an instruction by using the Mod and R/M fields of the ModR/M byte. The first 24 options provide
ways of specifying a memory location; the last eight (Mod = 11B) provide ways of specifying general-purpose, MMX
technology and XMM registers.
The Mod and R/M columns in Table 2-1 and Table 2-2 give the binary encodings of the Mod and R/M fields required
to obtain the effective address listed in the first column. For example: see the row indicated by Mod = 11B, R/M =
000B. The row identifies the general-purpose registers EAX, AX or AL; MMX technology register MM0; or XMM
register XMM0. The register used is determined by the opcode byte and the operand-size attribute.
Now look at the seventh row in either table (labeled “REG =”). This row specifies the use of the 3-bit Reg/Opcode
field when the field is used to give the location of a second operand. The second operand must be a general-
purpose, MMX technology, or XMM register. Rows one through five list the registers that may correspond to the
value in the table. Again, the register used is determined by the opcode byte along with the operand-size attribute.
If the instruction does not require a second operand, then the Reg/Opcode field may be used as an opcode exten-
sion. This use is represented by the sixth row in the tables (labeled “/digit (Opcode)”). Note that values in row six
are represented in decimal form.
The body of Table 2-1 and Table 2-2 (under the label “Value of ModR/M Byte (in Hexadecimal)”) contains a 32 by
8 array that presents all of 256 values of the ModR/M byte (in hexadecimal). Bits 3, 4, and 5 are specified by the
column of the table in which a byte resides. The row specifies bits 0, 1, and 2; and bits 6 and 7. The figure below
demonstrates interpretation of one table value.
Mod
11
RM
000
/digit (Opcode); REG =
001
C8H
11001000
Figure 2-2. Table Interpretation of ModR/M Byte (C8H)
2-4
Vol. 2A
INSTRUCTION FORMAT
Table 2-1. 16-Bit Addressing Forms with the ModR/M Byte
r8(/r)
AL
CL
DL
BL
AH
CH
DH
BH
r16(/r)
AX
CX
DX
BX
SP
BP1
SI
DI
r32(/r)
EAX
ECX
EDX
EBX
ESP
EBP
ESI
EDI
mm(/r)
MM0
MM1
MM2
MM3
MM4
MM5
MM6
MM7
xmm(/r)
XMM0
XMM1
XMM2
XMM3
XMM4
XMM5
XMM6
XMM7
(In decimal) /digit (Opcode)
0
1
2
3
4
5
6
7
(In binary) REG =
000
001
010
011
100
101
110
111
Effective Address
Mod
R/M
Value of ModR/M Byte (in Hexadecimal)
[BX+SI]
00
000
00
08
10
18
20
28
30
38
[BX+DI]
001
01
09
11
19
21
29
31
39
[BP+SI]
010
02
0A
12
1A
22
2A
32
3A
[BP+DI]
011
03
0B
13
1B
23
2B
33
3B
[SI]
100
04
0C
14
1C
24
2C
34
3C
[DI]
101
05
0D
15
1D
25
2D
35
3D
disp162
110
06
0E
16
1E
26
2E
36
3E
[BX]
111
07
0F
17
1F
27
2F
37
3F
[BX+SI]+disp83
01
000
40
48
50
58
60
68
70
78
[BX+DI]+disp8
001
41
49
51
59
61
69
71
79
[BP+SI]+disp8
010
42
4A
52
5A
62
6A
72
7A
[BP+DI]+disp8
011
43
4B
53
5B
63
6B
73
7B
[SI]+disp8
100
44
4C
54
5C
64
6C
74
7C
[DI]+disp8
101
45
4D
55
5D
65
6D
75
7D
[BP]+disp8
110
46
4E
56
5E
66
6E
76
7E
[BX]+disp8
111
47
4F
57
5F
67
6F
77
7F
[BX+SI]+disp16
10
000
80
88
90
98
A0
A8
B0
B8
[BX+DI]+disp16
001
81
89
91
99
A1
A9
B1
B9
[BP+SI]+disp16
010
82
8A
92
9A
A2
AA
B2
BA
[BP+DI]+disp16
011
83
8B
93
9B
A3
AB
B3
BB
[SI]+disp16
100
84
8C
94
9C
A4
AC
B4
BC
[DI]+disp16
101
85
8D
95
9D
A5
AD
B5
BD
[BP]+disp16
110
86
8E
96
9E
A6
AE
B6
BE
[BX]+disp16
111
87
8F
97
9F
A7
AF
B7
BF
EAX/AX/AL/MM0/XMM0
11
000
C0
C8
D0
D8
E0
E8
F0
F8
ECX/CX/CL/MM1/XMM1
001
C1
C9
D1
D9
E1
E9
F1
F9
EDX/DX/DL/MM2/XMM2
010
C2
CA
D2
DA
E2
EA
F2
FA
EBX/BX/BL/MM3/XMM3
011
C3
CB
D3
DB
E3
EB
F3
FB
ESP/SP/AHMM4/XMM4
100
C4
CC
D4
DC
E4
EC
F4
FC
EBP/BP/CH/MM5/XMM5
101
C5
CD
D5
DD
E5
ED
F5
FD
ESI/SI/DH/MM6/XMM6
110
C6
CE
D6
DE
E6
EE
F6
FE
EDI/DI/BH/MM7/XMM7
111
C7
CF
D7
DF
E7
EF
F7
FF
NOTES:
1. The default segment register is SS for the effective addresses containing a BP index, DS for other effective addresses.
2. The disp16 nomenclature denotes a 16-bit displacement that follows the ModR/M byte and that is added to the index.
3. The disp8 nomenclature denotes an 8-bit displacement that follows the ModR/M byte and that is sign-extended and added to the
index.
Vol. 2A
2-5
INSTRUCTION FORMAT
Table 2-2. 32-Bit Addressing Forms with the ModR/M Byte
r8(/r)
AL
CL
DL
BL
AH
CH
DH
BH
r16(/r)
AX
CX
DX
BX
SP
BP
SI
DI
r32(/r)
EAX
ECX
EDX
EBX
ESP
EBP
ESI
EDI
mm(/r)
MM0
MM1
MM2
MM3
MM4
MM5
MM6
MM7
xmm(/r)
XMM0
XMM1
XMM2
XMM3
XMM4
XMM5
XMM6
XMM7
(In decimal) /digit (Opcode)
0
1
2
3
4
5
6
7
(In binary) REG =
000
001
010
011
100
101
110
111
Effective Address
Mod
R/M
Value of ModR/M Byte (in Hexadecimal)
[EAX]
00
000
00
08
10
18
20
28
30
38
[ECX]
001
01
09
11
19
21
29
31
39
[EDX]
010
02
0A
12
1A
22
2A
32
3A
[EBX]
011
03
0B
13
1B
23
2B
33
3B
[--][--]1
100
04
0C
14
1C
24
2C
34
3C
disp322
101
05
0D
15
1D
25
2D
35
3D
[ESI]
110
06
0E
16
1E
26
2E
36
3E
[EDI]
111
07
0F
17
1F
27
2F
37
3F
[EAX]+disp83
01
000
40
48
50
58
60
68
70
78
[ECX]+disp8
001
41
49
51
59
61
69
71
79
[EDX]+disp8
010
42
4A
52
5A
62
6A
72
7A
[EBX]+disp8
011
43
4B
53
5B
63
6B
73
7B
[--][--]+disp8
100
44
4C
54
5C
64
6C
74
7C
[EBP]+disp8
101
45
4D
55
5D
65
6D
75
7D
[ESI]+disp8
110
46
4E
56
5E
66
6E
76
7E
[EDI]+disp8
111
47
4F
57
5F
67
6F
77
7F
[EAX]+disp32
10
000
80
88
90
98
A0
A8
B0
B8
[ECX]+disp32
001
81
89
91
99
A1
A9
B1
B9
[EDX]+disp32
010
82
8A
92
9A
A2
AA
B2
BA
[EBX]+disp32
011
83
8B
93
9B
A3
AB
B3
BB
[--][--]+disp32
100
84
8C
94
9C
A4
AC
B4
BC
[EBP]+disp32
101
85
8D
95
9D
A5
AD
B5
BD
[ESI]+disp32
110
86
8E
96
9E
A6
AE
B6
BE
[EDI]+disp32
111
87
8F
97
9F
A7
AF
B7
BF
EAX/AX/AL/MM0/XMM0
11
000
C0
C8
D0
D8
E0
E8
F0
F8
ECX/CX/CL/MM/XMM1
001
C1
C9
D1
D9
E1
E9
F1
F9
EDX/DX/DL/MM2/XMM2
010
C2
CA
D2
DA
E2
EA
F2
FA
EBX/BX/BL/MM3/XMM3
011
C3
CB
D3
DB
E3
EB
F3
FB
ESP/SP/AH/MM4/XMM4
100
C4
CC
D4
DC
E4
EC
F4
FC
EBP/BP/CH/MM5/XMM5
101
C5
CD
D5
DD
E5
ED
F5
FD
ESI/SI/DH/MM6/XMM6
110
C6
CE
D6
DE
E6
EE
F6
FE
EDI/DI/BH/MM7/XMM7
111
C7
CF
D7
DF
E7
EF
F7
FF
NOTES:
1. The [--][--] nomenclature means a SIB follows the ModR/M byte.
2. The disp32 nomenclature denotes a 32-bit displacement that follows the ModR/M byte (or the SIB byte if one is present) and that is
added to the index.
3. The disp8 nomenclature denotes an 8-bit displacement that follows the ModR/M byte (or the SIB byte if one is present) and that is
sign-extended and added to the index.
Table 2-3 is organized to give 256 possible values of the SIB byte (in hexadecimal). General purpose registers used
as a base are indicated across the top of the table, along with corresponding values for the SIB byte’s base field.
Table rows in the body of the table indicate the register used as the index (SIB byte bits 3, 4, and 5) and the scaling
factor (determined by SIB byte bits 6 and 7).
2-6
Vol. 2A
INSTRUCTION FORMAT
Table 2-3. 32-Bit Addressing Forms with the SIB Byte
r32
EAX
ECX
EDX
EBX
ESP
[*]
ESI
EDI
(In decimal) Base =
0
1
2
3
4
5
6
7
(In binary) Base =
000
001
010
011
100
101
110
111
Scaled Index
SS
Index
Value of SIB Byte (in Hexadecimal)
[EAX]
00
000
00
01
02
03
04
05
06
07
[ECX]
001
08
09
0A
0B
0C
0D
0E
0F
[EDX]
010
10
11
12
13
14
15
16
17
[EBX]
011
18
19
1A
1B
1C
1D
1E
1F
none
100
20
21
22
23
24
25
26
27
[EBP]
101
28
29
2A
2B
2C
2D
2E
2F
[ESI]
110
30
31
32
33
34
35
36
37
[EDI]
111
38
39
3A
3B
3C
3D
3E
3F
[EAX*2]
01
000
40
41
42
43
44
45
46
47
[ECX*2]
001
48
49
4A
4B
4C
4D
4E
4F
[EDX*2]
010
50
51
52
53
54
55
56
57
[EBX*2]
011
58
59
5A
5B
5C
5D
5E
5F
none
100
60
61
62
63
64
65
66
67
[EBP*2]
101
68
69
6A
6B
6C
6D
6E
6F
[ESI*2]
110
70
71
72
73
74
75
76
77
[EDI*2]
111
78
79
7A
7B
7C
7D
7E
7F
[EAX*4]
10
000
80
81
82
83
84
85
86
87
[ECX*4]
001
88
89
8A
8B
8C
8D
8E
8F
[EDX*4]
010
90
91
92
93
94
95
96
97
[EBX*4]
011
98
99
9A
9B
9C
9D
9E
9F
none
100
A0
A1
A2
A3
A4
A5
A6
A7
[EBP*4]
101
A8
A9
AA
AB
AC
AD
AE
AF
[ESI*4]
110
B0
B1
B2
B3
B4
B5
B6
B7
[EDI*4]
111
B8
B9
BA
BB
BC
BD
BE
BF
[EAX*8]
11
000
C0
C1
C2
C3
C4
C5
C6
C7
[ECX*8]
001
C8
C9
CA
CB
CC
CD
CE
CF
[EDX*8]
010
D0
D1
D2
D3
D4
D5
D6
D7
[EBX*8]
011
D8
D9
DA
DB
DC
DD
DE
DF
none
100
E0
E1
E2
E3
E4
E5
E6
E7
[EBP*8]
101
E8
E9
EA
EB
EC
ED
EE
EF
[ESI*8]
110
F0
F1
F2
F3
F4
F5
F6
F7
[EDI*8]
111
F8
F9
FA
FB
FC
FD
FE
FF
NOTES:
1. The [*] nomenclature means a disp32 with no base if the MOD is 00B. Otherwise, [*] means disp8 or disp32 + [EBP]. This provides the
following address modes:
MOD bits Effective Address
00
[scaled index] + disp32
01
[scaled index] + disp8 + [EBP]
10
[scaled index] + disp32 + [EBP]
2.2
IA-32E MODE
IA-32e mode has two sub-modes. These are:
Compatibility Mode. Enables a 64-bit operating system to run most legacy protected mode software
unmodified.
64-Bit Mode. Enables a 64-bit operating system to run applications written to access 64-bit address space.
2.2.1
REX Prefixes
REX prefixes are instruction-prefix bytes used in 64-bit mode. They do the following:
Specify GPRs and SSE registers.
Vol. 2A
2-7
INSTRUCTION FORMAT
Specify 64-bit operand size.
Specify extended control registers.
Not all instructions require a REX prefix in 64-bit mode. A prefix is necessary only if an instruction references one
of the extended registers or uses a 64-bit operand. If a REX prefix is used when it has no meaning, it is ignored.
Only one REX prefix is allowed per instruction. If used, the REX prefix byte must immediately precede the opcode
byte or the escape opcode byte (0FH). When a REX prefix is used in conjunction with an instruction containing a
mandatory prefix, the mandatory prefix must come before the REX so the REX prefix can be immediately preceding
the opcode or the escape byte. For example, CVTDQ2PD with a REX prefix should have REX placed between F3 and
0F E6. Other placements are ignored. The instruction-size limit of 15 bytes still applies to instructions with a REX
prefix. See Figure 2-3.
Legacy
REX
Opcode
ModR/M
SIB
Displacement
Immediate
Prefixes
Prefix
(optional)
1-, 2-, or
1 byte
Address
Immediate data
1 byte
Grp 1,
Grp
3-byte
(if required)
displacement of
of 1, 2, or 4
(if required)
2, Grp 3,
Grp 4
opcode
1, 2, or 4 bytes
bytes or none
(optional)
Figure 2-3. Prefix Ordering in 64-bit Mode
2.2.1.1
Encoding
Intel 64 and IA-32 instruction formats specify up to three registers by using 3-bit fields in the encoding, depending
on the format:
ModR/M: the reg and r/m fields of the ModR/M byte.
ModR/M with SIB: the reg field of the ModR/M byte, the base and index fields of the SIB (scale, index, base)
byte.
Instructions without ModR/M: the reg field of the opcode.
In 64-bit mode, these formats do not change. Bits needed to define fields in the 64-bit context are provided by the
addition of REX prefixes.
2.2.1.2
More on REX Prefix Fields
REX prefixes are a set of 16 opcodes that span one row of the opcode map and occupy entries 40H to 4FH. These
opcodes represent valid instructions (INC or DEC) in IA-32 operating modes and in compatibility mode. In 64-bit
mode, the same opcodes represent the instruction prefix REX and are not treated as individual instructions.
The single-byte-opcode forms of the INC/DEC instructions are not available in 64-bit mode. INC/DEC functionality
is still available using ModR/M forms of the same instructions (opcodes FF/0 and FF/1).
See Table 2-4 for a summary of the REX prefix format. Figure 2-4 though Figure 2-7 show examples of REX prefix
fields in use. Some combinations of REX prefix fields are invalid. In such cases, the prefix is ignored. Some addi-
tional information follows:
Setting REX.W can be used to determine the operand size but does not solely determine operand width. Like
the 66H size prefix, 64-bit operand size override has no effect on byte-specific operations.
For non-byte operations: if a 66H prefix is used with prefix (REX.W = 1), 66H is ignored.
If a 66H override is used with REX and REX.W = 0, the operand size is 16 bits.
REX.R modifies the ModR/M reg field when that field encodes a GPR, SSE, control or debug register. REX.R is
ignored when ModR/M specifies other registers or defines an extended opcode.
REX.X bit modifies the SIB index field.
2-8
Vol. 2A
INSTRUCTION FORMAT
REX.B either modifies the base in the ModR/M r/m field or SIB base field; or it modifies the opcode reg field
used for accessing GPRs.
Table 2-4. REX Prefix Fields [BITS: 0100WRXB]
Field Name
Bit Position
Definition
-
7:4
0100
W
3
0 = Operand size determined by CS.D
1 = 64 Bit Operand Size
R
2
Extension of the ModR/M reg field
X
1
Extension of the SIB index field
B
0
Extension of the ModR/M r/m field, SIB base field, or Opcode reg field
ModRM Byte
REX PREFIX
Opcode
mod
reg
r/m
0100WR0B
11
rrr
bbb
Rrrr
Bbbb
OM17Xfig1-3
Figure 2-4. Memory Addressing Without an SIB Byte; REX.X Not Used
ModRM Byte
REX PREFIX
Opcode
mod
reg
r/m
0100WR0B
11
rrr
bbb
Rrrr
Bbbb
OM17Xfig1-4
Figure 2-5. Register-Register Addressing (No Memory Operand); REX.X Not Used
Vol. 2A
2-9
INSTRUCTION FORMAT
ModRM Byte
SIB Byte
REX PREFIX
Opcode
mod
reg
r/m
scale
index
base
0100WRXB
11
rrr
100
ss
xxx
bbb
Rrrr
Xxxx
Bbbb
OM17Xfig1-5
Figure 2-6. Memory Addressing With a SIB Byte
REX PREFIX
Opcode
reg
0100W00B
bbb
Bbbb
OM17Xfig1-6
Figure 2-7. Register Operand Coded in Opcode Byte; REX.X & REX.R Not Used
In the IA-32 architecture, byte registers (AH, AL, BH, BL, CH, CL, DH, and DL) are encoded in the ModR/M byte’s
reg field, the r/m field or the opcode reg field as registers 0 through 7. REX prefixes provide an additional
addressing capability for byte-registers that makes the least-significant byte of GPRs available for byte operations.
Certain combinations of the fields of the ModR/M byte and the SIB byte have special meaning for register encod-
ings. For some combinations, fields expanded by the REX prefix are not decoded. Table 2-5 describes how each
case behaves.
2-10
Vol. 2A
INSTRUCTION FORMAT
Table 2-5. Special Cases of REX Encodings
ModR/M or
Sub-field
Compatibility Mode
Compatibility Mode
SIB
Encodings
Operation
Implications
Additional Implications
ModR/M Byte
mod ? 11
SIB byte present.
SIB byte required for
REX prefix adds a fourth bit (b) which is not decoded
ESP-based addressing.
(don't care).
r/m =
b*100(ESP)
SIB byte also required for R12-based addressing.
ModR/M Byte
mod = 0
Base register not
EBP without a
REX prefix adds a fourth bit (b) which is not decoded
used.
displacement must be
(don't care).
r/m =
done using
b*101(EBP)
Using RBP or R13 without displacement must be done
mod = 01 with
using mod = 01 with a displacement of 0.
displacement of 0.
SIB Byte
index =
Index register not
ESP cannot be used as
REX prefix adds a fourth bit (b) which is decoded.
0100(ESP)
used.
an index register.
There are no additional implications. The expanded
index field allows distinguishing RSP from R12,
therefore R12 can be used as an index.
SIB Byte
base =
Base register is
Base register depends
REX prefix adds a fourth bit (b) which is not decoded.
0101(EBP)
unused if mod = 0.
on mod encoding.
This requires explicit displacement to be used with
EBP/RBP or R13.
NOTES:
* Don’t care about value of REX.B
2.2.1.3
Displacement
Addressing in 64-bit mode uses existing 32-bit ModR/M and SIB encodings. The ModR/M and SIB displacement
sizes do not change. They remain 8 bits or 32 bits and are sign-extended to 64 bits.
2.2.1.4
Direct Memory-Offset MOVs
In 64-bit mode, direct memory-offset forms of the MOV instruction are extended to specify a 64-bit immediate
absolute address. This address is called a moffset. No prefix is needed to specify this 64-bit memory offset. For
these MOV instructions, the size of the memory offset follows the address-size default (64 bits in 64-bit mode). See
Table 2-6.
Table 2-6. Direct Memory Offset Form of MOV
Opcode
Instruction
A0
MOV AL, moffset
A1
MOV EAX, moffset
A2
MOV moffset, AL
A3
MOV moffset, EAX
2.2.1.5
Immediates
In 64-bit mode, the typical size of immediate operands remains 32 bits. When the operand size is 64 bits, the
processor sign-extends all immediates to 64 bits prior to their use.
Support for 64-bit immediate operands is accomplished by expanding the semantics of the existing move (MOV
reg, imm16/32) instructions. These instructions (opcodes B8H - BFH) move 16-bits or 32-bits of immediate data
(depending on the effective operand size) into a GPR. When the effective operand size is 64 bits, these instructions
can be used to load an immediate into a GPR. A REX prefix is needed to override the 32-bit default operand size to
a 64-bit operand size.
For example:
48 B8
8877665544332211 MOV RAX,1122334455667788H
Vol. 2A
2-11
INSTRUCTION FORMAT
2.2.1.6
RIP-Relative Addressing
A new addressing form, RIP-relative (relative instruction-pointer) addressing, is implemented in 64-bit mode. An
effective address is formed by adding displacement to the 64-bit RIP of the next instruction.
In IA-32 architecture and compatibility mode, addressing relative to the instruction pointer is available only with
control-transfer instructions. In 64-bit mode, instructions that use ModR/M addressing can use RIP-relative
addressing. Without RIP-relative addressing, all ModR/M modes address memory relative to zero.
RIP-relative addressing allows specific ModR/M modes to address memory relative to the 64-bit RIP using a signed
32-bit displacement. This provides an offset range of ±2GB from the RIP. Table 2-7 shows the ModR/M and SIB
encodings for RIP-relative addressing. Redundant forms of 32-bit displacement-addressing exist in the current
ModR/M and SIB encodings. There is one ModR/M encoding and there are several SIB encodings. RIP-relative
addressing is encoded using a redundant form.
In 64-bit mode, the ModR/M Disp32 (32-bit displacement) encoding is re-defined to be RIP+Disp32 rather than
displacement-only. See Table 2-7.
Table 2-7. RIP-Relative Addressing
ModR/M and SIB Sub-field Encodings
Compatibility Mode
64-bit Mode
Additional Implications in 64-bit mode
Operation
Operation
ModR/M Byte
mod = 00
Disp32
RIP + Disp32
In 64-bit mode, if one wants to use a Disp32
without specifying a base register, one can use a
r/m = 101 (none)
SIB byte encoding (indicated by ModR/M.r/m=100)
as described in the next row.
SIB Byte
base = 101 (none)
If mod = 00, Disp32
Same as legacy
None
index = 100 (none)
scale = 0, 1, 2, 4
The ModR/M encoding for RIP-relative addressing does not depend on using a prefix. Specifically, the r/m bit field
encoding of 101B (used to select RIP-relative addressing) is not affected by the REX prefix. For example, selecting
R13 (REX.B = 1, r/m = 101B) with mod = 00B still results in RIP-relative addressing. The 4-bit r/m field of REX.B
combined with ModR/M is not fully decoded. In order to address R13 with no displacement, software must encode
R13 + 0 using a 1-byte displacement of zero.
RIP-relative addressing is enabled by 64-bit mode, not by a 64-bit address-size. The use of the address-size prefix
does not disable RIP-relative addressing. The effect of the address-size prefix is to truncate and zero-extend the
computed effective address to 32 bits.
2.2.1.7
Default 64-Bit Operand Size
In 64-bit mode, two groups of instructions have a default operand size of 64 bits (do not need a REX prefix for this
operand size). These are:
Near branches.
All instructions, except far branches, that implicitly reference the RSP.
2.2.2
Additional Encodings for Control and Debug Registers
In 64-bit mode, more encodings for control and debug registers are available. The REX.R bit is used to modify the
ModR/M reg field when that field encodes a control or debug register (see Table 2-4). These encodings enable the
processor to address CR8-CR15 and DR8- DR15. An additional control register (CR8) is defined in 64-bit mode. CR8
becomes the Task Priority Register (TPR).
In the first implementation of IA-32e mode, CR9-CR15 and DR8-DR15 are not implemented. Any attempt to access
unimplemented registers results in an invalid-opcode exception (#UD).
2-12
Vol. 2A
INSTRUCTION FORMAT
2.3
INTEL® ADVANCED VECTOR EXTENSIONS (INTEL® AVX)
Intel AVX instructions are encoded using an encoding scheme that combines prefix bytes, opcode extension field,
operand encoding fields, and vector length encoding capability into a new prefix, referred to as VEX. In the VEX
encoding scheme, the VEX prefix may be two or three bytes long, depending on the instruction semantics. Despite
the two-byte or three-byte length of the VEX prefix, the VEX encoding format provides a more compact represen-
tation/packing of the components of encoding an instruction in Intel 64 architecture. The VEX encoding scheme
also allows more headroom for future growth of Intel 64 architecture.
2.3.1
Instruction Format
Instruction encoding using VEX prefix provides several advantages:
Instruction syntax support for three operands and up-to four operands when necessary. For example, the third
source register used by VBLENDVPD is encoded using bits 7:4 of the immediate byte.
Encoding support for vector length of 128 bits (using XMM registers) and 256 bits (using YMM registers).
Encoding support for instruction syntax of non-destructive source operands.
Elimination of escape opcode byte (0FH), SIMD prefix byte (66H, F2H, F3H) via a compact bit field represen-
tation within the VEX prefix.
Elimination of the need to use REX prefix to encode the extended half of general-purpose register sets (R8-
R15) for direct register access, memory addressing, or accessing XMM8-XMM15 (including YMM8-YMM15).
Flexible and more compact bit fields are provided in the VEX prefix to retain the full functionality provided by
REX prefix. REX.W, REX.X, REX.B functionalities are provided in the three-byte VEX prefix only because only a
subset of SIMD instructions need them.
Extensibility for future instruction extensions without significant instruction length increase.
Figure 2-8 shows the Intel 64 instruction encoding format with VEX prefix support. Legacy instruction without a
VEX prefix is fully supported and unchanged. The use of VEX prefix in an Intel 64 instruction is optional, but a VEX
prefix is required for Intel 64 instructions that operate on YMM registers or support three and four operand syntax.
VEX prefix is not a constant-valued, “single-purpose” byte like 0FH, 66H, F2H, F3H in legacy SSE instructions. VEX
prefix provides substantially richer capability than the REX prefix.
# Bytes
2,3
1
1
0,1
0,1,2,4
0,1
[Prefixes]
[VEX]
OPCODE
ModR/M
[SIB]
[DISP]
[IMM]
Figure 2-8. Instruction Encoding Format with VEX Prefix
2.3.2
VEX and the LOCK prefix
Any VEX-encoded instruction with a LOCK prefix preceding VEX will #UD.
2.3.3
VEX and the 66H, F2H, and F3H prefixes
Any VEX-encoded instruction with a 66H, F2H, or F3H prefix preceding VEX will #UD.
2.3.4
VEX and the REX prefix
Any VEX-encoded instruction with a REX prefix proceeding VEX will #UD.
Vol. 2A
2-13
INSTRUCTION FORMAT
2.3.5
The VEX Prefix
The VEX prefix is encoded in either the two-byte form (the first byte must be C5H) or in the three-byte form (the
first byte must be C4H). The two-byte VEX is used mainly for 128-bit, scalar, and the most common 256-bit AVX
instructions; while the three-byte VEX provides a compact replacement of REX and 3-byte opcode instructions
(including AVX and FMA instructions). Beyond the first byte of the VEX prefix, it consists of a number of bit fields
providing specific capability, they are shown in Figure 2-9.
The bit fields of the VEX prefix can be summarized by its functional purposes:
Non-destructive source register encoding (applicable to three and four operand syntax): This is the first source
operand in the instruction syntax. It is represented by the notation, VEX.vvvv. This field is encoded using 1’s
complement form (inverted form), i.e., XMM0/YMM0/R0 is encoded as 1111B, XMM15/YMM15/R15 is encoded
as 0000B.
Vector length encoding: This 1-bit field represented by the notation VEX.L. L= 0 means vector length is 128 bits
wide, L=1 means 256 bit vector. The value of this field is written as VEX.128 or VEX.256 in this document to
distinguish encoded values of other VEX bit fields.
REX prefix functionality: Full REX prefix functionality is provided in the three-byte form of VEX prefix. However
the VEX bit fields providing REX functionality are encoded using 1’s complement form, i.e., XMM0/YMM0/R0 is
encoded as 1111B, XMM15/YMM15/R15 is encoded as 0000B.
- Two-byte form of the VEX prefix only provides the equivalent functionality of REX.R, using 1’s complement
encoding. This is represented as VEX.R.
- Three-byte form of the VEX prefix provides REX.R, REX.X, REX.B functionality using 1’s complement
encoding and three dedicated bit fields represented as VEX.R, VEX.X, VEX.B.
- Three-byte form of the VEX prefix provides the functionality of REX.W only to specific instructions that need
to override default 32-bit operand size for a general purpose register to 64-bit size in 64-bit mode. For
those applicable instructions, VEX.W field provides the same functionality as REX.W. VEX.W field can
provide completely different functionality for other instructions.
Consequently, the use of REX prefix with VEX encoded instructions is not allowed. However, the intent of the
REX prefix for expanding register set is reserved for future instruction set extensions using VEX prefix
encoding format.
Compaction of SIMD prefix: Legacy SSE instructions effectively use SIMD prefixes (66H, F2H, F3H) as an
opcode extension field. VEX prefix encoding allows the functional capability of such legacy SSE instructions
(operating on XMM registers, bits 255:128 of corresponding YMM unmodified) to be encoded using the VEX.pp
field without the presence of any SIMD prefix. The VEX-encoded 128-bit instruction will zero-out bits 255:128
of the destination register. VEX-encoded instruction may have 128 bit vector length or 256 bits length.
Compaction of two-byte and three-byte opcode: More recently introduced legacy SSE instructions employ two
and three-byte opcode. The one or two leading bytes are: 0FH, and 0FH 3AH/0FH 38H. The one-byte escape
(0FH) and two-byte escape (0FH 3AH, 0FH 38H) can also be interpreted as an opcode extension field. The
VEX.mmmmm field provides compaction to allow many legacy instruction to be encoded without the constant
byte sequence, 0FH, 0FH 3AH, 0FH 38H. These VEX-encoded instruction may have 128 bit vector length or 256
bits length.
The VEX prefix is required to be the last prefix and immediately precedes the opcode bytes. It must follow any other
prefixes. If VEX prefix is present a REX prefix is not supported.
The 3-byte VEX leaves room for future expansion with 3 reserved bits. REX and the 66h/F2h/F3h prefixes are
reclaimed for future use.
VEX prefix has a two-byte form and a three byte form. If an instruction syntax can be encoded using the two-byte
form, it can also be encoded using the three byte form of VEX. The latter increases the length of the instruction by
one byte. This may be helpful in some situations for code alignment.
The VEX prefix supports 256-bit versions of floating-point SSE, SSE2, SSE3, and SSE4 instructions. Note, certain
new instruction functionality can only be encoded with the VEX prefix.
The VEX prefix will #UD on any instruction containing MMX register sources or destinations.
2-14
Vol. 2A
INSTRUCTION FORMAT
Byte 0
Byte 1
Byte 2
(Bit Position) 7
0
7
6
5
4
0
7
6
3
2
1
0
3-byte VEX
11000100
R X B
m-mmmm
W
vvvv
L
pp
7
0
7
6
3
2
1
0
2-byte VEX
11000101
R
vvvv
L
pp
R: REX.R in 1’s complement (inverted) form
1: Same as REX.R=0 (must be 1 in 32-bit mode)
0: Same as REX.R=1 (64-bit mode only)
X: REX.X in 1’s complement (inverted) form
1: Same as REX.X=0 (must be 1 in 32-bit mode)
0: Same as REX.X=1 (64-bit mode only)
B: REX.B in 1’s complement (inverted) form
1: Same as REX.B=0 (Ignored in 32-bit mode).
0: Same as REX.B=1 (64-bit mode only)
W: opcode specific (use like REX.W, or used for opcode
extension, or ignored, depending on the opcode byte)
m-mmmm:
00000: Reserved for future use (will #UD)
00001: implied 0F leading opcode byte
00010: implied 0F 38 leading opcode bytes
00011: implied 0F 3A leading opcode bytes
00100-11111: Reserved for future use (will #UD)
vvvv: a register specifier (in 1’s complement form) or 1111 if unused.
L: Vector Length
0: scalar or 128-bit vector
1: 256-bit vector
pp: opcode extension providing equivalent functionality of a SIMD prefix
00: None
01: 66
10: F3
11: F2
Figure 2-9. VEX bit fields
The following subsections describe the various fields in two or three-byte VEX prefix.
2.3.5.1
VEX Byte 0, bits[7:0]
VEX Byte 0, bits [7:0] must contain the value 11000101b (C5h) or 11000100b (C4h). The 3-byte VEX uses the C4h
first byte, while the 2-byte VEX uses the C5h first byte.
2.3.5.2
VEX Byte 1, bit [7] - ‘R’
VEX Byte 1, bit [7] contains a bit analogous to a bit inverted REX.R. In protected and compatibility modes the bit
must be set to ‘1’ otherwise the instruction is LES or LDS.
Vol. 2A
2-15
INSTRUCTION FORMAT
This bit is present in both 2- and 3-byte VEX prefixes.
The usage of WRXB bits for legacy instructions is explained in detail section 2.2.1.2 of Intel 64 and IA-32 Architec-
tures Software developer’s manual, Volume 2A.
This bit is stored in bit inverted format.
2.3.5.3
3-byte VEX byte 1, bit[6] - ‘X’
Bit[6] of the 3-byte VEX byte 1 encodes a bit analogous to a bit inverted REX.X. It is an extension of the SIB Index
field in 64-bit modes. In 32-bit modes, this bit must be set to ‘1’ otherwise the instruction is LES or LDS.
This bit is available only in the 3-byte VEX prefix.
This bit is stored in bit inverted format.
2.3.5.4
3-byte VEX byte 1, bit[5] - ‘B’
Bit[5] of the 3-byte VEX byte 1 encodes a bit analogous to a bit inverted REX.B. In 64-bit modes, it is an extension
of the ModR/M r/m field, or the SIB base field. In 32-bit modes, this bit is ignored.
This bit is available only in the 3-byte VEX prefix.
This bit is stored in bit inverted format.
2.3.5.5
3-byte VEX byte 2, bit[7] - ‘W’
Bit[7] of the 3-byte VEX byte 2 is represented by the notation VEX.W. It can provide following functions, depending
on the specific opcode.
• For AVX instructions that have equivalent legacy SSE instructions (typically these SSE instructions have a
general-purpose register operand with its operand size attribute promotable by REX.W), if REX.W promotes
the operand size attribute of the general-purpose register operand in legacy SSE instruction, VEX.W has same
meaning in the corresponding AVX equivalent form. In 32-bit modes for these instructions, VEX.W is silently
ignored.
• For AVX instructions that have equivalent legacy SSE instructions (typically these SSE instructions have oper-
ands with their operand size attribute fixed and not promotable by REX.W), if REX.W is don’t care in legacy
SSE instruction, VEX.W is ignored in the corresponding AVX equivalent form irrespective of mode.
• For new AVX instructions where VEX.W has no defined function (typically these meant the combination of the
opcode byte and VEX.mmmmm did not have any equivalent SSE functions), VEX.W is reserved as zero and
setting to other than zero will cause instruction to #UD.
2.3.5.6
2-byte VEX Byte 1, bits[6:3] and 3-byte VEX Byte 2, bits [6:3]- ‘vvvv’ the Source or Dest
Register Specifier
In 32-bit mode the VEX first byte C4 and C5 alias onto the LES and LDS instructions. To maintain compatibility with
existing programs the VEX 2nd byte, bits [7:6] must be 11b. To achieve this, the VEX payload bits are selected to
place only inverted, 64-bit valid fields (extended register selectors) in these upper bits.
The 2-byte VEX Byte 1, bits [6:3] and the 3-byte VEX, Byte 2, bits [6:3] encode a field (shorthand VEX.vvvv) that
for instructions with 2 or more source registers and an XMM or YMM or memory destination encodes the first source
register specifier stored in inverted (1’s complement) form.
VEX.vvvv is not used by the instructions with one source (except certain shifts, see below) or on instructions with
no XMM or YMM or memory destination. If an instruction does not use VEX.vvvv then it should be set to 1111b
otherwise instruction will #UD.
In 64-bit mode all 4 bits may be used. See Table 2-8 for the encoding of the XMM or YMM registers. In 32-bit and
16-bit modes bit 6 must be 1 (if bit 6 is not 1, the 2-byte VEX version will generate LDS instruction and the 3-byte
VEX version will ignore this bit).
2-16
Vol. 2A
INSTRUCTION FORMAT
Table 2-8. VEX.vvvv to register name mapping
General-Purpose Register
Valid in Legacy/Compatibility
VEX.vvvv
Dest Register
(If Applicable)1
32-bit modes?2
1111B
XMM0/YMM0
RAX/EAX
Valid
1110B
XMM1/YMM1
RCX/ECX
Valid
1101B
XMM2/YMM2
RDX/EDX
Valid
1100B
XMM3/YMM3
RBX/EBX
Valid
1011B
XMM4/YMM4
RSP/ESP
Valid
1010B
XMM5/YMM5
RBP/EBP
Valid
1001B
XMM6/YMM6
RSI/ESI
Valid
1000B
XMM7/YMM7
RDI/EDI
Valid
0111B
XMM8/YMM8
R8/R8D
Invalid
0110B
XMM9/YMM9
R9/R9D
Invalid
0101B
XMM10/YMM10
R10/R10D
Invalid
0100B
XMM11/YMM11
R11/R11D
Invalid
0011B
XMM12/YMM12
R12/R12D
Invalid
0010B
XMM13/YMM13
R13/R13D
Invalid
0001B
XMM14/YMM14
R14/R14D
Invalid
0000B
XMM15/YMM15
R15/R15D
Invalid
NOTES:
1. See Section 2.6, “VEX Encoding Support for GPR Instructions” for additional details.
2. Only the first eight General-Purpose Registers are accessible/encodable in 16/32b modes.
The VEX.vvvv field is encoded in bit inverted format for accessing a register operand.
2.3.6
Instruction Operand Encoding and VEX.vvvv, ModR/M
VEX-encoded instructions support three-operand and four-operand instruction syntax. Some VEX-encoded
instructions have syntax with less than three operands, e.g., VEX-encoded pack shift instructions support one
source operand and one destination operand).
The roles of VEX.vvvv, reg field of ModR/M byte (ModR/M.reg), r/m field of ModR/M byte (ModR/M.r/m) with
respect to encoding destination and source operands vary with different type of instruction syntax.
The role of VEX.vvvv can be summarized to three situations:
VEX.vvvv encodes the first source register operand, specified in inverted (1’s complement) form and is valid for
instructions with 2 or more source operands.
VEX.vvvv encodes the destination register operand, specified in 1’s complement form for certain vector shifts.
The instructions where VEX.vvvv is used as a destination are listed in Table 2-9. The notation in the “Opcode”
column in Table 2-9 is described in detail in section 3.1.1.
VEX.vvvv does not encode any operand, the field is reserved and should contain 1111b.
Table 2-9. Instructions with a VEX.vvvv destination
Opcode
Instruction mnemonic
VEX.128.66.0F 73 /7 ib
VPSLLDQ xmm1, xmm2, imm8
VEX.128.66.0F 73 /3 ib
VPSRLDQ xmm1, xmm2, imm8
VEX.128.66.0F 71 /2 ib
VPSRLW xmm1, xmm2, imm8
VEX.128.66.0F 72 /2 ib
VPSRLD xmm1, xmm2, imm8
VEX.128.66.0F 73 /2 ib
VPSRLQ xmm1, xmm2, imm8
VEX.128.66.0F 71 /4 ib
VPSRAW xmm1, xmm2, imm8
Vol. 2A
2-17
INSTRUCTION FORMAT
Opcode
Instruction mnemonic
VEX.128.66.0F 72 /4 ib
VPSRAD xmm1, xmm2, imm8
VEX.128.66.0F 71 /6 ib
VPSLLW xmm1, xmm2, imm8
VEX.128.66.0F 72 /6 ib
VPSLLD xmm1, xmm2, imm8
VEX.128.66.0F 73 /6 ib
VPSLLQ xmm1, xmm2, imm8
The role of ModR/M.r/m field can be summarized to two situations:
ModR/M.r/m encodes the instruction operand that references a memory address.
For some instructions that do not support memory addressing semantics, ModR/M.r/m encodes either the
destination register operand or a source register operand.
The role of ModR/M.reg field can be summarized to two situations:
ModR/M.reg encodes either the destination register operand or a source register operand.
For some instructions, ModR/M.reg is treated as an opcode extension and not used to encode any instruction
operand.
For instruction syntax that support four operands, VEX.vvvv, ModR/M.r/m, ModR/M.reg encodes three of the four
operands. The role of bits 7:4 of the immediate byte serves the following situation:
Imm8[7:4] encodes the third source register operand.
2.3.6.1
3-byte VEX byte 1, bits[4:0] - “m-mmmm”
Bits[4:0] of the 3-byte VEX byte 1 encode an implied leading opcode byte (0F, 0F 38, or 0F 3A). Several bits are
reserved for future use and will #UD unless 0.
Table 2-10. VEX.m-mmmm interpretation
VEX.m-mmmm
Implied Leading Opcode Bytes
00000B
Reserved
00001B
0F
00010B
0F 38
00011B
0F 3A
00100-11111B
Reserved
(2-byte VEX)
0F
VEX.m-mmmm is only available on the 3-byte VEX. The 2-byte VEX implies a leading 0Fh opcode byte.
2.3.6.2
2-byte VEX byte 1, bit[2], and 3-byte VEX byte 2, bit [2]- “L”
The vector length field, VEX.L, is encoded in bit[2] of either the second byte of 2-byte VEX, or the third byte of 3-
byte VEX. If “VEX.L = 1”, it indicates 256-bit vector operation. “VEX.L = 0” indicates scalar and 128-bit vector
operations.
The instruction VZEROUPPER is a special case that is encoded with VEX.L = 0, although its operation zero’s bits
255:128 of all YMM registers accessible in the current operating mode.
See the following table.
2-18
Vol. 2A
INSTRUCTION FORMAT
Table 2-11. VEX.L interpretation
VEX.L
Vector Length
0
128-bit (or 32/64-bit scalar)
1
256-bit
2.3.6.3
2-byte VEX byte 1, bits[1:0], and 3-byte VEX byte 2, bits [1:0]- “pp”
Up to one implied prefix is encoded by bits[1:0] of either the 2-byte VEX byte 1 or the 3-byte VEX byte 2. The prefix
behaves as if it was encoded prior to VEX, but after all other encoded prefixes.
See the following table.
Table 2-12. VEX.pp interpretation
pp
Implies this prefix after other prefixes but before VEX
00B
None
01B
66
10B
F3
11B
F2
2.3.7
The Opcode Byte
One (and only one) opcode byte follows the 2 or 3 byte VEX. Legal opcodes are specified in Appendix B, in color.
Any instruction that uses illegal opcode will #UD.
2.3.8
The ModR/M, SIB, and Displacement Bytes
The encodings are unchanged but the interpretation of reg_field or rm_field differs (see above).
2.3.9
The Third Source Operand (Immediate Byte)
VEX-encoded instructions can support instruction with a four operand syntax. VBLENDVPD, VBLENDVPS, and
PBLENDVB use imm8[7:4] to encode one of the source registers.
2.3.10 Intel® AVX Instructions and the Upper 128-bits of YMM registers
If an instruction with a destination XMM register is encoded with a VEX prefix, the processor zeroes the upper bits
(above bit 128) of the equivalent YMM register. Legacy SSE instructions without VEX preserve the upper bits.
2.3.10.1 Vector Length Transition and Programming Considerations
An instruction encoded with a VEX.128 prefix that loads a YMM register operand operates as follows:
Data is loaded into bits 127:0 of the register
Bits above bit 127 in the register are cleared.
Thus, such an instruction clears bits 255:128 of a destination YMM register on processors with a maximum vector-
register width of 256 bits. In the event that future processors extend the vector registers to greater widths, an
instruction encoded with a VEX.128 or VEX.256 prefix will also clear any bits beyond bit 255. (This is in contrast
with legacy SSE instructions, which have no VEX prefix; these modify only bits 127:0 of any destination register
operand.)
Programmers should bear in mind that instructions encoded with VEX.128 and VEX.256 prefixes will clear any
future extensions to the vector registers. A calling function that uses such extensions should save their state before
calling legacy functions. This is not possible for involuntary calls (e.g., into an interrupt-service routine). It is
recommended that software handling involuntary calls accommodate this by not executing instructions encoded
Vol. 2A
2-19
INSTRUCTION FORMAT
with VEX.128 and VEX.256 prefixes. In the event that it is not possible or desirable to restrict these instructions,
then software must take special care to avoid actions that would, on future processors, zero the upper bits of vector
registers.
Processors that support further vector-register extensions (defining bits beyond bit 255) will also extend the
XSAVE and XRSTOR instructions to save and restore these extensions. To ensure forward compatibility, software
that handles involuntary calls and that uses instructions encoded with VEX.128 and VEX.256 prefixes should first
save and then restore the vector registers (with any extensions) using the XSAVE and XRSTOR instructions with
save/restore masks that set bits that correspond to all vector-register extensions. Ideally, software should rely on
a mechanism that is cognizant of which bits to set. (E.g., an OS mechanism that sets the save/restore mask bits
for all vector-register extensions that are enabled in XCR0.) Saving and restoring state with instructions other than
XSAVE and XRSTOR will, on future processors with wider vector registers, corrupt the extended state of the vector
registers - even if doing so functions correctly on processors supporting 256-bit vector registers. (The same is true
if XSAVE and XRSTOR are used with a save/restore mask that does not set bits corresponding to all supported
extensions to the vector registers.)
2.3.11 Intel® AVX Instruction Length
The Intel AVX instructions described in this document (including VEX and ignoring other prefixes) do not exceed 11
bytes in length, but may increase in the future. The maximum length of an Intel 64 and IA-32 instruction remains
15 bytes.
2.3.12 Vector SIB (VSIB) Memory Addressing
In Intel® Advanced Vector Extensions 2 (Intel® AVX2), an SIB byte that follows the ModR/M byte can support VSIB
memory addressing to an array of linear addresses. VSIB addressing is only supported in a subset of Intel AVX2
instructions. VSIB memory addressing requires 32-bit or 64-bit effective address. In 32-bit mode, VSIB addressing
is not supported when address size attribute is overridden to 16 bits. In 16-bit protected mode, VSIB memory
addressing is permitted if address size attribute is overridden to 32 bits. Additionally, VSIB memory addressing is
supported only with VEX prefix.
In VSIB memory addressing, the SIB byte consists of:
The scale field (bit 7:6) specifies the scale factor.
The index field (bits 5:3) specifies the register number of the vector index register, each element in the vector
register specifies an index.
The base field (bits 2:0) specifies the register number of the base register.
Table 2-3 shows the 32-bit VSIB addressing form. It is organized to give 256 possible values of the SIB byte (in
hexadecimal). General purpose registers used as a base are indicated across the top of the table, along with corre-
sponding values for the SIB byte’s base field. The register names also include R8D-R15D applicable only in 64-bit
mode (when address size override prefix is used, but the value of VEX.B is not shown in Table 2-3). In 32-bit mode,
R8D-R15D does not apply.
Table rows in the body of the table indicate the vector index register used as the index field and each supported
scaling factor shown separately. Vector registers used in the index field can be XMM or YMM registers. The left-
most column includes vector registers VR8-VR15 (i.e., XMM8/YMM8-XMM15/YMM15), which are only available in
64-bit mode and does not apply if encoding in 32-bit mode.
2-20
Vol. 2A
INSTRUCTION FORMAT
Table 2-13. 32-Bit VSIB Addressing Forms of the SIB Byte
r32
EAX/
ECX/
EDX/
EBX/
ESP/
EBP/
ESI/
EDI/
R8D
R9D
R10D
R11D
R12D
R13D1
R14D
R15D
(In decimal) Base =
0
1
2
3
4
5
6
7
(In binary) Base =
000
001
010
011
100
101
110
111
Scaled Index
SS
Index
Value of SIB Byte (in Hexadecimal)
VR0/VR8
*1
00
000
00
01
02
03
04
05
06
07
VR1/VR9
001
08
09
0A
0B
0C
0D
0E
0F
VR2/VR10
010
10
11
12
13
14
15
16
17
VR3/VR11
011
18
19
1A
1B
1C
1D
1E
1F
VR4/VR12
100
20
21
22
23
24
25
26
27
VR5/VR13
101
28
29
2A
2B
2C
2D
2E
2F
VR6/VR14
110
30
31
32
33
34
35
36
37
VR7/VR15
111
38
39
3A
3B
3C
3D
3E
3F
VR0/VR8
*2
01
000
40
41
42
43
44
45
46
47
VR1/VR9
001
48
49
4A
4B
4C
4D
4E
4F
VR2/VR10
010
50
51
52
53
54
55
56
57
VR3/VR11
011
58
59
5A
5B
5C
5D
5E
5F
VR4/VR12
100
60
61
62
63
64
65
66
67
VR5/VR13
101
68
69
6A
6B
6C
6D
6E
6F
VR6/VR14
110
70
71
72
73
74
75
76
77
VR7/VR15
111
78
79
7A
7B
7C
7D
7E
7F
VR0/VR8
*4
10
000
80
81
82
83
84
85
86
87
VR1/VR9
001
88
89
8A
8B
8C
8D
8E
8F
VR2/VR10
010
90
91
92
93
94
95
96
97
VR3/VR11
011
98
89
9A
9B
9C
9D
9E
9F
VR4/VR12
100
A0
A1
A2
A3
A4
A5
A6
A7
VR5/VR13
101
A8
A9
AA
AB
AC
AD
AE
AF
VR6/VR14
110
B0
B1
B2
B3
B4
B5
B6
B7
VR7/VR15
111
B8
B9
BA
BB
BC
BD
BE
BF
VR0/VR8
*8
11
000
C0
C1
C2
C3
C4
C5
C6
C7
VR1/VR9
001
C8
C9
CA
CB
CC
CD
CE
CF
VR2/VR10
010
D0
D1
D2
D3
D4
D5
D6
D7
VR3/VR11
011
D8
D9
DA
DB
DC
DD
DE
DF
VR4/VR12
100
E0
E1
E2
E3
E4
E5
E6
E7
VR5/VR13
101
E8
E9
EA
EB
EC
ED
EE
EF
VR6/VR14
110
F0
F1
F2
F3
F4
F5
F6
F7
VR7/VR15
111
F8
F9
FA
FB
FC
FD
FE
FF
NOTES:
1. If ModR/M.mod = 00b, the base address is zero, then effective address is computed as [scaled vector index] + disp32. Otherwise the
base address is computed as [EBP/R13]+ disp, the displacement is either 8 bit or 32 bit depending on the value of ModR/M.mod:
MOD
Effective Address
00b
[Scaled Vector Register] + Disp32
01b
[Scaled Vector Register] + Disp8 + [EBP/R13]
10b
[Scaled Vector Register] + Disp32 + [EBP/R13]
2.3.12.1
64-bit Mode VSIB Memory Addressing
In 64-bit mode VSIB memory addressing uses the VEX.B field and the base field of the SIB byte to encode one of
the 16 general-purpose register as the base register. The VEX.X field and the index field of the SIB byte encode one
of the 16 vector registers as the vector index register.
In 64-bit mode the top row of Table 2-13 base register should be interpreted as the full 64-bit of each register.
2.4
INTEL® ADVANCED MATRIX EXTENSIONS (INTEL® AMX)
Intel® AMX instructions follow the general documentation convention established in previous sections. Additionally,
Intel® Advanced Matrix Extensions use notation conventions as described below.
In the instruction encoding boxes, sibmem is used to denote an encoding where a ModR/M byte and SIB byte are
used to indicate a memory operation where the base and displacement are used to point to memory, and the index
Vol. 2A
2-21
INSTRUCTION FORMAT
register (if present) is used to denote a stride between memory rows. The index register is scaled by the sib.scale
field as usual. The base register is added to the displacement, if present.
In the instruction encoding, the ModR/M byte is represented several ways depending on the role it plays. The
ModR/M byte has 3 fields: 2-bit ModR/M.mod field, a 3-bit ModR/M.reg field and a 3-bit ModR/M.r/m field. When all
bits of the ModR/M byte have fixed values for an instruction, the 2-hex nibble value of that byte is presented after
the opcode in the encoding boxes on the instruction description pages. When only some fields of the ModR/M byte
must contain fixed values, those values are specified as follows:
If only the ModR/M.mod must be 0b11, and ModR/M.reg and ModR/M.r/m fields are unrestricted, this is
denoted as 11:rrr:bbb. The rrr correspond to the 3-bits of the ModR/M.reg field and the bbb correspond to the
3-bits of the ModR/M.r/m field.
If the ModR/M.mod field is constrained to be a value other than 0b11, i.e., it must be one of 0b00, 0b01, or
0b10, then the notation !(11) is used.
If the ModR/M.reg field had a specific required value, e.g., 0b101, that would be denoted as mm:101:bbb.
NOTE
Historically this document only specified the ModR/M.reg field restrictions with the notation /0 ... /7
and did not specify restrictions on the ModR/M.mod and ModR/M.r/m fields in the encoding boxes.
2.5
INTEL® AVX AND INTEL® SSE INSTRUCTION EXCEPTION CLASSIFICATION
To look up the exceptions of legacy 128-bit SIMD instruction, 128-bit VEX-encoded instructions, and 256-bit VEX-
encoded instruction, Table 2-14 summarizes the exception behavior into separate classes, with detailed exception
conditions defined in sub-sections 2.5.1 through 2.6.1. For example, ADDPS contains the entry:
“See Exceptions Type 2”
In this entry, Type2” can be looked up in Table 2-14.
The instruction’s corresponding CPUID feature flag can be identified in the fourth column of the Instruction
summary table.
Note: #UD on CPUID feature flags=0 is not guaranteed in a virtualized environment if the hardware supports the
feature flag.
NOTE
Instructions that operate only with MMX, X87, or general-purpose registers are not covered by the
exception classes defined in this section. For instructions that operate on MMX registers, see
Section 23.25.3, “Exception Conditions of Legacy SIMD Instructions Operating on MMX Registers”
in the Intel® 64 and IA-32 Architectures Software Developer’s Manual, Volume 3B.
2-22
Vol. 2A

 

 

 

 

 

 

 

Content      ..     13      14      15      16     ..