|
|
|
龙芯 3A3000/3B3000 处理器用户手册 y 下册
Access class instruction with index register
Instruction
Instruction
ISA
mnemonic
function
Compatible
description
Category
LBUX
The index address takes an unsigned byte
MIPS DSP
LHX
Index address takes half a word
MIPS DSP
LWX
Index address fetch word
MIPS DSP
Branch instruction
Instruction
Instruction
ISA
mnemonic
function
Compatible
description
Category
BPOSGE32
DSP control register POS value is greater than or equal to 32 jump
MIPS DSP
38
龙芯 3A3000/3B3000 处理器用户手册 y 下册
DSP instructions no longer supported by MIPS
The instructions listed in this section, although defined in r2.34 of The MIPS® Architecture for Programmers
VolumeIV-e: The MIPS® DSP Application-Specific Extension to The MIPS64® Architecture, have been deleted
after r2.40.
Although
the
Loongson
3A300/3B3000 processor implements these instructions, it is not
recommended for use.
Instruction
Instruction
mnemonic
function
description
ABSQ_S. OB
The vector takes (8 right-most) byte absolute values and saturates them, resulting in symbol
expansion
ABSQ_S. PW
Vector take (right 4) absolute value of half word, and do saturation operation, the result sign
expansion
ABSQ_S. QH
Vector take (rightmost 2) absolute value of the word, and do saturation operation, the result
symbol expansion
ADDQ. PW
The vector 2 to the right plus a small number
ADDQ supachai
Vector (right 2) small number saturation plus
panitchpakdi W
ADDQ. QH
Vector (right 4) decimal half word plus
ADDQ S.Q H
Vector (4 to the right) decimal halfword saturation plus
ADDU. OB
Vector (right 8) decimal byte unsigned plus
ADDU S.O B
Vector (right 8) decimal byte unsigned saturation plus
ADDU. QH
Vector (right 4) decimal halfword unsigned plus
ADDU S.Q H
Vector (right 4) decimal halfword unsigned saturation plus
ADDUH. OB
The vector is unsigned (8 right-most) bytes plus, the result is divided by 2, and the result is
extended
ADDUH_R. OB
The vector is unsigned (8 right-most) bytes rounded plus, the result is divided by 2, and the
result is extended
BPOSGE64
The DSP control register POS value is 64 or higher, then it jumps
CMP. EQ. PW
Vector (rightmost 2) word equal comparison, result set condition bit
CMP. LT. PW
Vector (rightmost 2) word is less than comparison, result sets condition bit
CMP. LE. PW
Vector (rightmost 2) words less than or equal to the comparison, the result set condition bit
CMP. EQ. QH
Vector (right 4) half - word equal comparison, the result set condition bit
CMP. LT. QH
Vector (right-most 4) half character is less than the comparison, the result set condition bit
CMP. LE. QH
Vector (right 4) half - word less than or equal to the comparison, the result set condition bit
CMPGDU. EQ. OB
Vector (rightmost 8) unsigned byte equal comparison, the result of both conditional and
general register
CMPGDU. LT. OB
Vector (8 right-most) unsigned bytes less than the comparison, the result is both conditional
and general register
CMPGDU. LE. OB
Vector (8 right-most) unsigned bytes less than or equal to the comparison, the result is both
conditional and general register
CMPGU. EQ. OB
Vector (rightmost 8) byte equal comparison, the result set general purpose register
CMPGU. LT. OB
Vector (rightmost 8) bytes less than comparison, result in general purpose register
CMPGU. LE. OB
Vector (8 right-most) bytes less than or equal to the comparison, the result of the general
purpose register
CMPU. EQ. OB
Vector (rightmost 8) unsigned byte equal comparison, the result set condition bit
CMPU. LT. OB
Vector (8 right-most) unsigned bytes less than comparison, result set condition bit
CMPU. LE. OB
Vector (8 right-most) unsigned bytes less than or equal to the comparison, the result set
condition bit
39
龙芯 3A3000/3B3000 处理器用户手册 y 下册
DAPPEND
Right shift and high splice
DBALIGN
Two registers high and low byte splice
DEXTP
Extracts a fixed length from any position of the accumulator to a general purpose register.
The length is specified by the immediate number
DEXTPDP
Extract the fixed length number from any position of the accumulator to the general purpose
register and subtract the POS value. The length is denoted by the immediate number
set
40
龙芯 3A3000/3B3000 处理器用户手册 y 下册
Instruction
Instruction
mnemonic
function
description
DEXTPDPV
Extract the fixed length number from any position of the accumulator to the general purpose
register and subtract the POS value. The length is indicated by the register
set
DEXTPV
Extracts a fixed length number from any position of the accumulator to a general purpose
register. The length is determined by the register instruction
DEXTR. L
After the accumulator moves to the right, the double word is intercepted and assigned to the
general purpose register. The shift value is specified by the immediate number
DEXTR_R. L
The accumulator is rounded to the right, intercepting the double word and assigning it to the
general purpose register. The shift value is specified by the immediate number
DEXTR_RS. L
The accumulator moves to the right after saturated rounding, intercepts the double word and
assigns it to the general purpose register. The shift value is specified by the immediate
number
DEXTR w.
After the accumulator is moved to the right, the truncated word is assigned to the general
purpose register. The shift value is specified by the immediate number
DEXTR_R w.
The accumulator is rounded to the right, the truncated word is assigned to the general purpose
register, and the shift value is specified by the immediate number
DEXTR_RS w.
The accumulator moves to the right after saturated rounding, intercepting words to the general
purpose register, and shifting values are specified by the immediate number
DEXTR_S. H
Move right from accumulator saturation, extract half word to general purpose register, shift
value specified by immediate number
DEXTRV. L
After the accumulator moves to the right, the double word is intercepted and assigned to the
general purpose register. The shift value is specified by the register
DEXTRV_R. L
The accumulator is rounded to the right, intercepting the double word and assigning it to the
general purpose register. The shift value is specified by the register
DEXTRV_RS. L
The accumulator moves right after saturated rounding, intercepts the double word and assigns
it to the general purpose register. The shift value is specified by the register
DEXTRV_S. H
After the accumulator moves to the right, the truncated word is assigned to the general
purpose register, and the shifted value is specified by the register
DEXTRV w.
The accumulator is rounded to the right, the truncated word is assigned to the general purpose
register, and the shifted value is specified by the register
DEXTRV_R w.
The accumulator moves right after saturated rounding, intercepting words to the general
purpose register, and shifting values are specified by the register
DEXTRV_RS w.
Move right from accumulator saturation, extract half word to general purpose register, shift
value specified by register
DINSV
Variable bit field insertion
DMADD
Double characters are multiplied by symbols
DMADDU
Double word unsigned multiplication and addition
DMSUB
Double characters are signed multiplied and subtracted
DMSUBU
Double word unsigned times minus
DMTHLIP
LO values are copied to HI, general register values are copied to LO, pos values are increased
by 64
The DPA. W.Q H
Vector integers (4 to the right) half-word multiply and sum, and finally sum with ACC
DPAQ_S W.Q H
The vector decimals (4 to the right) are multiplied by the half-word, and the result is
multiplied by acc
DPAQ_SA L.P W
Vector small number multiplication, the result after saturated rounding, and then summing
with ACC
DPAU. H.O BL
The vector integer (4 leftmost) bytes are unsigned multiplied and accumulated, and the result
is then accumulated with ACC
DPAU. H.O BR
Vector integer (4 to the right) bytes unsigned multiplied and accumulated, the result is then
accumulated with ACC
The DPS. W.Q H
Vector integers (4 to the right) are multiplied and accumulated, and the result is then
subtracted from ACC
DPSQ_S W.Q H
The vector decimals (4 to the right) are multiplied by half words and subtracted from ACC
41
龙芯 3A3000/3B3000 处理器用户手册 y 下册
DPSQ_SA L.P W
Multiply the vector by a small number, the result is accumulated after saturated rounding, and
the result is accumulated with ACC
DPSU. H.O BL
The vector integer (4 leftmost) bytes are unsigned multiplied and accumulated, and the result
is then accumulated with ACC
DPSU. H.O BR
Vector integer (4 to the right) bytes unsigned multiplied and accumulated, the result is then
accumulated with ACC
DSHILO
The accumulator value is shifted and then written back to the same accumulator. The shift
value is determined by the immediate number
DSHILOV
The accumulator value is shifted and then written back to the same accumulator. The shifted
value is determined by the register
LDX
Base address plus index fetch
MAQ_S. L.P WL
The vector decimal (to the right) is multiplied by acc
MAQ_S. L.P WR
The vector decimal (far right) is multiplied by acc
MAQ_S. W.Q HLL
The vector decimal (leftmost) is multiplied by half a word and the result is added to ACC
MAQ_SA. W.Q HLL
The vector decimal (leftmost) is multiplied by a half-word, and the result is summed with
ACC after saturated rounding
MAQ_S W.Q HLR
The vector decimal (second left) is multiplied by half a word, and the result is added to ACC
42
龙芯 3A3000/3B3000 处理器用户手册 y 下册
Instruction
Instruction
mnemonic
function
description
MAQ_SA W.Q HLR
The vector decimal (left) is multiplied by half word, and the result is summed with ACC after
saturated rounding
MAQ_S. W.Q HRL
The vector decimal (to the right) is multiplied by half a word, and the result is added to ACC
MAQ_SA. W.Q HRL
Vector decimal (second right) half-word multiplication, the result is to take the saturated
rounding and then add with ACC
MAQ_S. W.Q HRR
The vector decimal (far right) is multiplied by half a word and the result is added to ACC
MAQ_SA. W.Q HRR
Vector decimal (right-most) half-word multiplication, the result is to take the saturated
rounding and then add with ACC
MULEQ_S. PW. QHL
The left half of a sign is saturated by multiplication, and the result is a word
MULEQ_S. PW. QHR
The right half of a sign is saturated by multiplication, and the result is a word
MULEU_S. QH. With
Vector (4 to the right) bytes unsigned times (4 to the right) halfwords, resulting in two
halfwords
MULEU_S. QH. The
Vector (4 right-most) bytes unsigned times (4 right-most) halfwords, resulting in two
OBR
halfwords
MULQ_RS. QH
The vector (4 right-most) halfwords are saturated and rounded, resulting in half-words
MULSAQ_S L.P W
Multiply the small number of vector, subtract after saturation, and then add the result with
ACC
MULSAQ_S W.Q H
The vector decimals (4 to the right) are multiplied and subtracted, and the result is added to
ACC
PACKRL. PW
Package the right-most word of source 1 with the right-second word of source 2
PICK the OB
Conditional bit - based (8 right-most) bytes selection
PICK the PW
Conditional bit - based (4 right - most) halfword selection
PICK the QH
Conditional bit - based (rightmost 2) word selection
PRECEQ. L.P WL
Vector decimal precision has signed extension, from (second right) word to double word
PRECEQ. L.P WR
Vector decimal precision signed extension, from (far right) word to double word
PRECEQ. PW. QHL
Vector decimal precision has symbolic extension, from (two to the right) half word to two
words
PRECEQ. PW. QHR
Vector decimal precision has symbolic expansion from (rightmost two) half words to two
words
PRECEQ. PW. QHLA
Vector decimal precision has a signed extension, from half - word to two - word
PRECEQ. PW. QHRA
Vector decimal precision has a signed extension, from half - word to two - word
PRECEQU. QH. With
Vector decimal precision extension, from (4 left) unsigned bytes to 4 half-words
PRECEQU. QH. The
Vector decimal precision extension, from (4 on the right) unsigned bytes to 4 half-words
OBR
PRECEQU. QH. OBLA
Vector decimal precision extension, from unsigned bytes (4 to the left of the right-most word)
to 4 half-words
PRECEQU. QH. OBRA
Vector decimal precision extension, from unsigned bytes (4 to the right of the right-most
word) to 4 half-words
PRECEU. QH. With
Vector precision extension, from (4 left) unsigned bytes to 4 halfwords
PRECEU. QH. The OBR
Vector precision expansion from (4 on the right) unsigned bytes to 4 halfwords
PRECEU. QH. OBLA
Vector precision expansion from unsigned bytes (4 left crosses of right-most word) to 4 half-
words
PRECEU. QH. OBRA
Vector precision expansion from unsigned bytes (4 crosses the right side of the right-most
word) to 4 half-words
PRECR. OB. QH
Vector integer precision reduced from (4 right-most) halfwords to 8 bytes
PRECR_SRA. QH. PW
Vector right shift integer precision reduction, from (rightmost 2) words to 4 half words,
resulting symbol expansion
PRECR_SRA_R. QH.
Vector right shift integer precision reduction from (rightmost 2) words to 4 half words, and do
PW
rounding, resulting symbol expansion
43
龙芯 3A3000/3B3000 处理器用户手册 y 下册
PRECRQ. OB. QH
Vector decimal precision reduced from (right-most 4) halfwords to 8 bytes
PRECRQ. PW. L
Vector decimal precision reduced from (2 words to the right) to 4 half words
PRECRQ. QH. PW
Vector decimal precision reduced from (2 words to the right) to (4 words)
PRECRQ_RS. QH. PW
Vector decimal precision is reduced from (right 2) words to (4) half words, and done with
saturation and rounding
PRECRQU_S. OB. QH
Vector decimal precision reduced from (right-most 4) halfwords to 8 unsigned bytes
PREPENDD
Double word right shift and high splice
PREPENDW
The word is shifted to the right and spliced high
RADDU L.O B
8 bytes unsigned accumulation
44
龙芯 3A3000/3B3000 处理器用户手册 y 下册
Instruction
Instruction
mnemonic
function
description
REPL. OB
Vector copies immediate number (integer) to 8 bytes, resulting symbol extension
REPL. PW
Vector copies immediate number (integer) to 2 words, resulting symbol extension
REPL. QH
Vector copies immediate Numbers (integers) to 4 halfwords, resulting in symbol expansion
REPLV. OB
Vector copies bytes to 8 bytes, resulting symbol extension
REPLV. PW
Vector copies word to 2 words, resulting symbol extension
REPLV. QH
Vector copy half word to 4 half word, resulting symbol expansion
SHLL. OB
The vector logic shifts 8 bytes to the right, the shift value is specified by the immediate
number, and the resulting symbol expands
SHLL. QH
Vector logic moves left (four right-most) halfwords, the shift value is specified by the
immediate number, and the resulting symbol expands
SHLL. PW
Vector logic moves left (rightmost 2) halfwords, the shift value is specified by the immediate
number, the result symbol expands
SHLLV. OB
The vector logic shifts 8 bytes to the left, the shift value is specified by the register, and the
result symbol expands
SHLLV. QH
Vector logic moves left (four right-most) halfwords, the shift value is specified by the
register, and the result symbol expands
SHLLV. PW
Vector logic moves left (rightmost 2) halfwords, the shift value is specified by the register,
and the result symbol expands
SHLL_S. QH
Vector saturation logic moves left (4 right-most) halfwords, the shift value is specified by the
immediate number, the result symbol expands
SHLL_S. PW
Vector saturation logic moves left (rightmost 2) halfwords, the shift value is specified by the
immediate number, the result symbol expands
SHLLV_S. QH
Vector saturation logic moves left (four right most) halfwords, the shift value is specified by
the register, the result symbol expands
SHLLV_S. PW
Vector saturation logic moves left (rightmost 2) halfwords, the shift value is specified by the
register, the result symbol expands
SHRA. OB
Vector arithmetic is shifted to the right (8 right-most) bytes, the shifted value is specified by
the immediate number, and the resulting symbol is extended
SHRA. QH
Vector arithmetic is right-shifted (4 right-most) halfwords, the shifted value is specified by
the immediate number, and the resulting symbol is extended
SHRA. PW
Vector arithmetic is right-shifted (right-most 2) halfwords, the shifted value is specified by
the immediate number, and the resulting symbol is extended
SHRAV. OB
Vector arithmetic is shifted to the right (8 rarest) bytes, the shifted value is specified by the
register, and the resulting symbol is extended
SHRAV. QH
Vector arithmetic is right-shifted (four right-most) halfwords, the shifted value is specified by
the register, and the resulting symbol is extended
SHRAV. PW
Vector arithmetic is right-shifted (rightmost 2) halfwords, the shifted value is specified by the
register, the result symbol is extended
SHRA_R. OB
Vector arithmetic is right-shifted (8 right-most) bytes, with the shift value specified by the
immediate number, rounded, and the resulting symbol extended
SHRA_R. QH
Vector arithmetic is right-shifted (4 right-most) halfwords, with the shift value specified by
the immediate number, rounded, and the resulting symbol extended
SHRA_R. PW
Vector arithmetic right-shift (rightmost two) words, the value of the shift is specified by the
immediate number, and done rounding, the result symbol expansion
SHRAV_R. OB
Vector arithmetic is shifted to the right (8 right-most) bytes, the shifted value is specified by
the register, and is rounded, resulting in symbol expansion
SHRAV_R. QH
Vector arithmetic right-shift (4 right-most) halfwords, shift values specified by registers, and
done rounding, resulting symbol expansion
SHRAV_R. PW
Vector arithmetic right-shift (rightmost two) words, shift values specified by registers, and
done rounding, resulting symbol expansion
SHRL. OB
Vector logic shifts 8 bytes to the right, the shift value is specified by the immediate number,
and the resulting symbol expands
SHRL. QH
Vector logic moves right (4 right-most) halfwords, the shift value is specified by the
immediate number, the result symbol expands
45
龙芯 3A3000/3B3000 处理器用户手册 y 下册
SHRLV. OB
Vector logic shifts 8 bytes to the right, the shifted value is specified by the register, and the
resulting symbol expands
SHRLV. QH
Vector logic moves right (4 right-most) halfwords, the shift value is specified by the register,
and the result symbol expands
SUBQ. PW
Vector (2 to the right) small number minus
SUBQ. QH
Vector (4 to the right) decimal half word minus
SUBQ_S. PW
Vector (rightmost 2) small number saturation minus
SUBQ_S. QH
Vector (4 to the right) decimal half word saturation minus
SUBU. OB
Vector (8 right most) decimal bytes unsigned minus
SUBU. QH
Vector (4 to the right) decimal halfword unsigned minus
SUBU_S. OB
Vector (right 8) decimal byte unsigned saturation minus
SUBU_S. QH
Vector (4 to the right) decimal halfword unsigned saturation minus
SUBUH. OB
The vector is unsigned (8 to the right), the result is subtracted, the result is divided by 2, the
result is extended
46
龙芯 3A3000/3B3000 处理器用户手册 y 下册
Instruction
Instruction
mnemonic
function
description
SUBUH_R. OB
Vector unsigned (8 right-most) bytes rounded down, result divided by 2, result symbol
extended
2.3.2 Supplementary instructions to MIPS DSP instruction manual
The definition of MTHI and MTLO in MIPS DSP instruction manual is unreasonable. In a 64-bit architecture,
the execution of the MTHI and MTLO instructions does not require writing to the HI/LO register after a high 32-bit
symbol extension of the value of the general purpose register.
YingCaiYong MIPS
TongYongZhiLingShouCeZhongDuiYu MTHI、MTLO 指令的定义形式,即 HI GPR[rs],LO GPR[rs]。
2.4 MIPS64 compatible instruction implementation definition
This section describes the implementation-related parts of the MIPS64 compatibility instructions implemented
by the GS464E, or those that differ from the MIPS64 specification.
2.4.1 The load instruction that targets the no. 0 general purpose register
All load instructions, except LDPTE and LWPTE, whose target is the universal register, perform as well as
PREF(hint=0) when their target bit is the zero universal register. However, for the GSLQ instruction of the godson
custom extension, it is necessary to set both of its target registers as general purpose register 0.
2.4.2 PREF instruction
The PREF directive only implements the hint=0 (prefetch for Load) and hint=1 (Prefetch for store) modes
according to the MIPS specification.
2.4.3 RDHWR instruction
RDHWR instruction not only implements rd value of 0, 1, 2, 3, 29 according to MIPS specification, but also
implements RD value of 30 and 31. The specific definition is as follows:
z Rd =30: Reads the external counter value. The counter is 64-bit and increases by 1 per beat at a frequency
consistent with the internal bus clock frequency of the chip. The characteristic of the clock frequency is
that it does not vary with the clock frequency of a processor core and can be considered as a constant
frequency.
z
Rd =31: Read the ratio between the current frequency of processor core and the frequency of bus in the
chip (M/N). Where M is in the [15:8] bit of the return value and N is in the [7:0] bit of the return value.
2.4.4 PREFX instruction
The PREFX directive implements five hints of 0, 1, 26, 27, and 28. Specific implementation details are as
follows:
z Hint =0, 1: Here the PREFX instruction prefetches in accordance with the MIPS specification, that is, it
prefetches one cacheline at a time. The difference is that the GS464E only extends the low 16-bit value
47
龙芯 3A3000/3B3000 处理器用户手册 y 下册
symbol in the index register as an index.
z Hint =26: Enter the first-level data cache according to the configuration of continuous prefetch. A
prefetch instruction can automatically prefetch the block of block_NUM with the distance of stride as the
length of block_size from the base address. At this point, part of the information in the index register is
used to configure the generation mode of the prefetch address. Specifically, the [15:0] bit of GPR[index]
is the
16bit signed address offset added to GPR[base]; The [16] bit of GPR[index] is the address
ascending and descending prefetch mark, 0 represents address ascending prefetch, and 1 represents
address descending prefetch. The [25:20] bit of GPR[index] is the pre-taken block_sid-1. The basic unit
of block_size is 128bit, so the longest block can contain 64×16=1KB of data. The [39:32] bit of
GPR[index] is block_num-1, so it supports pre-fetching of up to 256 blocks. Index [59:44] is the stride
between adjacent blocks, is the signed number, and the unit is byte.
48
龙芯 3A3000/3B3000 处理器用户手册 y 下册
z Hint =27: The hint is configured to continuously pre-fetch exclusive data into a level 1 data cache. In this mode,
the contents of the index register parse and
Hint =26 is exactly the same, except that this mode requires that the retrieved data be in an exclusive state.
z Hint =28: Configure the sequential pre-fetch data to enter the Shared cache. Parsing the contents of
the index register in this mode is exactly the same as hint=26, except that in this mode
the prefetch data is stored in a Shared cache.
If a WAIT_CACHE instruction is executed, the configuration prefetch is stopped. These include but are not
limited to cache, SYNC and SYNCI, and JRHB.
If you execute two PREFX (hint=25,26) in a row, the previous one will be covered with the new one. If you
execute two PREFX (hint=27) in a row, the previous one will be covered with the new one. However, there is no
conflict between the pre-fetch modes of hint=25,26, and hint=27.
2.4.5 WAIT instruction
There is only one mode supported by the WAIT instruction, which executes the same effect no matter what
value the software inserts into the Implementation Dependent Code field in the instruction code.
After the WAIT instruction is submitted in the processor pipeline, the processor core stops pointing and goes
into low-power mode, which continues until it is interrupted by NMI, internal clock interrupt, internal performance
counter interrupt, or external hardware interrupt.
It is important to note that after the WAIT instruction is executed, the data paths of the processor core's first-
level instruction cache, first-level data cache, and second-level sacrifice cache can still handle cache conformance
requests from outside. If the processor core's clock needs to be turned off completely, the software needs to do
more to ensure that it doesn't cause other cores to crash due to unresponsive consistency requests.
2.4.6 SYNC instructions
This processor implements only the SYNC instruction of Stype =0, and stype will leave exceptions for other
values. This instruction ACTS as a memory barrier to ensure that the access operation before SYNC has been
completed (for example, the data of store instruction has been written to dcache, the read and write of uncached has
been completed, and load has retrieved the value to register), and that the access operation after SYNC instruction
has not started yet.
2.4.7 SYNCI instruction
Since GS464E is maintained by the hardware for data consistency between the first-level instruction cache and
the first-level data cache, the implementation of the SYNCI instruction is adjusted. In the existing implementation,
the SYNCI directive, on the one hand, waits until all previous accesses have been executed (the same effect as the
SYNC directive). On the other hand, SYNCI will force subsequent instructions to repoint after execution to ensure
that the pointing unit can also see the execution effect of all accesses before the SYNCI instruction. At the same
time, the address information carried by the SYNCI directive is ignored, so tiB-related exceptions, address error
exceptions, and Cache error exceptions are no longer triggered.
2.4.8 TLBINV and TLBINVF directives
49
龙芯 3A3000/3B3000 处理器用户手册 y 下册
Although the config4.ie field in this processor is set to 0, the TLBINVF instruction (which is executed as
config4.ie =3 as defined in the MIPS specification) is implemented in the processor, that is, when a TLBINVF
instruction is executed, the hardware will invalidate the entire TLB table entry. In addition, the TLBINV directive
is also implemented in GS464E, but its execution effect is different from the MIPS specification definition and is
equivalent to the TLBINVF directive.
2.4.9 CACHE directives
There are three major differences between the CACHE instructions implemented by this processor and the
MIPS64 specification:
1. Relationship between instruction OP
and cache hierarchy
The correspondence between the [1:0] bits of the CACHE instruction OP field in GS464E and the CACHE
hierarchy is different from the definition of the MIPS64 specification, as shown in the table
2-26:
50
龙芯 3A3000/3B3000 处理器用户手册 y 下册
Table 2-26 CACHE instruction OP [1:0] corresponds to the CACHE hierarchy
The op (1-0)
GS464E
MIPS64
Simila
specification
rities
and
differ
ences
betwe
en
0 b00
First-order
First-order
consiste
instruction
instruction cache
nt
cache
0 b01
Level 1 data
Level 1 data
consiste
cache
cache
nt
0 b10
Secondary
Three levels of
Don't
sacrifice cache
cache
agree
0 bl1
Three-level
The second level
Don't
Shared cache
cache
agree
2. Address resolution method of Index class CACHE instruction
The address resolution of the index-class Cache instruction implemented by this processor is different from the
definition of the MIPS64 specification, as shown in Figure 2-8. It is important to note that because of GS464E
lowest several addresses are used to indicate to the operation of the Cache, sacrifice and secondary Cache
associated with level 3 group Shared Cache are 16 road, road, choose the lowest four need account address, so for
the secondary sacrifice Cache and tertiary Shared Cache on the Index Data Load and Index Data Store operation,
can only be read from the Cache line even piece of Data, the Cache line only odd-even two adjacent blocks to fill in
the same Data at the same time.
Figure 2-8. Address resolution format of the Index class CACHE instruction
Loongson 3A1500 chip processor
+ Log2 Log2 (CS/A) (L)
Log2 (L) Between (Log2 (A))
Block
Unused
The
Way
Index
Index
MIPS64 specification
Ceiling(Log2(A)) + Log2(CS/A) + Log2(L)
+ Log2 Log2 (CS/A)
Log2 (L)
(L)
Unused
Way
The
Byte Index
Index
Description:
Assume that the capacity of the Cache being operated is CS, the
degree of association is a-group association, and the size of the
Cache row is L bytes. Then the address number used to select the
path is: Ceiling(Log2(A))
Log2(CS/A) Log2(L) Log2(L) Log2(L)
3. Meaning of CACHE28, CACHE29, CACHE30 and CACHE31 Instructions
51
龙芯 3A3000/3B3000 处理器用户手册 y 下册
In GS464E, the CACHE28, CACHE29, CACHE30, and CACHE31 directives (i.e., the Cache directive of OP
[4:2]=0b111) have different meanings from the MIPS64 specification. These instructions in the godson processor
are all Index Store Data operations, while in MIPS64 specification, they are all Fetch and Lock operations.
4. Meaning of CACHE15 Instruction
In GS464E, CACHE15 instruction is a CACHE instruction related to implementation. Its function is to Fetch
and Lock Scache
Operations, specific functional definitions similar to the CACHE31 directive in the MIPS64 specification.
2.4.10 Madd.fmt, Msub.fmt, Nmadd.fmt, Nmsub.fMT instruction
In THE MIPS64 specification, if FIR.Has2008=0 or FCSR.MAC2008=0, then madd.fmt, msub.fmt,
nmadd.fmt
52
龙芯 3A3000/3B3000 处理器用户手册 y 下册
And Nmsub.fMT instruction to round the intermediate result after the multiplication operation, then conduct
subsequent addition and subtraction operation, and round the final result after addition and subtraction again.
Although GS464E does not support the ANSI/IEEE754-2008 binary floating point standard, the floating-point
multiplication plus class operation is performed only once at the final result, which is the Fuse-Multiply-Add
operation defined by the ANSI/IEEE754-2008 binary floating point standard.
2.4.11 EHB, SSNOP instructions
All execution correlation between GS464E instructions is handled by the hardware, so EHB and SSNOP
instructions are treated as NOP instructions in the processor.
2.4.12 DI and EI instructions
DI instruction can only be used in the 3A3000C/3B3000C model, which is compatible with the MIPS64
definition.
EI instruction can only be used in 3A3000C/3B3000C type, but the specific definition is slightly different
from MIPS64 definition: after EI execution, the IE position of the Status register can be set to 0, and the global
terminal enable can be turned off. However, only the lowest 3 bits in the return value are meaningful, which
correspond to the ERL, EXL and IE bits of the Status register before EI execution.
2.5 Loong core expansion command set
GS464E extension instructions include the following categories according to their functions:
y Access class instruction
y Arithmetic and logical operation instructions
y X86 binary acceleration instruction
y ARM binary acceleration instruction
y
64-bit multimedia instructions
y Misc
ellaneous
access class
instruction
Table 2-27 Instruction of longson extended access class
Instruction
Instruction
ISA
mnemonic
function
Compatible
description
Category
SETMEM
Access instruction extension prefix
LoongEXT32
LWDIR
32-bit page table directory entry access instruction
LoongEXT32
LWPTE
32-bit page table entry access instruction
LoongEXT32
LDDIR
64-bit page table directory entry access instruction
LoongEXT64
53
龙芯 3A3000/3B3000 处理器用户手册 y 下册
LDPTE
64-bit page table entry access instructions
LoongEXT64
GSLE
Exception if less than or equal to set address error
LoongEXT32
GSGT
Exception if greater than set address error
LoongEXT32
GSLWLC1 1
Take the left part of the word to the floating point register
LoongEXT32
GSLWRC1 2
Take the right part of the word to the floating point register
LoongEXT32
GSLDLC1
Take the left part of the double word to the floating point register
LoongEXT32
54
龙芯 3A3000/3B3000 处理器用户手册 y 下册
Instruction
Instruction
ISA
mnemonic
function
Compatible
description
Category
GSLDRC1
Take the right part of the double word to the floating point register
LoongEXT32
GSLBLE
Fetch byte with an out-of-bounds check
LoongEXT32
GSLBGT
The fetch byte with the next out-of-bounds check
LoongEXT32
GSLHLE
Take a half-character with an out-of-bounds check
LoongEXT32
GSLHGT
Take a half word with cross - boundary check below
LoongEXT32
GSLWLE
Take the word for cross - border inspection
LoongEXT32
GSLWGT
Take word with cross - boundary check
LoongEXT32
GSLDLE
Take double characters with cross - boundary check
LoongEXT64
GSLDGT
Take double characters with cross - boundary check below
LoongEXT64
GSLWLEC1
Fetch-word to the floating-point register with an out-of-bounds check
LoongEXT32
GSLWGTC1
Fetch-word to the floating-point register with the down - bound check
LoongEXT32
GSLDLEC1
Take a double word to the floating point register with an out - of - bounds
LoongEXT64
check
GSLDGTC1
Take a double word to the floating point register with the next cross check
LoongEXT64
GSLQ
Double target register take - point four word
LoongEXT64
GSLQC1
The dual target register takes a floating point quadword
LoongEXT64
GSLBX
Fetch byte with offset
LoongEXT32
GSLHX
Take half word with offset
LoongEXT32
GSLWX
Take word with offset
LoongEXT32
GSLDX
Take a double word with offset
LoongEXT64
GSLWXC1
Floating-point words with offset
LoongEXT32
GSLDXC1
Floating point double word with offset
LoongEXT32
GSSWLC1
Saves the left part of a word from a floating-point register
LoongEXT32
GSSWRC1
Saves the right part of the word from the floating point register
LoongEXT32
GSSDLC1
Saves the left part of a double word from a floating point register
LoongEXT32
GSSDRC1
Saves the right part of the double word from the floating point register
LoongEXT32
GSSBLE
Carries bytes that are checked out of bounds
LoongEXT32
GSSBGT
Save bytes with down - bound check
LoongEXT32
GSSHLE
Take the cross - border check of the storage half - word
LoongEXT32
GSSHGT
Take the next cross - check of the storage half - word
LoongEXT32
GSSWLE
Carry the word that crosses the line to check
LoongEXT32
GSSWGT
Bring down the word of cross - border check
LoongEXT32
GSSDLE
Take the cross - border check of the save double - character
LoongEXT64
GSSDGT
Bring down the cross - check of the save double - character
LoongEXT64
GSSWLEC1
Saves a word from a floating-point register with an out-of-bounds check
LoongEXT32
GSSWGTC1
Saves a word from a floating-point register with an overbounds check
LoongEXT32
GSSDLEC1
Save double words from the floating point register with an out - of -
LoongEXT32
bounds check
GSSDGTC1
Save a double word from a floating-point register with an overbounds
LoongEXT32
check
GSSQ
Dual source register saves fixed point four characters
LoongEXT64
GSSQC1
Dual source register saves fixed point four characters
LoongEXT64
55
龙芯 3A3000/3B3000 处理器用户手册 y 下册
GSSBX
Offset bytes
LoongEXT32
GSSHX
Offset memory halfwords
LoongEXT32
56
龙芯 3A3000/3B3000 处理器用户手册 y 下册
Instruction
Instruction
ISA
mnemonic
function
Compatible
description
Category
GSSWX
Offset memory words
LoongEXT32
GSSDX
Offset memory double word
LoongEXT64
GSSWXC1
Floating point words with offset
LoongEXT32
GSSDXC1
Floating point double word with offset
LoongEXT32
Arithmetic and
logical operation
Table 2-28 Loongson extended arithmetic and logic operation instructions
instructions
Instruction
Instruction
ISA
mnemonic
function
Compatible
description
Category
GSANDN
General purpose register logical bit non - and
LoongEXT32
GSORN
General purpose register logical bit non or
LoongEXT32
GSADC. D.
Double word with carry
LoongEXT64
GSADC w.
Carry word plus
LoongEXT32
GSADC. H
Carry half word plus
LoongEXT32
GSADC. B
Carry byte plus
LoongEXT32
GSADCU. D.
Unsigned double word plus with carry
LoongEXT64
GSADCU w.
Carry unsigned word addition
LoongEXT32
GSADCU. H
Unsigned halfword with carry
LoongEXT32
GSADCU. B
Unsigned byte plus with carry
LoongEXT32
GSSBC. D.
With borrow double word minus
LoongEXT64
GSSBC w.
Take the debit word minus
LoongEXT32
GSSBC. H
Take the debit half word minus
LoongEXT32
GSSBC. B
Subtract with borrow
LoongEXT32
GSSBCU. D.
Unsigned double word subtraction with debit
LoongEXT64
GSSBCU w.
Minus unsigned word with debit
LoongEXT32
GSSBCU. H
Unsigned half-word subtraction with debit
LoongEXT32
GSSBCU. B
Unsigned byte with debit
LoongEXT32
GSMULT
A sign word is multiplied, and the result is written in the general purpose
LoongEXT32
register
GSDMULT
Sign double word multiplication, the result is written to the general
LoongEXT64
purpose register
GSMULTU
Unsigned word multiplication, the result is written to the general purpose
LoongEXT32
register
GSDMULTU
Unsigned double word multiplication, the result is written to the general
LoongEXT64
purpose register
GSDIV
Sign word division, quotient write general register
LoongEXT32
GSDDIV
Sign double - word division, quotient - write general - purpose register
LoongEXT64
GSDIVU
Unsigned word division, quotient write general register
LoongEXT32
GSDDIVU
Unsigned double - word division, quotient write general - purpose register
LoongEXT64
GSMOD
Sign word division, remainder write general purpose register
LoongEXT32
GSDMOD
Sign double - word division, remainder write general - purpose register
LoongEXT64
GSMODU
Unsigned word division, remainder write general purpose register
LoongEXT32
57
GSDMODU
Unsigned double - word division, remainder write general - purpose
LoongEXT64
register
GSROTR. H
Move the half-word loop to the right龙芯 3A3000/3B3000 处理器用户LoongEXT32
GSROTR. B
The byte cycle moves right
LoongEXT32
GSROTRV. H
Variable shift halfword circulates to the right
LoongEXT32
58
龙芯 3A3000/3B3000 处理器用户手册 y 下册
Instruction
Instruction
ISA
mnemonic
function
Compatible
description
Category
GSROTRV. B
The variable shift byte cycles right
LoongEXT32
GSRCR. D.
Double word loop with CF bit moves right
LoongEXT64
GSRCR w.
The word loop with CF bit moves right
LoongEXT32
GSRCR. H
A half-word loop with CF bits moves to the right
LoongEXT32
GSRCR. B
The byte loop with CF bits is moved right
LoongEXT32
GSDRCR32
Shift + 32 double word cycle with CF bit shifts to the right
LoongEXT64
GSRCRV. D.
A double word loop with a variable shift with CF bits shifted right
LoongEXT64
GSRCRV w.
A word with a variable shift with CF bits circulates to the right
LoongEXT32
GSRCRV. H
A half-word cycle with a variable shift with CF bits moves right
LoongEXT32
GSRCRV. B
A byte cycle with a variable shift with CF bits shifts right
LoongEXT32
Binary translation Acceleration instruction (X86)
Table 2-29 Longson extension X86 binary translation acceleration instructions
Instruction
Instruction
ISA
mnemonic
function
Compatible
description
Category
SETX86FLAG. D.
The EFLAG two-word mode prefix instruction is set in x86 mode
LoongEXT64
SETX86FLAG w.
The EFLAG prefix instruction is placed in x86 mode
LoongEXT32
SETX86FLAG. H
EFLAG halfword mode prefix instruction is set in x86 mode
LoongEXT32
SETX86FLAG. B
The EFLAG byte mode prefix instruction is set in x86 mode
LoongEXT32
X86AND. D.
Only the two-word logical bits and of EFLAG are set in x86 mode
LoongEXT64
X86AND w.
Only the logical bits and words of EFLAG are set in x86 mode
LoongEXT32
X86AND. H
Only the half-word logical bits and of EFLAG are set in x86 mode
LoongEXT32
X86AND. B
Only the byte logical bits and bytes of EFLAG are set in x86 mode
LoongEXT32
X86OR. D.
Only the two-word logical bits or of EFLAG are set in x86 mode
LoongEXT64
X86OR w.
Only the word logical bits or of EFLAG are set in x86 mode
LoongEXT32
X86OR. H
Only the half-word logical bits or of EFLAG are set in x86 mode
LoongEXT32
X86OR. B
Only the byte logical bits or of EFLAG are set in x86 mode
LoongEXT32
X86XOR. D.
Only the two-word logical bit xor of EFLAG is set in x86 mode
LoongEXT64
X86XOR w.
Only the logical bit or of EFLAG is set in x86 mode
LoongEXT32
X86XOR. H
Only half-word logical bits or of EFLAG are set in x86 mode
LoongEXT32
X86XOR. B
Only the byte logical bit xor of EFLAG is set in x86 mode
LoongEXT32
X86ADD. D.
Only EFLAG double characters are set in x86 mode
LoongEXT64
X86ADD w.
Only EFLAG characters are set in x86 mode
LoongEXT32
X86ADD. H
Only EFLAG halfwords are set in x86 mode
LoongEXT32
X86ADD. B
Only the bytes of EFLAG are set in x86 mode
LoongEXT32
X86ADC. D.
In x86 mode, only EFLAG's carry double - word addition is set
LoongEXT64
X86ADC w.
Only EFLAG's carry characters are set in x86 mode
LoongEXT32
X86ADC. H
Only EFLAG's carry halfwords are set in x86 mode
LoongEXT32
X86ADC. B
Only the carry bytes of EFLAG are set in x86 mode
LoongEXT32
X86ADDU. D.
Set EFLAG in x86 only with no exception of double characters
LoongEXT64
59
龙芯 3A3000/3B3000 处理器用户手册 y 下册
X86ADDU w.
Set EFLAG in x86 only without exception
LoongEXT32
X86SUB. D.
Only double-word subtractions of EFLAG are set in x86 mode
LoongEXT64
60
龙芯 3A3000/3B3000 处理器用户手册 y 下册
Instruction
Instruction
ISA
mnemonic
function
Compatible
description
Category
X86SUB w.
Only the word subtracts of EFLAG are set in x86 mode
LoongEXT32
X86SUB. H
Only half - character subtractions of EFLAG are set in x86 mode
LoongEXT32
X86SUB. B
Only the byte reductions of EFLAG are set in x86 mode
LoongEXT32
X86SBC. D.
In x86 mode, only EFLAG's band borrow double-word subtraction is set
LoongEXT64
X86SBC w.
Only EFLAG's loanword subtraction is set in x86 mode
LoongEXT32
X86SBC. H
Only EFLAG's band debit halfwords are set in x86 mode
LoongEXT32
X86SBC. B
Only EFLAG's band borrow bytes are set in x86 mode
LoongEXT32
X86SUBU. D.
Only the double-word subtraction of EFLAG is set in x86 mode without
LoongEXT64
exception
X86SUBU w.
Set EFLAG in x86 only with no exception word subtractions
LoongEXT32
X86INC. D.
Only double - character augmentation 1 of EFLAG is set in x86 mode
LoongEXT64
X86INC w.
Only EFLAG's word augmentation of 1 is set in x86 mode
LoongEXT32
X86INC. H
Only half - word augmentation of EFLAG is set in x86 mode
LoongEXT32
X86INC. B
In x86 only the byte augmentation of EFLAG is set to 1
LoongEXT32
X86DEC. D.
Only the two-word autodecrement of EFLAG is set in x86 mode
LoongEXT64
X86DEC w.
Only the characters of EFLAG are subtracted by 1 in x86 mode
LoongEXT32
X86DEC. H
Only EFLAG halfwords are subtracted by 1 in x86 mode
LoongEXT32
X86DEC. B
In x86 only the bytes of EFLAG are subtracted by 1
LoongEXT32
X86SLL. D.
Only EFLAG's two-word left shift is set in x86 mode
LoongEXT64
X86SLL w.
Only the left shift of EFLAG words is set in x86 mode
LoongEXT32
X86SLL. H
Only the left shift of EFLAG halfwords is set in x86 mode
LoongEXT32
X86SLL. B
Only the bytes of EFLAG are set left in x86 mode
LoongEXT32
X86DSLL32
Only the shift amount of EFLAG plus the left shift of 32 two-word logic
LoongEXT64
is set in x86 mode
X86SLLV. D.
Only the two-word variable shift of EFLAG is set to the left in x86 mode
LoongEXT64
X86SLLV w.
Only the variable word shift of EFLAG is set left in x86 mode
LoongEXT32
X86SLLV. H
Only the half-word variable shift of EFLAG is set left in x86 mode
LoongEXT32
X86SLLV. B
In x86 only the byte variable shift of EFLAG is set to move left
LoongEXT32
X86SRL. D.
Only EFLAG's two-word logic is set to right shift in x86 mode
LoongEXT64
X86SRL w.
Only the right shift of EFLAG's word logic is set in x86 mode
LoongEXT32
X86SRL. H
Only EFLAG's half-word logic is set to right shift in x86 mode
LoongEXT32
X86SRL. B
Only the byte logic of EFLAG is set right in x86 mode
LoongEXT32
X86DSRL32
Only the shift of EFLAG plus the right shift of 32 double word logic is set
LoongEXT64
in x86 mode
X86SRLV. D.
Only the two-word variable shift logical right shift of EFLAG is set in
LoongEXT64
x86 mode
X86SRLV w.
Only the logical right shift of EFLAG's variable word shift is set in x86
LoongEXT32
mode
X86SRLV. H
Only the half-word variable shift logical right shift of EFLAG is set in
LoongEXT32
x86 mode
X86SRLV. B
Only the byte variable shift logical right shift of EFLAG is set in x86
LoongEXT32
mode
X86SRA. D.
Only EFLAG's two-word arithmetic right-shift is set in x86 mode
LoongEXT64
X86SRA w.
In x86 only the word arithmetic shift of EFLAG is set
LoongEXT32
61
龙芯 3A3000/3B3000 处理器用户手册 y 下册
X86SRA. H
Only the half-word arithmetic right shift of EFLAG is set in x86 mode
LoongEXT32
X86SRA. B
Only the byte arithmetic right shift of EFLAG is set in x86 mode
LoongEXT32
X86DSRA32
Only the shift of EFLAG plus the right shift of 32 two-word logical
LoongEXT64
arithmetic is set in x86 mode
X86SRAV. D.
In x86 only the two-word variable shift of EFLAG is set to the arithmetic
LoongEXT64
right shift
62
龙芯 3A3000/3B3000 处理器用户手册 y 下册
Instruction
Instruction
ISA
mnemonic
function
Compatible
description
Category
X86SRAV w.
In x86 only the variable word shift of EFLAG is set to the arithmetic right
LoongEXT32
shift
X86SRAV. H
In x86 only the half-word variable shift of EFLAG is set to the arithmetic
LoongEXT32
right shift
X86SRAV. B
In x86 only the byte variable shift of EFLAG is set to the arithmetic right
LoongEXT32
shift
X86ROTR. D.
Only the right shift of EFLAG's two-word loop is set in x86 mode
LoongEXT64
X86ROTR w.
Only EFLAG's word loops are set right in x86 mode
LoongEXT32
X86ROTR. H
Only EFLAG's half-word loop is set to right shift in x86 mode
LoongEXT32
X86ROTR. B
In x86 only the byte cycle of EFLAG is set to move right
LoongEXT32
X86DROTR32
In x86 only the shift amount of EFLAG plus the right shift of 32 two-
LoongEXT64
word logic loops
X86ROTL. D.
Only EFLAG double-word loops are set left in x86 mode
LoongEXT64
X86ROTL w.
Only the left shift of EFLAG's word loop is set in x86 mode
LoongEXT32
X86ROTL. H
Only EFLAG's half-word loop is set left in x86 mode
LoongEXT32
X86ROTL. B
Only the byte loop of EFLAG is set left in x86 mode
LoongEXT32
X86DROTL32
Only the shift amount of EFLAG plus the left shift of 32 two-word logic
LoongEXT64
loop is set in x86 mode
X86RCR. D.
In x86 only the CF bits of EFLAG are set to right shift
LoongEXT64
X86RCR w.
In x86 only the CF bits of EFLAG are set to right shift
LoongEXT32
X86RCR. H
In x86 only the CF bits of EFLAG are set to right shift
LoongEXT32
X86RCR. B
In x86, the byte cycle with CF bits of EFLAG is set right
LoongEXT32
Only the shift of EFLAG plus 32 double - word logic with CF bit is set in x86 mode
X86DRCR32
LoongEXT64
Ring moves to the right
X86RCL. D.
In x86 only the CF bits of EFLAG are set to the left of a two-word loop
LoongEXT64
X86RCL w.
In x86 only the CF bits of EFLAG are set to loop left
LoongEXT32
X86RCL. H
In x86 only the CF bit of EFLAG's half-word loop is set to move left
LoongEXT32
X86RCL. B
In x86 mode, only the byte loops with CF bits of EFLAG are set left
LoongEXT32
Only the shift of EFLAG plus 32 double - word logic with CF bit is set in x86 mode
X86DRCL32
LoongEXT64
Ring left
X86ROTRV. D.
In x86 mode, only the variable shift amount of EFLAG is set for a two-
LoongEXT64
word cycle right shift
X86ROTRV w.
In x86, a word cycle that only sets the variable shift amount of EFLAG
LoongEXT32
moves right
X86ROTRV. H
In x86 mode, only the variable shift amount of EFLAG is set for a half-
LoongEXT32
word loop right shift
X86ROTRV. B
In x86 mode, the byte cycle that sets only the variable shift amount of
LoongEXT32
EFLAG moves right
X86ROTLV. D.
In x86 mode, only the variable shift of EFLAG's two-word loop is set to
LoongEXT64
the left
X86ROTLV w.
In x86 mode, only the variable shift amount of EFLAG is set to the left of
LoongEXT32
the word loop
X86ROTLV. H
In x86 mode, only the variable shift amount of EFLAG is set to the left of
LoongEXT32
a halfword loop
X86ROTLV. B
In x86 mode, only the variable shift amount of EFLAG is set to the left of
LoongEXT32
the byte cycle
X86RCRV. D.
In x86 only the variable shift of EFLAG with CF bits is set to the right
LoongEXT64
X86RCRV w.
In x86 only a CF bit of a word with a variable shift of EFLAG is set right
LoongEXT32
63
龙芯 3A3000/3B3000 处理器用户手册 y 下册
X86RCRV. H
In x86 only a CF bit halfword loop with a variable shift of EFLAG is set
LoongEXT32
right
X86RCRV. B
In x86 only the variable shift of EFLAG's byte cycle with CF bits is set to
LoongEXT32
the right
X86RCLV. D.
In x86 only a CF bit double-word loop with a variable shift of EFLAG is
LoongEXT64
set to the left
X86RCLV w.
In x86 only the CF bits of a word with a variable shift of EFLAG are set
LoongEXT32
left
X86RCLV. H
In x86 only a CF bit halfword loop with a variable shift of EFLAG is set
LoongEXT32
left
X86RCLV. B
In x86 only the variable shift of EFLAG's byte cycle with CF bits is set to
LoongEXT32
the left
64
龙芯 3A3000/3B3000 处理器用户手册 y 下册
Instruction
Instruction
ISA
mnemonic
function
Compatible
description
Category
X86MFFLAG
The value of the EFLAG bit is extracted in x86 mode
LoongEXT32
X86MTFLAG
Modify the value of the EFLAG flag bit in an x86 manner
LoongEXT32
X86J
Jump in X86 mode based on EFLAG values
LoongEXT32
X86LOOP
The loop is based on EFLAG values in X86 mode
LoongEXT32
SETTM
The x86 floating-point stack mode is set
LoongEXT32
CLRTM
X86 floating-point stack mode cleared
LoongEXT32
INCTOP
Pointer to the top of the x86 floating-point stack plus 1
LoongEXT32
DECTOP
The pointer at the top of the x86 floating-point stack minus 1
LoongEXT32
MTTOP
Write the x86 floating-point top pointer
LoongEXT32
MFTOP
Read the pointer to the top of the x86 floating-point stack
LoongEXT32
SETTAG
Determines the collocation register
LoongEXT32
The CVT
Extended double to double precision
LoongEXT32
transmission. D.L
D
The CVT
Double precision converted to extended double precision low position
LoongEXT32
transmission. LD.
D
The CVT
Double precision converted to extended double precision high order
LoongEXT32
transmission. UD.
D
Binary translation Acceleration instruction (ARM)
Table 2-30 Acceleration instructions for ARM binary translation extension
Instruction
Instruction
ISA
mnemonic
function
Compatible
description
Category
ARMADC
Only EFLAG is set with carry word and ARM conditional execution
LoongEXT32
ARMADD
Word plus, ARM conditional execution only set EFLAG
LoongEXT32
ARMSBC
With carry subtracting, only EFLAG is set for ARM conditional execution LoongEXT32
ARMSUB
Subtracting, only EFLAG is set for ARM conditional execution
LoongEXT32
those
Only EFLAG is set in ARM mode conditional execution
LoongEXT32
ARMBIC
Only EFLAG is set in ARM mode conditional execution
LoongEXT32
ARMMOV
The word moves and only EFLAG is set for ARM mode conditional
LoongEXT32
execution
ARMMVN
The word is moved backwards, and only EFLAG is set for ARM
LoongEXT32
conditional execution
Low 32 bit LO register moves to general purpose register, ARM mode
ARMMVLO32
LoongEXT32
conditional execution only
Buy EFLAG
HI register low 32 bit moves to general purpose register, ARM mode
ARMMVHI32
LoongEXT32
conditional execution only Settings
EFLAG
The low 32 bits of the HI register and the low 32 bits of the LO register
ARMMVACC64
LoongEXT32
are combined into 64 bits in ARM mode
Conditional execution sets only EFLAG
ARMOR
Only EFLAG is set in ARM mode conditional execution
LoongEXT32
65
龙芯 3A3000/3B3000 处理器用户手册 y 下册
ARMORN
Only EFLAG is set in ARM mode conditional execution
LoongEXT32
ARMXOR
Only EFLAG is set in ARM mode conditional execution
LoongEXT32
ARMSLL
Move the word left and set only EFLAG in ARM mode conditional
LoongEXT32
execution
ARMSLLV
The variable shift quantity word moves left, and only EFLAG is set in
LoongEXT32
ARM mode conditional execution
ARMSRL
The word logic is shifted to the right, and only EFLAG is set for ARM
LoongEXT32
mode conditional execution
ARMSRLV
Variable shift quantity word logic moves right, and only EFLAG is set in
LoongEXT32
ARM mode conditional execution
ARMSRA
The word arithmetic is shifted to the right, and only EFLAG is set for
LoongEXT32
ARM conditional execution
ARMSRAV
The variable shift quantity word arithmetic moves right, and only EFLAG
LoongEXT32
is set in ARM mode conditional execution
66
龙芯 3A3000/3B3000 处理器用户手册 y 下册
Instruction
Instruction
ISA
mnemonic
function
Compatible
description
Category
ARMROTR
The word loops right, and only EFLAG is set for ARM conditional
LoongEXT32
execution
ARMROTRV
The variable shifted quantity word circularly moves right, and only
LoongEXT32
EFLAG is set in ARM mode conditional execution
ARMRRX
ARM conditional performs a right shift of a carry word loop that sets
LoongEXT32
EFLAG only
ARMFCMP. F32
ARM FCMP.f32 instruction is used for floating point comparison and
LoongEXT32
FCR1 EFLAGS are set
ARMFCMP. F64
ARM FCMP.f64 instruction is used for floating point comparison and
LoongEXT32
FCR1 EFLAGS are set
ARMFCMPE. F32
ARM FCMP.f32 instruction is used for floating point comparison and
LoongEXT32
FCR1 EFLAGS are set
ARMFCMPE. F64
ARM FCMP.f64 instruction is used for floating point comparison and
LoongEXT32
FCR1 EFLAGS are set
ARMMOVE
The general purpose registers move between each other in ARM mode
LoongEXT32
according to EFLAG conditions
ARMMFHI
HI moves to the general register in ARM mode according to EFLAG
LoongEXT32
conditions
ARMMFLO
LO moves to the general purpose register in ARM mode according to
LoongEXT32
EFLAG conditions
ARMMFFCR
Move FCR1 EFLAGS to EFLAGS by ARM according to EFLAG
LoongEXT32
conditions
ARMFMOV. S
Move single precision Numbers between floating point registers in ARM
LoongEXT32
mode according to EFLAG conditions
ARMFMOV. D.
Floating point registers move doubles in ARM mode according to EFLAG
LoongEXT32
conditions
ARMJ
Jump according to EFLAG value in ARM mode
LoongEXT32
ARMMFFLAG
The value of EFLAG bit is extracted by ARM
LoongEXT32
ARMMTFLAG
Modify the value of the EFLAG bit by ARM
LoongEXT32
64-bit multimedia acceleration instructions
Table 2-31 Loong core extension 64-bit multimedia acceleration instructions
Instruction
Instruction
ISA
mnemonic
function
Compatible
description
Category
PADDSH
Four 16-bit signed integers plus, sign saturation
LoongEXT32
PADDUSH
Four 16-bit unsigned integers plus, unsigned saturation
LoongEXT32
PADDH
Four 16-digit Numbers plus
LoongEXT32
PADDW
Two 32 digits plus
LoongEXT32
PADDSB
Eight 8-bit signed integers plus, sign saturation
LoongEXT32
PADDUSB
Eight 8-bit unsigned integers plus, unsigned saturation
LoongEXT32
PADDB
Eight 8-digit Numbers plus
LoongEXT32
PADDD
64 digits add
LoongEXT32
PSUBSH
Four 16-bit signed integers minus, sign saturation
LoongEXT32
PSUBUSH
Four 16-bit unsigned integers minus, unsigned saturation
LoongEXT32
PSUBH
Four 16-digit subtractions
LoongEXT32
PSUBW
Two 32-digit subtractions
LoongEXT32
PSUBSB
Eight 8-bit signed integers minus, sign saturation
LoongEXT32
67
|
|