ISA Viewer

SM100 (Blackwell) Instructions

163 base instructions, 625 total variants

ACQBULK
Wait for Bulk Release Status Warp State
1
ACQSHMINIT
Wait for Shared Memory Initialization Release Status Warp State
1
AL2P
Unknown op
unknown
ALD
Unknown op
unknown
ATOM
Atomic Operation on Generic Memory
10
ATOMG
Atomic Operation on Global Memory
10
ATOMS
Atomic Operation on Shared Memory
8
B2R
Move Barrier To Register
3
BMSK
Bitfield Mask
3
BREV
Bit Reverse
3
CCTL
Cache Control
12
CGAERRBAR
CGA Error Barrier
1
CREDUX
Coupled Reduction of a Vector Register into a Uniform Register
1
CS2R
Move Special Register to Register
1
CSMTEST
Unknown op
unknown
DADD
FP64 Add
3
DFMA
FP64 Fused Mutiply Add
5
DMMA
Matrix Multiply and Accumulate
2
DMUL
FP64 Multiply
3
DSETP
FP64 Compare And Set Predicate
3
ELECT
Elect a Leader Thread
2
ENDCOLLECTIVE
Reset the MCOLLECTIVE mask
1
ERRBAR
Error Barrier
1
F2FP
Unknown op
unknown
F2I
Floating Point To Integer Conversion
2
F2IP
FP32 Down-Convert to Integer and Pack
5
FADD
FP32 Add
3
FADD2
FP32 Add
3
FCHK
Floating-point Range Check
3
FFMA
FP32 Fused Multiply and Add
5
FFMA2
FP32 Fused Multiply and Add
5
FHADD
FP32 Addition
2
FHFMA
FP32 Fused Multiply and Add
3
FLO
Find Leading One
3
FMNMX
FP32 Minimum/Maximum
6
FMNMX3
3-Input Floating-point Minimum / Maximum
3
FMUL
FP32 Multiply
3
FMUL2
FP32 Multiply
3
FOOTPRINT
Unknown op
unknown
FRND
Round To Integer
3
FSEL
Floating Point Select
3
FSET
FP32 Compare And Set
3
FSETP
FP32 Compare And Set Predicate
3
FSWZADD
FP32 Swizzle Add
1
GETLMEMBASE
Get Local Memory Base Address
1
HADD2
FP16 Add
4
HFMA2
FP16 Fused Mutiply Add
10
HMMA
Matrix Multiply and Accumulate
4
HMNMX2
FP16 Minimum / Maximum
6
HMUL2
FP16 Multiply
3
HSET2
FP16 Compare And Set
3
HSETP2
FP16 Compare And Set Predicate
3
I2F
Integer To Floating Point Conversion
3
I2FP
Integer to FP32 Convert and Pack
5
I2I
Integer To Integer Conversion
3
I2IP
Integer To Integer Conversion and Packing
3
IABS
Integer Absolute Value
3
IADD
Integer Addition
6
IADD3
3-input Integer Addition
6
IDP
Integer Dot Product and Accumulate
2
IMAD
Integer Multiply And Add
26
IMMA
Integer Matrix Multiply and Accumulate
4
IMNMX
Integer Minimum/Maximum
3
IMUL
Integer Multiply
9
IPA
Unknown op
unknown
ISBERD
Unknown op
unknown
ISETP
Integer Compare And Set Predicate
6
LD
Load from generic Memory
8
LDC
Load Constant
4
LDG
Load from Global Memory
12
LDGDEPBAR
Global Load Dependency Barrier
1
LDGSTS
Asynchronous Global to Shared Memcopy
4
LDL
Load within Local Memory Window
6
LDS
Load within Shared Memory Window
4
LDSM
Load Matrix from Shared Memory with Element Size Expansion
4
LDTRAM
Unknown op
unknown
LEA
LOAD Effective Address
14
LEPC
Load Effective PC
2
LOP3
Logic Operation
3
MATCH
Match Register Values Across Thread Group
2
MOV
Move
4
MOVM
Move Matrix with Transposition or Expansion
1
MUFU
FP32 Multi Function Operation
3
NOP
No Operation
1
OUT
Unknown op
unknown
P2R
Move Predicate Register To Register
3
PIXLD
Unknown op
unknown
PLOP3
Predicate Logic Operation
8
PMTRIG
Performance Monitor Trigger
1
POPC
Population count
3
PREEXIT
Dependent Task Launch Hint
1
PRMT
Permute Register Pair
5
QADD4
Unknown op
unknown
QFMA4
Unknown op
unknown
QMUL4
Unknown op
unknown
QSPC
Query Space
4
R2P
Move Register To Predicate Register
3
R2UR
Move from Vector Register to a Uniform Register
1
REDAS
Asynchronous Reduction on Distributed Shared Memory With Explicit Synchronization
2
REDUX
Reduction of a Vector Register into a Uniform Register
1
RPCMOV
PC Register Move
6
S2R
Move Special Register to Register
1
S2UR
Move Special Register to Uniform Register
1
SEL
Select Source with Predicate
3
SETCTAID
Set CTA ID
1
SGXT
Sign Extend
3
SHF
Funnel Shift
5
SHFL
Warp Wide Register Shuffle
4
STAS
Asynchronous Store to Distributed Shared Memory With Explicit Synchronization
2
SUATOM
Atomic Op on Surface Memory
4
SULD
Surface Load
4
SUQUERY
Unknown op
unknown
SYNCS
Sync Unit
16
TEX
Texture Fetch
6
TLD
Texture Load
6
TLD4
Texture Load 4
6
TMML
Texture MipMap Level
6
TXD
Texture Fetch With Derivatives
6
TXQ
Texture Query
2
UBLKCP
Bulk Data Copy
2
UBLKPF
Bulk Data Prefetch
2
UBLKRED
Bulk Data Copy from Shared Memory with Reduction
2
UBMSK
Uniform Bitfield Mask
2
UBREV
Uniform Bit Reverse
2
UCGABAR_ARV
CGA Barrier Synchronization
1
UCGABAR_WAIT
CGA Barrier Synchronization
1
UCLEA
Load Effective Address for a Constant
2
UF2FP
Uniform FP32 Down-convert and Pack
4
UFLO
Uniform Find Leading One
2
UIADD3
Uniform Integer Addition
8
UIMAD
Uniform Integer Multiplication
10
UISETP
Uniform Integer Compare and Set Uniform Predicate
4
ULEA
Uniform Load Effective Address
10
ULEPC
Uniform Load Effective PC
2
ULOP3
Uniform Logic Operation
2
UMOV
Uniform Move
2
UP2UR
Uniform Predicate to Uniform Register
2
UPLOP3
Uniform Predicate Logic Operation
4
UPOPC
Uniform Population Count
2
UPRMT
Uniform Byte Permute
2
UR2UP
Uniform Register to Uniform Predicate
2
USEL
Uniform Select
2
USETMAXREG
Release, Deallocate and Allocate Registers
1
USETSHMSZ
Unknown op
unknown
USGXT
Uniform Sign Extend
2
USHF
Uniform Funnel Shift
3
UTCATOMSWS
Perform Atomic operation on SW State Register
3
UTMACCTL
TMA Cache Control
2
UTMACMDFLUSH
TMA Command Flush
1
UTMALDG
Tensor Load from Global to Shared Memory
4
UTMAPF
Tensor Prefetch
4
UTMAREDG
Tensor Store from Shared to Global Memory with Reduction
2
UTMASTG
Tensor Store from Shared to Global Memory
2
UVIRTCOUNT
Virtual Resource Management
2
VABSDIFF
Absolute Difference
5
VABSDIFF4
Absolute Difference
5
VHMNMX
SIMD FP16 3-Input Minimum / Maximum
3
VIADD
SIMD Integer Addition
3
VIADDMNMX
SIMD Integer Addition and Fused Min/Max Comparison
5
VIMNMX
SIMD Integer Minimum / Maximum
3
VIMNMX3
SIMD Integer 3-Input Minimum / Maximum
3
VOTE
Vote Across SIMT Thread Group
1
VOTEU
Voting across SIMD Thread Group with Results in Uniform Destination
1

Unfound Instructions

Our fuzzer has not found these 96 instructions. If you have a cubin that contains any of these instructions and would like to contribute it, message us at collab@sf-tensor.com

BAR
Barrier Synchronization
unfound
BMOV
Move Convergence Barrier State
unfound
BPT
BreakPoint/Trap
unfound
BRA
Relative Branch
unfound
BREAK
Break out of the Specified Convergence Barrier
unfound
BRX
Relative Branch Indirect
unfound
BRXU
Relative Branch with Uniform Register Based Offset
unfound
BSSY
Barrier Set Convergence Synchronization Point
unfound
BSYNC
Synchronize Threads on a Convergence Barrier
unfound
CALL
Call Function
unfound
CCTLL
Cache Control
unfound
CCTLT
Texture Cache Control
unfound
CS2UR
Load a Value from Constant Memory into a Uniform Register
unfound
DEPBAR
Dependency Barrier
unfound
EXIT
Exit Program
unfound
F2F
Floating Point To Floating Point Conversion
unfound
FADD32I
FP32 Add
unfound
FENCE
Memory Visibility Guarantee for Shared or Global Memory
unfound
FFMA32I
FP32 Fused Multiply and Add
unfound
FMUL32I
FP32 Multiply
unfound
HADD2_32I
FP16 Add
unfound
HFMA2_32I
FP16 Fused Mutiply Add
unfound
HMUL2_32I
FP16 Multiply
unfound
IADD32I
Integer Addition
unfound
IDP4A
Integer Dot Product and Accumulate
unfound
IMUL32I
Integer Multiply
unfound
ISCADD
Scaled Integer Addition
unfound
ISCADD32I
Scaled Integer Addition
unfound
JMP
Absolute Jump
unfound
JMX
Absolute Jump Indirect
unfound
JMXU
Absolute Jump with Uniform Register Based Offset
unfound
KILL
Kill Thread
unfound
LDCU
Load a Value from Constant Memory into a Uniform Register
unfound
LDGMC
Reducing Load
unfound
LDT
Load Matrix from Tensor Memory to Register File
unfound
LDTM
Load Matrix from Tensor Memory to Register File
unfound
LOP
Logic Operation
unfound
LOP32I
Logic Operation
unfound
MEMBAR
Memory Barrier
unfound
MOV32I
Move
unfound
NANOSLEEP
Suspend Execution
unfound
OMMA
FP4 Matrix Multiply and Accumulate Across a Warp
unfound
PSETP
Combine Predicates and Set Predicate
unfound
QMMA
FP8 Matrix Multiply and Accumulate Across a Warp
unfound
REDG
Reduction Operation on Generic Memory
unfound
RET
Return From Subroutine
unfound
SETLMEMBASE
Set Local Memory Base Address
unfound
SHL
Shift Left
unfound
SHR
Shift Right
unfound
ST
Store to Generic Memory
unfound
STG
Store to Global Memory
unfound
STL
Store to Local Memory
unfound
STS
Store to Shared Memory
unfound
STSM
Store Matrix to Shared Memory
unfound
STT
Store Matrix to Tensor Memory from Register File
unfound
STTM
Store Matrix to Tensor Memory from Register File
unfound
SURED
Reduction Op on Surface Memory
unfound
SUST
Surface Store
unfound
UF2F
Uniform Float-to-Float Conversion
unfound
UF2I
Uniform Float-to-Integer Conversion
unfound
UF2IP
Uniform FP32 Down-Convert to Integer and Pack
unfound
UFADD
Uniform Uniform FP32 Addition
unfound
UFFMA
Uniform FP32 Fused Multiply-Add
unfound
UFMNMX
Uniform Floating-point Minimum / Maximum
unfound
UFMUL
Uniform FP32 Multiply
unfound
UFRND
Uniform Round to Integer
unfound
UFSEL
Uniform Floating-Point Select
unfound
UFSET
Uniform Floating-Point Compare and Set
unfound
UFSETP
Uniform Floating-Point Compare and Set Predicate
unfound
UGETNEXTWORKID
Uniform Get Next Work ID
unfound
UI2F
Uniform Integer to Float conversion
unfound
UI2FP
Uniform Integer to FP32 Convert and Pack
unfound
UI2I
Uniform Saturating Integer-to-Integer Conversion
unfound
UI2IP
Uniform Dual Saturating Integer-to-Integer Conversion and Packing
unfound
UIABS
Uniform Integer Absolute Value
unfound
UIADD3.64
Uniform Integer Addition
unfound
UIMNMX
Uniform Integer Minimum / Maximum
unfound
ULOP
Uniform Logic Operation
unfound
ULOP32I
Uniform Logic Operation
unfound
UMEMSETS
Initialize Shared Memory
unfound
UPSETP
Uniform Predicate Logic Operation
unfound
UREDGR
Uniform Reduction on Global Memory with Release
unfound
USHL
Uniform Left Shift
unfound
USHR
Uniform Right Shift
unfound
USTGR
Uniform Store to Global Memory with Release
unfound
UTCBAR
Tensor Core Barrier
unfound
UTCCP
Asynchonous data copy from Shared Memory to Tensor Memory
unfound
UTCHMMA
Uniform Matrix Multiply and Accumulate
unfound
UTCIMMA
Uniform Matrix Multiply and Accumulate
unfound
UTCOMMA
Uniform Matrix Multiply and Accumulate
unfound
UTCQMMA
Uniform Matrix Multiply and Accumulate
unfound
UTCSHIFT
Shift elements in Tensor Memory
unfound
UVIADD
Uniform SIMD Integer Addition
unfound
UVIMNMX
Uniform SIMD Integer Minimum / Maximum
unfound
WARPSYNC
Synchronize Threads in Warp
unfound
YIELD
Yield Control
unfound