SM100 (Blackwell) Instructions
163 base instructions, 625 total variants
Unfound Instructions
Our fuzzer has not found these 96 instructions. If you have a cubin that contains any of these instructions and would like to contribute it, message us at collab@sf-tensor.com
BAR
Barrier Synchronization
unfound
BMOV
Move Convergence Barrier State
unfound
BPT
BreakPoint/Trap
unfound
BRA
Relative Branch
unfound
BREAK
Break out of the Specified Convergence Barrier
unfound
BRX
Relative Branch Indirect
unfound
BRXU
Relative Branch with Uniform Register Based Offset
unfound
BSSY
Barrier Set Convergence Synchronization Point
unfound
BSYNC
Synchronize Threads on a Convergence Barrier
unfound
CALL
Call Function
unfound
CCTLL
Cache Control
unfound
CCTLT
Texture Cache Control
unfound
CS2UR
Load a Value from Constant Memory into a Uniform Register
unfound
DEPBAR
Dependency Barrier
unfound
EXIT
Exit Program
unfound
F2F
Floating Point To Floating Point Conversion
unfound
FADD32I
FP32 Add
unfound
FENCE
Memory Visibility Guarantee for Shared or Global Memory
unfound
FFMA32I
FP32 Fused Multiply and Add
unfound
FMUL32I
FP32 Multiply
unfound
HADD2_32I
FP16 Add
unfound
HFMA2_32I
FP16 Fused Mutiply Add
unfound
HMUL2_32I
FP16 Multiply
unfound
IADD32I
Integer Addition
unfound
IDP4A
Integer Dot Product and Accumulate
unfound
IMUL32I
Integer Multiply
unfound
ISCADD
Scaled Integer Addition
unfound
ISCADD32I
Scaled Integer Addition
unfound
JMP
Absolute Jump
unfound
JMX
Absolute Jump Indirect
unfound
JMXU
Absolute Jump with Uniform Register Based Offset
unfound
KILL
Kill Thread
unfound
LDCU
Load a Value from Constant Memory into a Uniform Register
unfound
LDGMC
Reducing Load
unfound
LDT
Load Matrix from Tensor Memory to Register File
unfound
LDTM
Load Matrix from Tensor Memory to Register File
unfound
LOP
Logic Operation
unfound
LOP32I
Logic Operation
unfound
MEMBAR
Memory Barrier
unfound
MOV32I
Move
unfound
NANOSLEEP
Suspend Execution
unfound
OMMA
FP4 Matrix Multiply and Accumulate Across a Warp
unfound
PSETP
Combine Predicates and Set Predicate
unfound
QMMA
FP8 Matrix Multiply and Accumulate Across a Warp
unfound
REDG
Reduction Operation on Generic Memory
unfound
RET
Return From Subroutine
unfound
SETLMEMBASE
Set Local Memory Base Address
unfound
SHL
Shift Left
unfound
SHR
Shift Right
unfound
ST
Store to Generic Memory
unfound
STG
Store to Global Memory
unfound
STL
Store to Local Memory
unfound
STS
Store to Shared Memory
unfound
STSM
Store Matrix to Shared Memory
unfound
STT
Store Matrix to Tensor Memory from Register File
unfound
STTM
Store Matrix to Tensor Memory from Register File
unfound
SURED
Reduction Op on Surface Memory
unfound
SUST
Surface Store
unfound
UF2F
Uniform Float-to-Float Conversion
unfound
UF2I
Uniform Float-to-Integer Conversion
unfound
UF2IP
Uniform FP32 Down-Convert to Integer and Pack
unfound
UFADD
Uniform Uniform FP32 Addition
unfound
UFFMA
Uniform FP32 Fused Multiply-Add
unfound
UFMNMX
Uniform Floating-point Minimum / Maximum
unfound
UFMUL
Uniform FP32 Multiply
unfound
UFRND
Uniform Round to Integer
unfound
UFSEL
Uniform Floating-Point Select
unfound
UFSET
Uniform Floating-Point Compare and Set
unfound
UFSETP
Uniform Floating-Point Compare and Set Predicate
unfound
UGETNEXTWORKID
Uniform Get Next Work ID
unfound
UI2F
Uniform Integer to Float conversion
unfound
UI2FP
Uniform Integer to FP32 Convert and Pack
unfound
UI2I
Uniform Saturating Integer-to-Integer Conversion
unfound
UI2IP
Uniform Dual Saturating Integer-to-Integer Conversion and Packing
unfound
UIABS
Uniform Integer Absolute Value
unfound
UIADD3.64
Uniform Integer Addition
unfound
UIMNMX
Uniform Integer Minimum / Maximum
unfound
ULOP
Uniform Logic Operation
unfound
ULOP32I
Uniform Logic Operation
unfound
UMEMSETS
Initialize Shared Memory
unfound
UPSETP
Uniform Predicate Logic Operation
unfound
UREDGR
Uniform Reduction on Global Memory with Release
unfound
USHL
Uniform Left Shift
unfound
USHR
Uniform Right Shift
unfound
USTGR
Uniform Store to Global Memory with Release
unfound
UTCBAR
Tensor Core Barrier
unfound
UTCCP
Asynchonous data copy from Shared Memory to Tensor Memory
unfound
UTCHMMA
Uniform Matrix Multiply and Accumulate
unfound
UTCIMMA
Uniform Matrix Multiply and Accumulate
unfound
UTCOMMA
Uniform Matrix Multiply and Accumulate
unfound
UTCQMMA
Uniform Matrix Multiply and Accumulate
unfound
UTCSHIFT
Shift elements in Tensor Memory
unfound
UVIADD
Uniform SIMD Integer Addition
unfound
UVIMNMX
Uniform SIMD Integer Minimum / Maximum
unfound
WARPSYNC
Synchronize Threads in Warp
unfound
YIELD
Yield Control
unfound