Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
243df61a88 | ||
|
|
21d41b8f98 | ||
|
|
fe6e8e25cf | ||
|
|
047c32b40b | ||
|
|
8a615acc76 | ||
|
|
f6a2652ac7 | ||
|
|
bed8ad9815 | ||
|
|
d0e0d07817 | ||
|
|
ad14fd8081 | ||
|
|
8f0e2cfc94 | ||
|
|
7529766d33 | ||
|
|
f45bd5a570 | ||
|
|
3375a3a668 | ||
|
|
f6c4835da7 | ||
|
|
be0b93b8c6 |
@@ -1,8 +1,8 @@
|
||||
* @ROCm/rocm-documentation
|
||||
* @amd-aakash @jlgreathouse @samjwu @yhuiYH @ROCm/rocm-documentation
|
||||
# Documentation files
|
||||
docs/ @ROCm/rocm-documentation
|
||||
*.md @ROCm/rocm-documentation
|
||||
*.rst @ROCm/rocm-documentation
|
||||
docs/ @amd-aakash @jlgreathouse @samjwu @yhuiYH @ROCm/rocm-documentation
|
||||
*.md @amd-aakash @jlgreathouse @samjwu @yhuiYH @ROCm/rocm-documentation
|
||||
*.rst @amd-aakash @jlgreathouse @samjwu @yhuiYH @ROCm/rocm-documentation
|
||||
# External CI
|
||||
/.azuredevops/ @ROCm/external-ci
|
||||
tools/rocm-build/ @ROCm/rocm-devops
|
||||
|
||||
@@ -1,9 +0,0 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" width="1280" height="640" viewBox="0 0 1280 640" role="img" aria-label="ROCm">
|
||||
<rect width="1280" height="640" fill="#0A0A0A"/>
|
||||
<svg x="96" y="215" width="210" height="210" viewBox="0 0 67 67"><path d="M22.21 67V44.6369H0V67H22.21Z" fill="#fff"/><path d="M66.7038 22.3184H22.2534L0.0878906 44.6367H44.4634L66.7038 22.3184Z" fill="#fff"/><path d="M22.21 0H0V22.3184H22.21V0Z" fill="#fff"/><path d="M66.7198 0H44.5098V22.3184H66.7198V0Z" fill="#fff"/><path d="M66.7198 67V44.6369H44.5098V67H66.7198Z" fill="#fff"/></svg>
|
||||
<text x="378" y="276" font-family="Inter,system-ui,-apple-system,sans-serif" font-size="78" font-weight="800" letter-spacing="-2" fill="#ffffff">ROCm</text>
|
||||
<text x="378" y="322" font-family="Inter,system-ui,sans-serif" font-size="30" fill="#ffffff" opacity=".66">AMD ROCm™ Software - GitHub Home</text>
|
||||
<rect x="378" y="338" width="806" height="3" rx="1.5" fill="#ffffff" opacity=".9"/>
|
||||
<text x="378" y="390" font-family="Inter,system-ui,sans-serif" font-size="24" font-weight="600" fill="#ffffff" opacity=".5">github.com/hanzoai</text>
|
||||
<text x="1184" y="390" text-anchor="end" font-family="Inter,system-ui,sans-serif" font-size="24" font-weight="600" fill="#ffffff" opacity=".5">hanzo.ai</text>
|
||||
</svg>
|
||||
|
Before Width: | Height: | Size: 1.2 KiB |
@@ -17,6 +17,5 @@ __pycache__/
|
||||
|
||||
# avoid duplicating contributing.md due to conf.py
|
||||
docs/contribute/index.md
|
||||
docs/about/release-notes.md
|
||||
docs/release/changelog.md
|
||||
.claude/settings.local.json
|
||||
|
||||
@@ -4,23 +4,21 @@
|
||||
version: 2
|
||||
|
||||
sphinx:
|
||||
configuration: docs/conf.py
|
||||
configuration: docs/conf.py
|
||||
|
||||
formats: [htmlzip]
|
||||
formats: []
|
||||
|
||||
python:
|
||||
install:
|
||||
- requirements: docs/sphinx/requirements.txt
|
||||
install:
|
||||
- requirements: docs/sphinx/requirements.txt
|
||||
|
||||
build:
|
||||
os: ubuntu-22.04
|
||||
tools:
|
||||
python: "3.10"
|
||||
apt_packages:
|
||||
- "doxygen"
|
||||
- "gfortran" # For pre-processing fortran sources
|
||||
- "graphviz" # For dot graphs in doxygen
|
||||
os: ubuntu-24.04
|
||||
tools:
|
||||
python: "3.10"
|
||||
|
||||
search:
|
||||
ignore:
|
||||
- "**/previous-versions/**"
|
||||
- "**/archive/**"
|
||||
- "**/include/**"
|
||||
- "**/redirect/**"
|
||||
|
||||
@@ -1,16 +1,9 @@
|
||||
AAC
|
||||
ABI
|
||||
ACE
|
||||
ACEs
|
||||
ACS
|
||||
AITER
|
||||
ALU
|
||||
AllReduce
|
||||
AllToAll
|
||||
AGPR
|
||||
AGPRs
|
||||
AITER
|
||||
ALU
|
||||
AMD
|
||||
AMDGPU
|
||||
AMDGPUs
|
||||
@@ -31,6 +24,7 @@ ASICs
|
||||
ASan
|
||||
ASm
|
||||
ATI
|
||||
ATT
|
||||
AccVGPR
|
||||
AccVGPRs
|
||||
AddressSanitizer
|
||||
@@ -41,10 +35,13 @@ Arb
|
||||
Async
|
||||
Autocast
|
||||
BARs
|
||||
BDF
|
||||
BKC
|
||||
BLAS
|
||||
BLASLt
|
||||
BMC
|
||||
BNXT
|
||||
BRCM
|
||||
BSR
|
||||
BabelStream
|
||||
Backported
|
||||
BatchNorm
|
||||
@@ -68,11 +65,15 @@ CMakeLists
|
||||
CMakePackage
|
||||
CNP
|
||||
CP
|
||||
CPACK
|
||||
CPC
|
||||
CPF
|
||||
CPP
|
||||
CPU
|
||||
CPUs
|
||||
CPX
|
||||
CPack
|
||||
CQ
|
||||
CSC
|
||||
CSDATA
|
||||
CSE
|
||||
@@ -91,11 +92,13 @@ ChatGPT
|
||||
Cholesky
|
||||
CoRR
|
||||
Codespaces
|
||||
ComfyUI
|
||||
Commitizen
|
||||
CommonMark
|
||||
Concretized
|
||||
Conda
|
||||
ConnectX
|
||||
Conv
|
||||
CountOnes
|
||||
Cron
|
||||
CuPy
|
||||
@@ -110,13 +113,14 @@ DIMM
|
||||
DKMS
|
||||
DL
|
||||
DMA
|
||||
DMC
|
||||
DNN
|
||||
DNNL
|
||||
DOCA
|
||||
DOMContentLoaded
|
||||
DPM
|
||||
DPX
|
||||
DRI
|
||||
DSA
|
||||
DSCP
|
||||
DW
|
||||
DWORD
|
||||
@@ -137,13 +141,16 @@ Disaggregated
|
||||
Dockerfile
|
||||
Dockerized
|
||||
Doxygen
|
||||
Dyninst
|
||||
ELMo
|
||||
ENDPGM
|
||||
EP
|
||||
EPYC
|
||||
ESMI
|
||||
ESXi
|
||||
EoS
|
||||
FBGEMM
|
||||
FFN
|
||||
FFT
|
||||
FFTs
|
||||
FFmpeg
|
||||
@@ -151,9 +158,10 @@ FHS
|
||||
FIFOs
|
||||
FIXME
|
||||
FMA
|
||||
FNUZ
|
||||
FMHA
|
||||
FP
|
||||
FX
|
||||
FetchContent
|
||||
FiLM
|
||||
Filesystem
|
||||
FindDb
|
||||
@@ -161,6 +169,7 @@ Flang
|
||||
FlashAttention
|
||||
FlashInfer
|
||||
FlashInfer’s
|
||||
Flatmm
|
||||
FluxBenchmark
|
||||
Fortran
|
||||
Fuyu
|
||||
@@ -173,12 +182,15 @@ GCD
|
||||
GCDs
|
||||
GCN
|
||||
GCNN
|
||||
GDA
|
||||
GDB
|
||||
GDDR
|
||||
GDR
|
||||
GDS
|
||||
GEMM
|
||||
GEMMs
|
||||
GETRS
|
||||
GETRS
|
||||
GFLOPS
|
||||
GFXIP
|
||||
GFortran
|
||||
@@ -186,15 +198,12 @@ GGUF
|
||||
GID
|
||||
GIM
|
||||
GL
|
||||
Glibc
|
||||
GLM
|
||||
GIM
|
||||
GL
|
||||
GLXT
|
||||
GMI
|
||||
GNN
|
||||
GNNs
|
||||
GPG
|
||||
GPGPU
|
||||
GPR
|
||||
GPT
|
||||
GPU
|
||||
@@ -204,6 +213,7 @@ GPUVM
|
||||
GPUs
|
||||
GRBM
|
||||
GRE
|
||||
GSU
|
||||
GTT
|
||||
Gbps
|
||||
Gemma
|
||||
@@ -214,15 +224,18 @@ GitHub
|
||||
Gitpod
|
||||
Glibc
|
||||
Gloo
|
||||
Gluon
|
||||
GraphAPI
|
||||
GraphBolt
|
||||
GraphSage
|
||||
Graphbolt
|
||||
HBM
|
||||
HCA
|
||||
HEVC
|
||||
HGX
|
||||
HIPCC
|
||||
HIPExtension
|
||||
HIPIFY
|
||||
HIPOCProgram
|
||||
HIPification
|
||||
HIPify
|
||||
HLO
|
||||
@@ -234,27 +247,30 @@ HSA
|
||||
HW
|
||||
HWE
|
||||
HWS
|
||||
HX
|
||||
Haswell
|
||||
Higgs
|
||||
Huggingface
|
||||
Hunyuan
|
||||
HunyuanVideo
|
||||
HybridEngine
|
||||
Hyperparameters
|
||||
IB
|
||||
IC
|
||||
ICD
|
||||
InternVL
|
||||
ICT
|
||||
ICV
|
||||
IDE
|
||||
IDEs
|
||||
IFWI
|
||||
ILP
|
||||
ILU
|
||||
IMDb
|
||||
IOMMU
|
||||
IOP
|
||||
IOPM
|
||||
IOPS
|
||||
IOV
|
||||
IPC
|
||||
IPU
|
||||
IPs
|
||||
IRQ
|
||||
ISA
|
||||
@@ -271,7 +287,7 @@ Intersphinx
|
||||
Intra
|
||||
Ioffe
|
||||
JAX's
|
||||
JAXLIB
|
||||
JPG
|
||||
JSON
|
||||
Jinja
|
||||
Jupyter
|
||||
@@ -281,23 +297,21 @@ KMD
|
||||
KV
|
||||
KVM
|
||||
Karpathy's
|
||||
Kimi
|
||||
KiB
|
||||
Kineto
|
||||
Keras
|
||||
Khronos
|
||||
KiB
|
||||
Kineto
|
||||
LAPACK
|
||||
LASYF
|
||||
LCLK
|
||||
LDS
|
||||
LLM
|
||||
LLMs
|
||||
LLVM
|
||||
LLaMA
|
||||
LM
|
||||
LPDDR
|
||||
LRU
|
||||
LSE
|
||||
LSAN
|
||||
LSTMs
|
||||
LSan
|
||||
@@ -317,6 +331,7 @@ MIOpenGEMM
|
||||
MIVisionX
|
||||
MLA
|
||||
MLM
|
||||
MLX
|
||||
MMA
|
||||
MMIO
|
||||
MMIOH
|
||||
@@ -324,14 +339,16 @@ MMU
|
||||
MNIST
|
||||
MPI
|
||||
MPT
|
||||
MSI
|
||||
MSVC
|
||||
MTP
|
||||
MTU
|
||||
MVAPICH
|
||||
MVFFR
|
||||
MX
|
||||
MXFP
|
||||
Makefile
|
||||
Makefiles
|
||||
ManyLinux
|
||||
Matplotlib
|
||||
Matrox
|
||||
MaxText
|
||||
@@ -342,6 +359,7 @@ Mellanox
|
||||
Mellanox's
|
||||
Meta's
|
||||
MiB
|
||||
Microscaling
|
||||
Miniconda
|
||||
MirroredStrategy
|
||||
Mixtral
|
||||
@@ -351,9 +369,6 @@ Mooncake
|
||||
MosaicML
|
||||
Mpops
|
||||
Multicore
|
||||
Multimodal
|
||||
multimodal
|
||||
multihost
|
||||
Multithreaded
|
||||
MyEnvironment
|
||||
MyST
|
||||
@@ -362,6 +377,7 @@ NBIO
|
||||
NBIOs
|
||||
NCCL
|
||||
NCF
|
||||
NCHW
|
||||
NCS
|
||||
NFS
|
||||
NIC
|
||||
@@ -372,6 +388,8 @@ NN
|
||||
NOP
|
||||
NPKit
|
||||
NPS
|
||||
NPVT
|
||||
NSIS
|
||||
NSP
|
||||
NUMA
|
||||
NVCC
|
||||
@@ -379,7 +397,6 @@ NVIDIA
|
||||
NVLink
|
||||
NVPTX
|
||||
NaN
|
||||
NaNs
|
||||
Nano
|
||||
Navi
|
||||
NoReturn
|
||||
@@ -397,16 +414,17 @@ OMPI
|
||||
OMPT
|
||||
OMPX
|
||||
ONNX
|
||||
OSL
|
||||
OOM
|
||||
OSS
|
||||
OSU
|
||||
OOM
|
||||
OTF
|
||||
OpenCL
|
||||
OpenCV
|
||||
OpenFabrics
|
||||
OpenGL
|
||||
OpenMP
|
||||
OpenMPI
|
||||
OpenSHMEM
|
||||
OpenSSL
|
||||
OpenVX
|
||||
OpenXLA
|
||||
@@ -417,11 +435,15 @@ PCI
|
||||
PCIe
|
||||
PEFT
|
||||
PEQT
|
||||
PEs
|
||||
PID
|
||||
PIL
|
||||
PILImage
|
||||
PJRT
|
||||
PLDM
|
||||
POR
|
||||
POSIX
|
||||
POTF
|
||||
POTRF
|
||||
PRNG
|
||||
PRs
|
||||
PSID
|
||||
@@ -435,12 +457,10 @@ Pensando
|
||||
PerfDb
|
||||
Perfetto
|
||||
PipelineParallel
|
||||
Pipelining
|
||||
PnP
|
||||
Pollara
|
||||
PowerEdge
|
||||
PowerShell
|
||||
Preshuffled
|
||||
Pretrained
|
||||
Pretraining
|
||||
Primus
|
||||
@@ -448,12 +468,15 @@ Profiler's
|
||||
PyPi
|
||||
PyTorch
|
||||
Pytest
|
||||
QMCPACK
|
||||
QPS
|
||||
QPX
|
||||
Qcycles
|
||||
QoS
|
||||
Qwen
|
||||
RAII
|
||||
RAS
|
||||
RBT
|
||||
RCCL
|
||||
RDC
|
||||
RDC's
|
||||
@@ -478,10 +501,14 @@ ROCm
|
||||
ROCmCC
|
||||
ROCmSoftwarePlatform
|
||||
ROCmValidationSuite
|
||||
ROCprof
|
||||
ROCprofiler
|
||||
ROCr
|
||||
ROCtx
|
||||
RPP
|
||||
RST
|
||||
RTC
|
||||
RTC
|
||||
RW
|
||||
Radeon
|
||||
Radix
|
||||
@@ -495,10 +522,10 @@ RoCE
|
||||
Runfile
|
||||
Ryzen
|
||||
SALU
|
||||
safetensors
|
||||
SBIOS
|
||||
SCA
|
||||
SDK
|
||||
SDKs
|
||||
SDMA
|
||||
SDPA
|
||||
SDRAM
|
||||
@@ -519,24 +546,35 @@ SMEM
|
||||
SMFMA
|
||||
SMI
|
||||
SMT
|
||||
SPARSELt
|
||||
SPI
|
||||
SPIR
|
||||
SPX
|
||||
SQTT
|
||||
SQs
|
||||
SRAM
|
||||
SRAMECC
|
||||
SVD
|
||||
SVM
|
||||
SWE
|
||||
SYTF
|
||||
SYTRF
|
||||
SYTRS
|
||||
SageAttention
|
||||
ScaledGEMM
|
||||
SerDes
|
||||
Shardy
|
||||
ShareGPT
|
||||
Shlens
|
||||
SiLU
|
||||
Skylake
|
||||
Slurm
|
||||
Softmax
|
||||
Spack
|
||||
SplitK
|
||||
StreamingLLM
|
||||
Strix
|
||||
Supermicro
|
||||
SwiGLU
|
||||
Szegedy
|
||||
TCA
|
||||
TCC
|
||||
@@ -547,16 +585,23 @@ TCP
|
||||
TCR
|
||||
TF
|
||||
TFLOPS
|
||||
TGZ
|
||||
THREADGROUPS
|
||||
TP
|
||||
TPS
|
||||
TPU
|
||||
TPUs
|
||||
TPX
|
||||
TSME
|
||||
TT
|
||||
TTM
|
||||
TTY
|
||||
TUI
|
||||
TVM
|
||||
TagRAM
|
||||
Tagram
|
||||
Taichi
|
||||
Taichi's
|
||||
TensileLite
|
||||
TensorBoard
|
||||
TensorFloat
|
||||
@@ -575,11 +620,12 @@ TorchVision
|
||||
TransferBench
|
||||
TrapStatus
|
||||
UAC
|
||||
UBB
|
||||
UC
|
||||
UCC
|
||||
UCX
|
||||
UE
|
||||
UI
|
||||
UEK
|
||||
UIF
|
||||
UMC
|
||||
USM
|
||||
@@ -588,6 +634,7 @@ UTCL
|
||||
UTCL
|
||||
UTIL
|
||||
UTIL
|
||||
UUID
|
||||
UX
|
||||
UltraChat
|
||||
Uncached
|
||||
@@ -602,20 +649,25 @@ VM
|
||||
VMEM
|
||||
VMID
|
||||
VMIDs
|
||||
VMM
|
||||
VMWare
|
||||
VMs
|
||||
VMware
|
||||
VRAM
|
||||
VSIX
|
||||
VSkipped
|
||||
Vanhoucke
|
||||
Vulkan
|
||||
WDAG
|
||||
WGP
|
||||
WGPs
|
||||
WR
|
||||
WX
|
||||
WikiText
|
||||
Winograd
|
||||
Wojna
|
||||
Workgroups
|
||||
WrW
|
||||
Writebacks
|
||||
XCD
|
||||
XCDs
|
||||
@@ -634,6 +686,7 @@ YML
|
||||
YModel
|
||||
ZeRO
|
||||
ZenDNN
|
||||
abquant
|
||||
accuracies
|
||||
activations
|
||||
addEventListener
|
||||
@@ -644,8 +697,15 @@ alloc
|
||||
allocatable
|
||||
allocator
|
||||
allocators
|
||||
alltoall
|
||||
alltoallv
|
||||
amd
|
||||
amdclang
|
||||
amdgpu
|
||||
amdrocm
|
||||
amdsmi
|
||||
api
|
||||
aqlprofile
|
||||
async
|
||||
aten
|
||||
atmi
|
||||
@@ -654,12 +714,12 @@ atomics
|
||||
autogenerated
|
||||
autograd
|
||||
autotune
|
||||
autotuning
|
||||
avx
|
||||
awk
|
||||
az
|
||||
backend
|
||||
backends
|
||||
batchnorm
|
||||
bb
|
||||
benchmarked
|
||||
benchmarking
|
||||
@@ -668,11 +728,13 @@ bilinear
|
||||
bitcode
|
||||
bitsandbytes
|
||||
bitwise
|
||||
blas
|
||||
blit
|
||||
blockscale
|
||||
bootloader
|
||||
bootup
|
||||
boson
|
||||
bosons
|
||||
bottlenecked
|
||||
br
|
||||
btn
|
||||
buildable
|
||||
@@ -681,25 +743,31 @@ bzip
|
||||
cTDP
|
||||
cacheable
|
||||
carveout
|
||||
ccl
|
||||
cd
|
||||
centos
|
||||
centric
|
||||
cgroups
|
||||
changelog
|
||||
changelogs
|
||||
checkpointing
|
||||
chiplet
|
||||
cholesky
|
||||
classList
|
||||
cmake
|
||||
cmd
|
||||
coalescable
|
||||
codecs
|
||||
codename
|
||||
codenamed
|
||||
collater
|
||||
comfyui
|
||||
comgr
|
||||
compat
|
||||
completers
|
||||
composable
|
||||
composablekernel
|
||||
concretization
|
||||
cond
|
||||
config
|
||||
configs
|
||||
conformant
|
||||
@@ -709,7 +777,10 @@ convolutional
|
||||
convolves
|
||||
copyable
|
||||
cpp
|
||||
cron
|
||||
csn
|
||||
csv
|
||||
ctypes
|
||||
cuBLAS
|
||||
cuDNN
|
||||
cuFFT
|
||||
@@ -721,10 +792,10 @@ cuda
|
||||
cudnn
|
||||
customizable
|
||||
customizations
|
||||
cxx
|
||||
dGPU
|
||||
dGPUs
|
||||
da
|
||||
dataflows
|
||||
dataset
|
||||
datasets
|
||||
dataspace
|
||||
@@ -742,12 +813,14 @@ denoise
|
||||
denoised
|
||||
denoises
|
||||
denormalize
|
||||
depthwise
|
||||
dequantization
|
||||
dequantized
|
||||
dequantizes
|
||||
deserializers
|
||||
detections
|
||||
dev
|
||||
devel
|
||||
devicelibs
|
||||
devsel
|
||||
dgl
|
||||
@@ -756,9 +829,13 @@ disagg
|
||||
disaggregated
|
||||
disaggregation
|
||||
disambiguates
|
||||
discoverability
|
||||
distro
|
||||
distros
|
||||
dkms
|
||||
dll
|
||||
dnf
|
||||
dnn
|
||||
dropless
|
||||
dtype
|
||||
eb
|
||||
@@ -779,17 +856,22 @@ eth
|
||||
ethernet
|
||||
exascale
|
||||
executables
|
||||
factorizations
|
||||
fam
|
||||
fam
|
||||
fas
|
||||
ffmpeg
|
||||
fft
|
||||
filesystem
|
||||
flashinfer
|
||||
flang
|
||||
fmha
|
||||
forEach
|
||||
foreach
|
||||
fortran
|
||||
fp
|
||||
framebuffer
|
||||
gRPC
|
||||
galb
|
||||
gb
|
||||
gcc
|
||||
gdb
|
||||
gemm
|
||||
@@ -800,21 +882,27 @@ githooks
|
||||
github
|
||||
globals
|
||||
gnupg
|
||||
gpt
|
||||
gpu
|
||||
granularities
|
||||
grayscale
|
||||
gre
|
||||
gtest
|
||||
gx
|
||||
gz
|
||||
gzip
|
||||
hardcoded
|
||||
heterogenous
|
||||
hipBLAS
|
||||
hipBLASLt
|
||||
hipBLASLt's
|
||||
hipCUB
|
||||
hipDNN
|
||||
hipDataType
|
||||
hipFFT
|
||||
hipFFTW
|
||||
hipFORT
|
||||
hipLIB
|
||||
hipMemGetAddressRange
|
||||
hipModules
|
||||
hipRAND
|
||||
hipSOLVER
|
||||
hipSPARSE
|
||||
@@ -829,8 +917,11 @@ hipfft
|
||||
hipfort
|
||||
hipification
|
||||
hipify
|
||||
hipinfo
|
||||
hiprand
|
||||
hipsolver
|
||||
hipsparse
|
||||
hipsparselt
|
||||
hlist
|
||||
hostname
|
||||
hotspotting
|
||||
@@ -839,6 +930,7 @@ hpp
|
||||
href
|
||||
hsa
|
||||
hsakmt
|
||||
hx
|
||||
hyperparameter
|
||||
hyperparameters
|
||||
iDRAC
|
||||
@@ -862,6 +954,7 @@ invariants
|
||||
invocating
|
||||
ipo
|
||||
jax
|
||||
jpeg
|
||||
js
|
||||
json
|
||||
kdb
|
||||
@@ -870,9 +963,13 @@ kv
|
||||
lang
|
||||
latencies
|
||||
len
|
||||
libVA
|
||||
libdrm
|
||||
libelf
|
||||
libfabric
|
||||
libjpeg
|
||||
libs
|
||||
libva
|
||||
linalg
|
||||
linearized
|
||||
linter
|
||||
@@ -885,15 +982,20 @@ logits
|
||||
logsumexp
|
||||
loopback
|
||||
lossy
|
||||
lp
|
||||
lstsq
|
||||
mW
|
||||
macOS
|
||||
matchers
|
||||
maxtext
|
||||
megablocks
|
||||
megatron
|
||||
microarchitecture
|
||||
microscaling
|
||||
migraphx
|
||||
migratable
|
||||
milliwatt
|
||||
milliwatts
|
||||
miopen
|
||||
miopengemm
|
||||
mivisionx
|
||||
@@ -904,24 +1006,26 @@ mlirmiopen
|
||||
mtypes
|
||||
mul
|
||||
multihost
|
||||
multimodal
|
||||
mutex
|
||||
mvffr
|
||||
mx
|
||||
namespace
|
||||
namespaces
|
||||
nanoGPT
|
||||
netplan
|
||||
noncontiguous
|
||||
num
|
||||
numa
|
||||
numref
|
||||
ocl
|
||||
ol
|
||||
openai
|
||||
opencl
|
||||
opencv
|
||||
openmp
|
||||
openssl
|
||||
optimizers
|
||||
os
|
||||
oss
|
||||
oversubscription
|
||||
pageable
|
||||
pallas
|
||||
@@ -931,15 +1035,16 @@ param
|
||||
parameterization
|
||||
params
|
||||
passthrough
|
||||
pb
|
||||
pc
|
||||
pe
|
||||
perf
|
||||
perfcounter
|
||||
performant
|
||||
perl
|
||||
piecewise
|
||||
pipelined
|
||||
pipelining
|
||||
pkgman
|
||||
pmc
|
||||
ppt
|
||||
pragma
|
||||
pre
|
||||
prebuild
|
||||
@@ -977,6 +1082,7 @@ querySelectorAll
|
||||
queueing
|
||||
qwen
|
||||
radeon
|
||||
radix
|
||||
rc
|
||||
rccl
|
||||
rdc
|
||||
@@ -985,10 +1091,12 @@ reStructuredText
|
||||
reachability
|
||||
recommender
|
||||
recommenders
|
||||
redhat
|
||||
redirections
|
||||
refactorization
|
||||
reformats
|
||||
reinforcememt
|
||||
relocations
|
||||
repo
|
||||
repos
|
||||
representativeness
|
||||
@@ -1012,6 +1120,7 @@ rocMLIR
|
||||
rocPRIM
|
||||
rocPyDecode
|
||||
rocRAND
|
||||
rocSHMEM
|
||||
rocSOLVER
|
||||
rocSPARSE
|
||||
rocThrust
|
||||
@@ -1019,26 +1128,38 @@ rocWMMA
|
||||
rocalution
|
||||
rocblas
|
||||
rocclr
|
||||
rocdecode
|
||||
rocfft
|
||||
rocgdb
|
||||
rocjpeg
|
||||
rocm
|
||||
rocminfo
|
||||
rocpd
|
||||
rocprim
|
||||
rocprof
|
||||
rocprofiler
|
||||
rocprofv
|
||||
rocr
|
||||
rocrand
|
||||
rocshmem
|
||||
rocsolver
|
||||
rocsparse
|
||||
rocthrust
|
||||
roctracer
|
||||
rocwmma
|
||||
roofline
|
||||
rst
|
||||
runfile
|
||||
runtime
|
||||
runtimes
|
||||
rx
|
||||
ryzen
|
||||
sL
|
||||
scalability
|
||||
scalable
|
||||
scipy
|
||||
sdk
|
||||
seccomp
|
||||
seealso
|
||||
selectattr
|
||||
selectedTag
|
||||
@@ -1071,11 +1192,15 @@ submatrix
|
||||
submodule
|
||||
submodules
|
||||
subnet
|
||||
subnets
|
||||
supercomputing
|
||||
suse
|
||||
symlink
|
||||
symlinks
|
||||
sys
|
||||
syscall
|
||||
syscalls
|
||||
sysdeps
|
||||
sysfs
|
||||
tabindex
|
||||
targetContainer
|
||||
td
|
||||
@@ -1083,6 +1208,7 @@ tensorfloat
|
||||
tf
|
||||
th
|
||||
threadgroups
|
||||
toc
|
||||
tokenization
|
||||
tokenize
|
||||
tokenized
|
||||
@@ -1090,6 +1216,7 @@ tokenizer
|
||||
tokenizes
|
||||
toolchain
|
||||
toolchains
|
||||
toolkits
|
||||
toolset
|
||||
toolsets
|
||||
topk
|
||||
@@ -1102,9 +1229,14 @@ torchvision
|
||||
tp
|
||||
tqdm
|
||||
tracebacks
|
||||
tridiagonal
|
||||
txt
|
||||
typedef
|
||||
uarch
|
||||
ubuntu
|
||||
uclk
|
||||
ud
|
||||
udev
|
||||
uncacheable
|
||||
uncached
|
||||
uncorrectable
|
||||
@@ -1112,12 +1244,13 @@ underoptimized
|
||||
unfused
|
||||
unhandled
|
||||
uninstallation
|
||||
unmap
|
||||
unmapped
|
||||
unpadded
|
||||
unprefixed
|
||||
unsqueeze
|
||||
unstacking
|
||||
unswitching
|
||||
unswizzled
|
||||
untar
|
||||
untrusted
|
||||
untuned
|
||||
unwindowed
|
||||
@@ -1132,6 +1265,7 @@ vectorize
|
||||
vectorized
|
||||
vectorizer
|
||||
vectorizes
|
||||
ver
|
||||
verl
|
||||
verl's
|
||||
virtualize
|
||||
@@ -1141,7 +1275,6 @@ vllm
|
||||
voxel
|
||||
walkthrough
|
||||
walkthroughs
|
||||
warmup
|
||||
watchpoints
|
||||
wavefront
|
||||
wavefronts
|
||||
@@ -1159,7 +1292,9 @@ xPacked
|
||||
xargs
|
||||
xcc
|
||||
xdit
|
||||
xplane
|
||||
xe
|
||||
xt
|
||||
xtx
|
||||
xz
|
||||
yaml
|
||||
ysvmadyb
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
MIT License
|
||||
|
||||
Copyright (c) 2023 - 2025 Advanced Micro Devices, Inc. All rights reserved.
|
||||
Copyright (c) 2023 - 2026 Advanced Micro Devices, Inc. All rights reserved.
|
||||
|
||||
Permission is hereby granted, free of charge, to any person obtaining a copy
|
||||
of this software and associated documentation files (the "Software"), to deal
|
||||
|
||||
@@ -1,18 +0,0 @@
|
||||
Hanzo ROCm
|
||||
Copyright (c) 2026 Hanzo AI, Inc.
|
||||
|
||||
This product includes software from AMD ROCm (https://github.com/ROCm/ROCm),
|
||||
licensed under the MIT License:
|
||||
|
||||
Copyright (c) 2023 - 2025 Advanced Micro Devices, Inc. All rights reserved.
|
||||
|
||||
This repository is the AMD ROCm meta/manifest repository: documentation, build
|
||||
tooling, and the repo manifest (default.xml). This meta-repository is MIT-licensed
|
||||
and its upstream MIT license is retained in LICENSE.
|
||||
|
||||
The individual ROCm components referenced by the manifest and built by the tooling
|
||||
carry their own upstream licenses, including the MIT License, the Apache License 2.0,
|
||||
and the University of Illinois/NCSA Open Source License. Note that the ROCgdb
|
||||
component (AMD's fork of GNU GDB, packaged by tools/rocm-build/build_rocm-gdb.sh) is
|
||||
licensed under the GNU General Public License (GPL) — a copyleft license. Each
|
||||
component is governed by its own license; consult that component's repository.
|
||||
@@ -1,5 +1,3 @@
|
||||
<p align="center"><img src=".github/hero.svg" alt="ROCm" width="880"></p>
|
||||
|
||||
<div align="center">
|
||||
<img src="docs/data/amd-rocm-logo.png" width="200px" alt="ROCm logo">
|
||||
|
||||
|
||||
@@ -1,754 +0,0 @@
|
||||
<!-- Do not edit this file! -->
|
||||
<!-- This file is autogenerated with -->
|
||||
<!-- tools/autotag/tag_script.py -->
|
||||
<!-- Disable lints since this is an auto-generated file. -->
|
||||
<!-- markdownlint-disable blanks-around-headers -->
|
||||
<!-- markdownlint-disable no-duplicate-header -->
|
||||
<!-- markdownlint-disable no-blanks-blockquote -->
|
||||
<!-- markdownlint-disable ul-indent -->
|
||||
<!-- markdownlint-disable no-trailing-spaces -->
|
||||
<!-- markdownlint-disable reference-links-images -->
|
||||
<!-- markdownlint-disable no-missing-space-atx -->
|
||||
<!-- spellcheck-disable -->
|
||||
|
||||
# ROCm 7.2.4 release notes
|
||||
|
||||
ROCm 7.2.4 is a quality release focused on performance and stability fixes for AI inference workloads on AMD Instinct GPUs.
|
||||
|
||||
## Release highlights
|
||||
|
||||
The following are the notable changes in ROCm 7.2.4.
|
||||
|
||||
### Reduced hipGraphLaunch latency for multi-list graphs
|
||||
|
||||
The HIP runtime's graph dispatch mechanism has been optimized, reducing launch latency for workloads using `hipGraphLaunch` with multi-list graph topologies.
|
||||
|
||||
### Fixed H2D memory copy latency regression in CPX mode
|
||||
|
||||
HIP runtime synchronization behavior has been corrected on AMD Instinct MI300 Series GPUs in CPX mode, restoring latency to previous levels for inference workloads that run multiple HIP streams with concurrent memory copies.
|
||||
|
||||
### Reduced ROCprofiler-SDK profiling overhead
|
||||
|
||||
Profiling stability has been improved for vLLM workloads traced with PyTorch `torch.profiler` using the ROCprofiler-SDK backend. The large, sporadic idle gaps that previously appeared between GPU kernels in the trace have been substantially reduced in common configurations, and the traces now more accurately reflect actual runtime behavior. Coverage may vary depending on model and parallelism settings.
|
||||
|
||||
### Reduced copy overhead in MIGraphX concat operations
|
||||
|
||||
MIGraphX now recognizes ONNX models that concatenate the same tensor multiple times and avoids redundant device-side copies, improving inference throughput at small batch sizes for the affected model class on AMD Instinct MI300X GPUs.
|
||||
|
||||
### User space, driver, and firmware dependent changes
|
||||
|
||||
The software for AMD Data Center GPU products requires maintaining a hardware
|
||||
and software stack with interdependencies among the GPU and baseboard
|
||||
firmware, AMD GPU drivers, and the ROCm user space software. While AMD publishes drivers and ROCm user space components, your server or infrastructure provider publishes the GPU and baseboard firmware by bundling AMD’s firmware releases via the AMD Platform Level Data Model (PLDM) bundle, which includes the Integrated Firmware Image (IFWI).
|
||||
|
||||
GPU and baseboard firmware versioning might differ across GPU families.
|
||||
|
||||
<div class="pst-scrollable-table-container">
|
||||
<table class="table table--middle-left">
|
||||
<thead>
|
||||
<tr>
|
||||
<th class="head">
|
||||
<p>ROCm Version</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>GPU</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>PLDM Bundle (Firmware)</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>AMD GPU Driver (amdgpu)</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>AMD GPU <br>
|
||||
Virtualization Driver (GIM)</p>
|
||||
</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<style>
|
||||
tbody#virtualization-support-instinct tr:last-child {
|
||||
border-bottom: 2px solid var(--pst-color-primary);
|
||||
}
|
||||
</style>
|
||||
<tr>
|
||||
<td rowspan="9" style="vertical-align: middle;">ROCm 7.2.4</td>
|
||||
<td>MI355X</td>
|
||||
<td>
|
||||
01.26.00.02<br>
|
||||
01.25.17.07<br>
|
||||
01.25.16.03
|
||||
</td>
|
||||
<td>
|
||||
30.30.x where x (0-4)<br>
|
||||
30.20.x where x (0-1)<br>
|
||||
30.10.x where x (0-2)
|
||||
</td>
|
||||
<td rowspan="3" style="vertical-align: middle;">8.7.1.K</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>MI350X</td>
|
||||
<td>
|
||||
01.26.00.02<br>
|
||||
01.25.17.07<br>
|
||||
01.25.16.03
|
||||
</td>
|
||||
<td>
|
||||
30.30.x where x (0-4)<br>
|
||||
30.20.x where x (0-1)<br>
|
||||
30.10.x where x (0-2)
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>MI325X<a href="#footnote1"><sup>[1]</sup></a></td>
|
||||
<td>
|
||||
01.25.06.08<br>
|
||||
01.25.04.02
|
||||
</td>
|
||||
<td>30.30.x where x (0-4)<br>
|
||||
30.20.x where x (0-1)<a href="#footnote1"><sup>[1]</sup></a><br>
|
||||
30.10.x where x (0-2)<br>
|
||||
6.4.z where z (0-3)<br>
|
||||
6.3.3
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>MI300X<a href="#footnote2"><sup>[2]</sup></a></td>
|
||||
<td>01.25.06.04<br>
|
||||
01.25.03.12<br>
|
||||
01.25.02.04</td>
|
||||
<td rowspan="6" style="vertical-align: middle;">
|
||||
30.30.x where x (0-4)<br>
|
||||
30.20.x where x (0-1)<br>
|
||||
30.10.x where x (0-2)<br>
|
||||
6.4.z where z (0–3)<br>
|
||||
6.3.3
|
||||
</td>
|
||||
<td>8.7.1.K</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>MI300A</td>
|
||||
<td>BKC 26.1</td>
|
||||
<td rowspan="3" style="vertical-align: middle;">Not Applicable</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>MI250X</td>
|
||||
<td>IFWI 47 (or later)</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>MI250</td>
|
||||
<td>MU5 w/ IFWI 75 (or later)</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>MI210</td>
|
||||
<td>MU5 w/ IFWI 75 (or later)</td>
|
||||
<td>8.7.1.K</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>MI100</td>
|
||||
<td>VBIOS D3430401-037</td>
|
||||
<td>Not Applicable</td>
|
||||
</tr>
|
||||
</table>
|
||||
</div>
|
||||
|
||||
<p id="footnote1">[1]: For AMD Instinct MI325X KVM SR-IOV users, don't use AMD GPU driver (amdgpu) 30.20.0.</p>
|
||||
<p id="footnote2">[2]: AMD Instinct MI300X KVM SR-IOV with Multi-VF (8 VF) support requires a compatible firmware BKC bundle, which will be released in the coming months.</p>
|
||||
|
||||
```{note}
|
||||
ROCm 7.2.4 doesn't include any other significant changes or feature additions. For comprehensive changes, new features, and enhancements in ROCm 7.2.3, refer to the [ROCm 7.2.3 release notes](#rocm-7-2-3-release-notes) below.
|
||||
```
|
||||
|
||||
## ROCm 7.2.3 release notes
|
||||
|
||||
The release notes provide a summary of notable changes since the previous ROCm release.
|
||||
|
||||
- [Release highlights](#id1)
|
||||
|
||||
- [Supported hardware, operating system, and virtualization changes](#supported-hardware-operating-system-and-virtualization-changes)
|
||||
|
||||
- [User space, driver, and firmware dependent changes](#user-space-driver-and-firmware-dependent-changes)
|
||||
|
||||
- [ROCm components versioning](#rocm-components)
|
||||
|
||||
- [Detailed component changes](#detailed-component-changes)
|
||||
|
||||
- [ROCm known issues](#rocm-known-issues)
|
||||
|
||||
- [ROCm upcoming changes](#rocm-upcoming-changes)
|
||||
|
||||
```{note}
|
||||
If you’re using AMD Radeon™ GPUs or Ryzen™ for graphics workloads, see the [Use ROCm on Radeon and Ryzen](https://rocm.docs.amd.com/projects/radeon-ryzen/en/latest/index.html) documentation to verify compatibility and system requirements.
|
||||
```
|
||||
|
||||
### Release highlights
|
||||
|
||||
The following are notable new features and improvements in ROCm 7.2.3. For changes to individual components, see
|
||||
[Detailed component changes](#detailed-component-changes).
|
||||
|
||||
#### Supported hardware, operating system, and virtualization changes
|
||||
|
||||
Hardware, operating system, and virtualization support remains unchanged in this release.
|
||||
|
||||
For more information about:
|
||||
|
||||
* AMD hardware, see [Supported GPUs (Linux)](https://rocm.docs.amd.com/projects/install-on-linux/en/docs-7.2.3/reference/system-requirements.html#supported-gpus).
|
||||
|
||||
* Operating systems, see [Supported operating systems](https://rocm.docs.amd.com/projects/install-on-linux/en/docs-7.2.3/reference/system-requirements.html#supported-operating-systems) and [ROCm installation for Linux](https://rocm.docs.amd.com/projects/install-on-linux/en/docs-7.2.3/).
|
||||
|
||||
* Virtualization support, see [Virtualization support](https://rocm.docs.amd.com/projects/install-on-linux/en/docs-7.2.3/reference/system-requirements.html#virtualization-support).
|
||||
|
||||
#### User space, driver, and firmware dependent changes
|
||||
|
||||
The software for AMD Data Center GPU products requires maintaining a hardware
|
||||
and software stack with interdependencies among the GPU and baseboard
|
||||
firmware, AMD GPU drivers, and the ROCm user space software. While AMD publishes drivers and ROCm user space components, your server or infrastructure provider publishes the GPU and baseboard firmware by bundling AMD’s firmware releases via the AMD Platform Level Data Model (PLDM) bundle, which includes the Integrated Firmware Image (IFWI).
|
||||
|
||||
GPU and baseboard firmware versioning might differ across GPU families.
|
||||
|
||||
<div class="pst-scrollable-table-container">
|
||||
<table class="table table--middle-left">
|
||||
<thead>
|
||||
<tr>
|
||||
<th class="head">
|
||||
<p>ROCm Version</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>GPU</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>PLDM Bundle (Firmware)</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>AMD GPU Driver (amdgpu)</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>AMD GPU <br>
|
||||
Virtualization Driver (GIM)</p>
|
||||
</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<style>
|
||||
tbody#virtualization-support-instinct tr:last-child {
|
||||
border-bottom: 2px solid var(--pst-color-primary);
|
||||
}
|
||||
</style>
|
||||
<tr>
|
||||
<td rowspan="9" style="vertical-align: middle;">ROCm 7.2.3</td>
|
||||
<td>MI355X</td>
|
||||
<td>
|
||||
01.26.00.02<br>
|
||||
01.25.17.07<br>
|
||||
01.25.16.03
|
||||
</td>
|
||||
<td>
|
||||
30.30.x where x (0-3)<br>
|
||||
30.20.x where x (0-1)<br>
|
||||
30.10.x where x (0-2)
|
||||
</td>
|
||||
<td rowspan="3" style="vertical-align: middle;">8.7.1.K</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>MI350X</td>
|
||||
<td>
|
||||
01.26.00.02<br>
|
||||
01.25.17.07<br>
|
||||
01.25.16.03
|
||||
</td>
|
||||
<td>
|
||||
30.30.x where x (0-3)<br>
|
||||
30.20.x where x (0-1)<br>
|
||||
30.10.x where x (0-2)
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>MI325X<a href="#footnote1"><sup>[1]</sup></a></td>
|
||||
<td>
|
||||
01.25.06.08<br>
|
||||
01.25.04.02
|
||||
</td>
|
||||
<td>30.30.x where x (0-3)<br>
|
||||
30.20.x where x (0-1)<a href="#footnote1"><sup>[1]</sup></a><br>
|
||||
30.10.x where x (0-2)<br>
|
||||
6.4.z where z (0-3)<br>
|
||||
6.3.3
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>MI300X<a href="#footnote2"><sup>[2]</sup></a></td>
|
||||
<td>01.25.06.04<br>
|
||||
01.25.03.12<br>
|
||||
01.25.02.04</td>
|
||||
<td rowspan="6" style="vertical-align: middle;">
|
||||
30.30.x where x (0-3)<br>
|
||||
30.20.x where x (0-1)<br>
|
||||
30.10.x where x (0-2)<br>
|
||||
6.4.z where z (0–3)<br>
|
||||
6.3.3
|
||||
</td>
|
||||
<td>8.7.1.K</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>MI300A</td>
|
||||
<td>BKC 26.1</td>
|
||||
<td rowspan="3" style="vertical-align: middle;">Not Applicable</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>MI250X</td>
|
||||
<td>IFWI 47 (or later)</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>MI250</td>
|
||||
<td>MU5 w/ IFWI 75 (or later)</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>MI210</td>
|
||||
<td>MU5 w/ IFWI 75 (or later)</td>
|
||||
<td>8.7.1.K</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>MI100</td>
|
||||
<td>VBIOS D3430401-037</td>
|
||||
<td>Not Applicable</td>
|
||||
</tr>
|
||||
</table>
|
||||
</div>
|
||||
|
||||
<p id="footnote1">[1]: For AMD Instinct MI325X KVM SR-IOV users, don't use AMD GPU driver (amdgpu) 30.20.0.</p>
|
||||
<p id="footnote2">[2]: AMD Instinct MI300X KVM SR-IOV with Multi-VF (8 VF) support requires a compatible firmware BKC bundle, which will be released in the coming months.</p>
|
||||
|
||||
#### Improved profiling accuracy for vLLM workloads
|
||||
|
||||
ROCm 7.2.3 improves profiling stability for vLLM workloads traced with PyTorch `torch.profiler`. The large, sporadic idle gaps that previously appeared between GPU kernels in the trace have been substantially reduced in common configurations, and the traces now more accurately reflect actual runtime behavior. Coverage may vary depending on model and parallelism settings; additional improvements are in progress.
|
||||
|
||||
#### MIGraphX update
|
||||
|
||||
[MIGraphX](https://rocm.docs.amd.com/projects/AMDMIGraphX/en/docs-7.2.3/index.html) has the following enhancements:
|
||||
|
||||
##### Improved performance of the Gather operator
|
||||
|
||||
Performance for embedding‑heavy inference workloads is improved by merging multiple independent gather operations from similar embedding tables into a single batched operation. Multi‑gather workloads now run more efficiently with fewer kernel launches and reduced memory traffic by adding horizontal fusion for cross-embedding gather operators. These gather operators have been updated to use `transpose`/`reshape`/`broadcast`/`slice`, enabling better optimization across different backends and data layouts.
|
||||
|
||||
##### ONNX Runtime reliability improvement
|
||||
|
||||
ONNX Runtime workloads accelerated with MIGraphX now provide a more reliable experience through external stream support in the MIGraphX Execution Provider, with improved memory allocation and deallocation for multi-stream inference.
|
||||
|
||||
#### ROCm documentation updates
|
||||
|
||||
ROCm documentation has been updated with ROCm XIO documentation. ROCm XIO provides an API for Accelerator-Initiated IO (XIO) for an AMD GPU `__device__` code. It enables AMD GPUs to perform direct IO operations to hardware devices without CPU intervention. ROCm XIO was initially released in April 2026 as an early-access software technology preview. Running production workloads is not recommended.
|
||||
For more information, see the [ROCm XIO documentation](https://rocm.docs.amd.com/projects/rocm-xio/en/beta-0.1.0/index.html) and {fab}`github` [ROCm/rocm-xio](https://github.com/ROCm/rocm-xio) GitHub repository.
|
||||
|
||||
### ROCm components
|
||||
|
||||
The following table lists the versions of ROCm components for ROCm 7.2.3, including any version
|
||||
changes from 7.2.2/7.2.1 to 7.2.3. Click the component's updated version to go to a list of its changes.
|
||||
|
||||
Click {fab}`github` to go to the component's source code on GitHub.
|
||||
|
||||
<div class="pst-scrollable-table-container">
|
||||
<table id="rocm-rn-components" class="table">
|
||||
<thead>
|
||||
<tr>
|
||||
<th>Category</th>
|
||||
<th>Group</th>
|
||||
<th>Name</th>
|
||||
<th>Version</th>
|
||||
<th></th>
|
||||
</tr>
|
||||
</thead>
|
||||
<colgroup>
|
||||
<col span="1">
|
||||
<col span="1">
|
||||
</colgroup>
|
||||
<tbody class="rocm-components-libs rocm-components-ml">
|
||||
<tr>
|
||||
<th rowspan="9">Libraries</th>
|
||||
<th rowspan="9">Machine learning and computer vision</th>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/composable_kernel/en/docs-7.2.3/index.html">Composable Kernel</a></td>
|
||||
<td>1.2.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/composablekernel"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/AMDMIGraphX/en/docs-7.2.3/index.html">MIGraphX</a></td>
|
||||
<td>2.15.0 ⇒ <a href="#migraphx-2-15-0">2.15.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/AMDMIGraphX"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/MIOpen/en/docs-7.2.3/index.html">MIOpen</a></td>
|
||||
<td>3.5.1</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/miopen"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/MIVisionX/en/docs-7.2.3/index.html">MIVisionX</a></td>
|
||||
<td>3.5.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/MIVisionX"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocAL/en/docs-7.2.3/index.html">rocAL</a></td>
|
||||
<td>2.5.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocAL"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocDecode/en/docs-7.2.3/index.html">rocDecode</a></td>
|
||||
<td>1.7.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocDecode"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocJPEG/en/docs-7.2.3/index.html">rocJPEG</a></td>
|
||||
<td>1.4.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocJPEG"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocPyDecode/en/docs-7.2.3/index.html">rocPyDecode</a></td>
|
||||
<td>0.8.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocPyDecode"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rpp/en/docs-7.2.3/index.html">RPP</a></td>
|
||||
<td>2.2.1</a></td>
|
||||
<td><a href="https://github.com/ROCm/rpp"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
</tbody>
|
||||
<tbody class="rocm-components-libs rocm-components-communication tbody-reverse-zebra">
|
||||
<tr>
|
||||
<th rowspan="2"></th>
|
||||
<th rowspan="2">Communication</th>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rccl/en/docs-7.2.3/index.html">RCCL</a></td>
|
||||
<td>2.27.7</a></td>
|
||||
<td><a href="https://github.com/ROCm/rccl"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocSHMEM/en/docs-7.1.0/index.html">rocSHMEM</a></td>
|
||||
<td>3.2.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocSHMEM"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
</tbody>
|
||||
<tbody class="rocm-components-libs rocm-components-math tbody-reverse-zebra">
|
||||
<tr>
|
||||
<th rowspan="16"></th>
|
||||
<th rowspan="16">Math</th>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/hipBLAS/en/docs-7.2.3/index.html">hipBLAS</a></td>
|
||||
<td>3.2.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipblas"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/hipBLASLt/en/docs-7.2.3/index.html">hipBLASLt</a></td>
|
||||
<td>1.2.2</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipblaslt"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/hipFFT/en/docs-7.2.3/index.html">hipFFT</a></td>
|
||||
<td>1.0.22</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipfft"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/hipfort/en/docs-7.2.3/index.html">hipfort</a></td>
|
||||
<td>0.7.1</a></td>
|
||||
<td><a href="https://github.com/ROCm/hipfort"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/hipRAND/en/docs-7.2.3/index.html">hipRAND</a></td>
|
||||
<td>3.1.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hiprand"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/hipSOLVER/en/docs-7.2.3/index.html">hipSOLVER</a></td>
|
||||
<td>3.2.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipsolver"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/hipSPARSE/en/docs-7.2.3/index.html">hipSPARSE</a></td>
|
||||
<td>4.2.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipsparse"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/hipSPARSELt/en/docs-7.2.3/index.html">hipSPARSELt</a></td>
|
||||
<td>0.2.6</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipsparselt"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocALUTION/en/docs-7.2.3/index.html">rocALUTION</a></td>
|
||||
<td>4.1.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocALUTION"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocBLAS/en/docs-7.2.3/index.html">rocBLAS</a></td>
|
||||
<td>5.2.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocblas"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocFFT/en/docs-7.2.3/index.html">rocFFT</a></td>
|
||||
<td>1.0.36</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocfft"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocRAND/en/docs-7.2.3/index.html">rocRAND</a></td>
|
||||
<td>4.2.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocrand"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocSOLVER/en/docs-7.2.3/index.html">rocSOLVER</a></td>
|
||||
<td>3.32.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocsolver"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocSPARSE/en/docs-7.2.3/index.html">rocSPARSE</a></td>
|
||||
<td>4.2.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocsparse"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocWMMA/en/docs-7.2.3/index.html">rocWMMA</a></td>
|
||||
<td>2.2.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocwmma"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/Tensile/en/docs-7.2.3/src/index.html">Tensile</a></td>
|
||||
<td>4.45.0</td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/shared/tensile"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
</tbody>
|
||||
<tbody class="rocm-components-libs rocm-components-primitives tbody-reverse-zebra">
|
||||
<tr>
|
||||
<th rowspan="4"></th>
|
||||
<th rowspan="4">Primitives</th>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/hipCUB/en/docs-7.2.3/index.html">hipCUB</a></td>
|
||||
<td>4.2.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipcub"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/hipTensor/en/docs-7.2.3/index.html">hipTensor</a></td>
|
||||
<td>2.2.0</td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hiptensor"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocPRIM/en/docs-7.2.3/index.html">rocPRIM</a></td>
|
||||
<td>4.2.0</td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocprim"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocThrust/en/docs-7.2.3/index.html">rocThrust</a></td>
|
||||
<td>4.2.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocthrust"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
</tbody>
|
||||
<tbody class="rocm-components-tools rocm-components-system tbody-reverse-zebra">
|
||||
<tr>
|
||||
<th rowspan="7">Tools</th>
|
||||
<th rowspan="7">System management</th>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/amdsmi/en/docs-7.2.3/index.html">AMD SMI</a></td>
|
||||
<td>26.2.2</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/amdsmi"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rdc/en/docs-7.2.3/index.html">ROCm Data Center Tool</a></td>
|
||||
<td>1.2.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rdc/"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocminfo/en/docs-7.2.3/index.html">rocminfo</a></td>
|
||||
<td>1.0.0</td>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocminfo/"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocm_smi_lib/en/docs-7.2.3/index.html">ROCm SMI</a></td>
|
||||
<td>7.8.0</td>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocm-smi-lib/"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/ROCmValidationSuite/en/docs-7.2.3/index.html">ROCm Validation Suite</a></td>
|
||||
<td>1.3.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/ROCmValidationSuite"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
</tbody>
|
||||
<tbody class="rocm-components-tools rocm-components-perf">
|
||||
<tr>
|
||||
<th rowspan="6"></th>
|
||||
<th rowspan="6">Performance</th>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocm_bandwidth_test/en/docs-7.2.3/index.html">ROCm Bandwidth
|
||||
Test</a></td>
|
||||
<td>2.6.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm_bandwidth_test/"><i
|
||||
class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocprofiler-compute/en/docs-7.2.3/index.html">ROCm Compute Profiler</a></td>
|
||||
<td>3.4.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocprofiler-compute"><i
|
||||
class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocprofiler-systems/en/docs-7.2.3/index.html">ROCm Systems Profiler</a></td>
|
||||
<td>1.3.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocprofiler-systems/"><i
|
||||
class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocprofiler/en/docs-7.2.3/index.html">ROCProfiler</a></td>
|
||||
<td>2.0.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocprofiler/"><i
|
||||
class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/docs-7.2.3/index.html">ROCprofiler-SDK</a></td>
|
||||
<td>1.1.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocprofiler-sdk/"><i
|
||||
class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr >
|
||||
<td><a href="https://rocm.docs.amd.com/projects/roctracer/en/docs-7.2.3/index.html">ROCTracer</a></td>
|
||||
<td>4.1.0</td>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/roctracer/"><i
|
||||
class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
</tbody>
|
||||
<tbody class="rocm-components-tools rocm-components-dev">
|
||||
<tr>
|
||||
<th rowspan="5"></th>
|
||||
<th rowspan="5">Development</th>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/HIPIFY/en/docs-7.2.3/index.html">HIPIFY</a></td>
|
||||
<td>22.0.0</td>
|
||||
<td><a href="https://github.com/ROCm/HIPIFY/"><i
|
||||
class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/ROCdbgapi/en/docs-7.2.3/index.html">ROCdbgapi</a></td>
|
||||
<td>0.77.4</a></td>
|
||||
<td><a href="https://github.com/ROCm/ROCdbgapi/"><i
|
||||
class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/ROCmCMakeBuildTools/en/docs-7.2.3/index.html">ROCm CMake</a></td>
|
||||
<td>0.14.0</td>
|
||||
<td><a href="https://github.com/ROCm/rocm-cmake/"><i
|
||||
class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/ROCgdb/en/docs-7.2.3/index.html">ROCm Debugger (ROCgdb)</a>
|
||||
</td>
|
||||
<td>16.3</a></td>
|
||||
<td><a href="https://github.com/ROCm/ROCgdb/"><i
|
||||
class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/rocr_debug_agent/en/docs-7.2.3/index.html">ROCr Debug Agent</a>
|
||||
</td>
|
||||
<td>2.1.0</td>
|
||||
<td><a href="https://github.com/ROCm/rocr_debug_agent/"><i
|
||||
class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
</tbody>
|
||||
<tbody class="rocm-components-compilers tbody-reverse-zebra">
|
||||
<tr>
|
||||
<th rowspan="2" colspan="2">Compilers</th>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/HIPCC/en/docs-7.2.3/index.html">HIPCC</a></td>
|
||||
<td>1.1.1</td>
|
||||
<td><a href="https://github.com/ROCm/llvm-project/tree/amd-staging/amd/hipcc"><i
|
||||
class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/llvm-project/en/docs-7.2.3/index.html">llvm-project</a></td>
|
||||
<td>22.0.0</a></td>
|
||||
<td><a href="https://github.com/ROCm/llvm-project/"><i
|
||||
class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
</tbody>
|
||||
<tbody class="rocm-components-runtimes tbody-reverse-zebra">
|
||||
<tr>
|
||||
<th rowspan="2" colspan="2">Runtimes</th>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/HIP/en/docs-7.2.3/index.html">HIP</a></td>
|
||||
<td>7.2.1</a></td>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/hip"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://rocm.docs.amd.com/projects/ROCR-Runtime/en/docs-7.2.3/index.html">ROCr Runtime</a></td>
|
||||
<td>1.18.0</td>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocr-runtime"><i class="fab fa-github fa-lg"></i></a></td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
</div>
|
||||
|
||||
### Detailed component changes
|
||||
|
||||
The following sections describe key changes to ROCm components.
|
||||
|
||||
```{note}
|
||||
For a historical overview of ROCm component updates, see the {doc}`ROCm consolidated changelog </release/changelog>`.
|
||||
```
|
||||
|
||||
#### **MIGraphX** (2.15.0)
|
||||
|
||||
##### Added
|
||||
|
||||
* External stream support to the MIGraphX context, allowing external HIP streams to be used during execution.
|
||||
* Ability to return a vector for output alias, supporting operators like `make_tuple`.
|
||||
|
||||
##### Changed
|
||||
|
||||
* Refactored `move_output_instructions_after` into the module class.
|
||||
* Updated rocMLIR to fix `bert_squad` and `bert_tf` regressions.
|
||||
|
||||
##### Optimized
|
||||
|
||||
* Rewrote the `gather` operator to use `transpose`/`reshape`/`broadcast`/`slice` for improved performance.
|
||||
* Horizontally fuse cross-embedding `gather` operators.
|
||||
* Improved tuning for Split-K.
|
||||
* Removed extra assignments and inserts in `find_nop_reshapes` to reduce overhead.
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
The following issues have been fixed:
|
||||
|
||||
* `int` to `bf16`/`fp16` conversion errors.
|
||||
* Comparison logic in `find_concat_op` to match the correct I/O.
|
||||
* `shape_transform_descriptor::rebase` when flattening a broadcasted dimension.
|
||||
* An error with `rewrite_reshapes`.
|
||||
* A gather rewrite crash by validating the strided view element count.
|
||||
* A bug in gather rewrite with NHWC shapes.
|
||||
* A crash in rocMLIR with Inception v3 on RDNA3 architecture-based Radeon GPUs.
|
||||
* Filter zero-argument operators during ONNX parsing to prevent errors.
|
||||
* Conflict for missing `no_broadcast` parameter on ROCm 7.2.x.
|
||||
|
||||
### ROCm known issues
|
||||
|
||||
ROCm known issues are noted on {fab}`github` [GitHub](https://github.com/ROCm/ROCm/labels/Verified%20Issue). For known
|
||||
issues related to individual components, review the [Detailed component changes](#detailed-component-changes).
|
||||
|
||||
#### Minor performance regression for MIGraphX with int8-quantized models
|
||||
|
||||
You might observe a slight performance regression when running int8-quantized models with MIGraphX. This impact is generally minimal and does not affect correctness. However, workloads sensitive to peak throughput might have reduced performance when compared to non-quantized or alternative execution paths. This issue is currently under investigation and will be fixed in a future ROCm release. See [GitHub issue #6195](https://github.com/ROCm/ROCm/issues/6195).
|
||||
|
||||
### ROCm upcoming changes
|
||||
|
||||
The following changes to the ROCm software stack are anticipated for future releases.
|
||||
|
||||
#### ROCTracer, ROCProfiler, rocprof, and rocprofv2 deprecation
|
||||
|
||||
ROCTracer, ROCProfiler, `rocprof`, and `rocprofv2` are deprecated. It's strongly recommended to upgrade to the latest version of the [ROCprofiler-SDK](https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/latest/) library and the (`rocprofv3`) tool to ensure continued support and access to new features.
|
||||
|
||||
To learn about key feature improvements and benefits of ROCprofiler-SDK over the deprecated ROCProfiler and ROCTracer, see [Comparing ROCprofiler-SDK to legacy ROCm profiling tools](https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/latest/conceptual/comparing-with-legacy-tools.html).
|
||||
|
||||
It's anticipated that ROCTracer, ROCProfiler, `rocprof`, and `rocprofv2` will reach end of support (EoS) by the end of 2026 Q2.
|
||||
|
||||
#### ROCm SMI deprecation
|
||||
|
||||
[ROCm SMI](https://github.com/ROCm/rocm_smi_lib) will be phased out in an
|
||||
upcoming ROCm release and will enter maintenance mode. After this transition,
|
||||
only critical bug fixes will be addressed and no further feature development
|
||||
will take place.
|
||||
|
||||
It's strongly recommended to transition your projects to [AMD
|
||||
SMI](https://github.com/ROCm/rocm-systems/tree/develop/projects/amdsmi), the successor to ROCm SMI. AMD SMI
|
||||
includes all the features of the ROCm SMI and will continue to receive regular
|
||||
updates, new functionality, and ongoing support. For more information on AMD
|
||||
SMI, see the [AMD SMI documentation](https://rocm.docs.amd.com/projects/amdsmi/en/latest/).
|
||||
|
||||
#### Changes to ROCm Object Tooling
|
||||
|
||||
ROCm Object Tooling tools ``roc-obj-ls``, ``roc-obj-extract``, and ``roc-obj`` were
|
||||
deprecated in ROCm 6.4, and will be removed in a future release. Functionality
|
||||
has been added to the ``llvm-objdump --offloading`` tool option to extract all
|
||||
clang-offload-bundles into individual code objects found within the objects
|
||||
or executables passed as input. The ``llvm-objdump --offloading`` tool option also
|
||||
supports the ``--arch-name`` option, and only extracts code objects found with
|
||||
the specified target architecture. See [llvm-objdump](https://llvm.org/docs/CommandGuide/llvm-objdump.html)
|
||||
for more information.
|
||||
@@ -0,0 +1,15 @@
|
||||
root = true
|
||||
|
||||
[*]
|
||||
charset = utf-8
|
||||
end_of_line = lf
|
||||
indent_style = space
|
||||
indent_size = 4
|
||||
insert_final_newline = true
|
||||
trim_trailing_whitespace = true
|
||||
|
||||
[*.rst]
|
||||
indent_size = 3
|
||||
|
||||
[*.{html,md,yaml}]
|
||||
indent_size = 2
|
||||
@@ -0,0 +1,71 @@
|
||||
<table class="rocm-docs-table table">
|
||||
<thead>
|
||||
<tr>
|
||||
<th class="head">
|
||||
<p>Framework</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Supported versions</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Supported OS</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Supported Python versions</p>
|
||||
</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td rowspan="2" style="vertical-align: middle;">
|
||||
<p>PyTorch</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle;">
|
||||
<p>2.11.0, 2.10.0, 2.9.1</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>Linux</p>
|
||||
</td>
|
||||
<td rowspan="2">
|
||||
<p>3.14, 3.13, 3.12, 3.11</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style="vertical-align: middle;">
|
||||
<p>2.11.0</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle;">
|
||||
<p>Windows</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style="vertical-align: middle;">
|
||||
<p>JAX</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle;">
|
||||
<p>0.9.1, 0.8.2</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>Linux</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>3.14, 3.13, 3.12, 3.11</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style="vertical-align: middle;">
|
||||
<p>vLLM<br>(<a href="#release-supported-hw">gfx950, gfx942, gfx1200,<br>gfx1201, gfx1100,
|
||||
gfx1101,<br>gfx1102, gfx1151 GPUs only</a>)</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>0.19.1<br>(requires PyTorch 2.10.0)</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>Linux</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>3.13</p>
|
||||
</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
<table>
|
||||
@@ -0,0 +1,778 @@
|
||||
#### **AMD SMI (BM)** (26.4.0)
|
||||
|
||||
##### Added
|
||||
|
||||
* **Added APU metrics support (table versions 2.4 and 3.0)**.
|
||||
* New `amdsmi_apu_metrics_t` struct accessible via `amdsmi_gpu_metrics_t.apu_metrics` pointer (non-null when APU-specific metrics are available).
|
||||
* **v2.4 metrics**:
|
||||
* `temperature_gfx`, `temperature_soc`, `temperature_core[8]`, `temperature_l3[2]`
|
||||
* `average_gfx_activity`, `average_mm_activity`
|
||||
* `average_socket_power`, `average_cpu_power`, `average_soc_power`, `average_gfx_power`, `average_core_power[8]`
|
||||
* Average clocks: `gfxclk`, `socclk`, `uclk`, `fclk`, `vclk`, `dclk`
|
||||
* Current clocks: `gfxclk`, `socclk`, `uclk`, `fclk`, `vclk`, `dclk`, `coreclk[8]`, `l3clk[2]`
|
||||
* `average_temperature_gfx`, `average_temperature_soc`, `average_temperature_core[8]`, `average_temperature_l3[2]`
|
||||
* `average_cpu_voltage`, `average_soc_voltage`, `average_gfx_voltage`, `average_cpu_current`, `average_soc_current`, `average_gfx_current`
|
||||
* `throttle_status`, `indep_throttle_status`
|
||||
* `fan_pwm`
|
||||
* **v3.0 metrics**:
|
||||
* `temperature_core[16]`, `temperature_skin`
|
||||
* `average_vcn_activity`, `average_ipu_activity[8]`, `average_core_c0_activity[16]`
|
||||
* `average_dram_reads`, `average_dram_writes`, `average_ipu_reads`, `average_ipu_writes`
|
||||
* `average_apu_power`, `average_dgpu_power`, `average_all_core_power`, `average_ipu_power`, `average_sys_power`
|
||||
* `stapm_power_limit`, `current_stapm_power_limit`
|
||||
* `average_core_power[16]`, `current_coreclk[16]`
|
||||
* `current_core_maxfreq`, `current_gfx_maxfreq`
|
||||
* `average_vpeclk_frequency`, `average_ipuclk_frequency`, `average_mpipu_frequency`
|
||||
* `throttle_residency_prochot`, `throttle_residency_spl`, `throttle_residency_fppt`, `throttle_residency_sppt`, `throttle_residency_thm_core`, `throttle_residency_thm_gfx`, `throttle_residency_thm_soc`
|
||||
* `time_filter_alphavalue`
|
||||
* Fields not applicable to the current version are set to sentinel values: `0xFFFF` for `uint16_t`, `0xFFFFFFFF` for `uint32_t`, and `UINT64_MAX` for `uint64_t` fields.
|
||||
* Python bindings updated with `AmdSmiApuMetrics` ctypes structure.
|
||||
|
||||
* **Added `oam_id` to `amdsmi_enumeration_info_t`**.
|
||||
* `amd-smi list -e` now displays `OAM_ID` (Physical XGMI ID / OAM ID).
|
||||
* Added `--enumeration` as a long-form alias for `-e` in `amd-smi list`.
|
||||
|
||||
* **Added support for GPU metrics v1.9 new fields**.
|
||||
* Added new temperature fields to `amdsmi_gpu_metrics_t`:
|
||||
* `temperature_hbm_stacks` — per-stack HBM temperatures (°C)
|
||||
* `temperature_mid` — per-MID temperatures (°C)
|
||||
* `temperature_aid` — per-AID temperatures (°C)
|
||||
* `temperature_xcd` — per-XCC compute die temperatures (°C)
|
||||
* Added new per-die clock fields to `amdsmi_gpu_metrics_t`:
|
||||
* `current_uclk_aid` — per-AID uclk (MHz)
|
||||
* `current_socclks_mid` — per-MID SOC clock (MHz)
|
||||
* Added new constants:
|
||||
* `AMDSMI_MAX_NUM_HBM_STACKS` (12)
|
||||
* `AMDSMI_MAX_NUM_AID` (2)
|
||||
* `AMDSMI_MAX_NUM_MID` (2)
|
||||
* `AMDSMI_MAX_NUM_CLKS_PER_AID` (2)
|
||||
* `AMDSMI_MAX_NUM_CLKS_PER_MID` (2)
|
||||
|
||||
* **Added VRAM and GTT tuning interface**.
|
||||
* New `amd-smi static --mem-carveout` to view VRAM carveout options.
|
||||
* New `amd-smi set --mem-carveout` to change the VRAM carveout (APU).
|
||||
* New `amd-smi set --gtt` and `amd-smi reset --gtt` for system-wide GTT size tuning.
|
||||
* New APIs: `amdsmi_get_gpu_uma_carveout_info()`, `amdsmi_set_gpu_uma_carveout()`, `amdsmi_get_ttm_info()`, `amdsmi_set_ttm_pages_limit()`, `amdsmi_reset_ttm_pages_limit()`.
|
||||
|
||||
* **Added UBB power and power_limit fields to `amdsmi_power_info_t` and `amdsmi_npm_info_t`**.
|
||||
* `amd-smi metric --power` now displays `ubb_power` when available.
|
||||
* `amd-smi node -p` now displays UBB power threshold when available.
|
||||
|
||||
* **Added CPU support for family 1A Models 50h-57h**.
|
||||
* New APIs: `amdsmi_get_cpu_xgmi_pstate_range()`, `amdsmi_get_cpu_core_ccd_power()`, `amdsmi_get_cpu_tdelta()`, `amdsmi_get_cpu_dimm_sb_reg()`, `amdsmi_get_cpu_svi3_vr_controller_temp()`, `amdsmi_get_cpu_pc6_enable()`, `amdsmi_get_cpu_cc6_enable()`, `amdsmi_get_cpu_sdps_limit()`, `amdsmi_get_cpu_core_floor_freq_limit()`, `amdsmi_get_cpu_core_eff_floor_freq_limit()`, and corresponding set APIs.
|
||||
* **Note**: `amdsmi_get_dfc_ctrl()` renamed to `amdsmi_get_cpu_dfc_ctrl()` and `amdsmi_set_dfc_ctrl()` renamed to `amdsmi_set_cpu_dfc_ctrl()` for naming consistency.
|
||||
|
||||
* **Updated memory API documentation**
|
||||
Added note that the sum of per-process memory usage is not expected to equal total usage.
|
||||
|
||||
##### Changed
|
||||
|
||||
* **Renamed `processor_type_t` enum typedef to `amdsmi_processor_type_t`**.
|
||||
* The unprefixed typedef name did not follow the `amdsmi_*_t` convention used throughout `amdsmi.h` and was easy to collide with identifiers defined by other system-management libraries. New code should use `amdsmi_processor_type_t`. The old name is preserved as a backward-compatibility typedef alias, so existing callers continue to compile unchanged.
|
||||
|
||||
* **Package install no longer modifies the system-wide `logrotate` timer or cron schedule**.
|
||||
* Previously, installing `amd-smi-lib` overwrote `/lib/systemd/system/logrotate.timer` (or moved `/etc/cron.daily/logrotate` to `/etc/cron.hourly/`) to force hourly rotation, which affected every other package using `logrotate`.
|
||||
* The package now only ships `/etc/logrotate.d/amd_smi.conf`, which sets its own `hourly` + `size 1M` cadence. AMD-SMI logs still rotate at the same frequency; system-wide settings stay as the distribution configured them.
|
||||
|
||||
##### Optimized
|
||||
|
||||
* **Optimized `rsmi_dev_device_identifiers_get()` in the ROCm-SMI device layer**.
|
||||
* Removed unnecessary iteration by directly indexing the device list.
|
||||
* Added bounds checking for `device_id`, with clearer error handling/logging.
|
||||
* Improves performance for device identifier queries.
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
* **Fixed `amd-smi metric` crashing with `TypeError` on MI300A when no CPU flags are specified**.
|
||||
* When no CPU arguments are passed, `metric_cpu()` sets all boolean CPU args to `True` to display all available data. `--cpu-svi3-vr-controller-temp` takes a TYPE argument (and optional RAIL_INDEX) rather than a boolean flag — setting it to `True` caused a `TypeError` crash when the code tried to subscript it with `[0][0]`. Added `cpu_svi3_vr_controller_temp` to the show-all exclusion list, following the existing pattern for `cpu_lclk_dpm_level`, `cpu_io_bandwidth`, `cpu_dimm_sb_reg`, and similar argument-taking flags.
|
||||
|
||||
* **Fixed `amdsmi_get_gpu_accelerator_partition_profile()` returning incorrect `num_partitions` when `num_partition` is unavailable from GPU metrics**.
|
||||
* GPU metrics no longer always provides `num_partition`. The function now derives the partition count from the active partition type when `num_partition` is not available:
|
||||
* SPX → 1, DPX → 2, TPX → 3, QPX → 4
|
||||
* CPX → derived from the XCD counter via `amdsmi_get_gpu_xcd_counter()`
|
||||
|
||||
* **Fixed `amdsmi_topo_get_p2p_status()` returning a raw `ctypes.c_uint32` object instead of an integer for the `type` field**.
|
||||
* The `'type'` key in the returned dictionary now correctly returns `type_32.value` (an `int`) rather than the unwrapped ctypes object, consistent with the pattern used in `amdsmi_topo_get_link_type()`.
|
||||
|
||||
* **Adjusted KFD process caching to be more responsive**.
|
||||
* Updated process caching to allow cache duration adjustment via the `AMDSMI_PROCESS_INFO_CACHE_MS` environment variable for workflows with rapid metric polling.
|
||||
|
||||
* **Fixed CLI exit codes to use absolute values**.
|
||||
* Invalid GPU parameters now return positive error codes as documented.
|
||||
|
||||
* **Fixed CLI breakage when `amdgpu` driver is not present**.
|
||||
* Improved init to better catch driver loading issues.
|
||||
|
||||
* **Aligned `amdsmi_get_gpu_device_uuid()` with HIP/rocminfo UUID format**.
|
||||
* Modified `amdsmi_asic_info_t.asic_serial` to report per-socket serial using KFD's `unique_id`.
|
||||
|
||||
* **Fixed multiple bugs in NIC/switch code and `amdsmi_init()` NIC handling**.
|
||||
* Fixed `sizeof` operator precedence, `hw_mon` reset, NUMA=65535 handling, and several CLI function call errors.
|
||||
* Fixed `amdsmi_init()` to succeed when no NIC hardware is present.
|
||||
|
||||
* **Fixed shared mutex and self-heal**.
|
||||
* Improved self-heal logic to correctly identify and recover from corrupted or uninitialized mutex state.
|
||||
|
||||
* **Fixed `cu_occupancy` displaying `0%` instead of `N/A` when file is unavailable**.
|
||||
* Process `cu_occupancy` is now initialized to `INVALID` instead of zero, so `amd-smi process` displays `N/A` rather than a misleading `0%` when the sysfs file is not accessible.
|
||||
|
||||
* **Fixed CLI set commands silently succeeding on invalid input values**.
|
||||
* `amd-smi set --profile <INVALID>` now returns a non-zero exit code and lists available profiles in the error message; invalid profile names are rejected at parse time.
|
||||
* `amd-smi set --clk-level <CLK_TYPE>` (missing performance level indices) now returns a non-zero exit code with a usage hint instead of silently succeeding.
|
||||
* `amd-smi set --power-cap <OUT_OF_RANGE>` now returns a non-zero exit code.
|
||||
* `amd-smi set --fan <INVALID>%` no longer prompts the out-of-spec warning before validating the percentage range; invalid values are rejected immediately.
|
||||
|
||||
* **Fixed `amd-smi set --profile` help text omitting `BOOTUP_DEFAULT`**.
|
||||
* `BOOTUP_DEFAULT` was always accepted at runtime but was missing from the `--help` profile list. Auditing invalid-input handling exposed this gap. `amd-smi reset --profile` can also be used to return to the bootup default power profile.
|
||||
|
||||
* **Fixed `amd-smi monitor --brcm_nic` and `--brcm_switch` flags being registered on non-BRCM systems**.
|
||||
* These flags are now only registered when BRCM hardware is present, preventing spurious failures on AMD GPU-only systems.
|
||||
|
||||
* **Fixed `amd-smi` default command alignment**.
|
||||
* Updated default `amd-smi` output to align values to the left for improved readability.
|
||||
Several items were misaligned in the default output, and this change ensures a consistent left-aligned format across all fields.
|
||||
* *This change is purely cosmetic and does not affect any functionality.*
|
||||
|
||||
* **Renamed `lc_perf_other_end_recovery` to `lc_perf_other_end_recovery_count` in `amd-smi metric` CLI output for unification**.
|
||||
|
||||
* **Removed references to deprecated `amd-smi reset -r`**.
|
||||
* CLI help text and memory partition change warnings no longer reference `amd-smi reset -r` for driver reloading.
|
||||
* Users are now directed to use `sudo modprobe -r amdgpu && sudo modprobe amdgpu` to reload the driver after partition changes.
|
||||
|
||||
* **Changed CPU power APIs to return values in milliwatts (mW) for higher precision**.
|
||||
* Removed lossy integer rounding (`(mW + 500) / 1000`) from 6 CPU power get APIs. Values are now
|
||||
returned in milliwatts directly from the ESMI library, preserving sub-watt precision.
|
||||
* **C API**: Output parameter type remains `uint32_t*`, but the unit changed from watts to milliwatts (mW).
|
||||
* `amdsmi_get_cpu_socket_power`
|
||||
* `amdsmi_get_cpu_socket_power_cap`
|
||||
* `amdsmi_get_cpu_socket_power_cap_max`
|
||||
* `amdsmi_get_cpu_pwr_efficiency_mode` (ppt_limit field)
|
||||
* `amdsmi_get_cpu_core_ccd_power`
|
||||
* `amdsmi_get_cpu_sdps_limit`
|
||||
* **Python API (breaking)**: These functions now return `int` (milliwatts) instead of `str` (e.g., `"240 Watts"`).
|
||||
Callers that parsed the string output must update to handle the numeric return value.
|
||||
* **CLI output**: Power values now display with milliwatt precision (e.g., `240.500 Watts`).
|
||||
* Added missing null-pointer validation for output parameters in `amdsmi_get_cpu_socket_power_cap`
|
||||
and `amdsmi_get_cpu_socket_power_cap_max`.
|
||||
* Updated header documentation to specify milliwatt units for all affected get and set API parameters.
|
||||
|
||||
* **Changed power APIs to have consistent output parameter types**.
|
||||
* Modified 6 CPU power APIs to have consistent output power types. All set and get APIs have `uint32_t` output values.
|
||||
* Modified get and set APIs that had double output types to have `uint32_t` output types in milliwatts (mW).
|
||||
* `amdsmi_get_cpu_socket_power(amdsmi_processor_handle processor_handle, uint32_t* ppower)`
|
||||
* `amdsmi_get_cpu_socket_power_cap(amdsmi_processor_handle processor_handle, uint32_t* pcap)`
|
||||
* `amdsmi_get_cpu_socket_power_cap_max(amdsmi_processor_handle processor_handle, uint32_t* pmax)`
|
||||
* `amdsmi_get_cpu_pwr_efficiency_mode(amdsmi_processor_handle processor_handle, uint32_t* power_efficiency_mode, uint32_t* utilization, uint32_t* ppt_limit)`
|
||||
* `amdsmi_get_cpu_core_ccd_power(amdsmi_processor_handle processor_handle, uint32_t* power)`
|
||||
* `amdsmi_get_cpu_sdps_limit(amdsmi_processor_handle processor_handle, uint32_t* sdps_limit)`
|
||||
|
||||
#### **Composable Kernel** (1.3.0)
|
||||
|
||||
##### Added
|
||||
|
||||
* Added overload of `load_tile_transpose` that takes reference to output tensor as output parameter.
|
||||
* Use data type from LDS tensor view when determining tile distribution for transpose in the GEMM pipeline.
|
||||
* Added `eightwarps` support for abquant mode in blockscale GEMM.
|
||||
* Added `preshuffleB` support for abquant mode in blockscale GEMM.
|
||||
* Added support for explicit GEMM in `CK_TILE` grouped convolution forward and backward weight.
|
||||
* Added TF32 convolution support on gfx942 and gfx950 in CK. It can be enabled or disabled via `DTYPES` of `tf32`.
|
||||
* Added `streamingllm` sink support for FMHA FWD, include `qr_ks_vs`, `qr_async` and `splitkv` pipelines.
|
||||
* Added support for microscaling (MX) FP8/FP4 mixed data types to Flatmm pipeline.
|
||||
* Added support for fp8 dynamic tensor-wise quantization of FP8 fmha fwd kernel.
|
||||
* Added FP8 KV cache support for FMHA batch prefill.
|
||||
* Added FMHA batch prefill kernel support for several KV cache layouts, flexible page sizes, and different lookup table configurations.
|
||||
* Added gpt-oss sink support for FMHA FWD, include `qr_ks_vs`, `qr_async`, `qr_async_trload` and `splitkv` pipelines.
|
||||
* Added persistent async input scheduler for CK Tile universal GEMM kernels to support asynchronous input streaming.
|
||||
* Added FP8 block scale quantization for FMHA forward kernel.
|
||||
* Added gfx11xx support for FMHA.
|
||||
* Added microscaling (MX) FP8/FP4 support on gfx950 for FMHA forward kernel (`qr` pipeline only).
|
||||
* Added FP8 per-tensor quantization support for FMHA forward V3 pipeline on gfx950.
|
||||
|
||||
#### **HIP** (7.13)
|
||||
|
||||
##### Added
|
||||
|
||||
* New HIP APIs
|
||||
* `cooperative_groups::reduce()` allows calling reduce operators on `thread_block_tile` and `coalesced_threads`. The implementation is based on the `__reduce_*_sync` operations, so the macro `HIP_ENABLE_EXTRA_WARP_SYNC_TYPES` might be needed to unlock some optimizations.
|
||||
* New device attribute `hipDeviceAttributeGPUDirectRDMAWithHipVMMSupported`, indicating support for GPU Direct RDMA when using HIP VMM. This attribute corresponds to the CUDA `CU_DEVICE_ATTRIBUTE_GPU_DIRECT_RDMA_WITH_CUDA_VMM_SUPPORTED`.
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
* A segmentation fault that occurred in child graphs during the graph‑launch phase. The issue originated from the entire graph being launched solely according to the parent graph’s scheduling logic. The HIP runtime now introduces a per‑graph segment‑scheduling control flag and propagates the parent graph’s scheduling mode to its child graphs, ensuring consistent scheduling behavior (classic vs. segment) and preventing failures when the parent falls back to classic scheduling.
|
||||
* A segmentation fault caused by passing a null pointer to the hipMemGetAddressRange API. The function now handles null pointers correctly, matching the behavior of the corresponding CUDA API.
|
||||
|
||||
##### Changed
|
||||
|
||||
* `__reduce_and_sync()`, `__reduce_or_sync()` and `__reduce_xor_sync()` now provide a consistent behavior for all mask values and with CUDA. Previously, some masks were translated into bitwise operations, but others were not (such as those containing "holes"). Now, all masks cause bitwise instructions to be emitted. This is a change in behavior compared to previous versions.
|
||||
|
||||
##### Optimized
|
||||
|
||||
* Improved HIP runtime error logging when an application's fat binary does not include a compatible code object for the detected GPU architecture, offering clearer guidance to rebuild with the appropriate `--offload-arch=gfxXXXX` option.
|
||||
|
||||
* Enables in‑memory and background‑thread asynchronous logging in the HIP runtime by default to improve overall logging capability. This behavior can be disabled by setting the environment variable `AMD_LOG_ASYNC=0`.
|
||||
|
||||
#### **hipBLAS** (3.4.0)
|
||||
|
||||
##### Added
|
||||
|
||||
* gfx1250 and gfx90c support to clients.
|
||||
* Version and other properties to Windows `hipblas.dll`.
|
||||
* Support for `OpenBLAS` ILP64-based API usage in clients.
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
* Restored the fallback of using the deprecated rocBLAS API `rocblas_set_device_memory_size` if allocations are failing.
|
||||
|
||||
#### **hipBLASLt** (1.3.0)
|
||||
|
||||
##### Added
|
||||
|
||||
* General Batched GEMM support.
|
||||
|
||||
##### Changed
|
||||
|
||||
* Replaced `install.sh` with an invoke-based task runner (`tasks.py`) to support cross-platform builds including Windows (ROCm 7.0+).
|
||||
* `gtest` and `msgpack-cxx` are now fetched automatically using CMake FetchContent if not found on the system.
|
||||
|
||||
#### **hipCUB** (4.4.0)
|
||||
|
||||
##### Optimized
|
||||
|
||||
* Reduced build times for unit tests.
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
* Fixed more memory leak issues with some unit tests.
|
||||
|
||||
#### **hipFFT** (1.0.23)
|
||||
|
||||
##### Added
|
||||
|
||||
* hipFFTW plan creation functions for advanced and general plans:
|
||||
* `fftw_plan_many_dft`
|
||||
* `fftwf_plan_many_dft`
|
||||
* `fftw_plan_many_dft_r2c`
|
||||
* `fftwf_plan_many_dft_r2c`
|
||||
* `fftw_plan_many_dft_c2r`
|
||||
* `fftwf_plan_many_dft_c2r`
|
||||
* `fftw_plan_guru_dft`
|
||||
* `fftwf_plan_guru_dft`
|
||||
* `fftw_plan_guru_dft_r2c`
|
||||
* `fftwf_plan_guru_dft_r2c`
|
||||
* `fftw_plan_guru_dft_c2r`
|
||||
* `fftwf_plan_guru_dft_c2r`
|
||||
* `fftw_plan_guru64_dft`
|
||||
* `fftwf_plan_guru64_dft`
|
||||
* `fftw_plan_guru64_dft_r2c`
|
||||
* `fftwf_plan_guru64_dft_r2c`
|
||||
* `fftw_plan_guru64_dft_c2r`
|
||||
* `fftwf_plan_guru64_dft_c2r`
|
||||
* Support for gfx1150 architecture.
|
||||
|
||||
##### Changed
|
||||
|
||||
* Moved library to C++20 standard.
|
||||
* Removed Boost as a dependency for clients and samples.
|
||||
* Callback functions will be deprecated in a future release.
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
* Fixed potential launch failure of data generation kernels in test and benchmark programs.
|
||||
|
||||
#### **hipRAND** (3.3.0)
|
||||
|
||||
##### Added
|
||||
|
||||
* `hiprand.dll` now contains embedded file version metadata.
|
||||
|
||||
#### **hipSOLVER** (3.4.0)
|
||||
|
||||
##### Added
|
||||
|
||||
* Compatibility-only functions:
|
||||
* `geev`
|
||||
* `hipsolverDnXgeev_bufferSize`
|
||||
* `hipsolverDnXgeev`
|
||||
* `syevBatched`
|
||||
* `hipsolverDnXsyevBatched_bufferSize`
|
||||
* `hipsolverDnXsyevBatched`
|
||||
* `syevd`
|
||||
* `hipsolverDnXsyevd_bufferSize`
|
||||
* `hipsolverDnXsyevd`
|
||||
* `sytrs`
|
||||
* `hipsolverDnXsytrs_bufferSize`
|
||||
* `hipsolverDnXsytrs`
|
||||
|
||||
#### **hipSPARSELt** (0.2.8)
|
||||
|
||||
##### Added
|
||||
|
||||
* CTest and test categories support (`--smoke`, `--pre_checkin`, and `--nightly`).
|
||||
|
||||
##### Optimized
|
||||
|
||||
* Provided more kernels for the `FP16`, `BF16`, and `Int8` datatypes.
|
||||
* Improved the performance of the `HIPSPARSELT_PRUNE_SPMMA_TILE` function.
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
* Fixed incorrect behavior when retrieving the PCI chip ID.
|
||||
* Fixed LDS out-of-bounds read in `prune_tile_kernel`.
|
||||
* Fixed out-of-bounds access for compress function test cases.
|
||||
* Fixed missing null terminator in the return value of `hipsparseLtGetArchName()`.
|
||||
* Fixed incorrect CPU result when `bias_type` is `BF16` for spmm test cases.
|
||||
* Fixed double-free issue in the example code `example_prune_strip`.
|
||||
* Fixed symbol interposition in the hipSPARSELt library.
|
||||
|
||||
#### **MIOpen** (3.5.1)
|
||||
|
||||
##### Added
|
||||
|
||||
* Added `MIOPEN_LOG_BUFFER_SIZE` option: when set to non-zero, dumps recent MIOpen logs to file on error.
|
||||
* [Conv] Added `ConvDepthwiseFwd3D` solver for optimizing specific 3D depthwise convolutions.
|
||||
* [Conv] Added NHWC layout support for Winograd convolution solvers.
|
||||
* [Conv] Added regular GEMM solver support for Conv3D forward and backward-data with 1x1x1 filters.
|
||||
* [Conv] Added configurable problem size threshold (`MIOPEN_CONV_DIRECT_MAX_SIZE`) for direct solver.
|
||||
* [Softmax] Added tuning support via Generic Search.
|
||||
|
||||
##### Changed
|
||||
|
||||
* [Conv] Improved default kernel selection for Composable Kernel (CK) convolution solvers with ranked shortlists.
|
||||
* [Conv] Split CK grouped convolution kernels into per-architecture runtime-loaded dynamic libraries.
|
||||
|
||||
##### Optimized
|
||||
|
||||
* Optimized transpose operations with tiled and vectorized variants for NCHW/NHWC conversions.
|
||||
* [BatchNorm] Optimized batchnorm reduction using warp shuffle intrinsics.
|
||||
* [Conv] Added heuristic filtering of slow GEMM solver configurations during tuning.
|
||||
|
||||
##### Deprecated
|
||||
|
||||
* [Conv] Deprecated CK non-grouped convolution forward and backward solvers.
|
||||
* Deprecated `miopenConvolutionBackwardBias`: the underlying OpenCL kernel (`MIOpenConvBwdBias.cl`) has been removed. The function now returns `miopenStatusNotImplemented` and will be removed in a future release.
|
||||
|
||||
##### Removed
|
||||
|
||||
* Removed GraphAPI experimental feature and related code.
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
* [Conv] Fixed Winograd Fury grouped convolution correctness on gfx12xx when G > 1.
|
||||
* [Conv] Fixed bf16 WrW convolution precision loss in inter-batch accumulation.
|
||||
* [Conv] Fixed GPU memory fault in Winograd v3.0 WrW solver for large tensor shapes.
|
||||
* Fixed BF16 `abs` function precision error caused by unnecessary cast through FP16.
|
||||
* Fixed pooling kernel runtime compilation failure.
|
||||
* Fixed gfx1151 inline assembly compilation errors in batchnorm kernels.
|
||||
* Fixed use-after-free in HIPOCProgram binary loading.
|
||||
|
||||
#### **ROCm Data Center Tool (RDC)** (1.3.0)
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
* **Fixed broken partition metrics**.
|
||||
* Regardless of whether the GPU was partitioned, RDC only saw the GPU index and no instances due to upstream gpu_metrics changes.
|
||||
|
||||
#### **rocBLAS** (5.4.0)
|
||||
|
||||
##### Added
|
||||
|
||||
* gfx1250 and gfx90c enabled.
|
||||
* Trace logging using `ROCBLAS_LAYER=1` for `rocblas_gemm_ex_get_solutions`, `rocblas_gemm_batched_ex_get_solutions`, `rocblas_gemm_ex_get_solutions_by_type`, and `rocblas_gemm_batched_ex_get_solutions_by_type`.
|
||||
* Version and other properties to Windows `rocblas.dll`.
|
||||
* Support for `OpenBLAS` ILP64 API for host reference in clients.
|
||||
* Dockerfiles in the `docker` directory to assist in setting up development.
|
||||
|
||||
##### Optimized
|
||||
|
||||
* Improved the performance of Level 3 `geam` for pure transpose scale use cases.
|
||||
* Improved the performance of Level 2 `tpsv`.
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
* Fix for querying solutions when using the `hipBLASLt` backend with `rocblas_gemm_batched_ex_get_solutions` if using null data pointers.
|
||||
|
||||
#### **ROCdbgapi** (0.80.0)
|
||||
|
||||
##### Added
|
||||
|
||||
* `amd_dbgapi_process_get_info()` adds a new query to get a mask spanning
|
||||
over all the bits used by all the address spaces. The query is called
|
||||
`AMD_DBGAPI_PROCESS_INFO_SIGNIFICANT_ADDRESS_BITS`.
|
||||
|
||||
#### **rocDecode** (1.8.0)
|
||||
|
||||
##### Added
|
||||
|
||||
* Logging improvement: Added function entry and exit logs (at Info log level).
|
||||
* Logging improvement: Added duration to function exit logs and optimized log message formatting to reduce runtime overhead.
|
||||
* Logging improvement: Merged all logger instances into one global instance.
|
||||
* Logging improvement: Unified logging format in utility classes with core library logging format.
|
||||
* Logging improvement: Moved debug logging from a compile-time switch to the runtime logger level controlled by `ROCDEC_LOG_LEVEL` (debug = 4).
|
||||
* Added support for user-set output surface format.
|
||||
|
||||
##### Changed
|
||||
|
||||
* Removed CPack packaging (DEB/RPM/NSIS/TGZ/ZIP generation and all related CPACK variables).
|
||||
* Removed `rocDecode-setup.py` dependency installer script.
|
||||
* Removed Docker files.
|
||||
* Removed package install documentation; updated all documentation to reference TheRock for installation.
|
||||
* Simplified libva version check (single `>= 1.22` requirement).
|
||||
* Cleaned up CMake error messages.
|
||||
|
||||
#### **rocFFT** (1.0.37)
|
||||
|
||||
##### Optimized
|
||||
|
||||
* Allow plans to share hipModules if they use the same kernels. This reduces time spent and memory used when
|
||||
creating plans that exist concurrently.
|
||||
* Improved performance of unit-strided, interleaved, complex-to-complex and real-to-complex FFTs on gfx1201, gfx90a, gfx942, and gfx950.
|
||||
|
||||
Single-precision lengths:
|
||||
* (160,72,72)
|
||||
* (160,80,72)
|
||||
* (160,80,80)
|
||||
* (72,72,72)
|
||||
* (80,80,80)
|
||||
* (84,84,72)
|
||||
* (96,96,96)
|
||||
* (108,108,80)
|
||||
|
||||
Double-precision lengths:
|
||||
* (72,72,52)
|
||||
* (60,60,60)
|
||||
* (64,64,52)
|
||||
* (64,64,64)
|
||||
|
||||
##### Changed
|
||||
|
||||
* Moved library to C++20 standard.
|
||||
* Removed Boost as a dependency for clients and samples.
|
||||
* Split the precompiled kernel cache file (`rocfft_kernel_cache.db`) into per-architecture files (`rocfft_kernel_cache_gfx950.db`, `rocfft_kernel_cache_gfx1201.db`, etc).
|
||||
* `rocfft_plan_create` returns `rocfft_status_invalid_offset` for any usage of non-zero offsets in plan descriptions. The feature is not supported yet.
|
||||
* Callback functions will be deprecated in a future release.
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
* Potential issue with data generation for multi-dimensional transforms in rocfft-tests and rocfft-bench.
|
||||
* An issue that sometimes blocked complex-to-complex FFT plan creation when using noncontiguous strides in multiple dimensions.
|
||||
* An issue that sometimes blocked complex-to-real FFT plan creation when using noncontiguous strides in multiple dimensions.
|
||||
* An issue that sometimes blocked complex-to-real FFT plan creation when using noncontiguous strides with small lengths on the two fastest dimensions.
|
||||
* Potential launch failure of data generation kernels in test and benchmark programs.
|
||||
* Incorrect results on some strided real-complex FFTs on gfx90a.
|
||||
* Incorrect results on some even-length real FFTs that have odd-length strides on higher dimensions.
|
||||
* Callbacks on MPI transforms when not all ranks have the same number of data bricks.
|
||||
* Functional issues for multi-device, in-place real transforms.
|
||||
* Functional issues for multi-dimensional, multi-device transforms involving some unit length(s).
|
||||
* Functional issues for multi-device transforms involving data divisions along the slowest-varying axis (only) for some bricks but not all.
|
||||
* Functional issues for multi-device transforms setting no field on input or output.
|
||||
* Automatic allocation of work memory at plan execution time, when work memory is required on multiple devices.
|
||||
|
||||
#### **rocJPEG** (1.5.0)
|
||||
|
||||
##### Changed
|
||||
|
||||
* rocJPEG is now delivered as part of [TheRock](https://github.com/ROCm/TheRock). All core dependencies are provided by the TheRock build.
|
||||
* Removed CPack packaging (DEB/RPM/NSIS/TGZ/ZIP generation and all related CPACK variables).
|
||||
* Removed `rocJPEG-setup.py` dependency installer script.
|
||||
* Removed Docker files.
|
||||
* Removed package install documentation; updated all documentation to reference TheRock for installation.
|
||||
* Simplified libva version check (single `>= 1.22` requirement).
|
||||
* Cleaned up CMake error messages.
|
||||
|
||||
#### **ROCm Compute Profiler** (3.6.0)
|
||||
|
||||
##### Added
|
||||
|
||||
* Added L2 memory bandwidth derived metrics under `--membw-analysis` to allow L2 memory bandwidth specific profiling and analysis metric block 30.
|
||||
|
||||
* Added AMD Ryzen AI Max 300 series (gfx1151) support.
|
||||
* New memory hierarchy visualization for RDNA 3.5 (gfx115X) in analyze CLI mode.
|
||||
|
||||
* Introduced support for AMD Instinct MI350P GPU.
|
||||
|
||||
* ``--view table`` option in analyze mode to force all TTY output to plain tables and ignore ``cli_style`` from YAML config (for example, mem_chart, Roofline charts render as tables). The ``--view`` argument is reserved for future TTY views (for example, other chart styles).
|
||||
|
||||
* Added EA memory bandwidth derived metrics under `--membw-analysis` to allow EA memory bandwidth specific profiling and analysis metric block 30.
|
||||
|
||||
##### Changed
|
||||
|
||||
* Standalone roofline (`--roof-only` option) in profile mode now creates `roofline.csv` only. HTML roofline charts are generated via `rocprof-compute analyze`. The `calc_ai_profile()` function has been removed; `calc_ai_analyze()` is the single source of truth for arithmetic intensity calculation.
|
||||
* Roofline visualization options (`--sort`, `--mem-level`, `--roofline-data-type`) have moved from profile mode to analyze mode.
|
||||
|
||||
* Standardized unit naming in analysis configs and Python utilities: `pct`/`Pct` → `Percent`, `instr` → `Instructions`.
|
||||
|
||||
* Profile mode output format:
|
||||
* Profile mode now creates separate counter collection files for each application replay (pmc_perf_*.csv or results_*.csv).
|
||||
* Analyze mode automatically merges these files into a unified pmc_perf.csv containing information from all application replays during pre-processing.
|
||||
|
||||
* ROCm Compute Profiler now builds and runs profile mode with vanilla Python without requiring any Python dependencies to be installed via `pip`.
|
||||
* Note that analysis mode will still require Python dependencies and will report any missing packages.
|
||||
|
||||
##### Removed
|
||||
|
||||
* Removed HIP API tracing since it's out-of-scope for ROCm Compute Profiler and the trace files were not being analyzed.
|
||||
|
||||
##### Optimized
|
||||
|
||||
* Filtering for block 21 (`-b 21`) in profile mode now only performs pc sampling and skips unnecessary counter collection.
|
||||
* Filtering for block 21 in analysis mode now skips metrics calculations and only shows kernel/dispatch/system statistics and pc sampling table.
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
* Fixed roofline benchmark MFMA FP16/BF16/INT8 peaks for MI350.
|
||||
* Fixed an issue where pc sampling profiling failed with multi-argument commands and live process attachment.
|
||||
|
||||
##### Upcoming changes
|
||||
|
||||
* `--path` and `--subpath` options are deprecated and will be removed in a future release.
|
||||
* Intermediate CSV generation (`results_*.csv`) from rocpd databases during profiling is deprecated and will be removed in a future release. The analyze step will read `.db` files directly.
|
||||
* `--retain-rocpd-output` is deprecated and will be removed in a future release. `.db` files will be retained by default.
|
||||
|
||||
##### Known issues
|
||||
|
||||
* For AMD Ryzen AI Max 300 series, the roofline metrics table will have N/A values for "peak" field.
|
||||
* This is planned to be addressed by adding empirical benchmark support for AMD Ryzen AI Max 300 series in a future release.
|
||||
|
||||
#### **ROCm Systems Profiler** (1.6.0)
|
||||
|
||||
##### Added
|
||||
|
||||
* Kernel Fusion Driver (KFD) event tracing support to capture page faults, page migrations, queue evictions, GPU unmap events, and dropped events. Requires ROCprofiler-SDK 1.2.1 or later. Enable with `ROCPROFSYS_ROCM_DOMAINS=kfd_events`.
|
||||
* Support for pause and resume of profiling via `roctxProfilerPause` and `roctxProfilerResume`.
|
||||
* Support for selective region tracing via the `ROCPROFSYS_SELECTED_REGIONS` environment variable, limiting tracing to specified regions.
|
||||
* `--selected-regions` CLI argument to `rocprof-sys-sample`, `rocprof-sys-run`, and `rocprof-sys-instrument` for specifying selective region tracing from the command line.
|
||||
* Support for re-attaching to a previously profiled process. After detaching, `rocprof-sys-attach` can re-attach to the same PID for a new profiling session.
|
||||
* MPI-rank-based file output filtering feature controlled with two new CLI arguments: `--rank-filter-output` and `--rank-filter-id`.
|
||||
* JSON-based configurable preset system with `--preset=<name>` flag, replacing the old `--<preset-name>` flags. Presets are now loaded from JSON files in `source/bin/common/presets/`, making them extensible and exportable. Use `--list-presets` to see available presets and `--explain=<name>` for detailed preset information.
|
||||
* Domain flags for composable configuration: `--gpu[=metrics]`, `--rocm[=domains]`, `--cpu[=hz]`, `--parallel[=runtimes]`. Domain flags can be combined with presets to customize profiling without editing configuration files.
|
||||
* Configuration export via `--export-config[=file]` to save resolved settings as reusable JSON configuration files. Exported configs can be loaded back with `--preset=./config.json`.
|
||||
* Topic-based help system: `--help` now shows a compact summary with essential options and a list of help topics. Use `--help=<topic>` (e.g., `--help=sampling`, `--help=gpu`, `--help=tracing`) to see only relevant options. Use `--help=all` for the full option listing.
|
||||
* Post-run output summary during library finalization showing result file locations.
|
||||
* JSON schema file (`share/rocprofiler-systems/presets/schema.json`) for preset validation.
|
||||
* Documentation (`docs/how-to/instrumenting-rewriting-binary-application.rst`) describing what to do when Dyninst reports a "Failed to transform trace" error during instrumentation.
|
||||
|
||||
##### Changed
|
||||
|
||||
* `rocprof-sys-avail` no longer queries GPU devices or hardware counters unless `--hw-counters` or `--all` is requested, reducing startup time and allowing settings/component queries in environments without GPU/ROCm.
|
||||
* `rocprof-sys-instrument` diagnostic file dumps (available, instrumented, excluded, coverage, overlapping) are now gated behind the `--dump-info` flag instead of being generated unconditionally.
|
||||
* Preset flags changed from `--balanced` to `--preset=balanced` syntax. The old `--<preset-name>` flags are still supported and handled within `preset_registry`.
|
||||
* Removed the `ROCPROFSYS_USE_ROCM` CMake option. ROCm is now required for building the ROCm Systems Profiler.
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
* Fixed an issue where the `--rocm-domains` CLI option for `rocprof-sys-run` was not recognized.
|
||||
|
||||
#### **rocminfo** (1.0.0)
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
* Fixed BDF (Bus:Device.Function) ID truncation issue that caused incorrect display of PCI device identifiers. The `bdf_id` field was incorrectly declared as `uint16_t` instead of `uint32_t`, causing silent truncation when HSA runtime returned the full 32-bit BDF ID value. This has been corrected to properly display complete BDF information for all GPU agents.
|
||||
|
||||
#### **rocPRIM** (4.4.0)
|
||||
|
||||
##### Added
|
||||
|
||||
* Added type trait definitions for `__hip_bfloat16`. This should resolve issues where this type did not work with radix-based algorithms.
|
||||
* Unit tests for config_types.
|
||||
|
||||
##### Optimized
|
||||
|
||||
* Reduced build times for unit tests.
|
||||
* Reduced memory usage in unit tests.
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
* Fixed a silent overflow in `rocprim::device_segmented_reduce` where it could exceed the maximum number of HIP threads, resulting in missing output.
|
||||
* Certain large unit tests now properly detect if insufficient system memory is present and skip the test case accordingly.
|
||||
* Fixed out-of-bounds memory access in block run length decode.
|
||||
* Fixed memory leak in unit tests.
|
||||
|
||||
#### **ROCprofiler-SDK** (1.3.0)
|
||||
|
||||
##### Added
|
||||
|
||||
**API:**
|
||||
|
||||
* Late-start profiling support: Enables profiling when `rocprofiler-sdk` is loaded after HSA/HIP runtimes have already initialized.
|
||||
* `rocprofiler_force_configure()` now automatically detects and profiles runtimes initialized before the SDK loads.
|
||||
* Integrates with `rocprofiler-register` to retrieve the registered API tables.
|
||||
* Supports all runtime types (HSA, HIP, ROCTX, RCCL, rocDecode, rocJPEG, and more) automatically.
|
||||
* No explicit late-start API calls required; works transparently.
|
||||
|
||||
* KFD (Kernel Fusion Driver) event tracing support:
|
||||
* Buffer service configurations for each KFD buffer tracing type.
|
||||
* New type `tool_buffer_tracing_kfd_record_t` using `std::variant` to wrap 8 different KFD buffer tracing types.
|
||||
* Each KFD event generates `rocpd_info_pmc`, `rocpd_event`, `rocpd_region`, and `rocpd_pmc_event` rows.
|
||||
* Fixed handling for special SVM location in KFD prefetch location reporting.
|
||||
* Fixed parsing for queue restore events to handle both correct format (character '0') and broken driver format (NULL character '\0').
|
||||
|
||||
**rocprofv3 (CLI):**
|
||||
|
||||
* Multi-pass counter collection support: Support for multiple `--pmc` flags to define separate counter groups for different profiling passes.
|
||||
* Ability to combine command-line `--pmc` flags with input file counter groups.
|
||||
* Each pass generates output in a separate `pass_n` subdirectory.
|
||||
* Example: `rocprofv3 --pmc SQ_WAVES --pmc GRBM_COUNT -- <app>` creates two profiling passes.
|
||||
|
||||
* KFD (Kernel Fusion Driver) event tracing support:
|
||||
* KFD record dumping to `rocpd` with support for 8 main KFD event types.
|
||||
* Support for `rocpd` to Perfetto conversion for KFD events.
|
||||
* `--kfd-trace` flag to enable KFD event tracing.
|
||||
|
||||
* ROCTx support for ATT: Added ROCtx support to device thread trace when using `--att --selected-regions`.
|
||||
* Allows `roctxProfilerPause` and `roctxProfilerResume` to explicitly control when ATT data collection starts and stops.
|
||||
* Enables more precise, region-focused ATT tracing with reduced overhead and noise.
|
||||
* Supports multiple resume/pause cycles, each producing separate trace output files.
|
||||
* Incompatible with `--att-consecutive-kernels`.
|
||||
|
||||
* PC sampling support for dynamic attach: Allows users to attach to a running application and collect PC samples without restarting the workload.
|
||||
* Enables profiling long-running or production-style jobs at the point of interest.
|
||||
* Results integrate with the existing PC sampling analysis flow.
|
||||
|
||||
**Documentation:**
|
||||
|
||||
* Added marker-controlled thread tracing section to the thread trace how-to guide.
|
||||
* Added cross-reference from ROCTx documentation to ATT with `selected-regions`.
|
||||
|
||||
##### Changed
|
||||
|
||||
**Implementation:**
|
||||
|
||||
* Late-start architecture redesign: Removed direct runtime symbol access in favor of proper rocprofiler-register integration.
|
||||
* Replaced ~600 lines of `dlopen`/`dlsym` bypass logic with ~80 lines by using `rocprofiler_register_invoke_all_registrations()`.
|
||||
* Late-start now works by requesting `rocprofiler-register` to re-propagate stored API tables.
|
||||
* Extensible design. Automatically supports new runtimes without SDK code changes.
|
||||
* Provides a proper separation of concerns. `rocprofiler-register` manages the table storage while SDK manages the table wrapping.
|
||||
* Counter dimension encoding changed from fixed-width to variable-width allocation per dimension type.
|
||||
* Dimension selection and reduction logic now uses explicit dimension masks and single-index selection.
|
||||
* HSA queue interception extended to handle AMD extended kernel dispatch packets.
|
||||
|
||||
##### Removed
|
||||
|
||||
* Counter collection support for plain text (`.txt`) input files. Only structured file formats (JSON and YAML) with schema validation are now supported.
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
* Fixed rocpd OTF2 output to add `ACCELERATOR_DEVICE` as system tree node domain for AMD devices.
|
||||
* Fixed `rocprofv3` input file parsing where comment lines containing `pmc:` were incorrectly processed as valid counter collection directives, causing unintended profiling passes.
|
||||
|
||||
#### **rocRAND** (4.4.0)
|
||||
|
||||
##### Added
|
||||
|
||||
* gfx1150 and gfx1152 support.
|
||||
* rocrand.dll now contains embedded file version metadata.
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
* Fixed memory leak in unit tests.
|
||||
|
||||
#### **rocSHMEM** (3.4.0)
|
||||
|
||||
##### Added
|
||||
|
||||
* Added new APIs:
|
||||
* `rocshmem_quiet_on_stream`
|
||||
* `rocshmem_sync_all_on_stream`
|
||||
* `rocshmem_TYPENAME_alltoall_wg`
|
||||
* `rocshmem_TYPENAME_alltoallv_wg`
|
||||
* `rocshmem_team_my_pe`
|
||||
* `rocshmem_team_n_pes`
|
||||
* `rocshmem_barrier`
|
||||
* `rocshmem_barrier_wave`
|
||||
* `rocshmem_barrier_wg`
|
||||
* `rocshmem_buffer_register`
|
||||
* `rocshmem_buffer_unregister`
|
||||
* `rocshmem_info_get_version`
|
||||
* `rocshmem_info_get_name`
|
||||
* `rocshmem_vendor_get_version_info`
|
||||
* Added library constants: `ROCSHMEM_MAJOR_VERSION`, `ROCSHMEM_MINOR_VERSION`,
|
||||
`ROCSHMEM_MAX_NAME_LEN`, `ROCSHMEM_VENDOR_STRING`, `ROCSHMEM_VERSION`,
|
||||
`ROCSHMEM_VENDOR_MAJOR_VERSION`, `ROCSHMEM_VENDOR_MINOR_VERSION`,
|
||||
`ROCSHMEM_VENDOR_PATCH_VERSION`.
|
||||
* Added vendor string and backend metadata to the `rocshmem_info` output.
|
||||
* Added `ROCSHMEM_TEAM_WORLD` for device code.
|
||||
* Added `ROCSHMEM_TEAM_SHARED` predefined team for PEs sharing a common memory domain (same node).
|
||||
* Added new environment variables:
|
||||
* `ROCSHMEM_GDA_OVERRIDE_NIC_FIRMWARE_CHECK`
|
||||
* `ROCSHMEM_GDA_NUM_QPS_PER_PE_DEFAULT_CTX`
|
||||
* `ROCSHMEM_GDA_NUM_QPS_PER_PE_USR_CTX`
|
||||
* Added VMM POSIX memory allocator (`USE_HEAP_DEVICE_VMM_POSIX`):
|
||||
* Uses HIP Virtual Memory Management (VMM) APIs for fine-grained memory control.
|
||||
* Requires ROCm 7.0+ and Linux kernel 5.6+.
|
||||
* Not compatible with MPI-based initialization (use `ROCSHMEM_INIT_WITH_UNIQUEID` instead).
|
||||
|
||||
##### Changed
|
||||
|
||||
* Use CQ collapsing for the Mellanox MLX5 GDA conduit.
|
||||
|
||||
#### **rocSOLVER** (3.34.0)
|
||||
|
||||
##### Added
|
||||
|
||||
* Computation of solution for LU factorization without pivoting:
|
||||
* GETRS_NPVT (with batched and strided\_batched versions)
|
||||
* GETRS_NPVT_64 (with batched and strided\_batched versions)
|
||||
* Linear solver routines for symmetric matrices:
|
||||
* SYTRS (with batched and strided\_batched versions)
|
||||
* SYTRS_64 (with batched and strided\_batched versions)
|
||||
|
||||
##### Optimized
|
||||
|
||||
* Improved the performance of POTF2 and downstream functions such as POTRF.
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
* Fixed a memory access error in SYTRF and synchronization issues in LASYF and SYTF2.
|
||||
|
||||
#### **rocSPARSE** (4.6.0)
|
||||
|
||||
##### Added
|
||||
|
||||
* `rocsparse_create_const_bsr_descr` routine for creating a const sparse BSR matrix descriptor.
|
||||
* `rocsparse_spic0` and `rocsparse_spilu0` routines for incomplete factorizations, with strided batched computations enabled.
|
||||
* `rocsparse_sptrsv_descr_create` and `rocsparse_sptrsv_descr_destroy` routines.
|
||||
* `rocsparse_singularity` enumeration.
|
||||
* `rocsparse_sptrsv_output_singularity` and `rocsparse_sptrsv_output_singularity_position` in `rocsparse_sptrsv_output`.
|
||||
* Strided batched computations for `rocsparse_sptrsv`.
|
||||
|
||||
##### Optimized
|
||||
|
||||
* Significant performance improvement for `rocsparse_Xgtsv_no_pivot_strided_batch`.
|
||||
* Significant performance improvement for `rocsparse_Xgtsv_no_pivot`.
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
* Fixed incorrect usage of `__syncthreads` in `bsrmm`, `csrmm` (row_split), and `csritilu0x`.
|
||||
* Fixed incorrect usage of `__syncthreads` in `csx2dense`, `dense2csx`, `prune_dense2csr`, `csrcolor`, and `csrmm` (`nnz_split`).
|
||||
* Fixed `rocsparse_[s|d|c|z]csric0` where `rocsparse_status_invalid_value` was being returned when the maximum number of non-zeros in any row is between 513 and 1024.
|
||||
* Fixed compilation when using `--rocsparse_ILP64`.
|
||||
* Fixed off-by-one heap-buffer-overflow in temporary buffer allocation for `rocsparse_csrsort`, `rocsparse_check_matrix_csr`, and `rocsparse_check_matrix_gebsr` (and their delegating routines `rocsparse_cscsort`, `rocsparse_coosort`, `rocsparse_check_matrix_csc`, and `rocsparse_check_matrix_gebsc`) where the `shift_offsets_kernel` temp buffer was sized for `m` elements instead of `m+1`.
|
||||
|
||||
##### Removed
|
||||
|
||||
* The deprecated C++14 support, which is no longer supported by the rocPRIM dependency.
|
||||
|
||||
#### **rocThrust** (4.4.0)
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
* Fixed memory leak in unit test.
|
||||
* Fixed unit test compatibility with ASAN.
|
||||
|
||||
#### **rocWMMA** (2.2.1)
|
||||
|
||||
##### Added
|
||||
|
||||
* Added the following community samples for external contributions, with build support and documentation:
|
||||
* `simple_gemm_silu`: demonstrates a GEMM + SiLU fused operator using the rocWMMA API.
|
||||
* `simple_gemm_fusion`: demonstrates block-tile-level dual-GEMM fusion using the rocWMMA API.
|
||||
* `simple_gemm_swiglu`: demonstrates a SwiGLU fused dual-GEMM kernel (LLaMA/Mistral FFN gate layer) using the rocWMMA API.
|
||||
|
||||
##### Changed
|
||||
|
||||
* Updated the `find_package` search for OpenMP to prefer the `openmp-config.cmake` provided by ROCm, with a fallback to module search mode.
|
||||
* Updated `INSTALL_RPATH` and added `BUILD_RPATH` for OpenMP.
|
||||
|
||||
##### Resolved issues
|
||||
|
||||
* Improved HIP RTC regression test portability when deployed outside the default path.
|
||||
@@ -0,0 +1,193 @@
|
||||
<table class="rocm-docs-table table">
|
||||
<thead>
|
||||
<tr>
|
||||
<th class="head"><p>Component group</p></th>
|
||||
<th class="head"><p>Component name</p></th>
|
||||
<th class="head"><p>Version</p></th>
|
||||
<th class="head"><p>Supported platforms</p></th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td rowspan="18" style="vertical-align: middle;">
|
||||
<p>Math and compute libraries</p>
|
||||
</td>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipblas">hipBLAS</a></td>
|
||||
<td><a href="#hipblas-3-4-0">3.4.0</a></td>
|
||||
<td rowspan="16" style="vertical-align: middle;">Linux/Windows · Instinct/Radeon/Ryzen</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipblaslt">hipBLASLt</a></td>
|
||||
<td><a href="#hipblaslt-1-3-0">1.3.0</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipcub">hipCUB</a></td>
|
||||
<td><a href="#hipcub-4-4-0">4.4.0</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipfft">hipFFT</a></td>
|
||||
<td><a href="#hipfft-1-0-23">1.0.23</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hiprand">hipRAND</a></td>
|
||||
<td><a href="#hiprand-3-3-0">3.3.0</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipsolver">hipSOLVER</a></td>
|
||||
<td><a href="#hipsolver-3-4-0">3.4.0</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipsparse">hipSPARSE</a></td>
|
||||
<td>4.5.0</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/miopen">MIOpen</a></td>
|
||||
<td><a href="#miopen-3-5-1">3.5.1</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocblas">rocBLAS</a></td>
|
||||
<td><a href="#rocblas-5-4-0">5.4.0</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocfft">rocFFT</a></td>
|
||||
<td><a href="#rocfft-1-0-37">1.0.37</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocrand">rocRAND</a></td>
|
||||
<td><a href="#rocrand-4-4-0">4.4.0</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocsolver">rocSOLVER</a></td>
|
||||
<td><a href="#rocsolver-3-34-0">3.34.0</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocsparse">rocSPARSE</a></td>
|
||||
<td><a href="#rocsparse-4-6-0">4.6.0</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocprim">rocPRIM</a></td>
|
||||
<td><a href="#rocprim-4-4-0">4.4.0</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocthrust">rocThrust</a></td>
|
||||
<td><a href="#rocthrust-4-4-0">4.4.0</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocwmma">rocWMMA</a></td>
|
||||
<td><a href="#rocwmma-2-2-1">2.2.1</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/composablekernel">Composable
|
||||
Kernel</a></td>
|
||||
<td><a href="#composable-kernel-1-3-0">1.3.0</a></td>
|
||||
<td>Linux/Windows · Instinct/Radeon</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipsparselt">hipSPARSELt</a></td>
|
||||
<td><a href="#hipsparselt-0-2-8">0.2.8</a></td>
|
||||
<td>Linux/Windows · Instinct (gfx950/gfx942)</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="2" style="vertical-align: middle;">
|
||||
<p>Communication libraries</p>
|
||||
</td>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rccl">RCCL</a></td>
|
||||
<td>2.28.3</td>
|
||||
<td>Linux · Instinct/Radeon/Ryzen</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocshmem">rocSHMEM</a></td>
|
||||
<td><a href="#rocshmem-3-4-0">3.4.0</a></td>
|
||||
<td>Linux · Instinct (gfx950/gfx942/gfx90a) · Radeon (gfx1201/gfx1200/gfx1100/gfx1101/gfx1102)</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="2" style="vertical-align: middle;">
|
||||
<p>Media libraries</p>
|
||||
</td>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocdecode">rocDecode</a></td>
|
||||
<td><a href="#rocdecode-1-8-0">1.8.0</a></td>
|
||||
<td rowspan="2" style="vertical-align: middle;">Linux · Instinct/Radeon · Ryzen (gfx1150/gfx1151/gfx1152)</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocjpeg">rocJPEG</a></td>
|
||||
<td><a href="#rocjpeg-1-5-0">1.5.0</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="5" style="vertical-align: middle;">
|
||||
<p>Runtimes and compilers</p>
|
||||
</td>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/hip">HIP</a></td>
|
||||
<td><a href="#hip-7-13">7.13</a></td>
|
||||
<td rowspan="4" style="vertical-align: middle;">Linux/Windows · Instinct/Radeon/Ryzen</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/HIPIFY/tree/therock-7.13">HIPIFY</a></td>
|
||||
<td>7.13</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/llvm-project/tree/therock-7.13">LLVM</a></td>
|
||||
<td>23.0.0</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/SPIRV-LLVM-Translator/tree/therock-7.13">SPIRV-LLVM-Translator</a></td>
|
||||
<td>23.0.0</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocr-runtime">ROCr Runtime</a></td>
|
||||
<td>1.21.0</td>
|
||||
<td>Linux · Instinct/Radeon/Ryzen</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="6" style="vertical-align: middle;">
|
||||
<p>Profiling and debugging tools</p>
|
||||
</td>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocprofiler-compute">ROCm
|
||||
Compute Profiler (rocprofiler-compute)</a></td>
|
||||
<td><a href="#rocm-compute-profiler-3-6-0">3.6.0</a></td>
|
||||
<td rowspan="2" style="vertical-align: middle;">Linux · Instinct · Ryzen (gfx1150/gfx1151/gfx1152)</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocprofiler-systems">ROCm
|
||||
Systems Profiler (rocprofiler-systems)</a></td>
|
||||
<td><a href="#rocm-systems-profiler-1-6-0">1.6.0</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocprofiler-sdk">ROCprofiler-SDK</a></td>
|
||||
<td><a href="#rocprofiler-sdk-1-3-0">1.3.0</a></td>
|
||||
<td>Linux · Instinct/Radeon · Ryzen (gfx1150/gfx1151/gfx1152)</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocdbgapi">ROCdbgapi</a></td>
|
||||
<td><a href="#rocdbgapi-0-80-0">0.80.0</a></td>
|
||||
<td rowspan="3" style="vertical-align: middle;">Linux · Instinct/Radeon</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/ROCgdb/tree/therock-7.13">ROCm Debugger (ROCgdb)</a></td>
|
||||
<td>16.3</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocr-debug-agent">ROCr Debug
|
||||
Agent</a></td>
|
||||
<td>2.1.0</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="3" style="vertical-align: middle;">
|
||||
<p>Control and monitoring tools</p>
|
||||
</td>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/amdsmi">AMD SMI (BM)</a></td>
|
||||
<td><a href="#amd-smi-bm-26-4-0">26.4.0</a></td>
|
||||
<td>Linux · Instinct/Radeon</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocminfo">rocminfo</a></td>
|
||||
<td><a href="#rocminfo-1-0-0">1.0.0</a></td>
|
||||
<td>Linux · Instinct/Radeon/Ryzen</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rdc">ROCm Data Center Tool
|
||||
(RDC)</a></td>
|
||||
<td><a href="#rocm-data-center-tool-rdc-1-3-0">1.3.0</a></td>
|
||||
<td>Linux · Instinct</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
@@ -0,0 +1,268 @@
|
||||
::::{tab-set}
|
||||
:::{tab-item} Instinct
|
||||
:sync: instinct
|
||||
|
||||
<table class="rocm-docs-table table">
|
||||
<colgroup style="width: 25%;">
|
||||
<thead>
|
||||
<tr>
|
||||
<th class="head">
|
||||
<p>AMD device</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Firmware</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Linux driver</p>
|
||||
</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td>
|
||||
<p>Instinct MI355X</p>
|
||||
</td>
|
||||
<td rowspan="2" style="vertical-align: middle">
|
||||
<p>PLDM bundle 01.26.00.02</p>
|
||||
</td>
|
||||
<td rowspan="10" style="vertical-align: middle">
|
||||
<p>
|
||||
<strong>AMD GPU Driver (amdgpu)</strong><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/31.30.0-preview/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>31.30.0</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/31.20.0-preview/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>31.20.0</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/31.10.0-preview/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>31.10.0</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.3/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>30.30.3</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.2/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>30.30.2</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.1/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>30.30.1</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.0/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>30.30.0</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.20.1/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>30.20.1</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.20.0/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>30.20.0</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.10.2/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>30.10.2</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.10.1/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>30.10.1</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.10/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>30.10.0</a><br>
|
||||
</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>Instinct MI350X</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>Instinct MI350P</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>IFWI 00185129</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>Instinct MI325X</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>PLDM bundle 01.25.04.02</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>Instinct MI300X</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>PLDM bundle 01.26.00.02</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>Instinct MI300A</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>BKC 26.1</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>Instinct MI250X</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>IFWI 75 (or later)</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>Instinct MI250</p>
|
||||
</td>
|
||||
<td rowspan="2">
|
||||
<p>Maintenance update (MU) 5 with IFWI 75 (or later)</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>Instinct MI210</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>Instinct MI100</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>VBIOS D3430401-037</p>
|
||||
</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
:::
|
||||
|
||||
:::{tab-item} Radeon
|
||||
:sync: radeon
|
||||
|
||||
<table class="rocm-docs-table table">
|
||||
<colgroup style="width: 50%;">
|
||||
<thead>
|
||||
<tr>
|
||||
<th class="head">
|
||||
<p>Linux driver</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Windows driver</p>
|
||||
</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td style="vertical-align: middle">
|
||||
<p>
|
||||
<strong>AMD GPU Driver (amdgpu)</strong><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/31.30.0-preview/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>31.30.0</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/31.20.0-preview/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>31.20.0</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/31.10.0-preview/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>31.10.0</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.3/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>30.30.3</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.2/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>30.30.2</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.1/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>30.30.1</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.0/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>30.30.0</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.20.1/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>30.20.1</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.20.0/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>30.20.0</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.10.2/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>30.10.2</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.10.1/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>30.10.1</a><br>
|
||||
<a
|
||||
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.10/documentation/release-notes.html"
|
||||
target="_blank"
|
||||
>30.10.0</a><br>
|
||||
</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>
|
||||
<strong>AMD Software: Adrenalin Edition</strong>
|
||||
<a
|
||||
href="https://www.amd.com/en/resources/support-articles/release-notes/RN-RAD-WIN-26-5-1.html"
|
||||
target="_blank"
|
||||
>26.5.1</a>
|
||||
</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
:::
|
||||
|
||||
:::{tab-item} Ryzen
|
||||
:sync: ryzen
|
||||
|
||||
<table class="rocm-docs-table table">
|
||||
<colgroup style="width: 50%;">
|
||||
<thead>
|
||||
<tr>
|
||||
<th class="head">
|
||||
<p>Linux driver</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Windows driver</p>
|
||||
</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Inbox kernel driver in Ubuntu 26.04 or 24.04.4</p>
|
||||
</td>
|
||||
<td rowspan="30" style="vertical-align: middle">
|
||||
<p>
|
||||
<strong>AMD Software: Adrenalin Edition</strong>
|
||||
<a
|
||||
href="https://www.amd.com/en/resources/support-articles/release-notes/RN-RAD-WIN-26-5-1.html"
|
||||
target="_blank"
|
||||
>26.5.1</a>
|
||||
</p>
|
||||
</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
:::
|
||||
::::
|
||||
@@ -0,0 +1,535 @@
|
||||
::::{tab-set}
|
||||
:::{tab-item} Instinct
|
||||
:sync: instinct
|
||||
|
||||
<table class="rocm-docs-table table">
|
||||
<thead>
|
||||
<colgroup style="width: 33%;">
|
||||
<colgroup style="width: 32%;">
|
||||
<tr>
|
||||
<th class="head">
|
||||
<p>Device series</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Device</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>LLVM target</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Architecture</p>
|
||||
</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td class="stub">
|
||||
<a href="https://www.amd.com/en/products/accelerators/instinct/mi350.html" target="_blank">AMD Instinct MI350
|
||||
Series</a>
|
||||
</td>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi350/mi355x.html" target="_blank">Instinct
|
||||
MI355X</a></p>
|
||||
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi350/mi350x.html" target="_blank">Instinct
|
||||
MI350X</a></p>
|
||||
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi350/mi350p.html" target="_blank">Instinct
|
||||
MI350P</a></p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx950</p>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://www.amd.com/en/technologies/cdna.html#cdna4" target="_blank">CDNA 4</a>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td class="stub">
|
||||
<a href="https://www.amd.com/en/products/accelerators/instinct/mi300.html" target="_blank">AMD Instinct MI300
|
||||
Series</a>
|
||||
</td>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi300/mi325x.html" target="_blank">Instinct
|
||||
MI325X</a></p>
|
||||
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi300/mi300x.html" target="_blank">Instinct
|
||||
MI300X</a></p>
|
||||
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi300/mi300a.html" target="_blank">Instinct
|
||||
MI300A</a></p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx942</p>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://www.amd.com/en/technologies/cdna.html#cdna3" target="_blank">CDNA 3</a>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td class="stub">
|
||||
<a href="https://www.amd.com/en/products/accelerators/instinct/mi200.html" target="_blank">AMD Instinct MI200
|
||||
Series</a>
|
||||
</td>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi200/mi250x.html" target="_blank">Instinct
|
||||
MI250X</a></p>
|
||||
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi200/mi250.html" target="_blank">Instinct
|
||||
MI250</a></p>
|
||||
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi200/mi210.html" target="_blank">Instinct
|
||||
MI210</a></p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx90a</p>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://www.amd.com/en/technologies/cdna.html#cdna2" target="_blank">CDNA 2</a>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td class="stub">
|
||||
<a href="https://www.amd.com/en/products/accelerators/instinct/mi100.html" target="_blank">AMD Instinct MI100
|
||||
Series</a>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://www.amd.com/en/products/accelerators/instinct/mi100.html" target="_blank">Instinct MI100</a>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx908</p>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://www.amd.com/en/technologies/cdna.html#cdna" target="_blank">CDNA</a>
|
||||
</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
:::
|
||||
|
||||
:::{tab-item} Radeon
|
||||
:sync: radeon
|
||||
|
||||
<table class="rocm-docs-table table">
|
||||
<thead>
|
||||
<colgroup style="width: 33%;">
|
||||
<colgroup style="width: 32%;">
|
||||
<tr>
|
||||
<th class="head">
|
||||
<p>Device series</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Device</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>LLVM target</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Architecture</p>
|
||||
</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td>
|
||||
<a href="https://www.amd.com/en/products/graphics/workstations/radeon-ai-pro.html#tabs-95fa144b96-item-b95ec9e1ca-tab"
|
||||
target="_blank">AMD Radeon AI PRO R9000 Series</a>
|
||||
</td>
|
||||
<td>
|
||||
<p><a
|
||||
href="https://www.amd.com/en/products/graphics/workstations/radeon-ai-pro/ai-9000-series/amd-radeon-ai-pro-r9700.html"
|
||||
target="_blank">Radeon AI PRO R9700</a></p>
|
||||
<p><a
|
||||
href="https://www.amd.com/en/products/graphics/workstations/radeon-ai-pro/ai-9000-series/amd-radeon-ai-pro-r9600d.html"
|
||||
target="_blank">Radeon AI PRO R9600D</a></p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1201</p>
|
||||
</td>
|
||||
<td rowspan="3">
|
||||
<a href="https://www.amd.com/en/technologies/rdna.html#tabs-1fabb91c39-item-330ee548f0-tab" target="_blank">RDNA
|
||||
4</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="2" class="stub">
|
||||
<a href="https://www.amd.com/en/products/graphics/desktops/radeon.html#tabs-ff9c5c3863-item-37fb38a236-tab"
|
||||
target="_blank">AMD Radeon RX 9000 Series</p>
|
||||
</td>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9070xt.html"
|
||||
target="_blank">Radeon RX 9070 XT</a></p>
|
||||
<p><a href="https://www.amd.com/en/support/downloads/drivers.html/graphics/radeon-rx/radeon-rx-9000-series/amd-radeon-rx-9070-gre.html"
|
||||
target="_blank">Radeon RX 9070 GRE</a></p>
|
||||
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9070.html"
|
||||
target="_blank">Radeon RX 9070</a></p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1201</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9060xt-lp.html"
|
||||
target="_blank">Radeon RX 9060 XT LP</a></p>
|
||||
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9060xt.html"
|
||||
target="_blank">Radeon RX 9060 XT</a></p>
|
||||
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9060.html"
|
||||
target="_blank">Radeon RX 9060</a></p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1200</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="2" class="stub">
|
||||
<a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro.html#tabs-990fdead92-item-20daa37284-tab"
|
||||
target="_blank">AMD Radeon PRO W7000 Series</a>
|
||||
</td>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w7900-dual-slot.html"
|
||||
target="_blank">Radeon PRO W7900 Dual Slot</a></p>
|
||||
<p><a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w7900.html" target="_blank">Radeon
|
||||
PRO W7900</a></p>
|
||||
<p><a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w7800-48gb.html"
|
||||
target="_blank">Radeon PRO W7800 48GB</a></p>
|
||||
<p><a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w7800.html" target="_blank">Radeon
|
||||
PRO W7800</a></p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1100</p>
|
||||
</td>
|
||||
<td rowspan="6">
|
||||
<a href="https://www.amd.com/en/technologies/rdna.html#tabs-1fabb91c39-item-05915f6044-tab" target="_blank">RDNA
|
||||
3</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w7700.html" target="_blank">Radeon
|
||||
PRO W7700</a></p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1101</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="3" class="stub">
|
||||
<a href="https://www.amd.com/en/products/graphics/desktops/radeon.html#tabs-ff9c5c3863-item-b55a56bf12-tab"
|
||||
target="_blank">AMD Radeon RX 7000 Series</p>
|
||||
</td>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7900xtx.html"
|
||||
target="_blank">Radeon RX 7900 XTX</a></p>
|
||||
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7900xt.html"
|
||||
target="_blank">Radeon RX 7900 XT</a></p>
|
||||
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7900-gre.html"
|
||||
target="_blank">Radeon RX 7900 GRE</a></p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1100</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7800-xt.html"
|
||||
target="_blank">Radeon RX 7800 XT</a></p>
|
||||
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7700-xt.html"
|
||||
target="_blank">Radeon RX 7700 XT</a></p>
|
||||
<p>Radeon RX 7700 XE</p>
|
||||
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7700.html"
|
||||
target="_blank">Radeon RX 7700</a></p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1101</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7600.html"
|
||||
target="_blank">Radeon RX 7600</a></p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1102</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="2">
|
||||
<a href="https://www.amd.com/en/products/accelerators/radeon-pro.html" target="_blank">AMD Radeon PRO V
|
||||
Series</a>
|
||||
</td>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/accelerators/radeon-pro/amd-radeon-pro-v710.html"
|
||||
target="_blank">Radeon PRO V710</a></p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1101</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/accelerators/radeon-pro/amd-radeon-pro-v620.html"
|
||||
target="_blank">Radeon PRO V620</a></p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1030</p>
|
||||
</td>
|
||||
<td rowspan="2">
|
||||
<a href="https://www.amd.com/en/technologies/rdna.html#tabs-1fabb91c39-item-9ed969eddf-tab" target="_blank">RDNA
|
||||
2</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td class="stub">
|
||||
<a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w6800.html" target="_blank">AMD Radeon
|
||||
PRO W6000 Series</a>
|
||||
</td>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w6800.html" target="_blank">Radeon
|
||||
PRO W6800</a></p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1030</p>
|
||||
</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
:::
|
||||
|
||||
:::{tab-item} Ryzen
|
||||
:sync: ryzen
|
||||
|
||||
<table class="rocm-docs-table table">
|
||||
<thead>
|
||||
<colgroup style="width: 26%;">
|
||||
<colgroup style="width: 40%;">
|
||||
<tr>
|
||||
<th class="head">
|
||||
<p>Device series</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Device</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>LLVM target</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Architecture</p>
|
||||
</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td class="stub">
|
||||
<a href="https://www.amd.com/en/products/processors/workstations/mobile.html#tabs-7f0c432fb2-item-5116ab7a74-tab"
|
||||
target="_blank">AMD Ryzen AI Max PRO<br>300 Series</a>
|
||||
</td>
|
||||
<td>
|
||||
<p><a
|
||||
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-max-pro-300-series/amd-ryzen-ai-max-plus-pro-395.html"
|
||||
target="_blank">Ryzen AI Max+ PRO 395</a> (Radeon 8060S)</p>
|
||||
<p><a
|
||||
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-max-pro-300-series/amd-ryzen-ai-max-pro-390.html"
|
||||
target="_blank">Ryzen AI Max PRO 390</a> (Radeon 8050S)</p>
|
||||
<p><a
|
||||
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-max-pro-300-series/amd-ryzen-ai-max-pro-385.html"
|
||||
target="_blank">Ryzen AI Max PRO 385</a> (Radeon 8050S)</p>
|
||||
<p><a
|
||||
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-max-pro-300-series/amd-ryzen-ai-max-pro-380.html"
|
||||
target="_blank">Ryzen AI Max PRO 380</a> (Radeon 8040S)</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1151</p>
|
||||
</td>
|
||||
<td rowspan="10">
|
||||
<p>RDNA 3.5</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td class="stub">
|
||||
<a href="https://www.amd.com/en/products/processors/laptop/ryzen.html#tabs-1181ea0b44-item-6ccfea5f65-tab"
|
||||
target="_blank">AMD Ryzen AI Max<br>300 Series</a>
|
||||
</td>
|
||||
<td>
|
||||
<p><a
|
||||
href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-plus-395.html"
|
||||
target="_blank">Ryzen AI Max+ 395</a> (Radeon 8060S)</p>
|
||||
<p><a
|
||||
href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-plus-392.html"
|
||||
target="_blank">Ryzen AI Max+ 392</a> (Radeon 8060S)</p>
|
||||
<p><a
|
||||
href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-plus-388.html"
|
||||
target="_blank">Ryzen AI Max+ 388</a> (Radeon 8060S)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-390.html"
|
||||
target="_blank">Ryzen AI Max 390</a> (Radeon 8050S)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-385.html"
|
||||
target="_blank">Ryzen AI Max 385</a> (Radeon 8050S)</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1151</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="2" class="stub">
|
||||
<a href="https://www.amd.com/en/products/processors/laptop/ryzen-for-business.html#tabs-0d174caf43-item-87690677fc-tab"
|
||||
target="_blank">AMD Ryzen AI PRO<br>400 Series</a>
|
||||
</td>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-400-series/amd-ryzen-ai-9-hx-pro-475.html"
|
||||
target="_blank">Ryzen AI 9 HX PRO 475</a> (Radeon 890M)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-400-series/amd-ryzen-ai-9-hx-pro-470.html"
|
||||
target="_blank">Ryzen AI 9 HX PRO 470</a> (Radeon 890M)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-400-series/amd-ryzen-ai-9-pro-465.html"
|
||||
target="_blank">Ryzen AI 9 PRO 465</a> (Radeon 880M)</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1150</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-400-series/amd-ryzen-ai-7-pro-450.html"
|
||||
target="_blank">Ryzen AI 7 PRO 450</a> (Radeon 860M)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-400-series/amd-ryzen-ai-5-pro-440.html"
|
||||
target="_blank">Ryzen AI 5 PRO 440</a> (Radeon 840M)</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1152</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="2" class="stub">
|
||||
<a href="https://www.amd.com/en/products/processors/consumer/ryzen-ai.html#tabs-f556098628-item-808b56dca3-tab"
|
||||
target="_blank">AMD Ryzen AI<br>400 Series</a>
|
||||
</td>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-400-series/amd-ryzen-ai-9-hx-475.html"
|
||||
target="_blank">Ryzen AI 9 HX 475</a> (Radeon 890M)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-400-series/amd-ryzen-ai-9-hx-470.html"
|
||||
target="_blank">Ryzen AI 9 HX 470</a> (Radeon 890M)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-400-series/amd-ryzen-ai-9-465.html"
|
||||
target="_blank">Ryzen AI 9 465</a> (Radeon 880M)</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1150</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-400-series/amd-ryzen-ai-7-450.html"
|
||||
target="_blank">Ryzen AI 7 450</a> (Radeon 860M)</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1152</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="2" class="stub">
|
||||
<a href="https://www.amd.com/en/products/processors/workstations/mobile.html#tabs-7f0c432fb2-item-387526c6cc-tab"
|
||||
target="_blank">AMD Ryzen AI PRO<br>300 Series</a>
|
||||
</td>
|
||||
<td>
|
||||
<p><a
|
||||
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-300-series/amd-ryzen-ai-9-hx-pro-375.html"
|
||||
target="_blank">Ryzen AI 9 HX PRO 375</a> (Radeon 890M)</p>
|
||||
<p><a
|
||||
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-300-series/amd-ryzen-ai-9-hx-pro-370.html"
|
||||
target="_blank">Ryzen AI 9 HX PRO 370</a> (Radeon 890M)</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1150</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p><a
|
||||
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-300-series/amd-ryzen-ai-7-pro-350.html"
|
||||
target="_blank">Ryzen AI 7 PRO 350</a> (Radeon 860M)</p>
|
||||
<p><a
|
||||
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-300-series/amd-ryzen-ai-5-pro-340.html"
|
||||
target="_blank">Ryzen AI 5 PRO 340</a> (Radeon 840M)</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1152</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="2" class="stub">
|
||||
<a href="https://www.amd.com/en/products/processors/consumer/ryzen-ai.html#tabs-f556098628-item-54e149d850-tab"
|
||||
target="_blank">AMD Ryzen AI<br>300 Series</a>
|
||||
</td>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-9-hx-375.html"
|
||||
target="_blank">Ryzen AI 9 HX 375</a> (Radeon 890M)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-9-hx-370.html"
|
||||
target="_blank">Ryzen AI 9 HX 370</a> (Radeon 890M)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-9-365.html"
|
||||
target="_blank">Ryzen AI 9 365</a> (Radeon 880M)</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1150</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-7-350.html"
|
||||
target="_blank">Ryzen AI 7 350</a> (Radeon 860M)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-7-345.html"
|
||||
target="_blank">Ryzen AI 7 345</a> (Radeon 840M)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-5-340.html"
|
||||
target="_blank">Ryzen AI 5 340</a> (Radeon 840M)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-5-330.html"
|
||||
target="_blank">Ryzen AI 5 330</a> (Radeon 820M)</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1152</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td class="stub">
|
||||
<a href="https://www.amd.com/en/products/processors/laptop/ryzen-for-business.html#tabs-0d174caf43-item-a8ec88d07e-tab"
|
||||
target="_blank">AMD Ryzen PRO<br>200 Series</a>
|
||||
</td>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/200-series/amd-ryzen-7-pro-250.html"
|
||||
target="_blank">Ryzen 7 PRO 250</a> (Radeon 780M)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/200-series/amd-ryzen-5-pro-230.html"
|
||||
target="_blank">Ryzen 5 PRO 230</a> (Radeon 760M)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/200-series/amd-ryzen-5-pro-220.html"
|
||||
target="_blank">Ryzen 5 PRO 220</a> (Radeon 740M)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/200-series/amd-ryzen-5-pro-215.html"
|
||||
target="_blank">Ryzen 5 PRO 215</a> (Radeon 740M)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/200-series/amd-ryzen-3-pro-210.html"
|
||||
target="_blank">Ryzen 3 PRO 210</a> (Radeon 740M)</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1103</p>
|
||||
</td>
|
||||
<td rowspan="2">
|
||||
<a href="https://www.amd.com/en/technologies/rdna.html#tabs-1fabb91c39-item-05915f6044-tab" target="_blank">RDNA
|
||||
3</a>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td class="stub">
|
||||
<a href="https://www.amd.com/en/products/processors/laptop/ryzen.html#tabs-1181ea0b44-item-895d56feed-tab"
|
||||
target="_blank">AMD Ryzen<br>200 Series</a>
|
||||
</td>
|
||||
<td>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-9-270.html"
|
||||
target="_blank">Ryzen 9 270</a> (Radeon 780M)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-7-260.html"
|
||||
target="_blank">Ryzen 7 260</a> (Radeon 780M)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-7-250.html"
|
||||
target="_blank">Ryzen 7 250</a> (Radeon 780M)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-5-240.html"
|
||||
target="_blank">Ryzen 5 240</a> (Radeon 760M)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-5-230.html"
|
||||
target="_blank">Ryzen 5 230</a> (Radeon 760M)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-5-220.html"
|
||||
target="_blank">Ryzen 5 220</a> (Radeon 740M)</p>
|
||||
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-3-210.html"
|
||||
target="_blank">Ryzen 3 210</a> (Radeon 740M)</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>gfx1103</p>
|
||||
</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
:::
|
||||
::::
|
||||
@@ -0,0 +1,311 @@
|
||||
::::{tab-set}
|
||||
:::{tab-item} Instinct
|
||||
:sync: instinct
|
||||
|
||||
<table class="rocm-docs-table table">
|
||||
<thead>
|
||||
<colgroup style="width: 33%;">
|
||||
<tr>
|
||||
<th class="head">
|
||||
<p>Linux distribution</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Supported versions</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Linux kernel version</p>
|
||||
</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<th rowspan="3" class="stub" style="vertical-align: middle">
|
||||
<p>Ubuntu</p>
|
||||
</th>
|
||||
<td>
|
||||
<p>26.04</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>GA 7.0</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>24.04.4</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>GA 6.8</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>22.04.5</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>GA 5.15</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<th rowspan="2" class="stub" style="vertical-align: middle">
|
||||
<p>Debian</p>
|
||||
</th>
|
||||
<td>
|
||||
<p>13</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>6.12</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>12</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>6.1.0</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<th rowspan="6" class="stub" style="vertical-align: middle">
|
||||
<p>Red Hat Enterprise Linux (RHEL)</p>
|
||||
</th>
|
||||
<td>
|
||||
<p>10.1</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>6.12.0-124</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>10.0</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>6.12.0-55</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>9.7</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>5.14.0-611</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>9.6</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>5.14.0-570</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>9.4</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>5.14.0-427</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>8.10</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>4.18.0-553</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<th rowspan="3" class="stub" style="vertical-align: middle">
|
||||
<p>Oracle Linux</p>
|
||||
</th>
|
||||
<td>
|
||||
<p>10</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>UEK 8.1</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>9</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>UEK 8</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>8</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>UEK 7</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<th class="stub" style="vertical-align: middle">
|
||||
<p>Rocky Linux</p>
|
||||
</th>
|
||||
<td>
|
||||
<p>9</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>5.14.0-570</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<th rowspan="2" class="stub" style="vertical-align: middle">
|
||||
<p>SUSE Linux Enterprise Server (SLES)</p>
|
||||
</th>
|
||||
<td>
|
||||
<p>16.0</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>6.12</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>15.7</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>6.4.0-150700.51</p>
|
||||
</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
:::
|
||||
|
||||
:::{tab-item} Radeon
|
||||
:sync: radeon
|
||||
|
||||
<table class="rocm-docs-table table">
|
||||
<thead>
|
||||
<colgroup style="width: 33%;">
|
||||
<tr>
|
||||
<th class="head">
|
||||
<p>Operating system</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Supported versions</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Linux kernel version</p>
|
||||
</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<th rowspan="3" class="stub" style="vertical-align: middle">
|
||||
<p>Ubuntu</p>
|
||||
</th>
|
||||
<td>
|
||||
<p>26.04</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>GA 7.0</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>24.04.4</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>GA 6.8</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>22.04.5</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>GA 5.15</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<th rowspan="2" class="stub" style="vertical-align: middle">
|
||||
<p>Red Hat Enterprise Linux (RHEL)</p>
|
||||
</th>
|
||||
<td>
|
||||
<p>10.1</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>6.12.0-124</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>9.7</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>5.14.0-611</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<th class="stub" style="vertical-align: middle">
|
||||
<p>Windows</p>
|
||||
</th>
|
||||
<td>
|
||||
<p>11 25H2</p>
|
||||
</td>
|
||||
<td>
|
||||
<p style="text-align: center;"> — </p>
|
||||
</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
:::
|
||||
|
||||
:::{tab-item} Ryzen
|
||||
:sync: ryzen
|
||||
|
||||
<table class="rocm-docs-table table">
|
||||
<thead>
|
||||
<colgroup style="width: 33%;">
|
||||
<tr>
|
||||
<th class="head">
|
||||
<p>Operating system</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Supported versions</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Linux kernel version</p>
|
||||
</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<th rowspan="2" class="stub" style="vertical-align: middle">
|
||||
<p>Ubuntu</p>
|
||||
</th>
|
||||
<td>
|
||||
<p>26.04</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>GA 7.0</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>24.04.4</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>HWE 6.17</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<th class="stub" style="vertical-align: middle">
|
||||
<p>Windows</p>
|
||||
</th>
|
||||
<td>
|
||||
<p>11 25H2</p>
|
||||
</td>
|
||||
<td>
|
||||
<p style="text-align: center;"> — </p>
|
||||
</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
:::
|
||||
::::
|
||||
@@ -0,0 +1,69 @@
|
||||
<table class="rocm-docs-table table">
|
||||
<thead>
|
||||
<tr>
|
||||
<th class="head">
|
||||
<p>Device</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Compute partition mode</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>NPS mode</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Deployment</p>
|
||||
</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td rowspan="3" style="vertical-align: middle;">
|
||||
<p>Instinct MI355X, MI350X</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>CPX</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>NPS 2</p>
|
||||
</td>
|
||||
<td rowspan="5" style="vertical-align: middle;">
|
||||
<p>Bare metal</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>DPX</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>NPS 2</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>QPX</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>NPS 2</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="2" style="vertical-align: middle;">
|
||||
<p>Instinct MI300X</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>CPX</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>NPS 4</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<p>DPX</p>
|
||||
</td>
|
||||
<td>
|
||||
<p>NPS 2</p>
|
||||
</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
@@ -0,0 +1,227 @@
|
||||
<table class="rocm-docs-table table">
|
||||
<colgroup style="width: 14%;">
|
||||
<colgroup style="width: 14%;">
|
||||
<colgroup style="width: 17%;">
|
||||
<colgroup style="width: 17%;">
|
||||
<colgroup style="width: 19%;">
|
||||
<colgroup style="width: 19%;">
|
||||
<thead>
|
||||
<tr>
|
||||
<th class="head">
|
||||
<p>AMD GPU</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Hypervisor</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Virtualization technology</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<a>Virtualization driver</a>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Host OS</p>
|
||||
</th>
|
||||
<th class="head">
|
||||
<p>Guest OS</p>
|
||||
</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td rowspan="5" style="vertical-align: middle">
|
||||
<p>Instinct MI355X</p>
|
||||
</td>
|
||||
<td rowspan="4" style="vertical-align: middle">
|
||||
<p>KVM</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Passthrough</p>
|
||||
</td>
|
||||
<td>
|
||||
<p style="text-align: center">—</p>
|
||||
</td>
|
||||
<td rowspan="4" style="vertical-align: middle">
|
||||
<p>Ubuntu 24.04</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Ubuntu 24.04</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="3" style="vertical-align: middle">
|
||||
<p>SR-IOV</p>
|
||||
</td>
|
||||
<td rowspan="3" style="vertical-align: middle">
|
||||
<a href="https://github.com/amd/MxGPU-Virtualization/releases/tag/9.0.0.K" target="_blank">GIM 9.0.0K
|
||||
</a>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Ubuntu 24.04</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style="vertical-align: middle">
|
||||
<p>RHEL 10.0</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style="vertical-align: middle">
|
||||
<p>RHEL 9.6</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style="vertical-align: middle">
|
||||
<p>ESXi</p>
|
||||
</td>
|
||||
<td>
|
||||
<p style="text-align: center">—</p>
|
||||
</td>
|
||||
<td>
|
||||
<p style="text-align: center">—</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>VMware ESXi 9.1</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Ubuntu 24.04</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="3" style="vertical-align: middle">
|
||||
<p>Instinct MI350X</p>
|
||||
</td>
|
||||
<td rowspan="3" style="vertical-align: middle">
|
||||
<p>KVM</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Passthrough</p>
|
||||
</td>
|
||||
<td>
|
||||
<p style="text-align: center">—</p>
|
||||
</td>
|
||||
<td rowspan="3" style="vertical-align: middle">
|
||||
<p>Ubuntu 24.04</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Ubuntu 24.04</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="2" style="vertical-align: middle">
|
||||
<p>SR-IOV</p>
|
||||
</td>
|
||||
<td rowspan="2" style="vertical-align: middle">
|
||||
<a href="https://github.com/amd/MxGPU-Virtualization/releases/tag/9.0.0.K" target="_blank">GIM 9.0.0K
|
||||
</a>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Ubuntu 24.04</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style="vertical-align: middle">RHEL 9.6</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Instinct MI325X</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>KVM</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>SR-IOV</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<a href="https://github.com/amd/MxGPU-Virtualization/releases/tag/9.0.0.K" target="_blank">GIM 9.0.0K
|
||||
</a>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Ubuntu 22.04</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Ubuntu 22.04</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="3" style="vertical-align: middle">
|
||||
<p>Instinct MI300X</p>
|
||||
</td>
|
||||
<td rowspan="3" style="vertical-align: middle">
|
||||
<p>KVM</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Passthrough</p>
|
||||
</td>
|
||||
<td>
|
||||
<p style="text-align: center">—</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Ubuntu 22.04</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Ubuntu 22.04</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="2" style="vertical-align: middle">
|
||||
<p>SR-IOV</p>
|
||||
</td>
|
||||
<td rowspan="2" style="vertical-align: middle">
|
||||
<a href="https://github.com/amd/MxGPU-Virtualization/releases/tag/9.0.0.K" target="_blank">GIM 9.0.0K
|
||||
</a>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Ubuntu 24.04</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Ubuntu 24.04</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Ubuntu 22.04</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Ubuntu 22.04</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="3" style="vertical-align: middle">
|
||||
<p>Instinct MI210</p>
|
||||
</td>
|
||||
<td rowspan="3" style="vertical-align: middle">
|
||||
<p>KVM</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Passthrough</p>
|
||||
</td>
|
||||
<td>
|
||||
<p style="text-align: center">—</p>
|
||||
</td>
|
||||
<td rowspan="3" style="vertical-align: middle">
|
||||
<p>RHEL 9.4</p>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Ubuntu 22.04</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td rowspan="2" style="vertical-align: middle">
|
||||
<p>SR-IOV</p>
|
||||
</td>
|
||||
<td rowspan="2" style="vertical-align: middle">
|
||||
<a href="https://github.com/amd/MxGPU-Virtualization/releases/tag/9.0.0.K" target="_blank">GIM 9.0.0K
|
||||
</a>
|
||||
</td>
|
||||
<td style="vertical-align: middle">
|
||||
<p>Ubuntu 22.04</p>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style="vertical-align: middle">
|
||||
<p>RHEL 9.4</p>
|
||||
</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
@@ -4,7 +4,7 @@
|
||||
<meta name="keywords" content="license, licensing terms">
|
||||
</head>
|
||||
|
||||
# ROCm license
|
||||
# ROCm licenses
|
||||
|
||||
```{include} ../../LICENSE
|
||||
```
|
||||
|
||||
@@ -0,0 +1,608 @@
|
||||
# ROCm Core SDK {{ ROCM_VERSION }} release notes
|
||||
|
||||
ROCm Core SDK {{ ROCM_VERSION }} continues the technology preview release stream
|
||||
that began with ROCm 7.9.0, advancing the transition to the new
|
||||
[TheRock](https://github.com/rocm/therock) build and release system. To learn
|
||||
more, see the [transition guide](/about/transition-guide-TheRock).
|
||||
|
||||
(preview-stream-note)=
|
||||
:::{important}
|
||||
ROCm {{ ROCM_VERSION }} follows the
|
||||
<a href="https://rocm.docs.amd.com/en/7.9.0-preview/about/release-notes.html#preview-stream-note"
|
||||
target="_blank">versioning discontinuity that began with the 7.9.0 preview release</a>
|
||||
and remains separate from the 7.0 to 7.2 production releases. For the latest
|
||||
production stream release, see the
|
||||
<a href="https://rocm.docs.amd.com/en/latest/">ROCm documentation</a>.
|
||||
|
||||
Maintaining parallel release streams -- preview and production -- gives
|
||||
users ample time to evaluate and adopt the new build system and dependency
|
||||
changes. The technology preview stream is planned to continue through
|
||||
mid-2026, after which it will replace the current production stream.
|
||||
|
||||
For previous preview releases, see the
|
||||
<a target="_blank" href="https://rocm.docs.amd.com/en/7.12.0-preview/release/versions.html">release history</a>.
|
||||
:::
|
||||
|
||||
## Release highlights
|
||||
|
||||
ROCm Core SDK {{ ROCM_VERSION }} with TheRock builds upon the [7.12.0 preview
|
||||
release](https://rocm.docs.amd.com/en/7.12.0-preview/about/release-notes.html).
|
||||
|
||||
This release expands support for AI inference, distributed workloads, and
|
||||
profiling workflows across AMD Instinct™, Radeon™, and Ryzen™ AI platforms.
|
||||
ROCm 7.13.0 adds inference-ready vLLM containers, expands GPU virtualization
|
||||
and partitioning support, introduces new profiling and tracing capabilities,
|
||||
and improves AI kernel, sparse math, and communication libraries.
|
||||
|
||||
### Platform and hardware support
|
||||
|
||||
This release expands GPU, operating system, virtualization, and partitioning support.
|
||||
|
||||
#### Expanded AMD GPU support
|
||||
|
||||
ROCm 7.13.0 adds support for the following AMD GPUs and APUs:
|
||||
|
||||
* AMD Instinct MI350P (gfx950)
|
||||
* AMD Radeon PRO W6800 (gfx1030)
|
||||
* AMD Radeon PRO V620 (gfx1030)
|
||||
* AMD Ryzen AI 7 PRO 360 (gfx1152)
|
||||
* AMD Ryzen AI 7 PRO 350 (gfx1152)
|
||||
* AMD Ryzen AI 5 PRO 340 (gfx1152)
|
||||
* AMD Ryzen AI 7 350 (gfx1152)
|
||||
* AMD Ryzen AI 7 345 (gfx1152)
|
||||
* AMD Ryzen AI 5 340 (gfx1152)
|
||||
* AMD Ryzen AI 5 330 (gfx1152)
|
||||
|
||||
For the complete list of supported AMD hardware, see [AMD hardware support](#amd-hardware-support).
|
||||
|
||||
#### Expanded Ubuntu support
|
||||
|
||||
ROCm 7.13.0 adds support for Ubuntu 26.04 on Instinct, Radeon, and Ryzen
|
||||
devices.
|
||||
|
||||
24.04.4 is now the validated Ubuntu 24 release instead of Ubuntu 24.04.3.
|
||||
|
||||
For the full list of supported Linux distributions, see [Operating system support](#operating-system-support).
|
||||
|
||||
#### Expanded GPU virtualization support for Instinct GPUs
|
||||
|
||||
ROCm 7.13.0 adds support for the following virtualization configurations on AMD Instinct GPUs.
|
||||
|
||||
* On MI355X: VMware ESXi 9.1 with Ubuntu 24.04 guest OS.
|
||||
|
||||
* On MI300X: KVM SR-IOV with Ubuntu 24.04 host OS and Ubuntu 24.04 guest OS.
|
||||
|
||||
* On MI210:
|
||||
|
||||
* KVM passthrough with RHEL 9.4 host OS and Ubuntu 22.04 guest OS.
|
||||
|
||||
* KVM SR-IOV with RHEL 9.4 host OS and Ubuntu 22.04 guest OS.
|
||||
|
||||
* KVM SR-IOV with RHEL 9.4 host OS and RHEL 9.4 guest OS.
|
||||
|
||||
Supported SR-IOV configurations require the [GIM Driver
|
||||
9.0.0K](https://github.com/amd/MxGPU-Virtualization/releases/tag/9.0.0.K). For
|
||||
details, see [GPU virtualization support](#gpu-virtualization-support).
|
||||
|
||||
#### Expanded Instinct GPU partitioning support
|
||||
|
||||
ROCm 7.13.0 enables the QPX compute + NPS 2 memory partition combination in
|
||||
bare metal deployments.
|
||||
|
||||
For details, see [GPU partitioning support](#gpu-partitioning-support).
|
||||
|
||||
### AI inference and frameworks
|
||||
|
||||
This release adds inference-ready container images and improves multi-node communication for distributed workloads.
|
||||
|
||||
#### vLLM 0.19.1 Docker images and pip packages
|
||||
|
||||
With ROCm 7.13.0, Docker images for running vLLM inference workloads are
|
||||
available. Images include vLLM 0.19.1, PyTorch 2.10, and Python 3.13 on Ubuntu 24.04.
|
||||
|
||||
Architecture-specific images are available for:
|
||||
|
||||
* AMD Instinct GPUs: gfx942 (MI325X, MI300X, MI300A) and gfx950 (MI355X, MI350X, MI350P)
|
||||
* AMD Radeon GPUs: gfx1100, gfx1101, gfx1102, gfx1200, gfx1201
|
||||
* AMD Ryzen AI APUs: gfx1150, gfx1151, gfx1152
|
||||
|
||||
See [](../ai-inference/vllm) to get started.
|
||||
|
||||
#### RCCL multi-node optimization for AMD Ryzen AI Max 300 series
|
||||
|
||||
RCCL improves multi-node clustering performance on systems with AMD Ryzen AI
|
||||
Max 300 series connected over Ethernet. Building on the initial
|
||||
multi-node enablement in ROCm 7.12.0, this release optimizes collective
|
||||
communication for distributed AI inference workloads using tensor parallelism
|
||||
(TP) and expert parallelism (EP) across up to 4 Ethernet-connected nodes.
|
||||
|
||||
#### RCCL GDA-based alltoall via rocSHMEM integration (experimental)
|
||||
|
||||
RCCL adds experimental support for GPU Direct Async (GDA)-based alltoall and
|
||||
alltoallv collective operations through rocSHMEM integration. When enabled,
|
||||
RCCL invokes rocSHMEM operations that use GDA to reduce latency for small
|
||||
message alltoall patterns.
|
||||
|
||||
This feature requires building RCCL with the `--rocshmem` flag and setting
|
||||
`RCCL_ROCSHMEM_ENABLE=1` at runtime. GDA support currently requires Broadcom
|
||||
NICs with GDA capability.
|
||||
|
||||
### Developer tools and profiling
|
||||
|
||||
This release adds new profiling capabilities, introduces the open-source ROCprof Trace Decoder, and extends HIP programming APIs.
|
||||
|
||||
#### ROCprof Trace Decoder open source release
|
||||
|
||||
ROCprof Trace Decoder, previously delivered as a closed-source
|
||||
component within ROCprofiler-SDK, is now available as the open-source
|
||||
rocprof-trace-decoder library. The decoder converts raw SQTT data from AMD GPUs
|
||||
into structured execution traces for performance analysis and debugging. It
|
||||
supports a wide range of AMD GPUs spanning Instinct, Radeon, and Ryzen
|
||||
architectures, with unit and integration tests across all supported hardware.
|
||||
See [AMD hardware support](#amd-hardware-support) for the complete list.
|
||||
|
||||
<!-- For more information, see [ROCprof Trace Decoder and thread trace APIs](https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/docs-7.13.0/api-reference/thread_trace.html) and [Using thread trace](https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/docs-7.13.0/how-to/using-thread-trace.html) in the ROCprofiler-SDK documentation. -->
|
||||
|
||||
#### HIP cooperative groups reduce operations
|
||||
|
||||
HIP adds `cooperative_groups::reduce()` for performing reduction operations
|
||||
across `thread_block_tile` and `coalesced_threads` groups. The implementation
|
||||
is based on `__reduce_*_sync` operations, and the
|
||||
`HIP_ENABLE_EXTRA_WARP_SYNC_TYPES` macro might be required to enable some
|
||||
optimizations.
|
||||
|
||||
Additionally, `__reduce_and_sync()`, `__reduce_or_sync()`, and
|
||||
`__reduce_xor_sync()` now provide consistent behavior for all mask values. All
|
||||
masks now emit bitwise instructions, aligning behavior with NVIDIA CUDA. This
|
||||
is a change from previous versions, where some masks were translated to bitwise
|
||||
operations, and others were not.
|
||||
|
||||
#### ROCm Compute Profiler feature highlights
|
||||
|
||||
The following are notable enhancements to the ROCm Compute Profiler
|
||||
(rocprofiler-compute).
|
||||
|
||||
* **RDNA 3.5 support:** ROCm Compute Profiler now supports GPU performance
|
||||
profiling and analysis on AMD Ryzen AI Max 300 series processors.
|
||||
<!-- An [RDNA 3 -->
|
||||
<!-- section](https://rocm.docs.amd.com/projects/rocprofiler-compute/en/docs-7.13.0/conceptual/rdna/rdna-performance-model.html) -->
|
||||
<!-- has been added to the performance model documentation explaining the supported -->
|
||||
<!-- performance metrics for AMD Ryzen AI Max 300 series processors. A new memory -->
|
||||
<!-- chart visualization accommodates the architectural differences between -->
|
||||
<!-- RDNA 3.5 and CDNA GPUs. Roofline is not yet supported for AMD Ryzen AI -->
|
||||
<!-- Max 300 series processors. -->
|
||||
|
||||
* **Removed dependency requirements for profiling:** Building ROCm Compute
|
||||
Profiler and using profile mode no longer requires installing Python
|
||||
dependencies from the `requirements.txt` file. Analysis mode still requires
|
||||
Python dependencies.
|
||||
|
||||
This change moves several operations from profile mode to analysis mode,
|
||||
including roofline HTML generation, roofline-related options
|
||||
(`--sort`, `--mem-level`, `--roofline-data-type`), and creation of the
|
||||
combined `pmc_perf.csv` file. Profile mode now only runs the roofline
|
||||
empirical benchmark, creates a `roofline.csv` file, and creates per-replay
|
||||
CSV files without merging them.
|
||||
|
||||
#### ROCm Systems Profiler feature highlights
|
||||
|
||||
The following are notable enhancements to the ROCm Systems Profiler
|
||||
(rocprofiler-systems).
|
||||
|
||||
* **Pause and resume profiling:** ROCm Systems Profiler now supports pausing
|
||||
and resuming profiling at runtime through the `roctxProfilerPause` and
|
||||
`roctxProfilerResume` APIs. This allows you to capture profiling data only
|
||||
during specific execution phases, reducing overhead and minimizing output size
|
||||
for long-running workloads.
|
||||
<!-- For more information, see [Configuring runtime -->
|
||||
<!-- options](https://rocm.docs.amd.com/projects/rocprofiler-systems/en/docs-7.13.0/how-to/configuring-runtime-options.html) -->
|
||||
<!-- in the ROCm Systems Profiler documentation. -->
|
||||
|
||||
* **Selective region tracing:** You can now restrict tracing to defined regions
|
||||
of interest using the `ROCPROFSYS_SELECTED_REGIONS` environment variable,
|
||||
reducing noise and limiting data collection to relevant workload segments.
|
||||
<!-- For more -->
|
||||
<!-- information, see -->
|
||||
<!-- [ROCPROFSYS_SELECTED_REGIONS](https://rocm.docs.amd.com/projects/rocprofiler-systems/en/docs-7.13.0/how-to/configuring-runtime-options.html#rocprofsys-selected-regions) -->
|
||||
<!-- in the ROCm Systems Profiler documentation. -->
|
||||
|
||||
* **KFD event tracing:** Kernel Fusion Driver (KFD) event tracing is now
|
||||
available for GPU memory management analysis, including page faults, page
|
||||
migrations, queue evictions, GPU unmap events, and dropped events. Requires
|
||||
an XNACK-capable GPU and ROCprofiler-SDK 1.2.1 or later.
|
||||
<!-- For more -->
|
||||
<!-- information, see [Configuring runtime -->
|
||||
<!-- options](https://rocm.docs.amd.com/projects/rocprofiler-systems/en/docs-7.13.0/how-to/configuring-runtime-options.html#exploring-gpu-metrics) -->
|
||||
<!-- in the ROCm Systems Profiler documentation. -->
|
||||
|
||||
* **MPI file-output filtering:** You can now filter profiler output files based
|
||||
on MPI rank using the `--rank-filter-output` CLI option or the
|
||||
`ROCPROFSYS_RANK_FILTER_OUTPUT` configuration setting, suppressing output
|
||||
from all other ranks. An optional `--rank-filter-id` option
|
||||
(`ROCPROFSYS_RANK_FILTER_ID`) allows specifying a custom environment variable
|
||||
for rank identification.
|
||||
<!-- For more information, see [Selective rank -->
|
||||
<!-- profiling](https://rocm.docs.amd.com/projects/rocprofiler-systems/en/docs-7.13.0/how-to/communication-runtime-profiling.html#selective-rank-profiling) -->
|
||||
<!-- in the ROCm Systems Profiler documentation. -->
|
||||
|
||||
* **JSON-based profiling presets and domain flags:** You can now configure
|
||||
common profiling workflows using JSON-based presets and a single
|
||||
`--preset=<name>` flag instead of manually setting multiple `ROCPROFSYS_*`
|
||||
environment variables. Eleven built-in presets cover common profiling scenarios, including GPU
|
||||
tracing, HPC workloads, and API-level analysis. Composable domain flags
|
||||
(`--gpu`, `--rocm`, `--cpu`, `--parallel`) and a topic-based
|
||||
`--help=<topic>` system further simplify configuration and discoverability.
|
||||
<!-- For more information, see [Using preset profiles and domain -->
|
||||
<!-- flags](https://rocm.docs.amd.com/projects/rocprofiler-systems/en/docs-7.13.0/how-to/using-preset-profiles.html) -->
|
||||
<!-- in the ROCm Systems Profiler documentation. -->
|
||||
|
||||
#### AMD SMI feature highlights
|
||||
|
||||
* **APU metrics and memory tuning**: New APU telemetry provides per-core
|
||||
temperature, power, clock, voltage, current, and throttle monitoring, with
|
||||
additional support for IPU activity and DRAM bandwidth metrics. New VRAM
|
||||
carveout and GTT tuning controls enable configurable memory allocation on
|
||||
supported APU platforms.
|
||||
|
||||
* **Per-component GPU temperature and clock monitoring**: GPU metrics table
|
||||
version 1.9 adds HBM stack temperatures, per-die temperature monitoring, and
|
||||
per-die memory and SOC clock reporting for data center deployments.
|
||||
|
||||
* **CPU power APIs report in milliwatts (breaking change)**: CPU power APIs now
|
||||
return values in milliwatts (mW) instead of watts. Python bindings now return
|
||||
numeric integer values instead of formatted strings. Existing applications
|
||||
that parse previous string-based outputs must be updated.
|
||||
|
||||
For more information, see the AMD SMI section in the [ROCm component changelogs](#rocm-component-changelogs).
|
||||
|
||||
### Libraries
|
||||
|
||||
This release adds new routines, data type support, and performance improvements across ROCm math and AI libraries.
|
||||
|
||||
#### Composable Kernel adds quantization and attention kernel capabilities
|
||||
|
||||
Composable Kernel adds several capabilities for AI and large language model
|
||||
workloads:
|
||||
|
||||
* **Microscaling (MX) FP8/FP4 support:** Mixed data type support for MX FP8 and
|
||||
FP4 in GEMM and Flash Multi-Head Attention (FMHA) forward kernels on AMD
|
||||
Instinct MI350 Series GPUs.
|
||||
|
||||
* **FP8 quantization for FMHA:** FMHA forward kernels now support multiple FP8
|
||||
quantization modes, including dynamic tensor-wise quantization, block scale
|
||||
quantization, per-tensor quantization, and FP8 KV cache support for batch
|
||||
prefill.
|
||||
|
||||
* **StreamingLLM and long-context inference:** Sink token support for FMHA
|
||||
forward enables StreamingLLM-style long-context inference.
|
||||
|
||||
* **Batch prefill enhancements:** FMHA batch prefill kernels now support
|
||||
multiple KV cache layouts, flexible page sizes, and configurable lookup table
|
||||
configurations.
|
||||
|
||||
* **RDNA 3 FMHA support:** Flash Attention kernels are now available on RDNA 3
|
||||
architectures.
|
||||
|
||||
* **SageAttention v2 forward kernel:** Multi-granularity quantization for Q, K,
|
||||
and V tensors with FP8, INT8, and INT4 data types and per-tensor, per-block,
|
||||
per-warp, and per-thread scale granularities on AMD Instinct MI300 Series and
|
||||
MI350 Series GPUs.
|
||||
|
||||
#### General Batched GEMM support in hipBLASLt
|
||||
|
||||
hipBLASLt adds native support for General Batched GEMM, where all matrices in
|
||||
a batch share the same problem dimensions but can have independent leading
|
||||
dimensions and strides. This replaces the previous implementation through the
|
||||
`hipblaslt_ext` Grouped GEMM APIs, which had known limitations.
|
||||
|
||||
The new implementation includes support for Global Split-U (GSU) to improve
|
||||
performance at large problem sizes. General Batched GEMM is important for
|
||||
inference workloads that dispatch batches of same-shape GEMM operations.
|
||||
|
||||
<!-- For more information, see the [hipBLASLt -->
|
||||
<!-- documentation](https://rocm.docs.amd.com/projects/hipBLASLt/en/docs-7.13.0/index.html). -->
|
||||
|
||||
#### rocSOLVER adds new solver routines and matrix analysis functions
|
||||
|
||||
rocSOLVER adds the following new routines, all with 64-bit index support:
|
||||
|
||||
* **GETRS_NPVT:** Solution of linear systems using LU factorization without
|
||||
pivoting. Batched and strided-batched variants are available.
|
||||
|
||||
* **SYTRS:** Solution of linear systems for symmetric matrices. Batched and
|
||||
strided-batched variants are available.
|
||||
|
||||
Additionally, POTF2 and downstream POTRF Cholesky factorization performance
|
||||
have been improved.
|
||||
<!-- For more information, see the [rocSOLVER -->
|
||||
<!-- documentation](https://rocm.docs.amd.com/projects/rocSOLVER/en/docs-7.13.0/index.html). -->
|
||||
|
||||
#### rocSPARSE adds sparse factorization routines
|
||||
|
||||
rocSPARSE adds new generic API routines for sparse incomplete factorization and
|
||||
triangular solve:
|
||||
|
||||
* `rocsparse_spic0` and `rocsparse_spilu0`: Generic incomplete Cholesky (IC0)
|
||||
and incomplete LU (ILU0) factorization routines with strided-batched
|
||||
computation support.
|
||||
|
||||
* `rocsparse_sptrsv`: Extended with strided-batched computation support and
|
||||
singularity detection through the new `rocsparse_singularity` enumeration.
|
||||
|
||||
Performance of tridiagonal solvers `rocsparse_Xgtsv_no_pivot` and
|
||||
`rocsparse_Xgtsv_no_pivot_strided_batch` has been improved.
|
||||
<!-- For more -->
|
||||
<!-- information, see the [rocSPARSE -->
|
||||
<!-- documentation](https://rocm.docs.amd.com/projects/rocSPARSE/en/docs-7.13.0/index.html). -->
|
||||
|
||||
#### Added rocDecode and rocJPEG libraries to the ROCm Core SDK
|
||||
|
||||
rocDecode provides hardware-accelerated video decoding for H.264, H.265/HEVC,
|
||||
AV1, and VP9 codecs, while rocJPEG provides hardware-accelerated JPEG decoding
|
||||
on AMD GPUs. Together, they enable
|
||||
efficient GPU-based media processing pipelines for data-intensive workloads
|
||||
such as AI training.
|
||||
|
||||
Both libraries are supported on Linux on AMD Instinct, Radeon, and Ryzen AI. See
|
||||
the projects in [ROCm/rocm-systems](https://github.com/ROCm/rocm-systems) for
|
||||
more information.
|
||||
|
||||
#### Added ROCm Data Center Tool to the ROCm Core SDK
|
||||
|
||||
ROCm Data Center Tool (RDC) provides telemetry collection, health monitoring,
|
||||
and job-level GPU statistics for data center deployments with AMD Instinct
|
||||
accelerators. RDC enables system administrators and cluster managers to monitor
|
||||
GPU health, collect telemetry data, and track per-job GPU usage across
|
||||
multi-node environments.
|
||||
|
||||
RDC is supported on Linux with AMD Instinct GPUs.
|
||||
<!-- See the -->
|
||||
<!-- [RDC documentation](https://rocm.docs.amd.com/projects/rdc/en/docs-7.13.0/index.html) -->
|
||||
<!-- for more information. -->
|
||||
|
||||
(release-supported-hw)=
|
||||
|
||||
## AMD hardware support
|
||||
|
||||
The following table lists supported AMD Instinct GPUs, Radeon GPUs, and Ryzen
|
||||
APUs. Each supported device is listed with its corresponding GPU
|
||||
microarchitecture and LLVM target.
|
||||
|
||||
:::{note}
|
||||
|
||||
If your GPU is not listed, it might be community-enabled through TheRock
|
||||
nightly builds. For more information, see [TheRock supported
|
||||
GPUs](https://github.com/ROCm/TheRock/blob/main/SUPPORTED_GPUS.md). For
|
||||
installation guidance, see [TheRock
|
||||
releases](https://github.com/ROCm/TheRock/blob/main/RELEASES.md).
|
||||
:::
|
||||
|
||||
```{include} ./include/hardware-support-table.md
|
||||
:parser: myst
|
||||
```
|
||||
|
||||
(release-supported-os)=
|
||||
|
||||
## Operating system support
|
||||
|
||||
ROCm supports the following Linux distribution and Microsoft Windows versions.
|
||||
If you're running ROCm on Linux, ensure your system is using a supported kernel
|
||||
version.
|
||||
|
||||
:::{important}
|
||||
The following table is a general overview of supported OSes. Actual support
|
||||
might vary by AMD GPU or APU. Use the {doc}`Compatibility matrix
|
||||
</compatibility/compatibility-matrix>` to verify support for your specific
|
||||
setup before installation.
|
||||
:::
|
||||
|
||||
```{include} ./include/os-support-table.md
|
||||
:parser: myst
|
||||
```
|
||||
|
||||
## Installation updates
|
||||
|
||||
ROCm 7.13.0 introduces several improvements to the Runfile Installer:
|
||||
|
||||
* Performance improvements for installing and uninstalling gfx architectures.
|
||||
* ROCm component tests are now included.
|
||||
* Support for prerequisite OEM kernel installation as part of the dependency install on Ryzen systems. You no longer need to install it manually.
|
||||
* Auto-detection of the GPU when using the GUI or when the `gfx=` argument is not provided on the command line. If the installer cannot detect the GPU, you must specify the gfx architecture using the GUI or the `gfx=` argument.
|
||||
|
||||
(release-supported-fw)=
|
||||
|
||||
## Kernel driver and firmware bundle support
|
||||
|
||||
ROCm requires a coordinated stack of compatible firmware, driver, and user
|
||||
space components. Maintaining version alignment between these layers ensures
|
||||
correct GPU operation and performance, especially for AMD data center products.
|
||||
While AMD publishes the AMD GPU driver and ROCm user space components, your
|
||||
server OEM (original equipment manufacturer) or infrastructure provider
|
||||
distributes the firmware packages. AMD supplies those firmware images (PLDM
|
||||
bundles), which the OEM integrates and distributes.
|
||||
|
||||
```{include} ./include/driver-firmware-support-table.md
|
||||
:parser: myst
|
||||
```
|
||||
|
||||
(release-virtualization-support)=
|
||||
|
||||
## GPU virtualization support
|
||||
|
||||
AMD Instinct data center GPUs support virtualization in the following
|
||||
configurations. Supported SR-IOV configurations require the AMD GPU
|
||||
Virtualization Driver (GIM) 9.0.0K -- see the [AMD Instinct Virtualization
|
||||
Driver
|
||||
documentation](https://instinct.docs.amd.com/projects/virt-drv/en/mainline-9.0.0.k/)
|
||||
for more information.
|
||||
|
||||
```{include} ./include/virtualization-support-table.html
|
||||
:parser: myst
|
||||
```
|
||||
|
||||
(release-gpu-partitioning-support)=
|
||||
|
||||
## GPU partitioning support
|
||||
|
||||
The following compute partition and NUMA-per-socket (NPS) configurations are
|
||||
available on AMD Instinct GPUs in bare metal deployments.
|
||||
|
||||
```{include} ./include/partitioning-support-table.html
|
||||
:parser: myst
|
||||
```
|
||||
|
||||
See the [AMD GPU partitioning](https://instinct.docs.amd.com/projects/amdgpu-docs/en/latest/gpu-partitioning/index.html)
|
||||
topic in the AMD GPU Driver documentation to learn more.
|
||||
|
||||
(release-ai-ecosystem)=
|
||||
|
||||
## AI ecosystem support
|
||||
|
||||
ROCm 7.13.0 provides optimized support for popular deep learning frameworks and
|
||||
AI inference engines. The following table lists supported frameworks and
|
||||
libraries, their compatible operating systems, and validated versions.
|
||||
|
||||
```{include} ./include/ai-ecosystem-support-table.html
|
||||
:parser: myst
|
||||
```
|
||||
|
||||
(release-components)=
|
||||
|
||||
## ROCm Core SDK components
|
||||
|
||||
The following table lists core tools and libraries included in the ROCm 7.13.0
|
||||
release.
|
||||
|
||||
:::{important}
|
||||
The following table is a general overview of ROCm Core SDK components. Actual
|
||||
support for these libraries and tools can vary by GPU and OS. Use the
|
||||
{doc}`Compatibility matrix </compatibility/compatibility-matrix>` to verify
|
||||
support for your specific setup.
|
||||
:::
|
||||
|
||||
```{include} ./include/core-sdk-components-table.html
|
||||
:parser: myst
|
||||
```
|
||||
|
||||
### ROCm component changelogs
|
||||
|
||||
The following sections describe key changes to ROCm Core SDK components.
|
||||
|
||||
```{include} ./include/core-sdk-components-aggregated-changelog.md
|
||||
:parser: myst
|
||||
```
|
||||
|
||||
## ROCm known issues
|
||||
|
||||
ROCm known issues are noted on {fab}`github` [GitHub](https://github.com/ROCm/ROCm/labels/Verified%20Issue). These issues will be fixed in a future ROCm release. For known issues related to individual components, review the [ROCm component changelogs](#rocm-component-changelogs).
|
||||
|
||||
### ROCm Compute Profiler might fail when profiling bash script or command
|
||||
|
||||
Running a bash script or command as a target for ROCm Compute Profiler might fail because bash overwrites the required environment variables. As a workaround, use `--no-native-tool` option in the profile mode. Note that this will disable iteration multiplexing.
|
||||
|
||||
### hipFFT and rocFFT callback examples fail to build on Windows
|
||||
|
||||
The hipFFT and rocFFT callback examples in [rocm-examples](https://github.com/rocm/rocm-examples) fail to build on a Windows operating system due to a linker error. CMake configuration and HIP object compilation will complete successfully, but the final link step fails with `clang: error: invalid linker name in argument '-fuse-ld=lld-link'` This issue affects all Windows configurations using Relocatable Device Code (RDC) mode. Linux builds are not affected. As a workaround, skip the hipFFT and rocFFT callback examples on Windows, and refer to the Linux builds or [callback](https://github.com/ROCm/rocm-examples/tree/amd-staging/Libraries/rocFFT/callback/) functionality documentation.
|
||||
|
||||
### QMCPACK might become unresponsive during DMC simulation on AMD Instinct MI300A GPUs
|
||||
|
||||
QMCPACK might become unresponsive when running Diffusion Monte Carlo (DMC) simulations with certain inputs on AMD Instinct MI300A GPUs. The application stops making progress after initialization and must be terminated manually.
|
||||
|
||||
### Resource-intensive workloads might result in GPU memory faults
|
||||
|
||||
Applications that pass large, complex data structures between device functions using scratch memory, and particularly rely on compiler optimization to minimize the number of copy operations, might encounter GPU memory access faults and become unresponsive.
|
||||
|
||||
### Increased binary size for multi-target GPU builds
|
||||
|
||||
Applications targeting multiple AMD GPU architectures might observe significantly larger binary sizes. Multi-target builds can produce binaries up to 54 percent larger. Single-target builds add approximately 8 MB of additional size per GPU target. As a workaround, reduce the number of GPU targets in multi-target builds, or strip the resource-usage symbols from release binaries.
|
||||
|
||||
### HIP cooperative groups might fail when compiled using the SPIR-V path
|
||||
|
||||
HIP applications that use cooperative groups might fail at kernel launch when compiled with `--offload-arch=amdgcnspirv`. The application fails at runtime with `LLVM ERROR: Cannot select: intrinsic %llvm.amdgcn.s.wait.asynccnt` error message. This
|
||||
affects all GPU architectures when using the SPIR-V compilation path. As a workaround, compile using a direct GPU architecture target (for example, `--offload-arch=gfx942`) instead of `--offload-arch=amdgcnspirv`.
|
||||
|
||||
### Illegal memory address error when using placement new with device function returns
|
||||
|
||||
HIP kernels that use the placement new operators to construct objects in the `hipMalloc` device memory might crash with `hipErrorIllegalAddress` error message when you pass a `__device__` function return value as the constructor argument. This only affects non-trivially-copyable types (for example, types with user-defined or deleted copy/move constructors). Trivially-copyable types are not affected. As a workaround, assign the device function return value to a local variable before passing it to placement new.
|
||||
|
||||
### LLVM-based compilers might fail when compiling half-precision vector operations
|
||||
|
||||
LLVM-based compilers might fail, returning `Failed to find subregs!` error message in `SIInstrInfo::copyPhysReg`, when compiling half-precision vector operations with optimization enabled. The issue was observed at optimization levels `-O1` to `-O3`.
|
||||
|
||||
### hipBLAS test suites failure on Windows
|
||||
|
||||
When using hipBLAS on Windows, the test suites might return non-zero exit codes, even when all mathematical correctness tests pass. This issue can affect CI/CD pipeline validation and block automated testing workflows on Windows systems, because the test framework might fail to detect successful test completion.
|
||||
|
||||
### ROCm Systems Profiler overwrites ROCPD output after process re-attachment
|
||||
|
||||
When you use `rocprof-sys-attach` to re-attach to a previously profiled process, the `ROCPD` output database files (.db) are written to the initial session's output directory instead of a new timestamped directory. This makes it difficult to distinguish profiling data between sessions. Perfetto trace files are not affected. As a workaround, back up your output directory before re-attaching to a previously profiled process.
|
||||
|
||||
### Missing dependencies when installing ROCm Core SDK
|
||||
|
||||
Installing the ROCm Core SDK using `amdrocm-core-sdk` or `amdrocm-core-dev/devel` might succeed, but some dependencies from the dev/devel meta packages might not be installed. As a workaround, install the dev packages manually:
|
||||
|
||||
```bash
|
||||
sudo apt install amdrocm-*
|
||||
```
|
||||
|
||||
### Issues related to AddressSanitizer
|
||||
|
||||
Multiple issues associated with AddressSanitizer (ASAN) `-fsanitize=address` being enabled have been observed including:
|
||||
|
||||
#### ASAN reports false errors for GPU kernels using shared memory
|
||||
|
||||
When you compile GPU kernels with ASAN enabled, kernels that use `__shared__` memory might produce false heap-buffer-overflow errors or GPU memory faults. As a workaround, disable ASAN by removing `-fsanitize=address` setting for affected kernels.
|
||||
|
||||
#### GPU kernels fail to launch in ASAN builds with large thread counts
|
||||
|
||||
When you build GPU libraries with ASAN enabled, kernels configured with large thread counts might fail to launch with `HSA_STATUS_ERROR_INVALID_ISA` error. As a workaround, reduce the thread block sizes to 256 threads or fewer for ASAN builds. The issue is currently under investigation.
|
||||
|
||||
#### ASAN breaks multi-architecture HIP binary builds
|
||||
|
||||
HIP applications built with ASAN enabled, targeting multiple GPU architectures, might fail to launch with `RuntimeError: .hipFatBinSegment size N is not a multiple of wrapper size (24)` and `RuntimeError: Unexpected magic 0x00000000 at wrapper i` error messages. Single-architecture builds are not affected. As a workaround, build single-architecture binaries using `--offload-arch` targeting only one GPU architecture, or disable ASAN by removing `-fsanitize=address` for HIP compilation.
|
||||
|
||||
#### ASAN produces incorrect results with ternary operators on struct kernel arguments
|
||||
|
||||
When you compile GPU kernels with ASAN enabled, ternary operators with struct kernel arguments might produce incorrect results. This can mask real bugs and produce false-positive results during memory-safety validation. The issue doesn't occur when the kernel arguments are first copied to local variables, or when compiled without ASAN. As a workaround, copy kernel arguments to local variables before using them in ternary expressions:
|
||||
|
||||
```cpp
|
||||
auto local_arg = kernel_arg;
|
||||
result = condition ? local_arg : other_arg;
|
||||
```
|
||||
|
||||
Alternatively, disable ASAN by removing `-fsanitize=address` when compiling GPU kernels.
|
||||
|
||||
## ROCm resolved issues
|
||||
|
||||
The following notable issues have been fixed in ROCm 7.13.0.
|
||||
|
||||
### Multi-ROCm installation failed on RPM-based distributions
|
||||
|
||||
Previously, installing multiple ROCm versions side by side on RPM-based distributions (RHEL and SLES) failed due to `.build-id` file conflicts between versioned packages.
|
||||
|
||||
### vLLM server failed to launch in ROCm Docker images
|
||||
|
||||
Previously, the vLLM server failed to start in ROCm 7.12.0 Docker images with an `ImportError` for `librocm_smi64.so.1` due to missing library path configuration.
|
||||
|
||||
### vLLM server failed to launch with tensor parallelism
|
||||
|
||||
Previously, the vLLM server failed to start with an invalid device pointer error when launching models with tensor parallelism set to 8 on AMD Instinct MI300 and MI355X GPUs.
|
||||
|
||||
### PyTorch DDP Gloo backend test failed on AMD GPUs
|
||||
|
||||
Previously, the PyTorch Distributed Data Parallel (DDP) test `test_ddp_apply_optim_in_backward_grad_as_bucket_view_false` failed when using the Gloo backend.
|
||||
|
||||
### rocWMMA header produced unknown type errors in HIP RTC
|
||||
|
||||
Previously, HIP RTC programs that included the `rocwmma/rocwmma.hpp` header failed to compile with unknown type name errors.
|
||||
|
||||
## ROCm upcoming changes
|
||||
|
||||
Future releases will add support for:
|
||||
|
||||
* Additional ROCm Core SDK components
|
||||
|
||||
* Domain-specific expansion toolkits (data science, life science, finance,
|
||||
simulation, and other HPC domains)
|
||||
|
||||
* More AMD hardware support
|
||||
@@ -0,0 +1,338 @@
|
||||
# Transition guide from legacy ROCm release stream
|
||||
|
||||
[ROCm Core SDK 7.13.0](https://rocm-stg.amd.com/en/docs-7.13.0/index.html#rocm-core-sdk) marks a step change from the ROCm legacy release stream. It is a preview release built on our new build system, TheRock.
|
||||
|
||||
## Major changes
|
||||
|
||||
<table class="rocm-docs-table table">
|
||||
<thead>
|
||||
<tr>
|
||||
<th class="head">Feature</th>
|
||||
<th class="head">ROCm Core SDK</th>
|
||||
<th class="head">ROCm Legacy</th>
|
||||
<th class="head">Description</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td>Installation directory</td>
|
||||
<td><code class="docutils literal notranslate"><span class="pre">/opt/rocm/core</span></code></td>
|
||||
<td><code class="docutils literal notranslate"><span class="pre">/opt/rocm/</span></code></td>
|
||||
<td>To support additional release streams downstream of the ROCm Core SDK</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>Package names</td>
|
||||
<td><code class="docutils literal notranslate"><span class="pre">amdrocm-[$component]</span></code></td>
|
||||
<td><code class="docutils literal notranslate"><span class="pre">rocm-[$component]</span></code> or <code class="docutils literal notranslate"><span class="pre">roc[$component]</span></code> or <code class="docutils literal notranslate"><span class="pre">hip[$component]</span></code></td>
|
||||
<td>Unique package prefix to avoid conflicts with upstream packages</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>Extras directory</td>
|
||||
<td><code class="docutils literal notranslate"><span class="pre">/opt/rocm/extras-7/</span></code></td>
|
||||
<td>N/A</td>
|
||||
<td>Shared install prefix scoped to each ROCm major version for projects built on the ROCm Core SDK</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
## Paths and linking
|
||||
|
||||
ROCm Core SDK 7.13.0 maintains ABI and API compatibility with the ROCm 7.2
|
||||
legacy releases, so recompilation is not required. For installations using your
|
||||
Linux distribution's package manager, the `amdrocm` meta package configures
|
||||
`update-alternatives` and provides backward-compatible symlinks for
|
||||
`/opt/rocm/bin`, `/opt/rocm/lib`, and other `/opt/rocm/` directories. For
|
||||
tarball installs, update `PATH`, `LD_LIBRARY_PATH`, `ROCM_PATH`, or other
|
||||
environment variables to reflect the new installation path (`/opt/rocm/core`).
|
||||
|
||||
## Software packages
|
||||
|
||||
ROCm Core SDK packages are more consolidated than the legacy ROCm release
|
||||
stream. For example, hipBLAS and rocBLAS are now combined into one package,
|
||||
`amdrocm-blas`. The table below lists new packages, their contents, and the
|
||||
corresponding legacy packages.
|
||||
|
||||
> **Note:** ASAN packages are not available in 7.13.0 and are planned for a future release.
|
||||
|
||||
(linux-packages-available-in-rocm-7-13-0)=
|
||||
|
||||
### Linux packages available in ROCm 7.13.0
|
||||
|
||||
<table class="rocm-docs-table table">
|
||||
<thead>
|
||||
<tr>
|
||||
<th class="head">ROCm Core SDK Package</th>
|
||||
<th class="head">Package Contents</th>
|
||||
<th class="head">ROCm Legacy Package</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td>amdrocm-amdsmi</td>
|
||||
<td>amd-smi</td>
|
||||
<td>amd-smi-lib, rocm-smi-lib</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-llvm</td>
|
||||
<td>amdclang++, hipcc, flang</td>
|
||||
<td>rocm-llvm, rocm-llvm-dev, Fortran compiler (included in rocm-llvm OpenMP runtime)</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-runtime</td>
|
||||
<td>HIP, ROCR, runtime compilation</td>
|
||||
<td>hip-runtime-amd, rocm-hip-runtime, rocm-language-runtime, hsa-rocr, comgr</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-fft</td>
|
||||
<td>rocFFT, hipFFT, hipFFTW</td>
|
||||
<td>rocfft, hipfft</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-blas</td>
|
||||
<td>rocBLAS, hipBLAS, hipBLASLt, hipSPARSELt</td>
|
||||
<td>rocblas, hipblas, hipblaslt, hipsparselt</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-sparse</td>
|
||||
<td>rocSPARSE, hipSPARSE</td>
|
||||
<td>rocsparse, hipsparse</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-solver</td>
|
||||
<td>rocSOLVER, hipSOLVER</td>
|
||||
<td>rocsolver, hipsolver, rocalution</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-dnn</td>
|
||||
<td>hipDNN, MIOpen</td>
|
||||
<td>miopen-hip</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-rand</td>
|
||||
<td>rocRAND, hipRAND</td>
|
||||
<td>rocrand, hiprand</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-ccl</td>
|
||||
<td>rocPRIM, rocThrust, hipCUB</td>
|
||||
<td>rocprim, rocthrust, hipcub, rocwmma</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-profiler</td>
|
||||
<td>rocprofiler-systems, rocprofiler-compute, rocprofiler-sdk, roctracer</td>
|
||||
<td>rocprofiler, rocprofiler-compute, rocprofiler-systems, rocprofiler-sdk, roctracer</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-profiler-base</td>
|
||||
<td>rocprofiler-sdk, roctracer</td>
|
||||
<td>rocprofiler-register, roctracer, hsa-amd-aqlprofile</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-base</td>
|
||||
<td>rocminfo, rocm-core</td>
|
||||
<td>rocm-core, rocminfo, rocm-cmake, half</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-ck</td>
|
||||
<td>Composable Kernel</td>
|
||||
<td>composablekernel</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-debugger</td>
|
||||
<td>rocgdb, ROCdbgapi, ROCr Debug Agent</td>
|
||||
<td>rocm-gdb, rocm-dbgapi, rocm-debug-agent</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-hipify</td>
|
||||
<td>HIPIFY</td>
|
||||
<td>hipify-clang</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-opencl</td>
|
||||
<td>OpenCL runtime and ICD loader</td>
|
||||
<td>rocm-opencl-runtime, rocm-opencl, hip-opencl</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-decode</td>
|
||||
<td>rocDecode (newly included in the ROCm Core SDK)</td>
|
||||
<td>rocdecode</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-jpeg</td>
|
||||
<td>rocJPEG (newly included in the ROCm Core SDK)</td>
|
||||
<td>rocjpeg</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-rccl</td>
|
||||
<td>rccl</td>
|
||||
<td>rccl</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-rocshmem</td>
|
||||
<td>rocSHMEM</td>
|
||||
<td>rocshmem</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-rdc</td>
|
||||
<td>ROCm Data Center Tool (newly included in the ROCm Core SDK)</td>
|
||||
<td>rdc</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>amdrocm-sysdeps</td>
|
||||
<td>Bundled third-party dependencies (libdrm, libelf, numa, libVA)</td>
|
||||
<td>System dependencies</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
Packages are offered in the following variants:
|
||||
|
||||
- **For all supported GPUs** -- works across all GPUs supported by ROCm (for example, `apt install amdrocm-core-sdk7.13`).
|
||||
- **For a specific GPU architecture** -- smaller install size, but requires you to know the GPU installed in your system (for example, `apt install amdrocm-core-sdk7.13-gfx110x`).
|
||||
|
||||
Installing all GPU architectures is not required. You can install packages for a specific architecture, multiple architectures side by side, or all supported GPU architectures.
|
||||
|
||||
When redistributing software built on the ROCm Core SDK (for example, via containers), we recommend the all GPU package variant for broad hardware support. If disk footprint is a concern, you can use a single GPU architecture package variant instead.
|
||||
|
||||
### Architecture-specific packages available in ROCm 7.13.0
|
||||
|
||||
<table class="rocm-docs-table table">
|
||||
<thead>
|
||||
<tr>
|
||||
<th class="head">Architecture Family</th>
|
||||
<th class="head">Package Suffix</th>
|
||||
<th class="head">Product Name (Not Exhaustive)</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td>CDNA4</td>
|
||||
<td>-gfx950</td>
|
||||
<td>AMD Instinct MI355X / MI350X</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>CDNA3</td>
|
||||
<td>-gfx94x</td>
|
||||
<td>AMD Instinct MI325X / MI300X / MI300A</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>CDNA2</td>
|
||||
<td>-gfx90a</td>
|
||||
<td>AMD Instinct MI250X / MI250 / MI210</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>CDNA</td>
|
||||
<td>-gfx908</td>
|
||||
<td>AMD Instinct MI100</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>RDNA4</td>
|
||||
<td>-gfx120x</td>
|
||||
<td>AMD Radeon RX 9070 / AMD Radeon RX 9060 / AMD Radeon RX 9070 XT / AMD Radeon RX 9060 XT / AMD Radeon RX 9070 GRE / AMD Radeon AI PRO R9700 / AMD Radeon AI PRO R9600D / AMD Radeon RX 9060 XT LP</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>RDNA3.5</td>
|
||||
<td>-gfx1150<br>-gfx1151<br>-gfx1152</td>
|
||||
<td>AMD Ryzen AI 9 465 / AMD Ryzen AI 9 365 / AMD Ryzen AI 9 HX 475 / AMD Ryzen AI 9 HX 470 / AMD Ryzen AI 9 HX 375 / AMD Ryzen AI 9 HX 370 / AMD Ryzen AI 9 PRO 465 / AMD Ryzen AI 9 PRO HX 475 / AMD Ryzen AI 9 PRO HX 470 / AMD Ryzen AI 9 HX PRO 375 / AMD Ryzen AI 9 HX PRO 370 / AMD Ryzen AI Max 390 / AMD Ryzen AI Max 385 / AMD Ryzen AI Max+ 395 / AMD Ryzen AI Max+ 392 / AMD Ryzen AI Max+ 388 / AMD Ryzen AI Max PRO 390 / AMD Ryzen AI Max PRO 385 / AMD Ryzen AI Max PRO 380 / AMD Ryzen AI Max+ PRO 395 / AMD Ryzen AI 7 450 / AMD Ryzen AI 7 350 / AMD Ryzen AI 7 345 / AMD Ryzen AI 5 340 / AMD Ryzen AI 5 330 / AMD Ryzen AI 7 PRO 450 / AMD Ryzen AI 5 PRO 440 / AMD Ryzen AI 7 PRO 350 / AMD Ryzen AI 5 PRO 340</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>RDNA3</td>
|
||||
<td>-gfx110x</td>
|
||||
<td>AMD Radeon RX 7700 / AMD Radeon RX 7600 / AMD Radeon PRO V710 / AMD Radeon PRO W7900 / AMD Radeon PRO W7800 / AMD Radeon PRO W7700 / AMD Radeon RX 7900 XT / AMD Radeon RX 7800 XT / AMD Radeon RX 7700 XT / AMD Radeon RX 7700 XE / AMD Radeon RX 7900 XTX / AMD Radeon RX 7900 GRE / AMD Radeon PRO W7800 48GB / AMD Radeon PRO W7900 Dual Slot</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>RDNA2</td>
|
||||
<td>-gfx1030</td>
|
||||
<td>AMD Radeon PRO V620 / AMD Radeon PRO W6800</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
## ROCm Core SDK component changes (moved or removed)
|
||||
|
||||
### Planned for future releases
|
||||
|
||||
- ROCm Core SDK: RPP
|
||||
- ROCm-Extras: hipfort, rocALUTION, rocPyDecode, rocAL, MIVisionX
|
||||
|
||||
### Moved to ROCm-Extras
|
||||
|
||||
- ROCm Validation Suite
|
||||
- ROCm Bandwidth Test
|
||||
- TransferBench
|
||||
- MIGraphX
|
||||
|
||||
### Moved to Standalone/ONNX
|
||||
|
||||
- ONNX runtime
|
||||
|
||||
### Removed
|
||||
|
||||
- [ROCm SMI](https://rocm.docs.amd.com/en/latest/about/release-notes.html#rocm-smi-deprecation) (replaced by AMD SMI)
|
||||
|
||||
## Notable package relocations
|
||||
|
||||
- rocMLIR (now included in MIGraphX)
|
||||
- HIPCC (now included in `amdrocm-llvm`)
|
||||
- FLANG (now included in `amdrocm-llvm`)
|
||||
- ROCm CMake (now in `amdrocm-base`)
|
||||
- ROCTracer (now in `amdrocm-profiler-base`)
|
||||
- ROCProfiler (functionality in `amdrocm-profiler`)
|
||||
|
||||
## Components available in the ROCm Core SDK, ROCm-Extras, and Standalone/ONNX
|
||||
|
||||
<table class="rocm-docs-table table">
|
||||
<thead>
|
||||
<tr>
|
||||
<th class="head"></th>
|
||||
<th class="head">Category</th>
|
||||
<th class="head">Present</th>
|
||||
<th class="head">Absent/Moved</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td rowspan="6" class="stub" style="vertical-align: middle"><strong>ROCm Core SDK</strong></td>
|
||||
<td>Math and compute libraries</td>
|
||||
<td>CK, hipBLAS, hipBLASLt, hipCUB, hipFFT, hipRAND, hipSOLVER, hipSPARSE/SPARSELt, MIOpen, rocBLAS, rocFFT, rocRAND, rocSOLVER, rocSPARSE, rocPRIM, rocThrust, rocWMMA</td>
|
||||
<td>hipfort, rocALUTION</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>Communication libraries</td>
|
||||
<td>RCCL, rocSHMEM</td>
|
||||
<td>—</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>Media libraries</td>
|
||||
<td>rocDecode, rocJPEG, ROCm Performance Primitives (RPP planned for a future release)</td>
|
||||
<td>rocPyDecode, rocAL, MIVisionX, MIGraphX, CK (moved to math and compute)</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>Runtime, compilers, build tools</td>
|
||||
<td>HIP, HIPIFY, LLVM</td>
|
||||
<td>HIPCC, FLANG, ROCm CMake</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>Profiling and debugging tools</td>
|
||||
<td>ROCm Compute Profiler, ROCm Systems Profiler, ROCprofiler-SDK, ROCdbgapi, ROCm Debugger, ROCr Debug Agent</td>
|
||||
<td>ROCTracer, ROCProfiler</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>Control and monitoring tools</td>
|
||||
<td>AMD SMI, ROCm Data Center Tool, rocminfo, hipinfo</td>
|
||||
<td>ROCm SMI (removed), ROCm Validation Suite, ROCm Bandwidth Test</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style="vertical-align: middle"><strong>ROCm-Extras</strong></td>
|
||||
<td>—</td>
|
||||
<td>ROCm Validation Suite, ROCm Bandwidth Test, TransferBench, MIGraphX</td>
|
||||
<td>—</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style="vertical-align: middle"><strong>Standalone/ONNX</strong></td>
|
||||
<td>—</td>
|
||||
<td>rocMLIR, ONNX runtime</td>
|
||||
<td>—</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
@@ -8,7 +8,7 @@ docker:
|
||||
- "Parallel VAE decode support for Wan models"
|
||||
- "Batch inference and data parallel support"
|
||||
components:
|
||||
TheRock:
|
||||
TheRock:
|
||||
version: 9b611c6
|
||||
url: https://github.com/ROCm/TheRock
|
||||
rocm-libraries:
|
||||
@@ -75,7 +75,7 @@ docker:
|
||||
- '--guidance_scale 6.0 \'
|
||||
- '--use_torch_compile \'
|
||||
- '--attention_backend aiter \'
|
||||
- '--output_directory results'
|
||||
- "--output_directory results"
|
||||
- model: Hunyuan Video 1.5
|
||||
model_repo: hunyuanvideo-community/HunyuanVideo-1.5-Diffusers-720p_t2v
|
||||
url: https://huggingface.co/hunyuanvideo-community/HunyuanVideo-1.5-Diffusers-720p_t2v
|
||||
@@ -97,7 +97,7 @@ docker:
|
||||
- '--enable_tiling --enable_slicing \'
|
||||
- '--use_torch_compile \'
|
||||
- '--attention_backend aiter \'
|
||||
- '--output_directory results'
|
||||
- "--output_directory results"
|
||||
- group: Wan-AI
|
||||
js_tag: wan
|
||||
models:
|
||||
@@ -123,7 +123,7 @@ docker:
|
||||
- '--num_inference_steps 40 \'
|
||||
- '--use_torch_compile \'
|
||||
- '--attention_backend aiter \'
|
||||
- '--output_directory results'
|
||||
- "--output_directory results"
|
||||
- model: Wan2.2
|
||||
model_repo: Wan-AI/Wan2.2-I2V-A14B-Diffusers
|
||||
url: https://huggingface.co/Wan-AI/Wan2.2-I2V-A14B-Diffusers
|
||||
@@ -146,7 +146,7 @@ docker:
|
||||
- '--num_inference_steps 40 \'
|
||||
- '--use_torch_compile \'
|
||||
- '--attention_backend aiter \'
|
||||
- '--output_directory results'
|
||||
- "--output_directory results"
|
||||
- group: FLUX
|
||||
js_tag: flux
|
||||
models:
|
||||
@@ -172,7 +172,7 @@ docker:
|
||||
- '--guidance_scale 0.0 \'
|
||||
- '--num_iterations 50 \'
|
||||
- '--attention_backend aiter \'
|
||||
- '--output_directory results'
|
||||
- "--output_directory results"
|
||||
- model: FLUX.1 Kontext
|
||||
model_repo: black-forest-labs/FLUX.1-Kontext-dev
|
||||
url: https://huggingface.co/black-forest-labs/FLUX.1-Kontext-dev
|
||||
@@ -196,7 +196,7 @@ docker:
|
||||
- '--guidance_scale 2.5 \'
|
||||
- '--num_iterations 25 \'
|
||||
- '--attention_backend aiter \'
|
||||
- '--output_directory results'
|
||||
- "--output_directory results"
|
||||
- model: FLUX.2
|
||||
model_repo: black-forest-labs/FLUX.2-dev
|
||||
url: https://huggingface.co/black-forest-labs/FLUX.2-dev
|
||||
@@ -220,7 +220,7 @@ docker:
|
||||
- '--guidance_scale 4.0 \'
|
||||
- '--num_iterations 25 \'
|
||||
- '--attention_backend aiter \'
|
||||
- '--output_directory results'
|
||||
- "--output_directory results"
|
||||
- model: FLUX.2 Klein
|
||||
model_repo: black-forest-labs/FLUX.2-klein-9B
|
||||
url: https://huggingface.co/black-forest-labs/FLUX.2-klein-9B
|
||||
@@ -242,7 +242,7 @@ docker:
|
||||
- '--guidance_scale 1.0 \'
|
||||
- '--num_iterations 25 \'
|
||||
- '--attention_backend aiter \'
|
||||
- '--output_directory results'
|
||||
- "--output_directory results"
|
||||
- group: StableDiffusion
|
||||
js_tag: stablediffusion
|
||||
models:
|
||||
@@ -263,7 +263,7 @@ docker:
|
||||
- '--use_cfg_parallel \'
|
||||
- '--use_torch_compile \'
|
||||
- '--attention_backend aiter \'
|
||||
- '--output_directory results'
|
||||
- "--output_directory results"
|
||||
- group: Z-Image
|
||||
js_tag: z_image
|
||||
models:
|
||||
@@ -289,7 +289,7 @@ docker:
|
||||
- '--guidance_scale 4.0 \'
|
||||
- '--num_iterations 25 \'
|
||||
- '--attention_backend aiter \'
|
||||
- '--output_directory results'
|
||||
- "--output_directory results"
|
||||
- group: LTX
|
||||
js_tag: ltx
|
||||
models:
|
||||
@@ -313,7 +313,7 @@ docker:
|
||||
- '--guidance_scale 4.0 \'
|
||||
- '--num_iterations 1 \'
|
||||
- '--attention_backend aiter \'
|
||||
- '--output_directory results'
|
||||
- "--output_directory results"
|
||||
- group: Qwen-Image
|
||||
js_tag: qwen_image
|
||||
models:
|
||||
@@ -336,7 +336,7 @@ docker:
|
||||
- '--use_torch_compile \'
|
||||
- '--num_iterations 1 \'
|
||||
- '--attention_backend aiter \'
|
||||
- '--output_directory results'
|
||||
- "--output_directory results"
|
||||
- model: Qwen-Image-Edit
|
||||
model_repo: Qwen/Qwen-Image-Edit
|
||||
url: https://huggingface.co/Qwen/Qwen-Image-Edit
|
||||
@@ -358,4 +358,4 @@ docker:
|
||||
- '--use_torch_compile \'
|
||||
- '--num_iterations 1 \'
|
||||
- '--attention_backend aiter \'
|
||||
- '--output_directory results'
|
||||
- "--output_directory results"
|
||||
@@ -17,7 +17,7 @@ vLLM inference performance testing
|
||||
|
||||
.. _vllm-benchmark-unified-docker-812:
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.0_20250812-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.10.0-20250812.yaml
|
||||
|
||||
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
|
||||
{% set model_groups = data.vllm_benchmark.model_groups %}
|
||||
@@ -64,7 +64,7 @@ The following is summary of notable changes since the :doc:`previous ROCm/vLLM D
|
||||
Supported models
|
||||
================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.0_20250812-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.10.0-20250812.yaml
|
||||
|
||||
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
|
||||
{% set model_groups = data.vllm_benchmark.model_groups %}
|
||||
@@ -123,8 +123,6 @@ Supported models
|
||||
|
||||
vLLM is a toolkit and library for LLM inference and serving. AMD implements
|
||||
high-performance custom kernels and modules in vLLM to enhance performance.
|
||||
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
|
||||
more information.
|
||||
|
||||
.. _vllm-benchmark-performance-measurements-812:
|
||||
|
||||
@@ -157,7 +155,7 @@ To test for optimal performance, consult the recommended :ref:`System health ben
|
||||
<rocm-for-ai-system-health-bench>`. This suite of tests will help you verify and fine-tune your
|
||||
system's configuration.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.0_20250812-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.10.0-20250812.yaml
|
||||
|
||||
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
|
||||
{% set model_groups = data.vllm_benchmark.model_groups %}
|
||||
@@ -296,21 +294,21 @@ system's configuration.
|
||||
|
||||
* -
|
||||
- ``configs/extended.csv``
|
||||
-
|
||||
-
|
||||
|
||||
* -
|
||||
- ``configs/performance.csv``
|
||||
-
|
||||
-
|
||||
|
||||
* - ``--benchmark``
|
||||
- ``throughput``
|
||||
- Measure offline end-to-end throughput.
|
||||
|
||||
* -
|
||||
* -
|
||||
- ``serving``
|
||||
- Measure online serving performance.
|
||||
|
||||
* -
|
||||
* -
|
||||
- ``all``
|
||||
- Measure both throughput and serving.
|
||||
|
||||
@@ -433,9 +431,6 @@ Further reading
|
||||
- To learn how to run community models from Hugging Face on AMD GPUs, see
|
||||
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
|
||||
|
||||
- To learn how to fine-tune LLMs and optimize inference, see
|
||||
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
|
||||
|
||||
- For a list of other ready-made Docker images for AI with ROCm, see
|
||||
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
|
||||
|
||||
@@ -16,7 +16,7 @@ vLLM inference performance testing
|
||||
|
||||
.. _vllm-benchmark-unified-docker-909:
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20250909-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.10.1-20250909.yaml
|
||||
|
||||
{% set docker = data.dockers[0] %}
|
||||
|
||||
@@ -57,7 +57,7 @@ The following is summary of notable changes since the :doc:`previous ROCm/vLLM D
|
||||
Supported models
|
||||
================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20250909-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.10.1-20250909.yaml
|
||||
|
||||
{% set docker = data.dockers[0] %}
|
||||
{% set model_groups = data.model_groups %}
|
||||
@@ -146,7 +146,7 @@ To test for optimal performance, consult the recommended :ref:`System health ben
|
||||
<rocm-for-ai-system-health-bench>`. This suite of tests will help you verify and fine-tune your
|
||||
system's configuration.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20250909-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.10.1-20250909.yaml
|
||||
|
||||
{% set docker = data.dockers[0] %}
|
||||
{% set model_groups = data.model_groups %}
|
||||
@@ -433,9 +433,6 @@ Further reading
|
||||
- To learn more about system settings and management practices to configure your system for
|
||||
AMD Instinct MI300X Series accelerators, see `AMD Instinct MI300X system optimization <https://instinct.docs.amd.com/projects/amdgpu-docs/en/latest/system-optimization/mi300x.html>`_.
|
||||
|
||||
- See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
|
||||
a brief introduction to vLLM and optimization strategies.
|
||||
|
||||
- For application performance optimization strategies for HPC and AI workloads,
|
||||
including inference with vLLM, see :doc:`/how-to/rocm-for-ai/inference-optimization/workload`.
|
||||
|
||||
@@ -16,7 +16,7 @@ vLLM inference performance testing
|
||||
|
||||
.. _vllm-benchmark-unified-docker-930:
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20251006-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.10.1-20251006.yaml
|
||||
|
||||
{% set docker = data.dockers[0] %}
|
||||
|
||||
@@ -75,7 +75,7 @@ The following is summary of notable changes since the :doc:`previous ROCm/vLLM D
|
||||
Supported models
|
||||
================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20251006-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.10.1-20251006.yaml
|
||||
|
||||
{% set docker = data.dockers[0] %}
|
||||
{% set model_groups = data.model_groups %}
|
||||
@@ -178,7 +178,7 @@ system's configuration.
|
||||
Pull the Docker image
|
||||
=====================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20251006-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.10.1-20251006.yaml
|
||||
|
||||
{% set docker = data.dockers[0] %}
|
||||
|
||||
@@ -192,7 +192,7 @@ Pull the Docker image
|
||||
Benchmarking
|
||||
============
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20251006-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.10.1-20251006.yaml
|
||||
|
||||
{% set docker = data.dockers[0] %}
|
||||
{% set model_groups = data.model_groups %}
|
||||
@@ -441,7 +441,7 @@ To reproduce this ROCm-enabled vLLM Docker image release, follow these steps:
|
||||
|
||||
2. Use the following command to build the image directly from the specified commit.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20251006-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.10.1-20251006.yaml
|
||||
|
||||
{% set docker = data.dockers[0] %}
|
||||
.. code-block:: shell
|
||||
@@ -467,9 +467,6 @@ Further reading
|
||||
- To learn more about system settings and management practices to configure your system for
|
||||
AMD Instinct MI300X Series GPUs, see `AMD Instinct MI300X system optimization <https://instinct.docs.amd.com/projects/amdgpu-docs/en/latest/system-optimization/mi300x.html>`_.
|
||||
|
||||
- See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
|
||||
a brief introduction to vLLM and optimization strategies.
|
||||
|
||||
- For application performance optimization strategies for HPC and AI workloads,
|
||||
including inference with vLLM, see :doc:`/how-to/rocm-for-ai/inference-optimization/workload`.
|
||||
|
||||
@@ -16,7 +16,7 @@ vLLM inference performance testing
|
||||
|
||||
.. _vllm-benchmark-unified-docker-1103:
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.11.1_20251103-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.11.1-20251103.yaml
|
||||
|
||||
{% set docker = data.dockers[0] %}
|
||||
|
||||
@@ -61,7 +61,7 @@ The following is summary of notable changes since the :doc:`previous ROCm/vLLM D
|
||||
Supported models
|
||||
================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.11.1_20251103-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.11.1-20251103.yaml
|
||||
|
||||
{% set docker = data.dockers[0] %}
|
||||
{% set model_groups = data.model_groups %}
|
||||
@@ -164,7 +164,7 @@ system's configuration.
|
||||
Pull the Docker image
|
||||
=====================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.11.1_20251103-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.11.1-20251103.yaml
|
||||
|
||||
{% set docker = data.dockers[0] %}
|
||||
|
||||
@@ -178,7 +178,7 @@ Pull the Docker image
|
||||
Benchmarking
|
||||
============
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.11.1_20251103-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.11.1-20251103.yaml
|
||||
|
||||
{% set docker = data.dockers[0] %}
|
||||
{% set model_groups = data.model_groups %}
|
||||
@@ -431,7 +431,7 @@ To reproduce this ROCm-enabled vLLM Docker image release, follow these steps:
|
||||
|
||||
2. Use the following command to build the image directly from the specified commit.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.11.1_20251103-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.11.1-20251103.yaml
|
||||
|
||||
{% set docker = data.dockers[0] %}
|
||||
.. code-block:: shell
|
||||
@@ -457,9 +457,6 @@ Further reading
|
||||
- To learn more about system settings and management practices to configure your system for
|
||||
AMD Instinct MI300X Series GPUs, see `AMD Instinct MI300X system optimization <https://instinct.docs.amd.com/projects/amdgpu-docs/en/latest/system-optimization/mi300x.html>`_.
|
||||
|
||||
- See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
|
||||
a brief introduction to vLLM and optimization strategies.
|
||||
|
||||
- For application performance optimization strategies for HPC and AI workloads,
|
||||
including inference with vLLM, see :doc:`/how-to/rocm-for-ai/inference-optimization/workload`.
|
||||
|
||||
@@ -44,9 +44,7 @@ optimizing performance with popular AI models.
|
||||
consumption and increases throughput by leveraging dynamic key and value
|
||||
allocation in GPU memory. vLLM also incorporates many LLM acceleration
|
||||
and quantization algorithms. In addition, AMD implements high-performance
|
||||
custom kernels and modules in vLLM to enhance performance further. See
|
||||
:ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for more
|
||||
information.
|
||||
custom kernels and modules in vLLM to enhance performance further.
|
||||
|
||||
Getting started
|
||||
===============
|
||||
@@ -277,7 +275,7 @@ options and their descriptions.
|
||||
|
||||
Latency benchmark example
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
|
||||
Use this command to benchmark the latency of the Llama 3.1 8B model on one GPU with the ``float16`` data type.
|
||||
|
||||
.. code-block::
|
||||
@@ -334,9 +332,6 @@ Further reading
|
||||
- To learn how to run community models from Hugging Face on AMD GPUs, see
|
||||
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
|
||||
|
||||
- To learn how to fine-tune LLMs and optimize inference, see
|
||||
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
|
||||
|
||||
- For a list of other ready-made Docker images for AI with ROCm, see
|
||||
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
|
||||
|
||||
@@ -68,8 +68,6 @@ optimizing performance with popular AI models.
|
||||
|
||||
vLLM is a toolkit and library for LLM inference and serving. AMD implements
|
||||
high-performance custom kernels and modules in vLLM to enhance performance.
|
||||
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
|
||||
more information.
|
||||
|
||||
Getting started
|
||||
===============
|
||||
@@ -342,7 +340,7 @@ options and their descriptions.
|
||||
|
||||
Example 1: latency benchmark
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
|
||||
Use this command to benchmark the latency of the Llama 3.1 8B model on one GPU with the ``float16`` and ``float8`` data types.
|
||||
|
||||
.. code-block::
|
||||
@@ -404,9 +402,6 @@ Further reading
|
||||
- To learn how to run community models from Hugging Face on AMD GPUs, see
|
||||
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
|
||||
|
||||
- To learn how to fine-tune LLMs and optimize inference, see
|
||||
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
|
||||
|
||||
- For a list of other ready-made Docker images for AI with ROCm, see
|
||||
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
|
||||
|
||||
@@ -41,8 +41,6 @@ and :ref:`standalone benchmarking <vllm-benchmark-standalone-v066-options>`.
|
||||
|
||||
vLLM is a toolkit and library for LLM inference and serving. AMD implements
|
||||
high-performance custom kernels and modules in vLLM to enhance performance.
|
||||
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
|
||||
more information.
|
||||
|
||||
Getting started
|
||||
===============
|
||||
@@ -387,7 +385,7 @@ options and their descriptions.
|
||||
|
||||
Example 1: latency benchmark
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
|
||||
Use this command to benchmark the latency of the Llama 3.1 70B model on eight GPUs with the ``float16`` and ``float8`` data types.
|
||||
|
||||
.. code-block::
|
||||
@@ -449,9 +447,6 @@ Further reading
|
||||
- To learn how to run community models from Hugging Face on AMD GPUs, see
|
||||
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
|
||||
|
||||
- To learn how to fine-tune LLMs and optimize inference, see
|
||||
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
|
||||
|
||||
- For a list of other ready-made Docker images for AI with ROCm, see
|
||||
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
|
||||
|
||||
@@ -17,7 +17,7 @@ vLLM inference performance testing
|
||||
|
||||
.. _vllm-benchmark-unified-docker:
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.7.3_20250325-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.7.3-20250325.yaml
|
||||
|
||||
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
|
||||
{% set model_groups = data.vllm_benchmark.model_groups %}
|
||||
@@ -93,8 +93,6 @@ vLLM inference performance testing
|
||||
|
||||
vLLM is a toolkit and library for LLM inference and serving. AMD implements
|
||||
high-performance custom kernels and modules in vLLM to enhance performance.
|
||||
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
|
||||
more information.
|
||||
|
||||
.. _vllm-benchmark-performance-measurements-v073:
|
||||
|
||||
@@ -317,9 +315,6 @@ Further reading
|
||||
- To learn how to run community models from Hugging Face on AMD GPUs, see
|
||||
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
|
||||
|
||||
- To learn how to fine-tune LLMs and optimize inference, see
|
||||
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
|
||||
|
||||
- For a list of other ready-made Docker images for AI with ROCm, see
|
||||
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
|
||||
|
||||
@@ -12,7 +12,7 @@ vLLM inference performance testing
|
||||
|
||||
.. _vllm-benchmark-unified-docker:
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.8.3_20250415-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.8.3-20250415.yaml
|
||||
|
||||
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
|
||||
{% set model_groups = data.vllm_benchmark.model_groups %}
|
||||
@@ -88,8 +88,6 @@ vLLM inference performance testing
|
||||
|
||||
vLLM is a toolkit and library for LLM inference and serving. AMD implements
|
||||
high-performance custom kernels and modules in vLLM to enhance performance.
|
||||
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
|
||||
more information.
|
||||
|
||||
.. _vllm-benchmark-performance-measurements-v083:
|
||||
|
||||
@@ -333,9 +331,6 @@ Further reading
|
||||
- To learn how to run community models from Hugging Face on AMD GPUs, see
|
||||
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
|
||||
|
||||
- To learn how to fine-tune LLMs and optimize inference, see
|
||||
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
|
||||
|
||||
- For a list of other ready-made Docker images for AI with ROCm, see
|
||||
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
|
||||
|
||||
@@ -17,7 +17,7 @@ vLLM inference performance testing
|
||||
|
||||
.. _vllm-benchmark-unified-docker:
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.8.5_20250513-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.8.5-20250513.yaml
|
||||
|
||||
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
|
||||
{% set model_groups = data.vllm_benchmark.model_groups %}
|
||||
@@ -97,8 +97,6 @@ vLLM inference performance testing
|
||||
|
||||
vLLM is a toolkit and library for LLM inference and serving. AMD implements
|
||||
high-performance custom kernels and modules in vLLM to enhance performance.
|
||||
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
|
||||
more information.
|
||||
|
||||
.. _vllm-benchmark-performance-measurements-v085-20250513:
|
||||
|
||||
@@ -342,9 +340,6 @@ Further reading
|
||||
- To learn how to run community models from Hugging Face on AMD GPUs, see
|
||||
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
|
||||
|
||||
- To learn how to fine-tune LLMs and optimize inference, see
|
||||
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
|
||||
|
||||
- For a list of other ready-made Docker images for AI with ROCm, see
|
||||
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
|
||||
|
||||
@@ -17,7 +17,7 @@ vLLM inference performance testing
|
||||
|
||||
.. _vllm-benchmark-unified-docker:
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.8.5_20250521-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.8.5-20250521.yaml
|
||||
|
||||
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
|
||||
{% set model_groups = data.vllm_benchmark.model_groups %}
|
||||
@@ -97,8 +97,6 @@ vLLM inference performance testing
|
||||
|
||||
vLLM is a toolkit and library for LLM inference and serving. AMD implements
|
||||
high-performance custom kernels and modules in vLLM to enhance performance.
|
||||
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
|
||||
more information.
|
||||
|
||||
.. _vllm-benchmark-performance-measurements-v085-20250521:
|
||||
|
||||
@@ -342,9 +340,6 @@ Further reading
|
||||
- To learn how to run community models from Hugging Face on AMD GPUs, see
|
||||
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
|
||||
|
||||
- To learn how to fine-tune LLMs and optimize inference, see
|
||||
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
|
||||
|
||||
- For a list of other ready-made Docker images for AI with ROCm, see
|
||||
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
|
||||
|
||||
@@ -17,7 +17,7 @@ vLLM inference performance testing
|
||||
|
||||
.. _vllm-benchmark-unified-docker:
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.9.0.1_20250605-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.9.0.1-20250605.yaml
|
||||
|
||||
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
|
||||
{% set model_groups = data.vllm_benchmark.model_groups %}
|
||||
@@ -97,8 +97,6 @@ vLLM inference performance testing
|
||||
|
||||
vLLM is a toolkit and library for LLM inference and serving. AMD implements
|
||||
high-performance custom kernels and modules in vLLM to enhance performance.
|
||||
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
|
||||
more information.
|
||||
|
||||
.. _vllm-benchmark-performance-measurements-v0901-20250605:
|
||||
|
||||
@@ -341,9 +339,6 @@ Further reading
|
||||
- To learn how to run community models from Hugging Face on AMD GPUs, see
|
||||
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
|
||||
|
||||
- To learn how to fine-tune LLMs and optimize inference, see
|
||||
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
|
||||
|
||||
- For a list of other ready-made Docker images for AI with ROCm, see
|
||||
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
|
||||
|
||||
@@ -17,7 +17,7 @@ vLLM inference performance testing
|
||||
|
||||
.. _vllm-benchmark-unified-docker-702:
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.9.1_20250702-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.9.1-20250702.yaml
|
||||
|
||||
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
|
||||
{% set model_groups = data.vllm_benchmark.model_groups %}
|
||||
@@ -97,8 +97,6 @@ vLLM inference performance testing
|
||||
|
||||
vLLM is a toolkit and library for LLM inference and serving. AMD implements
|
||||
high-performance custom kernels and modules in vLLM to enhance performance.
|
||||
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
|
||||
more information.
|
||||
|
||||
.. _vllm-benchmark-performance-measurements-20250702:
|
||||
|
||||
@@ -341,9 +339,6 @@ Further reading
|
||||
- To learn how to run community models from Hugging Face on AMD GPUs, see
|
||||
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
|
||||
|
||||
- To learn how to fine-tune LLMs and optimize inference, see
|
||||
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
|
||||
|
||||
- For a list of other ready-made Docker images for AI with ROCm, see
|
||||
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
|
||||
|
||||
@@ -17,7 +17,7 @@ vLLM inference performance testing
|
||||
|
||||
.. _vllm-benchmark-unified-docker-715:
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.9.1_20250715-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.9.1-20250715.yaml
|
||||
|
||||
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
|
||||
{% set model_groups = data.vllm_benchmark.model_groups %}
|
||||
@@ -70,7 +70,7 @@ The following is summary of notable changes since the :doc:`previous ROCm/vLLM D
|
||||
Supported models
|
||||
================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.9.1_20250715-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.9.1-20250715.yaml
|
||||
|
||||
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
|
||||
{% set model_groups = data.vllm_benchmark.model_groups %}
|
||||
@@ -129,8 +129,6 @@ Supported models
|
||||
|
||||
vLLM is a toolkit and library for LLM inference and serving. AMD implements
|
||||
high-performance custom kernels and modules in vLLM to enhance performance.
|
||||
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
|
||||
more information.
|
||||
|
||||
.. _vllm-benchmark-performance-measurements-715:
|
||||
|
||||
@@ -163,7 +161,7 @@ To test for optimal performance, consult the recommended :ref:`System health ben
|
||||
<rocm-for-ai-system-health-bench>`. This suite of tests will help you verify and fine-tune your
|
||||
system's configuration.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.9.1_20250715-benchmark-models.yaml
|
||||
.. datatemplate:yaml:: ./data/vllm-0.9.1-20250715.yaml
|
||||
|
||||
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
|
||||
{% set model_groups = data.vllm_benchmark.model_groups %}
|
||||
@@ -438,9 +436,6 @@ Further reading
|
||||
- To learn how to run community models from Hugging Face on AMD GPUs, see
|
||||
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
|
||||
|
||||
- To learn how to fine-tune LLMs and optimize inference, see
|
||||
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
|
||||
|
||||
- For a list of other ready-made Docker images for AI with ROCm, see
|
||||
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
|
||||
|
||||
@@ -6,13 +6,19 @@
|
||||
prebuilt and optimized docker images.
|
||||
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
|
||||
|
||||
************************
|
||||
xDiT diffusion inference
|
||||
************************
|
||||
******************************
|
||||
xDiT diffusion inference 25.10
|
||||
******************************
|
||||
|
||||
.. caution::
|
||||
|
||||
This documentation does not reflect the latest version of the xDiT diffusion
|
||||
inference performance documentation. See
|
||||
:doc:`/ai-inference/archive/xdit-history` for the latest version.
|
||||
|
||||
.. _xdit-video-diffusion-2510:
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.10-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.10.yaml
|
||||
|
||||
{% set docker = data.xdit_diffusion_inference.docker %}
|
||||
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
|
||||
@@ -59,7 +65,7 @@ The following models are supported for inference performance benchmarking.
|
||||
Some instructions, commands, and recommendations in this documentation might
|
||||
vary by model -- select one to get started.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.10-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.10.yaml
|
||||
|
||||
{% set docker = data.xdit_diffusion_inference.docker %}
|
||||
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
|
||||
@@ -122,7 +128,7 @@ guide to properly configure your system settings before starting.
|
||||
Pull the Docker image
|
||||
=====================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.10-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.10.yaml
|
||||
|
||||
{% set docker = data.xdit_diffusion_inference.docker %}
|
||||
|
||||
@@ -139,7 +145,7 @@ Validate and benchmark
|
||||
Once the image has been downloaded you can follow these steps to
|
||||
run benchmarks and generate outputs.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.10-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.10.yaml
|
||||
|
||||
{% set model_groups = data.xdit_diffusion_inference.model_groups %}
|
||||
{% for model_group in model_groups %}
|
||||
@@ -166,7 +172,7 @@ Prepare the model
|
||||
|
||||
You can either use an existing Hugging Face cache or download the model fresh inside the container.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.10-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.10.yaml
|
||||
|
||||
{% set docker = data.xdit_diffusion_inference.docker %}
|
||||
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
|
||||
@@ -264,7 +270,7 @@ Run inference
|
||||
You can benchmark models through `MAD <https://github.com/ROCm/MAD>`__-integrated automation or standalone
|
||||
torchrun commands.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.10-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.10.yaml
|
||||
|
||||
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
|
||||
{% for model_group in model_groups %}
|
||||
@@ -295,7 +301,7 @@ torchrun commands.
|
||||
--tags {{model.mad_tag}} \
|
||||
--keep-model-dir \
|
||||
--live-output
|
||||
|
||||
|
||||
MAD launches a Docker container with the name
|
||||
``container_ci-{{model.mad_tag}}``. The throughput and serving reports of the
|
||||
model are collected in the following paths: ``{{ model.mad_tag }}_throughput.csv``
|
||||
@@ -395,5 +401,5 @@ Further reading
|
||||
Previous versions
|
||||
=================
|
||||
|
||||
See :doc:`xdit-history` to find documentation for previous releases
|
||||
of xDiT diffusion inference performance testing.
|
||||
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
|
||||
releases of xDiT diffusion inference performance testing.
|
||||
@@ -6,20 +6,19 @@
|
||||
prebuilt and optimized docker images.
|
||||
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
|
||||
|
||||
************************
|
||||
xDiT diffusion inference
|
||||
************************
|
||||
******************************
|
||||
xDiT diffusion inference 25.11
|
||||
******************************
|
||||
|
||||
.. caution::
|
||||
|
||||
This documentation does not reflect the latest version of ROCm vLLM
|
||||
This documentation does not reflect the latest version of the xDiT diffusion
|
||||
inference performance documentation. See
|
||||
:doc:`/how-to/rocm-for-ai/inference/xdit-diffusion-inference` for the latest
|
||||
version.
|
||||
:doc:`/ai-inference/archive/xdit-history` for the latest version.
|
||||
|
||||
.. _xdit-video-diffusion-2511:
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
|
||||
|
||||
{% set docker = data.xdit_diffusion_inference.docker | selectattr("version", "equalto", "v25-11") | first %}
|
||||
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
|
||||
@@ -48,7 +47,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
|
||||
What's new
|
||||
==========
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
|
||||
|
||||
{% set docker = data.xdit_diffusion_inference.docker | selectattr("version", "equalto", "v25-11") | first %}
|
||||
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
|
||||
@@ -66,7 +65,7 @@ The following models are supported for inference performance benchmarking.
|
||||
Some instructions, commands, and recommendations in this documentation might
|
||||
vary by model -- select one to get started.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
|
||||
|
||||
{% set docker = data.xdit_diffusion_inference.docker | selectattr("version", "equalto", "v25-11") | first %}
|
||||
{% set model_groups = data.xdit_diffusion_inference.model_groups %}
|
||||
@@ -145,7 +144,7 @@ system's configuration.
|
||||
Pull the Docker image
|
||||
=====================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
|
||||
|
||||
{% set docker = data.xdit_diffusion_inference.docker | selectattr("version", "equalto", "v25-11") | first %}
|
||||
|
||||
@@ -162,7 +161,7 @@ Validate and benchmark
|
||||
Once the image has been downloaded you can follow these steps to
|
||||
run benchmarks and generate outputs.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
|
||||
|
||||
{% for model_group in model_groups %}
|
||||
{% for model in model_group.models %}
|
||||
@@ -180,7 +179,7 @@ Choose your setup method
|
||||
|
||||
You can either use an existing Hugging Face cache or download the model fresh inside the container.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
|
||||
|
||||
{% set docker = data.xdit_diffusion_inference.docker | selectattr("version", "equalto", "v25-11") | first %}
|
||||
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
|
||||
@@ -270,7 +269,7 @@ You can either use an existing Hugging Face cache or download the model fresh in
|
||||
Run inference
|
||||
=============
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
|
||||
|
||||
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
|
||||
{% for model_group in model_groups %}
|
||||
@@ -384,7 +383,5 @@ Run inference
|
||||
Previous versions
|
||||
=================
|
||||
|
||||
See
|
||||
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
|
||||
to find documentation for previous releases of xDiT diffusion inference
|
||||
performance testing.
|
||||
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
|
||||
releases of xDiT diffusion inference performance testing.
|
||||
@@ -6,20 +6,19 @@
|
||||
prebuilt and optimized docker images.
|
||||
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
|
||||
|
||||
************************
|
||||
xDiT diffusion inference
|
||||
************************
|
||||
******************************
|
||||
xDiT diffusion inference 25.12
|
||||
******************************
|
||||
|
||||
.. caution::
|
||||
|
||||
This documentation does not reflect the latest version of xDiT diffusion
|
||||
This documentation does not reflect the latest version of the xDiT diffusion
|
||||
inference performance documentation. See
|
||||
:doc:`/how-to/rocm-for-ai/inference/xdit-diffusion-inference` for the latest
|
||||
version.
|
||||
:doc:`/ai-inference/archive/xdit-history` for the latest version.
|
||||
|
||||
.. _xdit-video-diffusion-2512:
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -51,7 +50,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
|
||||
What's new
|
||||
==========
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -68,7 +67,7 @@ The following models are supported for inference performance benchmarking.
|
||||
Some instructions, commands, and recommendations in this documentation might
|
||||
vary by model -- select one to get started.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -133,7 +132,7 @@ system's configuration.
|
||||
Pull the Docker image
|
||||
=====================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -147,7 +146,7 @@ Pull the Docker image
|
||||
Validate and benchmark
|
||||
======================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -170,7 +169,7 @@ Choose your setup method
|
||||
|
||||
You can either use an existing Hugging Face cache or download the model fresh inside the container.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -262,7 +261,7 @@ You can either use an existing Hugging Face cache or download the model fresh in
|
||||
Run inference
|
||||
=============
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -406,7 +405,5 @@ Run inference
|
||||
Previous versions
|
||||
=================
|
||||
|
||||
See
|
||||
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
|
||||
to find documentation for previous releases of xDiT diffusion inference
|
||||
performance testing.
|
||||
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
|
||||
releases of xDiT diffusion inference performance testing.
|
||||
@@ -6,20 +6,19 @@
|
||||
prebuilt and optimized docker images.
|
||||
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
|
||||
|
||||
************************
|
||||
xDiT diffusion inference
|
||||
************************
|
||||
******************************
|
||||
xDiT diffusion inference 25.13
|
||||
******************************
|
||||
|
||||
.. caution::
|
||||
|
||||
This documentation does not reflect the latest version of the xDiT diffusion
|
||||
inference performance documentation. See
|
||||
:doc:`/how-to/rocm-for-ai/inference/xdit-diffusion-inference` for the latest
|
||||
version.
|
||||
:doc:`/ai-inference/archive/xdit-history` for the latest version.
|
||||
|
||||
.. _xdit-video-diffusion-2513:
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -52,7 +51,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
|
||||
What's new
|
||||
==========
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -69,7 +68,7 @@ The following models are supported for inference performance benchmarking.
|
||||
Some instructions, commands, and recommendations in this documentation might
|
||||
vary by model -- select one to get started.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -150,7 +149,7 @@ system's configuration.
|
||||
Pull the Docker image
|
||||
=====================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -164,7 +163,7 @@ Pull the Docker image
|
||||
Validate and benchmark
|
||||
======================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -187,7 +186,7 @@ Choose your setup method
|
||||
|
||||
You can either use an existing Hugging Face cache or download the model fresh inside the container.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -279,7 +278,7 @@ You can either use an existing Hugging Face cache or download the model fresh in
|
||||
Run inference
|
||||
=============
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -469,7 +468,5 @@ Run inference
|
||||
Previous versions
|
||||
=================
|
||||
|
||||
See
|
||||
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
|
||||
to find documentation for previous releases of xDiT diffusion inference
|
||||
performance testing.
|
||||
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
|
||||
releases of xDiT diffusion inference performance testing.
|
||||
@@ -6,20 +6,19 @@
|
||||
prebuilt and optimized docker images.
|
||||
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
|
||||
|
||||
************************
|
||||
xDiT diffusion inference
|
||||
************************
|
||||
*****************************
|
||||
xDiT diffusion inference 26.1
|
||||
*****************************
|
||||
|
||||
.. caution::
|
||||
|
||||
This documentation does not reflect the latest version of the xDiT diffusion
|
||||
inference performance documentation. See
|
||||
:doc:`/how-to/rocm-for-ai/inference/xdit-diffusion-inference` for the latest
|
||||
version.
|
||||
:doc:`/ai-inference/archive/xdit-history` for the latest version.
|
||||
|
||||
.. _xdit-video-diffusion-v261-v261:
|
||||
.. _xdit-video-diffusion-v261:
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -47,7 +46,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
|
||||
What's new
|
||||
==========
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -64,7 +63,7 @@ The following models are supported for inference performance benchmarking.
|
||||
Some instructions, commands, and recommendations in this documentation might
|
||||
vary by model -- select one to get started.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -129,7 +128,7 @@ system's configuration.
|
||||
Pull the Docker image
|
||||
=====================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -143,7 +142,7 @@ Pull the Docker image
|
||||
Validate and benchmark
|
||||
======================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -166,7 +165,7 @@ Choose your setup method
|
||||
|
||||
You can either use an existing Hugging Face cache or download the model fresh inside the container.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -258,7 +257,7 @@ You can either use an existing Hugging Face cache or download the model fresh in
|
||||
Run inference
|
||||
=============
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -317,7 +316,5 @@ Run inference
|
||||
Previous versions
|
||||
=================
|
||||
|
||||
See
|
||||
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
|
||||
to find documentation for previous releases of xDiT diffusion inference
|
||||
performance testing.
|
||||
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
|
||||
releases of xDiT diffusion inference performance testing.
|
||||
@@ -14,12 +14,11 @@ xDiT diffusion inference
|
||||
|
||||
This documentation does not reflect the latest version of the xDiT diffusion
|
||||
inference performance documentation. See
|
||||
:doc:`/how-to/rocm-for-ai/inference/xdit-diffusion-inference` for the latest
|
||||
version.
|
||||
:doc:`/ai-inference/archive/xdit-history` for the latest version.
|
||||
|
||||
.. _xdit-video-diffusion-262:
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -47,7 +46,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
|
||||
What's new
|
||||
==========
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -64,7 +63,7 @@ The following models are supported for inference performance benchmarking.
|
||||
Some instructions, commands, and recommendations in this documentation might
|
||||
vary by model -- select one to get started.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -129,7 +128,7 @@ system's configuration.
|
||||
Pull the Docker image
|
||||
=====================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -143,7 +142,7 @@ Pull the Docker image
|
||||
Validate and benchmark
|
||||
======================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -166,7 +165,7 @@ Choose your setup method
|
||||
|
||||
You can either use an existing Hugging Face cache or download the model fresh inside the container.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -258,7 +257,7 @@ You can either use an existing Hugging Face cache or download the model fresh in
|
||||
Run inference
|
||||
=============
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -315,7 +314,5 @@ Run inference
|
||||
Previous versions
|
||||
=================
|
||||
|
||||
See
|
||||
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
|
||||
to find documentation for previous releases of xDiT diffusion inference
|
||||
performance testing.
|
||||
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
|
||||
releases of xDiT diffusion inference performance testing.
|
||||
@@ -6,20 +6,19 @@
|
||||
prebuilt and optimized docker images.
|
||||
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
|
||||
|
||||
************************
|
||||
xDiT diffusion inference
|
||||
************************
|
||||
*****************************
|
||||
xDiT diffusion inference 26.3
|
||||
*****************************
|
||||
|
||||
.. caution::
|
||||
|
||||
This documentation does not reflect the latest version of the xDiT diffusion
|
||||
inference performance documentation. See
|
||||
:doc:`/how-to/rocm-for-ai/inference/xdit-diffusion-inference` for the latest
|
||||
version.
|
||||
:doc:`/ai-inference/archive/xdit-history` for the latest version.
|
||||
|
||||
.. _xdit-video-diffusion-263:
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -47,7 +46,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
|
||||
What's new
|
||||
==========
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -64,7 +63,7 @@ The following models are supported for inference performance benchmarking.
|
||||
Some instructions, commands, and recommendations in this documentation might
|
||||
vary by model -- select one to get started.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -129,7 +128,7 @@ system's configuration.
|
||||
Pull the Docker image
|
||||
=====================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -143,7 +142,7 @@ Pull the Docker image
|
||||
Validate and benchmark
|
||||
======================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -166,7 +165,7 @@ Choose your setup method
|
||||
|
||||
You can either use an existing Hugging Face cache or download the model fresh inside the container.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -258,7 +257,7 @@ You can either use an existing Hugging Face cache or download the model fresh in
|
||||
Run inference
|
||||
=============
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -315,7 +314,5 @@ Run inference
|
||||
Previous versions
|
||||
=================
|
||||
|
||||
See
|
||||
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
|
||||
to find documentation for previous releases of xDiT diffusion inference
|
||||
performance testing.
|
||||
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
|
||||
releases of xDiT diffusion inference performance testing.
|
||||
@@ -1,25 +1,26 @@
|
||||
:orphan:
|
||||
:no-search:
|
||||
:selector-toc2: Model
|
||||
:selector-toc2-icon: fa-solid fa-robot
|
||||
|
||||
.. meta::
|
||||
:description: Learn to validate diffusion model video generation on MI300X, MI350X and MI355X accelerators using
|
||||
prebuilt and optimized docker images.
|
||||
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
|
||||
|
||||
************************
|
||||
xDiT diffusion inference
|
||||
************************
|
||||
*****************************
|
||||
xDiT diffusion inference 26.4
|
||||
*****************************
|
||||
|
||||
.. caution::
|
||||
|
||||
This documentation does not reflect the latest version of the xDiT diffusion
|
||||
inference performance documentation. See
|
||||
:doc:`/how-to/rocm-for-ai/inference/xdit-diffusion-inference` for the latest
|
||||
version.
|
||||
:doc:`/ai-inference/archive/xdit-history` for the latest version.
|
||||
|
||||
.. _xdit-video-diffusion-264:
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.4-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -47,7 +48,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
|
||||
What's new
|
||||
==========
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.4-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -64,43 +65,37 @@ The following models are supported for inference performance benchmarking.
|
||||
Some instructions, commands, and recommendations in this documentation might
|
||||
vary by model -- select one to get started.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.4-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
.. raw:: html
|
||||
.. selector:: Model
|
||||
:key: model-group
|
||||
|
||||
<div id="vllm-benchmark-ud-params-picker" class="container-fluid">
|
||||
<div class="row gx-0">
|
||||
<div class="col-2 me-1 px-2 model-param-head">Model</div>
|
||||
<div class="row col-10 pe-0">
|
||||
{% for model_group in docker.supported_models %}
|
||||
<div class="col-6 px-2 model-param" data-param-k="model-group" data-param-v="{{ model_group.js_tag }}" tabindex="0">{{ model_group.group }}</div>
|
||||
{% endfor %}
|
||||
</div>
|
||||
</div>
|
||||
{% for model_group in docker.supported_models %}
|
||||
.. selector-option:: {{ model_group.group }}
|
||||
:value: {{ model_group.js_tag }}
|
||||
:width: 25%
|
||||
|
||||
<div class="row gx-0 pt-1">
|
||||
<div class="col-2 me-1 px-2 model-param-head">Variant</div>
|
||||
<div class="row col-10 pe-0">
|
||||
{% for model_group in docker.supported_models %}
|
||||
{% set models = model_group.models %}
|
||||
{% for model in models %}
|
||||
{% if models|length % 3 == 0 %}
|
||||
<div class="col-4 px-2 model-param" data-param-k="model" data-param-v="{{ model.js_tag }}" data-param-group="{{ model_group.js_tag }}" tabindex="0">{{ model.model }}</div>
|
||||
{% else %}
|
||||
<div class="col-6 px-2 model-param" data-param-k="model" data-param-v="{{ model.js_tag }}" data-param-group="{{ model_group.js_tag }}" tabindex="0">{{ model.model }}</div>
|
||||
{% endif %}
|
||||
{% endfor %}
|
||||
{% endfor %}
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
{% endfor %}
|
||||
|
||||
{% for model_group in docker.supported_models %}
|
||||
.. selector:: Variant
|
||||
:key: model
|
||||
:show-cond: model-group={{ model_group.js_tag }}
|
||||
|
||||
{% set models = model_group.models %}
|
||||
{% for model in models %}
|
||||
.. selector-option:: {{ model.model }}
|
||||
:value: {{ model.js_tag }}
|
||||
|
||||
{% endfor %}
|
||||
{% endfor %}
|
||||
|
||||
{% for model_group in docker.supported_models %}
|
||||
{% for model in model_group.models %}
|
||||
|
||||
.. container:: model-doc {{ model.js_tag }}
|
||||
.. selected:: model={{ model.js_tag }}
|
||||
|
||||
.. note::
|
||||
|
||||
@@ -129,7 +124,7 @@ system's configuration.
|
||||
Pull the Docker image
|
||||
=====================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.4-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -143,7 +138,7 @@ Pull the Docker image
|
||||
Validate and benchmark
|
||||
======================
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.4-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
@@ -153,7 +148,7 @@ Validate and benchmark
|
||||
{% for model_group in docker.supported_models %}
|
||||
{% for model in model_group.models %}
|
||||
|
||||
.. container:: model-doc {{model.js_tag}}
|
||||
.. selected:: model={{ model.js_tag }}
|
||||
|
||||
The following commands are written for {{ model.model }}.
|
||||
See :ref:`xdit-video-diffusion-supported-models-264` to switch to another available model.
|
||||
@@ -166,13 +161,13 @@ Choose your setup method
|
||||
|
||||
You can either use an existing Hugging Face cache or download the model fresh inside the container.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.4-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
{% for model_group in docker.supported_models %}
|
||||
{% for model in model_group.models %}
|
||||
.. container:: model-doc {{model.js_tag}}
|
||||
.. selected:: model={{model.js_tag}}
|
||||
|
||||
.. tab-set::
|
||||
|
||||
@@ -258,14 +253,14 @@ You can either use an existing Hugging Face cache or download the model fresh in
|
||||
Run inference
|
||||
=============
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.4-inference-models.yaml
|
||||
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
|
||||
|
||||
{% set docker = data.docker %}
|
||||
|
||||
{% for model_group in docker.supported_models %}
|
||||
{% for model in model_group.models %}
|
||||
|
||||
.. container:: model-doc {{ model.js_tag }}
|
||||
.. selected:: model={{ model.js_tag }}
|
||||
|
||||
.. tab-set::
|
||||
|
||||
@@ -315,7 +310,5 @@ Run inference
|
||||
Previous versions
|
||||
=================
|
||||
|
||||
See
|
||||
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
|
||||
to find documentation for previous releases of xDiT diffusion inference
|
||||
performance testing.
|
||||
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
|
||||
releases of xDiT diffusion inference performance testing.
|
||||
@@ -1,8 +1,8 @@
|
||||
:orphan:
|
||||
|
||||
************************************************************
|
||||
xDiT diffusion inference performance testing version history
|
||||
************************************************************
|
||||
****************************************
|
||||
xDiT diffusion inference version history
|
||||
****************************************
|
||||
|
||||
This table lists previous versions of the ROCm xDiT diffusion inference performance
|
||||
testing environment. For detailed information about available models for
|
||||
@@ -20,7 +20,7 @@ benchmarking, see the version-specific documentation.
|
||||
* ROCm 7.13.0
|
||||
* TheRock cbff3d1
|
||||
-
|
||||
* :doc:`Documentation </how-to/rocm-for-ai/inference/xdit-diffusion-inference>`
|
||||
* :doc:`Documentation <../xdit>`
|
||||
* `Docker Hub <https://hub.docker.com/layers/rocm/pytorch-xdit/v26.5/images/sha256-b8ad9fd4b41bc116ac2aff07c1066bf369cf7fc110b1a323f6302191985a51fd>`__
|
||||
|
||||
* - ``rocm/pytorch-xdit:v26.4``
|
||||
@@ -0,0 +1,124 @@
|
||||
********************************
|
||||
ComfyUI image generation on ROCm
|
||||
********************************
|
||||
|
||||
`ComfyUI <https://github.com/comfyanonymous/ComfyUI>`__ is an open-source,
|
||||
node-based interface for building and running image generation workflows with
|
||||
diffusion models such as Stable Diffusion. Its modular graph-based design lets
|
||||
you construct, customize, and share complex pipelines without writing code. This
|
||||
page walks through installing and running ComfyUI on AMD GPUs.
|
||||
|
||||
Prerequisites
|
||||
=============
|
||||
|
||||
Ensure your working environment is running ROCm-enabled PyTorch on
|
||||
a :ref:`supported system <compat-matrix>`. See :ref:`pytorch-install` for
|
||||
instructions.
|
||||
|
||||
.. important::
|
||||
|
||||
On Windows, ComfyUI might not start if Smart App Control is enabled in your
|
||||
Windows security settings.
|
||||
|
||||
Installation
|
||||
============
|
||||
|
||||
After installing ROCm and PyTorch in your Python environment, follow these
|
||||
steps to install ComfyUI.
|
||||
|
||||
1. Clone the ComfyUI repository.
|
||||
|
||||
.. code-block:: shell
|
||||
|
||||
git clone https://github.com/comfyanonymous/ComfyUI.git
|
||||
|
||||
2. Activate your Python virtual environment and install dependencies.
|
||||
|
||||
.. tab-set::
|
||||
|
||||
.. tab-item:: Linux
|
||||
:sync: linux
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
pip install -r ComfyUI/requirements.txt
|
||||
|
||||
.. tab-item:: Windows
|
||||
:sync: windows
|
||||
|
||||
.. code-block:: bat
|
||||
|
||||
pip install -r ComfyUI\requirements.txt
|
||||
|
||||
Run ComfyUI
|
||||
===========
|
||||
|
||||
Use the following steps for a simple example of running ComfyUI.
|
||||
|
||||
1. Start the ComfyUI server from the command line.
|
||||
|
||||
.. tab-set::
|
||||
|
||||
.. tab-item:: Linux
|
||||
:sync: linux
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
python ComfyUI/main.py
|
||||
|
||||
.. tab-item:: Windows
|
||||
:sync: windows
|
||||
|
||||
.. code-block:: bat
|
||||
|
||||
python ComfyUI\main.py
|
||||
|
||||
This starts the server, displaying a prompt like:
|
||||
|
||||
.. code-block:: text
|
||||
|
||||
To see the GUI go to: http://127.0.0.1:8188
|
||||
|
||||
2. Go to ``http://127.0.0.1:8188`` in your web browser. You might need to
|
||||
replace ``8188`` with the appropriate port.
|
||||
|
||||
.. image:: ./images/comfyui/comfyui-main.png
|
||||
:align: center
|
||||
|
||||
3. Search for one of the following templates and download any missing
|
||||
models.
|
||||
|
||||
.. tab-set::
|
||||
|
||||
.. tab-item:: SD3.5 Simple
|
||||
|
||||
Select **Template** → **Model Filter** → **SD3.5** → **SD3.5 Simple**
|
||||
|
||||
.. image:: ./images/comfyui/sd3_5-simple-card.png
|
||||
:align: center
|
||||
|
||||
Download required models, if missing.
|
||||
|
||||
.. image:: ./images/comfyui/sd3_5-missing-models.png
|
||||
:align: center
|
||||
|
||||
.. tab-item:: Chroma1 Radiance text to image
|
||||
|
||||
Select **Template** → **Model Filter** → **Chroma** → **Chroma1 Radiance text to image**
|
||||
|
||||
.. image:: ./images/comfyui/chroma1-radiance-tti-card.png
|
||||
:align: center
|
||||
|
||||
Download required models, if missing.
|
||||
|
||||
.. image:: ./images/comfyui/chroma1-radiance-tti-missing-models.png
|
||||
:align: center
|
||||
|
||||
4. Click the **Run** button.
|
||||
|
||||
The application will use your AMD GPU to convert the prompted text to an image.
|
||||
|
||||
.. seealso::
|
||||
|
||||
To learn more about the ComfyUI interface and workflows, see the `ComfyUI
|
||||
documentation <https://docs.comfy.org/development/core-concepts/workflow>`__.
|
||||
|
Before Width: | Height: | Size: 43 KiB After Width: | Height: | Size: 43 KiB |
|
Before Width: | Height: | Size: 44 KiB After Width: | Height: | Size: 44 KiB |
|
Before Width: | Height: | Size: 28 KiB After Width: | Height: | Size: 28 KiB |
|
Before Width: | Height: | Size: 112 KiB After Width: | Height: | Size: 112 KiB |
|
Before Width: | Height: | Size: 188 KiB After Width: | Height: | Size: 188 KiB |
|
Before Width: | Height: | Size: 138 KiB After Width: | Height: | Size: 138 KiB |
|
Before Width: | Height: | Size: 62 KiB After Width: | Height: | Size: 62 KiB |
|
Before Width: | Height: | Size: 27 KiB After Width: | Height: | Size: 27 KiB |
|
Before Width: | Height: | Size: 86 KiB After Width: | Height: | Size: 86 KiB |
|
Before Width: | Height: | Size: 49 KiB After Width: | Height: | Size: 49 KiB |
|
Before Width: | Height: | Size: 30 KiB After Width: | Height: | Size: 30 KiB |
|
Before Width: | Height: | Size: 129 KiB After Width: | Height: | Size: 129 KiB |
|
Before Width: | Height: | Size: 80 KiB After Width: | Height: | Size: 80 KiB |
|
Before Width: | Height: | Size: 153 KiB After Width: | Height: | Size: 153 KiB |
|
Before Width: | Height: | Size: 219 KiB After Width: | Height: | Size: 219 KiB |
|
Before Width: | Height: | Size: 310 KiB After Width: | Height: | Size: 310 KiB |
|
Before Width: | Height: | Size: 342 KiB After Width: | Height: | Size: 342 KiB |
@@ -2,35 +2,35 @@
|
||||
:description: How to Use ROCm for AI inference optimization
|
||||
:keywords: ROCm, LLM, AI inference, Optimization, GPUs, usage, tutorial
|
||||
|
||||
*******************************************
|
||||
Use ROCm for AI inference optimization
|
||||
*******************************************
|
||||
**********************
|
||||
Inference optimization
|
||||
**********************
|
||||
|
||||
AI inference optimization is the process of improving the performance of machine learning models and speeding up the inference process. It includes:
|
||||
|
||||
- **Quantization**: This involves reducing the precision of model weights and activations while maintaining acceptable accuracy levels. Reduced precision improves inference efficiency because lower precision data requires less storage and better utilizes the hardware's computation power.
|
||||
- **Quantization**: This involves reducing the precision of model weights and activations while maintaining acceptable accuracy levels. Reduced precision improves inference efficiency because lower precision data requires less storage and better utilizes the hardware's computation power.
|
||||
|
||||
- **Kernel optimization**: This technique involves optimizing computation kernels to exploit the underlying hardware capabilities. For example, the kernels can be optimized to use multiple GPU cores or utilize specialized hardware like tensor cores to accelerate the computations.
|
||||
- **Kernel optimization**: This technique involves optimizing computation kernels to exploit the underlying hardware capabilities. For example, the kernels can be optimized to use multiple GPU cores or utilize specialized hardware like tensor cores to accelerate the computations.
|
||||
|
||||
- **Libraries**: Libraries such as Flash Attention, xFormers, and PyTorch TunableOp are used to accelerate deep learning models and improve the performance of inference workloads.
|
||||
- **Libraries**: Libraries such as Flash Attention, xFormers, and PyTorch TunableOp are used to accelerate deep learning models and improve the performance of inference workloads.
|
||||
|
||||
- **Hardware acceleration**: Hardware acceleration techniques, like GPUs for AI inference, can significantly improve performance due to their parallel processing capabilities.
|
||||
- **Hardware acceleration**: Hardware acceleration techniques, like GPUs for AI inference, can significantly improve performance due to their parallel processing capabilities.
|
||||
|
||||
- **Pruning**: This involves removing unnecessary connections, layers, or weights from a pre-trained model while maintaining acceptable accuracy levels, resulting in a smaller model that requires fewer computational resources to run inference.
|
||||
- **Pruning**: This involves removing unnecessary connections, layers, or weights from a pre-trained model while maintaining acceptable accuracy levels, resulting in a smaller model that requires fewer computational resources to run inference.
|
||||
|
||||
Utilizing these optimization techniques with the ROCm™ software platform can significantly reduce inference time, improve performance, and reduce the cost of your AI applications.
|
||||
Utilizing these optimization techniques with the ROCm™ software platform can significantly reduce inference time, improve performance, and reduce the cost of your AI applications.
|
||||
|
||||
Throughout the following topics, this guide discusses optimization techniques for inference workloads.
|
||||
|
||||
- :doc:`Model quantization <model-quantization>`
|
||||
|
||||
- :doc:`Model acceleration libraries <model-acceleration-libraries>`
|
||||
- :doc:`Model acceleration libraries <model-acceleration-libs>`
|
||||
|
||||
- :doc:`Optimizing with Composable Kernel <optimizing-with-composable-kernel>`
|
||||
- :doc:`Optimizing with Composable Kernel <optimize-with-composable-kernel>`
|
||||
|
||||
- :doc:`Optimizing Triton kernels <optimizing-triton-kernel>`
|
||||
- :doc:`Optimizing Triton kernels <optimize-triton-kernels>`
|
||||
|
||||
- :doc:`Profiling and debugging <profiling-and-debugging>`
|
||||
- :doc:`Workload tuning <workload-optimization>`
|
||||
|
||||
- :doc:`Workload tuning <workload>`
|
||||
- :ref:`Profiling and debugging <mi300x-profiling-tools>`
|
||||
|
||||
@@ -7,9 +7,7 @@ LLM inference frameworks
|
||||
************************
|
||||
|
||||
This section discusses how to implement `vLLM <https://docs.vllm.ai/en/latest>`_ and `Hugging Face TGI
|
||||
<https://huggingface.co/docs/text-generation-inference/en/index>`_ using
|
||||
:doc:`single-accelerator <../fine-tuning/single-gpu-fine-tuning-and-inference>` and
|
||||
:doc:`multi-accelerator <../fine-tuning/multi-gpu-fine-tuning-and-inference>` systems.
|
||||
<https://huggingface.co/docs/text-generation-inference/en/index>`_.
|
||||
|
||||
.. _fine-tuning-llms-vllm:
|
||||
|
||||
@@ -68,7 +66,7 @@ Installing vLLM
|
||||
|
||||
The following log message is displayed in your command line indicates that the server is listening for requests.
|
||||
|
||||
.. image:: ../../../data/how-to/llm-fine-tuning-optimization/vllm-single-gpu-log.png
|
||||
.. image:: ./images/llm-inference-frameworks/vllm-single-gpu-log.png
|
||||
:alt: vLLM API server log message
|
||||
:align: center
|
||||
|
||||
@@ -141,7 +139,7 @@ Installing vLLM
|
||||
|
||||
ROCm provides a prebuilt optimized Docker image for validating the performance of LLM inference with vLLM
|
||||
on the MI300X GPU. The Docker image includes ROCm, vLLM, and PyTorch.
|
||||
For more information, see :doc:`/how-to/rocm-for-ai/inference/benchmark-docker/vllm`.
|
||||
For more information, see :doc:`/ai-inference/vllm`.
|
||||
|
||||
.. _fine-tuning-llms-tgi:
|
||||
|
||||
@@ -20,7 +20,7 @@ Attention (GQA), and Multi-Query Attention (MQA). This reduction in memory movem
|
||||
time-to-first-token (TTFT) latency for large batch sizes and long prompt sequences, thereby enhancing overall
|
||||
performance.
|
||||
|
||||
.. image:: ../../../data/how-to/llm-fine-tuning-optimization/attention-module.png
|
||||
.. image:: ./images/model-acceleration-libs/attention-module.png
|
||||
:alt: Attention module of a large language module utilizing tiling
|
||||
:align: center
|
||||
|
||||
@@ -36,7 +36,7 @@ These can be installed by following the official
|
||||
`PyTorch installation guide <https://pytorch.org/get-started/locally/>`_. Alternatively, for a simpler setup, you can use a preconfigured
|
||||
:ref:`ROCm PyTorch Docker image <using-docker-with-pytorch-pre-installed>`, which already includes the required libraries.
|
||||
|
||||
Installing Flash Attention 2
|
||||
Installing Flash Attention 2
|
||||
----------------------------
|
||||
|
||||
`Flash Attention <https://github.com/Dao-AILab/flash-attention>`_ supports two backend implementations on AMD GPUs.
|
||||
@@ -61,29 +61,29 @@ To install Flash Attention 2, use the following commands:
|
||||
pip install ninja
|
||||
|
||||
# To install the CK backend flash attention
|
||||
python setup.py install
|
||||
python setup.py install
|
||||
|
||||
# To install the Triton backend flash attention
|
||||
FLASH_ATTENTION_TRITON_AMD_ENABLE="TRUE" python setup.py install
|
||||
FLASH_ATTENTION_TRITON_AMD_ENABLE="TRUE" python setup.py install
|
||||
|
||||
# To install both CK and Triton backend flash attention
|
||||
FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE && FLASH_ATTENTION_SKIP_CK_BUILD=FALSE python setup.py install
|
||||
|
||||
For detailed installation instructions, see `Flash Attention <https://github.com/Dao-AILab/flash-attention>`_.
|
||||
|
||||
Benchmarking Flash Attention 2
|
||||
Benchmarking Flash Attention 2
|
||||
------------------------------
|
||||
|
||||
Benchmark scripts to evaluate the performance of Flash Attention 2 are stored in the ``flash-attention/benchmarks/`` directory.
|
||||
|
||||
To benchmark the CK backend
|
||||
To benchmark the CK backend
|
||||
|
||||
.. code-block:: shell
|
||||
|
||||
cd flash-attention/benchmarks
|
||||
pip install transformers einops ninja
|
||||
|
||||
python3 benchmark_flash_attention.py
|
||||
python3 benchmark_flash_attention.py
|
||||
|
||||
To benchmark the Triton backend
|
||||
|
||||
@@ -91,7 +91,7 @@ To benchmark the Triton backend
|
||||
|
||||
FLASH_ATTENTION_TRITON_AMD_ENABLE="TRUE" python3 benchmark_flash_attention.py
|
||||
|
||||
Using Flash Attention 2
|
||||
Using Flash Attention 2
|
||||
-----------------------
|
||||
|
||||
.. code-block:: python
|
||||
@@ -128,13 +128,13 @@ xFormers also improves the performance of attention modules. Although xFormers a
|
||||
similarly to Flash Attention 2 due to its tiling behavior of query, key, and value, it’s widely used for LLMs and
|
||||
Stable Diffusion models with the Hugging Face Diffusers library.
|
||||
|
||||
Installing CK xFormers
|
||||
Installing CK xFormers
|
||||
----------------------
|
||||
|
||||
Use the following commands to install CK xFormers.
|
||||
|
||||
.. code-block:: shell
|
||||
|
||||
|
||||
# Install from source
|
||||
git clone https://github.com/ROCm/xformers.git
|
||||
cd xformers/
|
||||
@@ -175,20 +175,20 @@ of the PyTorch compilation.
|
||||
os.environ["TOKENIZERS_PARALLELISM"] = "false"
|
||||
model_name = "NousResearch/Meta-Llama-3-8B"
|
||||
prompts = []
|
||||
|
||||
|
||||
for b in range(1):
|
||||
prompts.append("New york city is where "
|
||||
)
|
||||
|
||||
|
||||
tokenizer = AutoTokenizer.from_pretrained(model_name)
|
||||
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.float16).to(device).eval()
|
||||
inputs = tokenizer(prompts, return_tensors="pt").to(model.device)
|
||||
|
||||
|
||||
def decode_one_tokens(model, cur_token, input_pos, cache_position):
|
||||
logits = model(cur_token, position_ids=input_pos, cache_position=cache_position, return_dict=False, use_cache=True)[0]
|
||||
new_token = torch.argmax(logits[:, -1], dim=-1)[:, None]
|
||||
return new_token
|
||||
|
||||
|
||||
batch_size, seq_length = inputs["input_ids"].shape
|
||||
|
||||
# Static key-value cache
|
||||
@@ -198,16 +198,16 @@ of the PyTorch compilation.
|
||||
cache_position = torch.arange(seq_length, device=device)
|
||||
generated_ids = torch.zeros(batch_size, seq_length + max_new_tokens + 1, dtype=torch.int, device=device)
|
||||
generated_ids[:, cache_position] = inputs["input_ids"].to(device).to(torch.int)
|
||||
|
||||
|
||||
logits = model(**inputs, cache_position=cache_position, return_dict=False, use_cache=True)[0]
|
||||
next_token = torch.argmax(logits[:, -1], dim=-1)[:, None]
|
||||
|
||||
# torch compilation
|
||||
decode_one_tokens = torch.compile(decode_one_tokens, mode="max-autotune-no-cudagraphs",fullgraph=True)
|
||||
|
||||
|
||||
generated_ids[:, seq_length] = next_token[:, 0]
|
||||
cache_position = torch.tensor([seq_length + 1], device=device)
|
||||
|
||||
|
||||
with torch.no_grad():
|
||||
for _ in range(1, max_new_tokens):
|
||||
with torch.backends.cuda.sdp_kernel(enable_flash=False, enable_mem_efficient=False, enable_math=True):
|
||||
@@ -235,7 +235,7 @@ page describes the options.
|
||||
|
||||
# To turn on TunableOp, simply set this environment variable
|
||||
export PYTORCH_TUNABLEOP_ENABLED=1
|
||||
|
||||
|
||||
# Python
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
@@ -244,7 +244,7 @@ page describes the options.
|
||||
W = torch.rand(200, 20, device="cuda")
|
||||
Out = F.linear(A, W)
|
||||
print(Out.size())
|
||||
|
||||
|
||||
# tunableop_results0.csv
|
||||
Validator,PT_VERSION,2.4.0
|
||||
Validator,ROCM_VERSION,6.1.0.0-82-5fabb4c
|
||||
@@ -253,7 +253,7 @@ page describes the options.
|
||||
Validator,ROCBLAS_VERSION,4.1.0-cefa4a9b-dirty
|
||||
GemmTunableOp_float_TN,tn_200_100_20,Gemm_Rocblas_32323,0.00669595
|
||||
|
||||
.. image:: ../../../data/how-to/llm-fine-tuning-optimization/tunableop.png
|
||||
.. image:: ./images/model-acceleration-libs/tunableop.png
|
||||
:alt: GEMM and TunableOp
|
||||
:align: center
|
||||
|
||||
@@ -270,7 +270,7 @@ and as a back end for PyTorch quantized operators. FBGEMM offers optimized on-CP
|
||||
strong performance on native tensor formats, and the ability to generate
|
||||
high-performance shape- and size-specific kernels at runtime.
|
||||
|
||||
FBGEMM_GPU collects several high-performance PyTorch GPU operator libraries
|
||||
FBGEMM_GPU collects several high-performance PyTorch GPU operator libraries
|
||||
for use in training and inference. It provides efficient table-batched embedding functionality,
|
||||
data layout transformation, and quantization support.
|
||||
|
||||
@@ -288,7 +288,7 @@ Installing FBGEMM_GPU consists of the following steps:
|
||||
* Install ROCm using Docker or the :doc:`package manager <rocm-install-on-linux:install/install-methods/package-manager-index>`
|
||||
* Install the nightly `PyTorch <https://pytorch.org/>`_ build
|
||||
* Complete the pre-build and build tasks
|
||||
|
||||
|
||||
.. note::
|
||||
|
||||
FBGEMM_GPU doesn't require the installation of FBGEMM. To optionally install
|
||||
@@ -375,7 +375,7 @@ and run the ROCm Docker image, use this command:
|
||||
|
||||
You can also install ROCm using the package manager. FBGEMM_GPU requires the installation of the full ROCm package.
|
||||
For more information, see :doc:`the ROCm installation guide <rocm-install-on-linux:install/detailed-install>`.
|
||||
The ROCm package also requires the :doc:`MIOpen <miopen:index>` component as a dependency.
|
||||
The ROCm package also requires the :doc:`MIOpen <miopen:index>` component as a dependency.
|
||||
To install MIOpen, use the ``apt install`` command.
|
||||
|
||||
.. code-block:: shell
|
||||
@@ -407,7 +407,7 @@ Install `PyTorch <https://pytorch.org/>`_ using ``pip`` for the most reliable an
|
||||
Perform the prebuild and build
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
#. Clone the FBGEMM repository and the relevant submodules. Use ``pip`` to install the
|
||||
#. Clone the FBGEMM repository and the relevant submodules. Use ``pip`` to install the
|
||||
components in ``requirements.txt``. Run the following commands inside the Miniconda environment.
|
||||
|
||||
.. code-block:: shell
|
||||
@@ -447,7 +447,7 @@ Perform the prebuild and build
|
||||
# Set the Python platform name for the Linux case
|
||||
export python_plat_name="manylinux2014_${ARCH}"
|
||||
|
||||
#. Build FBGEMM_GPU for the ROCm platform. Set ``ROCM_PATH`` to the path to your ROCm installation.
|
||||
#. Build FBGEMM_GPU for the ROCm platform. Set ``ROCM_PATH`` to the path to your ROCm installation.
|
||||
Run these commands from the ``fbgemm_gpu/`` directory inside the Miniconda environment.
|
||||
|
||||
.. code-block:: shell
|
||||
@@ -474,7 +474,7 @@ Perform the prebuild and build
|
||||
--package_variant=rocm \
|
||||
-DHIP_ROOT_DIR="${ROCM_PATH}" \
|
||||
-DCMAKE_C_FLAGS="-DTORCH_USE_HIP_DSA" \
|
||||
-DCMAKE_CXX_FLAGS="-DTORCH_USE_HIP_DSA"
|
||||
-DCMAKE_CXX_FLAGS="-DTORCH_USE_HIP_DSA"
|
||||
|
||||
Post-build validation
|
||||
----------------------
|
||||
@@ -533,8 +533,8 @@ follow these instructions:
|
||||
# Run the test
|
||||
python -m pytest -v -rsx -s -W ignore::pytest.PytestCollectionWarning split_table_batched_embeddings_test.py
|
||||
|
||||
To run the FBGEMM_GPU ``uvm`` test, use these commands. These tests only support the AMD MI210 and
|
||||
more recent GPUs.
|
||||
To run the FBGEMM_GPU ``uvm`` test, use these commands. These tests only support the AMD MI210 and
|
||||
more recent GPUs.
|
||||
|
||||
.. code-block:: shell
|
||||
|
||||
@@ -6,14 +6,14 @@
|
||||
Optimizing Triton kernels
|
||||
*************************
|
||||
|
||||
This section introduces the general steps for
|
||||
This section introduces the general steps for
|
||||
`Triton <https://openai.com/index/triton/>`_ kernel optimization. Broadly,
|
||||
Triton kernel optimization is similar to :doc:`HIP <hip:how-to/performance_guidelines>`
|
||||
and CUDA kernel optimization.
|
||||
|
||||
Refer to the
|
||||
:ref:`Triton kernel performance optimization <mi300x-triton-kernel-performance-optimization>`
|
||||
section of the :doc:`workload` guide
|
||||
section of the :doc:`workload-optimization` guide
|
||||
for detailed information.
|
||||
|
||||
Triton kernel performance optimization includes the following topics.
|
||||
@@ -29,11 +29,11 @@ The template parameters of the instance are grouped into four parameter types:
|
||||
- [Parameters for determining extra operations on matrix elements](matrix-element-operation)
|
||||
- [Performance-oriented tunable parameters](tunable-parameters)
|
||||
|
||||
<!--
|
||||
<!--
|
||||
================
|
||||
### Figure 2
|
||||
================ -->
|
||||
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-template_parameters.jpg
|
||||
```{figure} ./images/optimize-with-composable-kernel/ck-template_parameters.jpg
|
||||
The template parameters of the selected GEMM kernel are classified into four groups. These template parameter groups should be defined properly before running the instance.
|
||||
```
|
||||
|
||||
@@ -100,7 +100,7 @@ struct AddRelu
|
||||
|
||||
(tunable-parameters)=
|
||||
|
||||
#### Tunable parameters
|
||||
#### Tunable parameters
|
||||
|
||||
The CK instance includes a series of tunable template parameters to control the parallel granularity of the workload to achieve load balancing on different hardware platforms.
|
||||
|
||||
@@ -123,11 +123,11 @@ After determining the template parameters, we instantiate the kernel with actual
|
||||
|
||||
The row and column, and stride information of input matrices are also passed to the instance. For batched GEMM, you must pass in additional batch count and batch stride values. The extra operations for pre and post-processing are also passed with an actual argument; for example, α and β for GEMM scaling operations. Afterward, the instantiated kernel is launched by the invoker, as illustrated in Figure 3.
|
||||
|
||||
<!--
|
||||
<!--
|
||||
================
|
||||
### Figure 3
|
||||
================ -->
|
||||
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-kernel_launch.jpg
|
||||
```{figure} ./images/optimize-with-composable-kernel/ck-kernel_launch.jpg
|
||||
Templated kernel launching consists of kernel instantiation, making arguments by passing in actual application parameters, creating an invoker, and running the instance through the invoker.
|
||||
```
|
||||
|
||||
@@ -152,11 +152,11 @@ The following section discusses the analysis of the operation flow of `Linear_Re
|
||||
|
||||
The first operation in the process is to perform the multiplication of input matrices A and B. The resulting matrix C is then scaled with α to obtain T1. At the same time, the process performs a scaling operation on D elements to obtain T2. Afterward, the process performs matrix addition between T1 and T2, element activation calculation using ReLU, and element rounding sequentially. The operations to generate E1, E2, and E are encapsulated and completed by a user-defined template function in CK (given in the next sub-section). This template function is integrated into the fundamental instance directly during the compilation phase so that all these steps can be fused in a single GPU kernel.
|
||||
|
||||
<!--
|
||||
<!--
|
||||
================
|
||||
### Figure 4
|
||||
================ -->
|
||||
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-operation_flow.jpg
|
||||
```{figure} ./images/optimize-with-composable-kernel/ck-operation_flow.jpg
|
||||
Operation flow.
|
||||
```
|
||||
|
||||
@@ -168,11 +168,11 @@ Third, consider the platform for implementing CK instances. The instances suffix
|
||||
|
||||
Here, we use [DeviceBatchedGemmMultiD_Xdl](https://github.com/ROCm/composable_kernel/tree/develop/example/24_batched_gemm) as the fundamental instance to implement the functionalities in the previous table.
|
||||
|
||||
<!--
|
||||
<!--
|
||||
================
|
||||
### Figure 5
|
||||
================ -->
|
||||
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-root_instance.jpg
|
||||
```{figure} ./images/optimize-with-composable-kernel/ck-root_instance.jpg
|
||||
Use the ‘DeviceBatchedGemmMultiD_Xdl’ instance as a root.
|
||||
```
|
||||
|
||||
@@ -189,7 +189,7 @@ The inference of SQ quantized models relies on using PyTorch and Transformer lib
|
||||
In GEMM, the A and B inputs are two-dimensional matrices, and the required input matrices of the selected fundamental CK instance are three-dimensional matrices. Therefore, we must convert the input 2-D tensors to 3-D tensors, by using `tensor`'s `unsqueeze()` method before passing these matrices to the instance. For batched GEMM in the preceding table, ignore this step.
|
||||
|
||||
```c++
|
||||
// Function input and output
|
||||
// Function input and output
|
||||
torch::Tensor linear_relu_abde_i8(
|
||||
torch::Tensor A_,
|
||||
torch::Tensor B_,
|
||||
@@ -197,10 +197,10 @@ torch::Tensor linear_relu_abde_i8(
|
||||
float alpha,
|
||||
float beta)
|
||||
{
|
||||
// Convert torch::Tensor A_ (M, K) to torch::Tensor A (1, M, K)
|
||||
// Convert torch::Tensor A_ (M, K) to torch::Tensor A (1, M, K)
|
||||
auto A = A_.unsqueeze(0);
|
||||
|
||||
// Convert torch::Tensor B_ (K, N) to torch::Tensor A (1, K, N)
|
||||
// Convert torch::Tensor B_ (K, N) to torch::Tensor A (1, K, N)
|
||||
auto B = B_.unsqueeze(0);
|
||||
...
|
||||
```
|
||||
@@ -232,7 +232,7 @@ As shown in the following code block, we obtain M, N, and K values using input t
|
||||
auto D = D_.view({1,-1}).repeat({M, 1});
|
||||
|
||||
// Allocate memory for E
|
||||
auto E = torch::empty({batch_count, M, N},
|
||||
auto E = torch::empty({batch_count, M, N},
|
||||
torch::dtype(torch::kInt8).device(A.device()));
|
||||
```
|
||||
|
||||
@@ -241,7 +241,7 @@ In the following code block, `ADataType`, `BDataType` and `D0DataType` are used
|
||||
`AccDataType` determines the data precision used to represent the multiply-add results of A and B elements. Generally, a larger range data type is applied to store the multiply-add results of A and B to avoid result overflow; `I32` is applied in this case. The `CShuffleDataType I32` data type indicates that the multiply-add results continue to be stored in LDS as an `I32` data format. All of this is implemented through the following code block.
|
||||
|
||||
```c++
|
||||
// Data precision
|
||||
// Data precision
|
||||
using ADataType = I8;
|
||||
using BDataType = I8;
|
||||
using AccDataType = I32;
|
||||
@@ -265,7 +265,7 @@ Following the convention of various linear algebra libraries, row-major and colu
|
||||
In CK, `PassThrough` is a struct denoting if an operation is applied to the tensor it binds to. To fuse the operations between E1, E2, and E introduced in section [Operation flow analysis](#operation-flow-analysis), we define a custom C++ struct, `ScaleScaleAddRelu`, and bind it to `CDEELementOp`. It determines the operations that will be applied to `CShuffle` (A×B results), tensor D, α, and β.
|
||||
|
||||
```c++
|
||||
// No operations bound to the elements of A and B
|
||||
// No operations bound to the elements of A and B
|
||||
using AElementOp = PassThrough;
|
||||
using BElementOp = PassThrough;
|
||||
|
||||
@@ -290,17 +290,17 @@ struct ScaleScaleAddRelu {
|
||||
|
||||
// Perform addition operation
|
||||
F32 temp = c_scale + d_scale;
|
||||
|
||||
|
||||
// Perform RELU operation
|
||||
temp = temp > 0 ? temp : 0;
|
||||
|
||||
// Perform rounding operation
|
||||
// Perform rounding operation
|
||||
temp = temp > 127 ? 127 : temp;
|
||||
|
||||
|
||||
// Return to E
|
||||
e = ck::type_convert<I8>(temp);
|
||||
}
|
||||
|
||||
|
||||
F32 alpha;
|
||||
F32 beta;
|
||||
};
|
||||
@@ -315,16 +315,16 @@ static constexpr auto GemmDefault = ck::tensor_operation::device::GemmSpecializa
|
||||
The template parameters of the target fundamental instance are initialized with the above parameters and includes default tunable parameters. For specific tuning methods, see [Tunable parameters](#tunable-parameters).
|
||||
|
||||
```c++
|
||||
using DeviceOpInstance = ck::tensor_operation::device::DeviceBatchedGemmMultiD_Xdl<
|
||||
using DeviceOpInstance = ck::tensor_operation::device::DeviceBatchedGemmMultiD_Xdl<
|
||||
// Tensor layout
|
||||
ALayout, BLayout, DsLayout, ELayout,
|
||||
ALayout, BLayout, DsLayout, ELayout,
|
||||
// Tensor data type
|
||||
ADataType, BDataType, AccDataType, CShuffleDataType, DsDataType, EDataType,
|
||||
ADataType, BDataType, AccDataType, CShuffleDataType, DsDataType, EDataType,
|
||||
// Tensor operation
|
||||
AElementOp, BElementOp, CDEElementOp,
|
||||
// Padding strategy
|
||||
AElementOp, BElementOp, CDEElementOp,
|
||||
// Padding strategy
|
||||
GemmDefault,
|
||||
// Tunable parameters
|
||||
// Tunable parameters
|
||||
tunable parameters>;
|
||||
```
|
||||
|
||||
@@ -356,13 +356,13 @@ invoker.Run(argument, StreamConfig{nullptr, 0});
|
||||
The output of the fundamental instance is a calculated batched matrix E (batch, M, N). Before the return, it needs to be converted to a 2-D matrix if a normal GEMM result is required.
|
||||
|
||||
```c++
|
||||
// Convert (1, M, N) to (M, N)
|
||||
// Convert (1, M, N) to (M, N)
|
||||
return E.squeeze(0);
|
||||
```
|
||||
|
||||
### Binding to Python
|
||||
|
||||
Since these functions are written in C++ and `torch::Tensor`, you can use `pybind11` to bind the functions and import them as Python modules. For the example, the necessary binding code for exposing the functions in the table spans but a few lines.
|
||||
Since these functions are written in C++ and `torch::Tensor`, you can use `pybind11` to bind the functions and import them as Python modules. For the example, the necessary binding code for exposing the functions in the table spans but a few lines.
|
||||
|
||||
```c++
|
||||
#include <torch/extension.h>
|
||||
@@ -390,7 +390,7 @@ os.environ["CXX"] = "hipcc"
|
||||
sources = [
|
||||
'torch_int/kernels/linear.cpp',
|
||||
'torch_int/kernels/bmm.cpp',
|
||||
'torch_int/kernels/pybind.cpp',
|
||||
'torch_int/kernels/pybind.cpp',
|
||||
]
|
||||
|
||||
include_dirs = ['torch_int/kernels/include']
|
||||
@@ -418,11 +418,11 @@ setup(
|
||||
|
||||
Run `python setup.py install` to build and install the extension. It should look something like Figure 6:
|
||||
|
||||
<!--
|
||||
<!--
|
||||
================
|
||||
### Figure 6
|
||||
================ -->
|
||||
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-compilation.jpg
|
||||
```{figure} ./images/optimize-with-composable-kernel/ck-compilation.jpg
|
||||
Compilation and installation of the INT8 kernels.
|
||||
```
|
||||
|
||||
@@ -430,11 +430,11 @@ Compilation and installation of the INT8 kernels.
|
||||
|
||||
The implementation architecture of running SmoothQuant models on MI300X GPUs is illustrated in Figure 7, where (a) shows the decoder layer composition components of the target model, (b) shows the major implementation class for the decoder layer components, and \(c\) denotes the underlying GPU kernels implemented by CK instance.
|
||||
|
||||
<!--
|
||||
<!--
|
||||
================
|
||||
### Figure 7
|
||||
================ -->
|
||||
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-inference_flow.jpg
|
||||
```{figure} ./images/optimize-with-composable-kernel/ck-inference_flow.jpg
|
||||
The implementation architecture of running SmoothQuant models on AMD MI300X GPUs.
|
||||
```
|
||||
|
||||
@@ -456,11 +456,11 @@ Note that since the default values were used for the tunable parameters of the f
|
||||
|
||||
Figure 8 shows the performance comparisons between the original FP16 and the SmoothQuant-quantized INT8 models on a single MI300X GPU. The GPU memory footprints of SmoothQuant-quantized models are significantly reduced. It also indicates the per-sample inference latency is significantly reduced for all SmoothQuant-quantized OPT models (illustrated in (b)). Notably, the performance of the CK instance-based INT8 kernel steadily improves with an increase in model size.
|
||||
|
||||
<!--
|
||||
<!--
|
||||
================
|
||||
### Figure 8
|
||||
================ -->
|
||||
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-comparisons.jpg
|
||||
```{figure} ./images/optimize-with-composable-kernel/ck-comparisons.jpg
|
||||
Performance comparisons between the original FP16 and the SmoothQuant-quantized INT8 models on a single MI300X GPU.
|
||||
```
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
.. meta::
|
||||
:description: Learn about vLLM V1 inference tuning on AMD Instinct GPUs for optimal performance.
|
||||
:keywords: AMD, Instinct, MI300X, MI325X, MI350X, MI355X, HPC, tuning, BIOS settings, NBIO, ROCm,
|
||||
:keywords: AMD, Instinct, MI300X, HPC, tuning, BIOS settings, NBIO, ROCm,
|
||||
environment variable, performance, HIP, Triton, PyTorch TunableOp, vLLM, RCCL,
|
||||
MIOpen, GPU, resource utilization
|
||||
|
||||
@@ -25,18 +25,14 @@ Instinct MI300X, MI325X, MI350X, and MI355X GPUs. Learn how to:
|
||||
Performance environment variables
|
||||
=================================
|
||||
|
||||
The following variables are generally useful for Instinct MI300X/MI325X/MI350X/MI355X GPUs and vLLM:
|
||||
The following variables are generally useful for Instinct MI300X/MI355X GPUs and vLLM:
|
||||
|
||||
* **HIP and math libraries**
|
||||
|
||||
* ``export HIP_FORCE_DEV_KERNARG=1`` — improves kernel launch performance by
|
||||
forcing device kernel arguments. This is already set by default in
|
||||
:doc:`vLLM ROCm Docker images
|
||||
</how-to/rocm-for-ai/inference/benchmark-docker/vllm>`. Bare-metal users
|
||||
:doc:`vLLM ROCm Docker images </ai-inference/vllm>`. Bare-metal users
|
||||
should set this manually.
|
||||
* ``export SAFETENSORS_FAST_GPU=1`` — enables GPU-accelerated safetensors
|
||||
loading, significantly reducing model load time for large models. Already
|
||||
set in vLLM ROCm Docker images. Bare-metal users should set this manually.
|
||||
* ``export TORCH_BLAS_PREFER_HIPBLASLT=1`` — explicitly prefers hipBLASLt
|
||||
over hipBLAS for GEMM operations. By default, PyTorch uses heuristics to
|
||||
choose the best BLAS library. Setting this can improve linear layer
|
||||
@@ -45,7 +41,7 @@ The following variables are generally useful for Instinct MI300X/MI325X/MI350X/M
|
||||
* **RCCL (collectives for multi-GPU)**
|
||||
|
||||
* ``export NCCL_MIN_NCHANNELS=112`` — increases RCCL channels from default
|
||||
(typically 32-64) to 112 on the Instinct MI300X/MI325X. **Only beneficial for
|
||||
(typically 32-64) to 112 on the Instinct MI300X. **Only beneficial for
|
||||
multi-GPU distributed workloads** (tensor parallelism, pipeline
|
||||
parallelism). Single-GPU inference does not need this.
|
||||
|
||||
@@ -54,31 +50,31 @@ The following variables are generally useful for Instinct MI300X/MI325X/MI350X/M
|
||||
AITER (AI Tensor Engine for ROCm) switches
|
||||
==========================================
|
||||
|
||||
AITER (AI Tensor Engine for ROCm) provides ROCm-specific fused kernels optimized for Instinct MI350 Series and MI300X/MI325X GPUs in vLLM V1.
|
||||
AITER (AI Tensor Engine for ROCm) provides ROCm-specific fused kernels optimized for Instinct MI350 Series and MI300X GPUs in vLLM V1.
|
||||
|
||||
Enable all AITER optimizations with a single master switch:
|
||||
How AITER flags work:
|
||||
|
||||
* ``VLLM_ROCM_USE_AITER`` is the master switch (defaults to ``False``/``0``).
|
||||
* Individual feature flags (``VLLM_ROCM_USE_AITER_LINEAR``, ``VLLM_ROCM_USE_AITER_MOE``, and so on) default to ``True`` but only activate when the master switch is enabled.
|
||||
* To enable a specific AITER feature, you must set both ``VLLM_ROCM_USE_AITER=1`` and the specific feature flag to ``1``.
|
||||
|
||||
Quick start examples:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
# Enable all AITER optimizations (recommended for most workloads)
|
||||
export VLLM_ROCM_USE_AITER=1
|
||||
vllm serve MODEL_NAME
|
||||
|
||||
Most individual AITER sub-flags default to ``1`` when the master switch is on,
|
||||
while specialized features retain the defaults listed below. You rarely need to
|
||||
change them. To select a specific attention backend, use ``--attention-backend``
|
||||
(see :ref:`backend selection <vllm-optimization-aiter-backend-selection>`).
|
||||
# Enable AITER Fused MoE and enable Triton Prefill-Decode (split) attention
|
||||
export VLLM_ROCM_USE_AITER=1
|
||||
export VLLM_V1_USE_PREFILL_DECODE_ATTENTION=1
|
||||
export VLLM_ROCM_USE_AITER_MHA=0
|
||||
vllm serve MODEL_NAME
|
||||
|
||||
**Flags you might adjust:**
|
||||
|
||||
* ``VLLM_ROCM_SHUFFLE_KV_CACHE_LAYOUT=1`` — Set for high-concurrency MHA workloads (≥32 concurrent requests) with ``ROCM_AITER_FA``. Defaults to ``0``.
|
||||
* ``VLLM_ROCM_USE_AITER_MOE=0`` — Disable only if you hit ``RuntimeError: wrong! device_gemm ...``. Try ``AITER_ONLINE_TUNE=1`` first. See :ref:`AITER MoE requirements <vllm-optimization-aiter-moe-requirements>`.
|
||||
* ``VLLM_ROCM_USE_AITER=0`` — Disable AITER entirely to fall back to Triton kernels (for debugging).
|
||||
|
||||
Advanced: individual AITER flags
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
The following table lists AITER-related sub-flags for fine-grained control. Most
|
||||
users do not need to modify these; the default behavior for each flag is listed below.
|
||||
# Disable AITER entirely (i.e, use vLLM Triton Unified Attention Kernel)
|
||||
export VLLM_ROCM_USE_AITER=0
|
||||
vllm serve MODEL_NAME
|
||||
|
||||
.. list-table::
|
||||
:header-rows: 1
|
||||
@@ -88,135 +84,232 @@ users do not need to modify these; the default behavior for each flag is listed
|
||||
- Description (default behavior)
|
||||
|
||||
* - ``VLLM_ROCM_USE_AITER``
|
||||
- Master switch to enable AITER kernels (``0`` by default). All other ``VLLM_ROCM_USE_AITER_*`` flags require this to be set to ``1``.
|
||||
- Master switch to enable AITER kernels (``0``/``False`` by default). All other ``VLLM_ROCM_USE_AITER_*`` flags require this to be set to ``1``.
|
||||
|
||||
* - ``VLLM_ROCM_USE_AITER_LINEAR``
|
||||
- Use AITER quantization operators + GEMM for linear layers (defaults to ``1`` when AITER is on). Accelerates matrix multiplications in all transformer layers. **Recommended to keep enabled**.
|
||||
- Use AITER quantization operators + GEMM for linear layers (defaults to ``True`` when AITER is on). Accelerates matrix multiplications in all transformer layers. **Recommended to keep enabled**.
|
||||
|
||||
* - ``VLLM_ROCM_USE_AITER_MOE``
|
||||
- Use AITER fused-MoE kernels (defaults to ``1`` when AITER is on). Accelerates Mixture-of-Experts routing and computation. See the note on :ref:`AITER MoE requirements <vllm-optimization-aiter-moe-requirements>`.
|
||||
- Use AITER fused-MoE kernels (defaults to ``True`` when AITER is on). Accelerates Mixture-of-Experts routing and computation. See the note on :ref:`AITER MoE requirements <vllm-optimization-aiter-moe-requirements>`.
|
||||
|
||||
* - ``VLLM_ROCM_USE_AITER_RMSNORM``
|
||||
- Use AITER RMSNorm kernels (defaults to ``1`` when AITER is on). Accelerates normalization layers. **Recommended: keep enabled.**
|
||||
- Use AITER RMSNorm kernels (defaults to ``True`` when AITER is on). Accelerates normalization layers. **Recommended: keep enabled.**
|
||||
|
||||
* - ``VLLM_ROCM_USE_AITER_MLA``
|
||||
- Use AITER Multi-head Latent Attention for supported models, for example, DeepSeek-V3/R1 (defaults to ``1`` when AITER is on). See the section on :ref:`AITER MLA requirements <vllm-optimization-aiter-mla-requirements>`.
|
||||
- Use AITER Multi-head Latent Attention for supported models, for example, DeepSeek-V3/R1 (defaults to ``True`` when AITER is on). See the section on :ref:`AITER MLA requirements <vllm-optimization-aiter-mla-requirements>`.
|
||||
|
||||
* - ``VLLM_ROCM_USE_AITER_MHA``
|
||||
- Use AITER Multi-Head Attention kernels (defaults to ``1`` when AITER is on; set to ``0`` to use Triton attention backends or ``ROCM_ATTN`` backend instead). See :ref:`attention backend selection <vllm-optimization-aiter-backend-selection>`.
|
||||
- Use AITER Multi-Head Attention kernels (defaults to ``True`` when AITER is on; set to ``0`` to use Triton attention backends and Prefill-Decode attention backend instead). See :ref:`attention backend selection <vllm-optimization-aiter-backend-selection>`.
|
||||
|
||||
* - ``VLLM_ROCM_USE_AITER_UNIFIED_ATTENTION``
|
||||
- Enable AITER's optimized unified attention kernel (defaults to ``0``). Only takes effect when AITER is enabled and AITER MHA is disabled (``VLLM_ROCM_USE_AITER_MHA=0``). When set to ``0``, falls back to vLLM's Triton unified attention. Can also be enabled via ``--attention-backend ROCM_AITER_UNIFIED_ATTN``.
|
||||
- Enable AITER's optimized unified attention kernel (defaults to ``False``). Only takes effect when: AITER is enabled; unified attention mode is active (``VLLM_V1_USE_PREFILL_DECODE_ATTENTION=0``); and AITER MHA is disabled (``VLLM_ROCM_USE_AITER_MHA=0``). When disabled, falls back to vLLM's Triton unified attention.
|
||||
|
||||
* - ``VLLM_ROCM_USE_AITER_FP8BMM``
|
||||
- Use AITER ``FP8`` batched matmul (defaults to ``1`` when AITER is on). Fuses ``FP8`` per-token quantization with batched GEMM (used in MLA models like DeepSeek-V3).
|
||||
|
||||
* - ``VLLM_ROCM_USE_AITER_FP4BMM``
|
||||
- Use AITER ``FP4`` batched matmul (defaults to ``1`` when AITER is on). Fuses ``FP4`` per-token quantization with batched GEMM (used in MLA models like DeepSeek-V3). Requires an Instinct MI350X/MI355X GPU.
|
||||
|
||||
* - ``VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS``
|
||||
- Fuse shared expert computation into the AITER fused-MoE kernel (defaults to ``0``). Applies to MoE models with shared experts (for example, DeepSeek-V3/R1 with 1 shared expert). Requires SiLU/GELU activation (``is_act_and_mul``). Incompatible with `MoRI <https://github.com/ROCm/mori#mori>`__ scheduling — disable with ``VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS=0`` if using MoRI.
|
||||
|
||||
* - ``VLLM_ROCM_USE_AITER_FP4_ASM_GEMM``
|
||||
- Enable AITER assembly (HIP) FP4 GEMM kernels for MXFP4-quantized models (defaults to ``0``). When set to ``1``, uses hand-tuned ASM kernels instead of Triton for ``FP4×FP4`` weight GEMMs — faster at small batch sizes (``M`` ≤ 64). Requires Instinct MI350X/MI355X (``supports_mx()``). Combine with ``--quantization quark`` for MXFP4 models.
|
||||
- Use AITER ``FP8`` batched matmul (defaults to ``True`` when AITER is on). Fuses ``FP8`` per-token quantization with batched GEMM (used in MLA models like DeepSeek-V3). Requires an Instinct MI300X/MI355X GPU.
|
||||
|
||||
* - ``VLLM_ROCM_USE_SKINNY_GEMM``
|
||||
- Prefer skinny-GEMM kernel variants for small batch sizes (defaults to ``1``). Improves performance when ``M`` dimension is small. **Recommended to keep enabled**.
|
||||
- Prefer skinny-GEMM kernel variants for small batch sizes (defaults to ``True``). Improves performance when ``M`` dimension is small. **Recommended to keep enabled**.
|
||||
|
||||
* - ``VLLM_ROCM_FP8_PADDING``
|
||||
- Pad ``FP8`` linear weight tensors to improve memory locality (defaults to ``1``). Minor memory overhead for better performance.
|
||||
- Pad ``FP8`` linear weight tensors to improve memory locality (defaults to ``True``). Minor memory overhead for better performance.
|
||||
|
||||
* - ``VLLM_ROCM_MOE_PADDING``
|
||||
- Pad MoE weight tensors for better memory access patterns (defaults to ``1``). Same memory/performance tradeoff as ``FP8`` padding.
|
||||
|
||||
* - ``VLLM_ROCM_SHUFFLE_KV_CACHE_LAYOUT``
|
||||
- This only affects the ``ROCM_AITER_FA`` backend. When set to ``0``, it uses the HIP Paged Attention implementation ``torch.ops.aiter.paged_attention_v1`` from AITER. **Good for low concurrency (for example, concurrency ≤32).** When set to ``1``, it uses the ASM Paged Attention Kernel ``pa_fwd_asm`` from AITER. **Good for high concurrency (for example, concurrency ≥32).** (Defaults to ``0`` when AITER is on.)
|
||||
- Pad MoE weight tensors for better memory access patterns (defaults to ``True``). Same memory/performance tradeoff as ``FP8`` padding.
|
||||
|
||||
* - ``VLLM_ROCM_CUSTOM_PAGED_ATTN``
|
||||
- Use custom paged-attention decode kernel when ``ROCM_ATTN`` backend is selected (defaults to ``1``). See :ref:`Attention backend selection with AITER <vllm-optimization-aiter-backend-selection>`.
|
||||
- Use custom paged-attention decode kernel when Prefill-Decode attention backend is selected (defaults to ``True``). See :ref:`Attention backend selection with AITER <vllm-optimization-aiter-backend-selection>`.
|
||||
|
||||
.. note::
|
||||
|
||||
When ``VLLM_ROCM_USE_AITER=1``, most AITER component flags (``LINEAR``,
|
||||
``MOE``, ``RMSNORM``, ``MLA``, ``MHA``, ``FP8BMM``) automatically default to
|
||||
``True``. You typically only need to set the master switch
|
||||
``VLLM_ROCM_USE_AITER=1`` to enable all optimizations. ROCm provides a
|
||||
prebuilt optimized Docker image for validating the performance of LLM
|
||||
inference with vLLM on MI300X Series GPUs. The Docker image includes ROCm,
|
||||
vLLM, and PyTorch. For more information, see :doc:`/ai-inference/vllm`.
|
||||
|
||||
.. _vllm-optimization-aiter-moe-requirements:
|
||||
|
||||
AITER MoE requirements (Mixtral, DeepSeek-V2/V3, Qwen-MoE models)
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
``VLLM_ROCM_USE_AITER_MOE`` enables AITER's optimized Mixture-of-Experts kernels, such as expert routing (topk selection) and expert computation for better performance.
|
||||
|
||||
Applicable models:
|
||||
|
||||
* Mixtral series: for example, Mixtral-8x7B / Mixtral-8x22B
|
||||
* Llama-4 family: for example, Llama-4-Scout-17B-16E / Llama-4-Maverick-17B-128E
|
||||
* DeepSeek family: DeepSeek-V2 / DeepSeek-V3 / DeepSeek-R1
|
||||
* Qwen family: Qwen1.5-MoE / Qwen2-MoE / Qwen2.5-MoE series
|
||||
* Other MoE architectures
|
||||
|
||||
When to enable:
|
||||
|
||||
* **Enable (default):** For all MoE models on the Instinct MI300X/MI355X for best throughput
|
||||
* **Disable:** Only for debugging or if you encounter numerical issues
|
||||
|
||||
Example usage:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
# Standard MoE model (Mixtral)
|
||||
VLLM_ROCM_USE_AITER=1 vllm serve mistralai/Mixtral-8x7B-Instruct-v0.1
|
||||
|
||||
# Hybrid MoE+MLA model (DeepSeek-V3) - requires both MOE and MLA flags
|
||||
VLLM_ROCM_USE_AITER=1 vllm serve deepseek-ai/DeepSeek-V3 \
|
||||
--block-size 1 \
|
||||
--tensor-parallel-size 8
|
||||
|
||||
.. _vllm-optimization-aiter-mla-requirements:
|
||||
.. _vllm-optimization-aiter-mla-sparse-requirements:
|
||||
|
||||
AITER MLA requirements (DeepSeek-V3/R1 models)
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
``VLLM_ROCM_USE_AITER_MLA`` enables AITER MLA (Multi-head Latent Attention) optimization for supported models. Defaults to **True** when AITER is on.
|
||||
|
||||
Critical requirement:
|
||||
|
||||
* **Must** explicitly set ``--block-size 1``
|
||||
|
||||
.. important::
|
||||
|
||||
If you omit ``--block-size 1``, vLLM will raise an error rather than defaulting to 1.
|
||||
|
||||
Applicable models:
|
||||
|
||||
* DeepSeek-V3 / DeepSeek-R1
|
||||
* DeepSeek-V2
|
||||
* Other models using multi-head latent attention (MLA) architecture
|
||||
|
||||
Example usage:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
# DeepSeek-R1 with AITER MLA (requires 8 GPUs)
|
||||
VLLM_ROCM_USE_AITER=1 vllm serve deepseek-ai/DeepSeek-R1 \
|
||||
--block-size 1 \
|
||||
--tensor-parallel-size 8
|
||||
|
||||
.. _vllm-optimization-aiter-backend-selection:
|
||||
|
||||
Attention backend selection with AITER
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
Most models work out of the box with ``VLLM_ROCM_USE_AITER=1`` — vLLM auto-selects
|
||||
the optimal backend. Use ``--attention-backend`` to override the auto-selected backend.
|
||||
Understanding which attention backend to use helps optimize your deployment.
|
||||
|
||||
.. code-block:: bash
|
||||
Quick reference: Which attention backend will I get?
|
||||
|
||||
export VLLM_ROCM_USE_AITER=1
|
||||
vllm serve <your-model> --tensor-parallel-size <tp>
|
||||
Default behavior (no configuration)
|
||||
|
||||
.. note::
|
||||
Always set ``VLLM_ROCM_USE_AITER=1`` even when using ``--attention-backend`` explicitly.
|
||||
``--attention-backend`` only overrides the attention kernel; ``VLLM_ROCM_USE_AITER=1``
|
||||
is still required to enable AITER for GEMM, RMSNorm, and MoE kernels.
|
||||
The Radeon/fallback backends (``ROCM_ATTN``, ``TRITON_MLA``) are the exception —
|
||||
they do not use AITER and do not require the env var.
|
||||
Without setting any environment variables, vLLM uses:
|
||||
|
||||
The table below shows which backend is selected per model type and how to tune it.
|
||||
* **vLLM Triton Unified Attention** — A single Triton kernel handling both prefill and decode phases
|
||||
* Works on all ROCm platforms
|
||||
* Good baseline performance
|
||||
|
||||
**Recommended**: Enable AITER (set ``VLLM_ROCM_USE_AITER=1``)
|
||||
|
||||
When you enable AITER, the backend is automatically selected based on your model:
|
||||
|
||||
.. code-block:: text
|
||||
|
||||
Is your model using MLA architecture? (DeepSeek-V3/R1/V2)
|
||||
├─ YES → AITER MLA Backend
|
||||
│ • Requires --block-size 1
|
||||
│ • Best performance for MLA models
|
||||
│ • Automatically selected
|
||||
│
|
||||
└─ NO → AITER MHA Backend
|
||||
• For standard transformer models (Llama, Mistral, etc.)
|
||||
• Optimized for Instinct MI300X/MI355X
|
||||
• Automatically selected
|
||||
|
||||
**Advanced**: Manual backend selection
|
||||
|
||||
Most users won't need this, but you can override the defaults:
|
||||
|
||||
.. list-table::
|
||||
:widths: 40 60
|
||||
:header-rows: 1
|
||||
:widths: 15 18 35 32
|
||||
|
||||
* - Model type
|
||||
- Backend
|
||||
- How to enable
|
||||
- Tuning tips
|
||||
* - To use this backend
|
||||
- Set these flags
|
||||
|
||||
* - **MHA models** (Llama, Mistral, Qwen, Mixtral, MiniMax-M2.5)
|
||||
- **ROCM_AITER_FA** (recommended, auto-selected)
|
||||
- ``VLLM_ROCM_USE_AITER=1`` (auto-selected). To override: ``VLLM_ROCM_USE_AITER=1 --attention-backend ROCM_AITER_FA``. Add ``VLLM_ROCM_SHUFFLE_KV_CACHE_LAYOUT=1`` for shuffled KV cache.
|
||||
- **2.7–4.4x TPS** over legacy ``ROCM_ATTN``. Set ``VLLM_ROCM_SHUFFLE_KV_CACHE_LAYOUT=1`` for **15–20% decode improvement** at high concurrency. Low TTFT: ``--max-num-batched-tokens`` ≤ 8k–16k. High throughput: ≥ 32k with ``cudagraph_mode=FULL``.
|
||||
* - AITER MLA (MLA models only)
|
||||
- ``VLLM_ROCM_USE_AITER=1`` (auto-selects for DeepSeek-V3/R1)
|
||||
|
||||
* - **MLA models** (DeepSeek-V3/R1/V2, Kimi-K2.5, Mistral-Large-3-675B)
|
||||
- **ROCM_AITER_MLA** (recommended, auto-selected)
|
||||
- ``VLLM_ROCM_USE_AITER=1`` (auto-selected). To override: ``VLLM_ROCM_USE_AITER=1 --attention-backend ROCM_AITER_MLA``. ``--block-size 1`` is no longer mandatory (vLLM ≥0.14); still recommended for prefix-caching workloads.
|
||||
- **1.2–1.5x higher TPS** over ``TRITON_MLA``. Supports uniform-batch CUDA graphs and MTP. On MI300X/MI325X (gfx942), ``ROCM_AITER_TRITON_MLA`` may show 2–3% higher TPS. On MI355X (gfx950), ``ROCM_AITER_MLA`` is preferred (uses AITER assembly MHA for prefill).
|
||||
* - AITER MHA (standard models)
|
||||
- ``VLLM_ROCM_USE_AITER=1`` (auto-selects for non-MLA models)
|
||||
|
||||
* - **DSA models** (DeepSeek-V3.2, GLM-5)
|
||||
- **ROCM_AITER_MLA_SPARSE** (auto-selected)
|
||||
- ``VLLM_ROCM_USE_AITER=1`` — auto-detected from ``index_topk`` in model config. Requires ``--block-size 1``.
|
||||
- Instinct MI300X/MI325X/MI350X/MI355X only.
|
||||
* - vLLM Triton Unified (default)
|
||||
- ``VLLM_ROCM_USE_AITER=0`` (or unset)
|
||||
|
||||
* - **gpt-oss models** (gpt-oss-120b/20b)
|
||||
- **ROCM_AITER_UNIFIED_ATTN**
|
||||
- ``VLLM_ROCM_USE_AITER=1 --attention-backend ROCM_AITER_UNIFIED_ATTN``
|
||||
-
|
||||
* - Triton Prefill-Decode (split) without AITER
|
||||
- | ``VLLM_V1_USE_PREFILL_DECODE_ATTENTION=1``
|
||||
|
||||
* - **Radeon / fallback**
|
||||
- **ROCM_ATTN** (MHA) or **TRITON_MLA** (MLA)
|
||||
- ``--attention-backend ROCM_ATTN`` or ``TRITON_MLA``
|
||||
- ``ROCM_ATTN`` is preferred over ``TRITON_ATTN`` — it uses a custom HIP paged-attention kernel for decode when the model's KV head size is supported, and falls back to Triton only when not. If ``ROCM_ATTN`` is slow for your model (unsupported head size triggers Triton decode), try ``TRITON_ATTN``. Both work on Radeon GPUs.
|
||||
* - Triton Prefill-Decode (split) along with AITER Fused-MoE
|
||||
- | ``VLLM_ROCM_USE_AITER=1``
|
||||
| ``VLLM_ROCM_USE_AITER_MHA=0``
|
||||
| ``VLLM_V1_USE_PREFILL_DECODE_ATTENTION=1``
|
||||
|
||||
.. note::
|
||||
**MoE models** (Mixtral, Llama-4-Scout/Maverick, DeepSeek-V2/V3/R1, Kimi-K2.5, MiniMax-M2.5, GLM-5, Qwen-MoE): AITER MoE kernels activate automatically with ``VLLM_ROCM_USE_AITER=1`` — no extra attention backend flags needed. If you hit ``RuntimeError: wrong! device_gemm ...``, set ``AITER_ONLINE_TUNE=1`` and retry. Only disable MoE kernels (``VLLM_ROCM_USE_AITER_MOE=0``) if that also fails.
|
||||
|
||||
Once AITER is configured, see `Parallelism strategies (run vLLM on multiple GPUs)`_ for TP/DP/EP choices — especially for MLA and MoE models where the wrong strategy wastes memory or throughput.
|
||||
* - AITER Unified Attention
|
||||
- | ``VLLM_ROCM_USE_AITER=1``
|
||||
| ``VLLM_ROCM_USE_AITER_MHA=0``
|
||||
| ``VLLM_ROCM_USE_AITER_UNIFIED_ATTENTION=1``
|
||||
|
||||
**Quick start examples**:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
# DSA model (DeepSeek-V3.2) — backend auto-selected from model config
|
||||
VLLM_ROCM_USE_AITER=1 vllm serve deepseek-ai/DeepSeek-V3.2 \
|
||||
# Recommended: Standard model with AITER (Llama, Mistral, Qwen, etc.)
|
||||
VLLM_ROCM_USE_AITER=1 vllm serve meta-llama/Llama-3.3-70B-Instruct
|
||||
|
||||
# MLA model with AITER (DeepSeek-V3/R1)
|
||||
VLLM_ROCM_USE_AITER=1 vllm serve deepseek-ai/DeepSeek-R1 \
|
||||
--block-size 1 \
|
||||
--tensor-parallel-size 8
|
||||
|
||||
# Explicitly select a backend for MLA models
|
||||
VLLM_ROCM_USE_AITER=1 vllm serve deepseek-ai/DeepSeek-R1-0528 \
|
||||
--tensor-parallel-size 8 \
|
||||
--attention-backend ROCM_AITER_MLA
|
||||
# Advanced: Use Prefill-Decode split (for short input cases) with AITER Fused-MoE
|
||||
VLLM_ROCM_USE_AITER=1 \
|
||||
VLLM_ROCM_USE_AITER_MHA=0 \
|
||||
VLLM_V1_USE_PREFILL_DECODE_ATTENTION=1 \
|
||||
vllm serve meta-llama/Llama-4-Scout-17B-16E
|
||||
|
||||
# MHA model with shuffled KV cache layout for high concurrency
|
||||
VLLM_ROCM_USE_AITER=1 VLLM_ROCM_SHUFFLE_KV_CACHE_LAYOUT=1 \
|
||||
vllm serve meta-llama/Llama-3.3-70B-Instruct \
|
||||
--attention-backend ROCM_AITER_FA
|
||||
**Which backend should I choose?**
|
||||
|
||||
.. list-table::
|
||||
:widths: 30 70
|
||||
:header-rows: 1
|
||||
|
||||
* - Your use case
|
||||
- Recommended backend
|
||||
|
||||
* - **Standard transformer models** (Llama, Mistral, Qwen, Mixtral)
|
||||
- **AITER MHA** (``VLLM_ROCM_USE_AITER=1``) — **Recommended for most workloads** on Instinct MI300X/MI355X. Provides optimized attention kernels for both prefill and decode phases.
|
||||
|
||||
* - **MLA models** (DeepSeek-V3/R1/V2)
|
||||
- **AITER MLA** (auto-selected with ``VLLM_ROCM_USE_AITER=1``) — Required for optimal performance, must use ``--block-size 1``
|
||||
|
||||
* - **gpt-oss models** (gpt-oss-120b/20b)
|
||||
- **AITER Unified Attention** (``VLLM_ROCM_USE_AITER=1``, ``VLLM_ROCM_USE_AITER_MHA=0``, ``VLLM_ROCM_USE_AITER_UNIFIED_ATTENTION=1``) — Required for optimal performance
|
||||
|
||||
* - **Debugging or compatibility**
|
||||
- **vLLM Triton Unified** (default with ``VLLM_ROCM_USE_AITER=0``) — Generic fallback, works everywhere
|
||||
|
||||
**Important notes:**
|
||||
|
||||
* **AITER MHA and AITER MLA are mutually exclusive** — vLLM automatically detects MLA models and selects the appropriate backend
|
||||
* **For 95% of users:** Simply set ``VLLM_ROCM_USE_AITER=1`` and let vLLM choose the right backend
|
||||
* When in doubt, start with AITER enabled (the recommended configuration) and profile your specific workload
|
||||
|
||||
Backend choice quick recipes
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
* **Standard transformers (any prompt length):** Start with ``VLLM_ROCM_USE_AITER=1`` → AITER MHA. For CUDA graph modes, see architecture-specific guidance below (Dense vs MoE models have different optimal modes).
|
||||
* **Latency-sensitive chat (low TTFT):** keep ``--max-num-batched-tokens`` ≤ **8k–16k** with AITER.
|
||||
* **Streaming decode (low ITL):** raise ``--max-num-batched-tokens`` to **32k–64k**.
|
||||
* **Offline max throughput:** ``--max-num-batched-tokens`` ≥ **32k** with ``cudagraph_mode=FULL``.
|
||||
|
||||
**How to verify which backend is active**
|
||||
|
||||
@@ -225,13 +318,68 @@ Check vLLM's startup logs to confirm which attention backend is being used:
|
||||
.. code-block:: bash
|
||||
|
||||
# Start vLLM and check logs
|
||||
VLLM_ROCM_USE_AITER=1 vllm serve meta-llama/Llama-3.3-70B-Instruct 2>&1 | grep -i "using.*backend"
|
||||
VLLM_ROCM_USE_AITER=1 vllm serve meta-llama/Llama-3.3-70B-Instruct 2>&1 | grep -i attention
|
||||
|
||||
Look for ``Using <backend_name> backend.`` in the startup output — for example,
|
||||
``Using ROCM_AITER_FA backend.``
|
||||
**Expected log messages:**
|
||||
|
||||
For in-depth architecture and benchmarks of all 7 ROCm attention backends, see the
|
||||
`ROCm Attention Backend blog post <https://vllm.ai/blog/rocm-attention-backend>`_.
|
||||
* AITER MHA: ``Using Aiter Flash Attention backend on V1 engine.``
|
||||
* AITER MLA: ``Using AITER MLA backend on V1 engine.``
|
||||
* vLLM Triton MLA: ``Using Triton MLA backend on V1 engine.``
|
||||
* vLLM Triton Unified: ``Using Triton Attention backend on V1 engine.``
|
||||
* AITER Triton Unified: ``Using Aiter Unified Attention backend on V1 engine.``
|
||||
* AITER Triton Prefill-Decode: ``Using Rocm Attention backend on V1 engine.``
|
||||
|
||||
Attention backend technical details
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
This section provides technical details about vLLM's attention backends on ROCm.
|
||||
|
||||
vLLM V1 on ROCm provides these attention implementations:
|
||||
|
||||
1. **vLLM Triton Unified Attention** (default when AITER is **off**)
|
||||
|
||||
* Single unified Triton kernel handling both chunked prefill and decode phases
|
||||
* Generic implementation that works across all ROCm platforms
|
||||
* Good baseline performance
|
||||
* Automatically selected when ``VLLM_ROCM_USE_AITER=0`` (or unset)
|
||||
* Supports GPT-OSS
|
||||
|
||||
2. **AITER Triton Unified Attention** (advanced, requires manual configuration)
|
||||
|
||||
* The AMD optimized unified Triton kernel
|
||||
* Enable with ``VLLM_ROCM_USE_AITER=1``, ``VLLM_ROCM_USE_AITER_MHA=0``, and ``VLLM_ROCM_USE_AITER_UNIFIED_ATTENTION=1``.
|
||||
* Only useful for specific workloads. Most users should use AITER MHA instead.
|
||||
* Recommended this backend when running GPT-OSS.
|
||||
|
||||
3. **AITER Triton Prefill–Decode Attention** (hybrid, Instinct MI300X-optimized)
|
||||
|
||||
* Enable with ``VLLM_V1_USE_PREFILL_DECODE_ATTENTION=1``
|
||||
* Uses separate kernels for prefill and decode phases:
|
||||
|
||||
* **Prefill**: ``context_attention_fwd`` Triton kernel
|
||||
* **Primary decode**: ``torch.ops._rocm_C.paged_attention`` (custom ROCm kernel optimized for head sizes 64/128, block sizes 16/32, GQA 1–16, context ≤131k; sliding window not supported)
|
||||
* **Fallback decode**: ``kernel_paged_attention_2d`` Triton kernel when shapes don't meet primary decode requirements
|
||||
|
||||
* Usually better compared to unified Triton kernels
|
||||
* Performance vs AITER MHA varies: AITER MHA is typically faster overall, but Prefill-Decode split may win in short input scenarios
|
||||
* The custom paged attention decode kernel is controlled by ``VLLM_ROCM_CUSTOM_PAGED_ATTN`` (default **True**)
|
||||
|
||||
4. **AITER Multi-Head Attention (MHA)** (default when AITER is **on**)
|
||||
|
||||
* Controlled by ``VLLM_ROCM_USE_AITER_MHA`` (**1** = enabled)
|
||||
* Best all-around performance for standard transformer models
|
||||
* Automatically selected when ``VLLM_ROCM_USE_AITER=1`` and model is not MLA
|
||||
|
||||
5. **vLLM Triton Multi-head Latent Attention (MLA)** (for DeepSeek-V3/R1/V2)
|
||||
|
||||
* Automatically selected when ``VLLM_ROCM_USE_AITER=0`` (or unset)
|
||||
|
||||
6. **AITER Multi-head Latent Attention (MLA)** (for DeepSeek-V3/R1/V2)
|
||||
|
||||
* Controlled by ``VLLM_ROCM_USE_AITER_MLA`` (``1`` = enabled)
|
||||
* Required for optimal performance on MLA architecture models
|
||||
* Automatically selected when ``VLLM_ROCM_USE_AITER=1`` and model uses MLA
|
||||
* Requires ``--block-size 1``
|
||||
|
||||
Quick Reduce (large all-reduces on ROCm)
|
||||
========================================
|
||||
@@ -246,7 +394,7 @@ It supports FP16/BF16 as well as symmetric INT8/INT6/INT4 quantized all-reduce (
|
||||
Control via:
|
||||
|
||||
* ``VLLM_ROCM_QUICK_REDUCE_QUANTIZATION`` ∈ ``["NONE","FP","INT8","INT6","INT4"]`` (default ``NONE``).
|
||||
* ``VLLM_ROCM_QUICK_REDUCE_CAST_BF16_TO_FP16``: cast BF16 input to FP16 (``1`` by default for performance).
|
||||
* ``VLLM_ROCM_QUICK_REDUCE_CAST_BF16_TO_FP16``: cast BF16 input to FP16 (``1/True`` by default for performance).
|
||||
* ``VLLM_ROCM_QUICK_REDUCE_MAX_SIZE_BYTES_MB``: cap the preset buffer (default ``NONE`` ≈ ``2048`` MB).
|
||||
|
||||
Quick Reduce tends to help **throughput** at higher TP counts (for example, 4–8) with many concurrent requests.
|
||||
@@ -263,34 +411,12 @@ vLLM supports the following parallelism strategies:
|
||||
|
||||
For more details, see `Parallelism and scaling <https://docs.vllm.ai/en/stable/serving/parallelism_scaling.html>`_.
|
||||
|
||||
**Quick-reference decision table:**
|
||||
**Choosing the right strategy:**
|
||||
|
||||
.. list-table::
|
||||
:header-rows: 1
|
||||
:widths: 30 35 35
|
||||
|
||||
* - Model type
|
||||
- Low concurrency (≤128 requests)
|
||||
- High concurrency (≥512 requests)
|
||||
|
||||
* - **Dense** (for example, Llama, Qwen-dense, Mistral-dense)
|
||||
- TP only
|
||||
- TP + independent DP replicas (your own load balancer)
|
||||
|
||||
* - **MoE, standard density ≥3%** (for example, Qwen3-235B-A22B, DeepSeek-V3/R1)
|
||||
- TP + EP
|
||||
- DP + EP
|
||||
|
||||
* - **MoE, ultra-sparse <1%** (for example, Llama-4-Maverick at 0.78%)
|
||||
- TP only — **no EP** (AllToAll overhead exceeds benefit)
|
||||
- DP only — **no EP**
|
||||
|
||||
* - **MLA models** (for example, DeepSeek-V2/V3/R1, Kimi-K2.5, Mistral-Large-3-675B)
|
||||
- TP + EP
|
||||
- **DP + EP** — TP alone duplicates the full KV cache on every GPU; use DP Attention to partition it
|
||||
|
||||
EP = ``--enable-expert-parallel``. DP = ``--data-parallel-size N``.
|
||||
See `Data Parallel Attention (advanced)`_ for the MLA memory explanation and `Expert parallelism`_ for EP details.
|
||||
* **Tensor Parallelism (TP)**: Use when model doesn't fit on one GPU. Prefer staying within a single XGMI island (≤8 GPUs on the Instinct MI300X).
|
||||
* **Pipeline Parallelism (PP)**: Use for very large models across nodes. Set TP to GPUs per node, scale with PP across nodes.
|
||||
* **Data Parallelism (DP)**: Use when model fits on single GPU or TP group, and you need higher throughput. Combine with TP/PP for large models.
|
||||
* **Expert Parallelism (EP)**: Use for MoE models with ``--enable-expert-parallel``. More efficient than TP for MoE layers.
|
||||
|
||||
Tensor parallelism
|
||||
^^^^^^^^^^^^^^^^^^
|
||||
@@ -319,9 +445,6 @@ Tensor parallelism splits each layer of the model weights across multiple GPUs w
|
||||
.. tip::
|
||||
For structured data parallelism deployments with load balancing, see :ref:`data-parallelism-section`.
|
||||
|
||||
.. note::
|
||||
**MLA models (DeepSeek, Kimi-K2.5, Mistral-Large-3-675B):** TP alone replicates the full KV cache on every GPU, which wastes memory at high concurrency. See `Data Parallel Attention (advanced)`_ for the DP+EP configuration that partitions the KV cache instead.
|
||||
|
||||
Pipeline parallelism
|
||||
^^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
@@ -344,7 +467,7 @@ Pipeline parallelism splits the model's layers across multiple GPUs or nodes, wi
|
||||
--pipeline-parallel-size 2
|
||||
|
||||
.. note::
|
||||
**ROCm best practice**: On Instinct MI300X/MI325X/MI350X/MI355X, prefer staying within a single XGMI island (≤8 GPUs) using TP only. Use PP when scaling beyond eight GPUs or across nodes.
|
||||
**ROCm best practice**: On the Instinct MI300X, prefer staying within a single XGMI island (≤8 GPUs) using TP only. Use PP when scaling beyond eight GPUs or across nodes.
|
||||
|
||||
.. _data-parallelism-section:
|
||||
|
||||
@@ -413,21 +536,29 @@ For more technical details, see `vLLM Data Parallel Deployment <https://docs.vll
|
||||
Data Parallel Attention (advanced)
|
||||
""""""""""""""""""""""""""""""""""
|
||||
|
||||
For MLA models (DeepSeek V2/V3/R1, Kimi-K2.5), **DP+EP is the recommended configuration at high concurrency** (≥512 concurrent requests). Unlike traditional DP which replicates model weights, Data Parallel Attention uses inter-GPU AllToAll communication to partition KV cache across GPUs, avoiding the KV cache duplication that occurs with tensor parallelism.
|
||||
For models with Multi-head Latent Attention (MLA) architecture like DeepSeek V2, V3, and R1, vLLM supports **Data Parallel Attention**,
|
||||
which provides request-level parallelism instead of model replication. This avoids KV cache duplication across tensor parallel ranks,
|
||||
significantly reducing memory usage and enabling larger batch sizes.
|
||||
|
||||
* At **≤128 concurrent requests**, TP=8 provides 40–86% higher throughput
|
||||
* At **≥512 concurrent requests**, DP=8+EP provides 16–47% higher throughput
|
||||
* Crossover typically occurs around **256–512 concurrent requests**
|
||||
**Key benefits for MLA models:**
|
||||
|
||||
* Eliminates KV cache duplication when using tensor parallelism
|
||||
* Enables higher throughput for high-QPS serving scenarios
|
||||
* Better memory efficiency for large context windows
|
||||
|
||||
**Usage with Expert Parallelism:**
|
||||
|
||||
Data parallel attention works seamlessly with Expert Parallelism for MoE models:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
# DeepSeek-R1 with DP attention and expert parallelism (high concurrency)
|
||||
# DeepSeek-R1 with DP attention and expert parallelism
|
||||
VLLM_ALL2ALL_BACKEND="allgather_reducescatter" vllm serve deepseek-ai/DeepSeek-R1 \
|
||||
--data-parallel-size 8 \
|
||||
--enable-expert-parallel \
|
||||
--disable-nccl-for-dp-synchronization
|
||||
|
||||
For more technical details, see `vLLM RFC #16037 <https://github.com/vllm-project/vllm/issues/16037>`_ and the `vLLM MoE Playbook <https://rocm.blogs.amd.com/software-tools-optimization/vllm-moe-guide/README.html>`_.
|
||||
For more technical details, see `vLLM RFC #16037 <https://github.com/vllm-project/vllm/issues/16037>`_.
|
||||
|
||||
Expert parallelism
|
||||
^^^^^^^^^^^^^^^^^^
|
||||
@@ -435,47 +566,23 @@ Expert parallelism
|
||||
Expert parallelism (EP) distributes expert layers of Mixture-of-Experts (MoE) models across multiple GPUs,
|
||||
where tokens are routed to the GPUs holding the experts they need.
|
||||
|
||||
**Performance considerations:**
|
||||
|
||||
Expert parallelism is designed primarily for cross-node MoE deployments where high-bandwidth interconnects (like InfiniBand) between nodes make EP communication efficient. For single-node Instinct MI300X/MI355X deployments with XGMI connectivity, tensor parallelism typically provides better performance due to optimized all-to-all collectives on XGMI.
|
||||
|
||||
**When to use EP:**
|
||||
|
||||
.. list-table::
|
||||
:header-rows: 1
|
||||
:widths: 30 35 35
|
||||
* Multi-node MoE deployments with fast inter-node networking
|
||||
* Models with very large numbers of experts that benefit from expert distribution
|
||||
* Workloads where EP's reduced data movement outweighs communication overhead
|
||||
|
||||
* - Scenario
|
||||
- Recommended config
|
||||
- Rationale
|
||||
|
||||
* - **Low concurrency** (≤128 requests)
|
||||
- TP=8 (EP optional)
|
||||
- 40–86% higher throughput than DP at low concurrency.
|
||||
|
||||
* - **High concurrency** (≥512 requests)
|
||||
- DP=8 + EP
|
||||
- 16–47% higher throughput at scale (for example, 7,114 TPS for DeepSeek-R1 at 1024 concurrent requests).
|
||||
|
||||
* - **MLA/MQA models** (DeepSeek-V2/V3/R1, Kimi-K2.5)
|
||||
- DP + EP
|
||||
- Avoids KV cache duplication across TP ranks. Mandatory for optimal memory at high concurrency.
|
||||
|
||||
* - **Ultra-sparse MoE** (<1% activation density, for example, Llama-4-Maverick)
|
||||
- DP or TP **without** EP
|
||||
- EP adds AllToAll overhead that exceeds the benefit — EP is 7–12% *slower* for these models.
|
||||
|
||||
* - **Standard MoE** (≥3% activation density, for example, DeepSeek-R1, Qwen3-235B)
|
||||
- EP flag
|
||||
- Improves expert routing efficiency.
|
||||
**Single-node recommendation:** For Instinct MI300X/MI355X within a single node (≤8 GPUs), prefer tensor parallelism over expert parallelism for MoE models to leverage XGMI's high bandwidth and low latency.
|
||||
|
||||
**Basic usage:**
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
# DP + EP for MLA+MoE models (DeepSeek-R1, high concurrency)
|
||||
VLLM_ALL2ALL_BACKEND="allgather_reducescatter" vllm serve deepseek-ai/DeepSeek-R1 \
|
||||
--data-parallel-size 8 \
|
||||
--enable-expert-parallel \
|
||||
--disable-nccl-for-dp-synchronization
|
||||
|
||||
# TP + EP (low concurrency, non-MLA models)
|
||||
# Enable expert parallelism for MoE models (DeepSeek example with 8 GPUs)
|
||||
vllm serve deepseek-ai/DeepSeek-R1 \
|
||||
--tensor-parallel-size 8 \
|
||||
--enable-expert-parallel
|
||||
@@ -487,37 +594,17 @@ When EP is enabled alongside tensor parallelism:
|
||||
* Fused MoE layers use expert parallelism
|
||||
* Non-fused MoE layers use tensor parallelism
|
||||
|
||||
Multimodal model optimization (vision-language)
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
**Combining with Data Parallelism:**
|
||||
|
||||
For multimodal models (Qwen3-VL, InternVL, step3), use batch-level data parallelism
|
||||
for the vision encoder instead of the default tensor parallelism:
|
||||
EP works seamlessly with Data Parallel Attention for optimal memory efficiency in MLA+MoE models (for example, DeepSeek V3):
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
vllm serve Qwen/Qwen3-VL-235B-A22B-Instruct \
|
||||
--tensor-parallel-size 8 \
|
||||
--mm-encoder-tp-mode data \
|
||||
# DP attention + EP for DeepSeek-R1
|
||||
VLLM_ALL2ALL_BACKEND="allgather_reducescatter" vllm serve deepseek-ai/DeepSeek-R1 \
|
||||
--data-parallel-size 8 \
|
||||
--enable-expert-parallel \
|
||||
--max-model-len 32768
|
||||
|
||||
``--mm-encoder-tp-mode data`` replaces per-layer all-reduce synchronization (58–126 ops
|
||||
in TP mode) with a single all-gather after encoding, yielding **10–45% throughput
|
||||
improvement** with negligible memory overhead (0.2–2.3% model size increase).
|
||||
|
||||
**When it helps most:**
|
||||
|
||||
* High-resolution images (1024×1024 px): **+16% average** throughput
|
||||
* 1–3 images per request: **+13–16%** throughput
|
||||
* Deep vision encoders (for example, InternVL 45 blocks, step3 63 blocks)
|
||||
|
||||
**When to skip it:**
|
||||
|
||||
* Very small vision encoders (<1% of total model parameters)
|
||||
* 10+ small images per request (diminishing returns)
|
||||
* Memory-constrained deployments (encoder weights are replicated per GPU)
|
||||
|
||||
For more details, see the `vLLM Multimodal DP blog post <https://rocm.blogs.amd.com/software-tools-optimization/vllm-dp-vision/README.html>`_.
|
||||
--disable-nccl-for-dp-synchronization
|
||||
|
||||
Throughput benchmarking
|
||||
=======================
|
||||
@@ -556,22 +643,22 @@ Maximizing instances per node
|
||||
To maximize **per-node throughput**, run as many vLLM instances as model memory allows,
|
||||
balancing KV-cache capacity.
|
||||
|
||||
* **HBM capacities**: MI300X = 192 GB HBM3; MI325X = 256 GB HBM3E; MI350X/MI355X = 288 GB HBM3E.
|
||||
* **HBM capacities**: MI300X = 192 GB HBM3; MI355X = 288 GB HBM3E.
|
||||
|
||||
* Up to **eight** single-GPU vLLM instances can run in parallel on an 8×GPU node (one per GPU):
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
for i in $(seq 0 7); do
|
||||
CUDA_VISIBLE_DEVICES="$i" vllm bench throughput
|
||||
-tp 1 --model /path/to/model
|
||||
CUDA_VISIBLE_DEVICES="$i" vllm bench throughput
|
||||
-tp 1 --model /path/to/model
|
||||
--dataset /path/to/ShareGPT_V3_unfiltered_cleaned_split.json &
|
||||
done
|
||||
|
||||
Total throughput from **N** single-GPU instances usually exceeds one instance stretched across **N** GPUs (``-tp N``).
|
||||
|
||||
**Model coverage**: Llama 2 (7B/13B/70B), Llama 3 (8B/70B), Qwen2 (7B/72B), Mixtral-8x7B/8x22B, and others Llama2‑70B
|
||||
and Llama3‑70B can fit a single MI300X/MI325X/MI350X/MI355X; Llama3.1‑405B fits on a single 8×MI300X/MI325X/MI350X/MI355X node.
|
||||
and Llama3‑70B can fit a single MI300X/MI355X; Llama3.1‑405B fits on a single 8×MI300X/MI355X node.
|
||||
|
||||
Configure the gpu-memory-utilization parameter
|
||||
==================================================
|
||||
@@ -681,15 +768,11 @@ CUDA graphs reduce kernel launch overhead by capturing and replaying GPU operati
|
||||
|
||||
* - Attention backend
|
||||
- CUDA graph support
|
||||
* - ``TRITON_ATTN``
|
||||
* - vLLM/AITER Triton Unified Attention, vLLM Prefill-Decode Attention
|
||||
- Full support (prefill + decode)
|
||||
* - ``ROCM_ATTN``, ``ROCM_AITER_UNIFIED_ATTN``
|
||||
- Full support (prefill + decode)
|
||||
* - ``ROCM_AITER_FA``, ``ROCM_AITER_MLA``, ``ROCM_AITER_TRITON_MLA``
|
||||
* - AITER MHA, AITER MLA
|
||||
- Uniform batches only
|
||||
* - ``ROCM_AITER_MLA_SPARSE``
|
||||
- Uniform single-token decode only
|
||||
* - ``TRITON_MLA``
|
||||
* - vLLM Triton MLA
|
||||
- Must exclude attention from graph — ``PIECEWISE`` required
|
||||
|
||||
**Usage examples:**
|
||||
@@ -716,20 +799,20 @@ CUDA graphs reduce kernel launch overhead by capturing and replaying GPU operati
|
||||
Quantization support
|
||||
====================
|
||||
|
||||
vLLM supports FP4/FP8 (4-bit/8-bit floating point) weight and activation quantization using hardware acceleration on the Instinct MI300X, MI325X, MI350X, and MI355X.
|
||||
Quantization of models with FP4/FP8 allows for a **2x-4x** reduction in model memory requirements and up to a **1.6x**
|
||||
improvement in throughput with minimal impact on accuracy.
|
||||
vLLM supports FP4/FP8 (4-bit/8-bit floating point) weight and activation quantization using hardware acceleration on the Instinct MI300X and MI355X.
|
||||
Quantization of models with FP4/FP8 allows for a **2x-4x** reduction in model memory requirements and up to a **1.6x**
|
||||
improvement in throughput with minimal impact on accuracy.
|
||||
|
||||
vLLM ROCm supports a variety of quantization demands:
|
||||
vLLM ROCm supports a variety of quantization demands:
|
||||
|
||||
* On-the-fly quantization
|
||||
* On-the-fly quantization
|
||||
|
||||
* Pre-quantized model through Quark and llm-compressor
|
||||
* Pre-quantized model through Quark and llm-compressor
|
||||
|
||||
Supported quantization methods
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
vLLM on ROCm supports the following quantization methods for the AMD Instinct MI300 series and Instinct MI350 series GPUs:
|
||||
vLLM on ROCm supports the following quantization methods for the AMD Instinct MI300 series and Instinct MI355X GPUs:
|
||||
|
||||
.. list-table::
|
||||
:header-rows: 1
|
||||
@@ -851,23 +934,40 @@ For models without pre-quantization, vLLM can quantize ``FP16``/``BF16`` models
|
||||
GPTQ
|
||||
^^^^
|
||||
|
||||
GPTQ (4-bit/8-bit weight quantization) is fully supported on ROCm via HIP-compiled kernels.
|
||||
Pre-quantized GPTQ models from Hugging Face work out of the box. For better throughput on AMD Instinct GPUs,
|
||||
consider **AWQ with Triton kernels** or **FP8 quantization** instead.
|
||||
GPTQ is a 4-bit/8-bit weight quantization method that compresses models with minimal accuracy loss. GPTQ
|
||||
is fully supported on ROCm via HIP-compiled kernels in vLLM.
|
||||
|
||||
**ROCm support status**:
|
||||
|
||||
- **Fully supported** - GPTQ kernels compile and run on ROCm via HIP
|
||||
- **Pre-quantized models work** with standard GPTQ kernels
|
||||
|
||||
**Recommendation**: For the AMD Instinct MI300X, **AWQ with Triton kernels** or **FP8 quantization** might provide better
|
||||
performance due to ROCm-specific optimizations, but GPTQ is a viable alternative.
|
||||
|
||||
**Using pre-quantized GPTQ models**:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
# Using pre-quantized GPTQ model on ROCm
|
||||
vllm serve RedHatAI/Meta-Llama-3.1-70B-Instruct-quantized.w4a16 \
|
||||
--quantization gptq \
|
||||
--dtype auto \
|
||||
--tensor-parallel-size 1
|
||||
|
||||
**Important notes**:
|
||||
|
||||
- **Kernel support:** GPTQ uses standard HIP-compiled kernels on ROCm
|
||||
- **Performance:** AWQ with Triton kernels might offer better throughput on AMD GPUs due to ROCm optimizations
|
||||
- **Compatibility:** GPTQ models from Hugging Face work on ROCm with standard performance
|
||||
- **Use case:** GPTQ is suitable when pre-quantized GPTQ models are readily available
|
||||
|
||||
AWQ (Activation-aware Weight Quantization)
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
AWQ (Activation-aware Weight Quantization) is a 4-bit weight quantization technique that provides excellent
|
||||
model compression with minimal accuracy loss (<1%). ROCm supports AWQ quantization on the AMD Instinct MI300 series and
|
||||
MI350 series GPUs with vLLM.
|
||||
MI355X GPUs with vLLM.
|
||||
|
||||
**Using pre-quantized AWQ models:**
|
||||
|
||||
@@ -880,7 +980,7 @@ Many AWQ-quantized models are available on Hugging Face. Use them directly with
|
||||
vllm serve hugging-quants/Meta-Llama-3.1-70B-Instruct-AWQ-INT4 \
|
||||
--quantization awq \
|
||||
--tensor-parallel-size 1 \
|
||||
--dtype auto
|
||||
--dtype auto
|
||||
|
||||
**Important Notes:**
|
||||
|
||||
@@ -998,7 +1098,7 @@ Speculative decoding (experimental)
|
||||
===================================
|
||||
|
||||
Recent vLLM versions add support for speculative decoding backends (for example, Eagle‑v3). Evaluate for your model and latency/throughput goals.
|
||||
Speculative decoding is a technique to reduce latency when max number of concurrency is low.
|
||||
Speculative decoding is a technique to reduce latency when max number of concurrency is low.
|
||||
Depending on the methods, the effective concurrency varies, for example, from 16 to 64.
|
||||
|
||||
Example command:
|
||||
@@ -1026,7 +1126,7 @@ Example command:
|
||||
|
||||
It has been observed that more ``num_speculative_tokens`` causes less
|
||||
acceptance rate of draft model tokens and a decline in throughput. As a
|
||||
workaround, set ``num_speculative_tokens`` to <= 2.
|
||||
workaround, set ``num_speculative_tokens`` to <= 2.
|
||||
|
||||
|
||||
Multi-node checklist and troubleshooting
|
||||
@@ -1037,16 +1137,8 @@ Multi-node checklist and troubleshooting
|
||||
3. For GPUDirect RDMA, set ``RCCL_NET_GDR_LEVEL=2`` and verify links (``ibstat``). Requires supported NICs (for example, ConnectX‑6+).
|
||||
4. Collect RCCL logs: ``RCCL_DEBUG=INFO`` and optionally ``RCCL_DEBUG_SUBSYS=INIT,GRAPH`` for init/graph stalls.
|
||||
|
||||
Deprecated terms
|
||||
================
|
||||
|
||||
* **Prefill-Decode attention** has been renamed to **ROCM_ATTN** (ROCm attention). Use ``--attention-backend ROCM_ATTN`` to select this backend.
|
||||
|
||||
Further reading
|
||||
===============
|
||||
|
||||
* :doc:`workload`
|
||||
* :doc:`/how-to/rocm-for-ai/inference/benchmark-docker/vllm`
|
||||
* `ROCm Attention Backend deep-dive <https://vllm.ai/blog/rocm-attention-backend>`_ — architecture and benchmarks for all 7 backends
|
||||
* `vLLM MoE Playbook - A Practical Guide to TP, DP, PP and Expert Parallelism <https://rocm.blogs.amd.com/software-tools-optimization/vllm-moe-guide/README.html>`_ — DP+EP tuning for MoE models
|
||||
* `Multimodal DP optimization <https://rocm.blogs.amd.com/software-tools-optimization/vllm-dp-vision/README.html>`_ — batch-level DP for vision encoders
|
||||
* :doc:`workload-optimization`
|
||||
* :doc:`/ai-inference/vllm`
|
||||
@@ -7,7 +7,7 @@ docker:
|
||||
- "Support fp8 MLA for MI355"
|
||||
- "Block wise sparsity support for AMD triton FAv3 Sage attention"
|
||||
components:
|
||||
TheRock:
|
||||
TheRock:
|
||||
version: cbff3d1
|
||||
url: https://github.com/ROCm/TheRock
|
||||
rocm-libraries:
|
||||
|
After Width: | Height: | Size: 167 KiB |
|
After Width: | Height: | Size: 47 KiB |
|
After Width: | Height: | Size: 778 KiB |
|
After Width: | Height: | Size: 39 KiB |