Compare commits

...
37 Commits
Author SHA1 Message Date
peterjunparkandGitHub ae2f20be45 [docs/7.13.0] Fix dropdown state rendering issue (#6315) 2026-05-29 11:18:16 -04:00
pmoutsias-amdandGitHub 591cc4e484 Merge pull request #6309 from peterjunpark/docs/7.13.0
[docs/7.13] Fix selector state
2026-05-27 22:08:28 -04:00
Peter Park 6cc2f6a2bd js: fix persisted state on init
fix
2026-05-27 22:00:53 -04:00
Peter Park 768ae67199 fix selector 2026-05-27 21:53:29 -04:00
peterjunparkandGitHub 13a0037cd9 [docs/7.13.0] fix(selector.js): Fix dropdown initialization from localStorage data (#6308) 2026-05-27 15:54:14 -04:00
peterjunparkandGitHub a2293976f4 [docs/7.13.0] fix radeon/ryzen>mixed>windows (#6307) 2026-05-27 12:50:24 -04:00
Peter Parkandpeterjunpark 833a3e0160 [docs/7.13.0] fix: install rocdxg for wsl after installing rocm 2026-05-26 17:31:46 -04:00
Peter Parkandpeterjunpark 69daace269 [docs/7.13.0] Add Radeon AI PRO R9700S gpu specs 2026-05-26 13:07:19 -04:00
aa85dae7ea [docs/7.13.0] Add MI350P partitioning support note (#6292)
* [docs/7.13.0] Add MI350P partitioning support note

* fix table rowspan

* update

* Apply suggestions from code review

Co-authored-by: Pratik Basyal <pratik.basyal@amd.com>

---------

Co-authored-by: Pratik Basyal <pratik.basyal@amd.com>
2026-05-21 18:54:17 -04:00
peterjunparkandGitHub 89236dadcb [docs/7.13.0] Add WSL and mixed graphics + compute use case (via amdgpu-install) to installation options (#6291)
* add back graphics/compute use case selector

* add wsl (librocdxg) instructions

* add g++ prereq and update rocminfo output for wsl

* update wordlist.txt

* add ubuntu 26 for wsl

* add 26.04 prereq

* limit WSL to validated GPUs
2026-05-21 18:54:06 -04:00
peterjunparkandGitHub 23a0baa23b [docs/7.13.0] Update runfile to 7.13.0-3 (#6293) 2026-05-21 17:06:53 -04:00
pmoutsias-amdandGitHub c9d3b1a54e Merge pull request #6287 from peterjunpark/docs/7.13.0
[docs/7.13.0] Add AMD Radeon AI PRO R9700S
2026-05-21 09:47:59 -04:00
Peter Park e43f6ea95b [docs/7.13.0] Add AMD Radeon AI PRO R9700S 2026-05-21 09:23:48 -04:00
peterjunparkandGitHub 905ff81ce3 [docs/7.13.0] Add Windows multi-arch installation (#6284)
* [docs/7.13.0] Add Windows multi-arch installation

* fix
2026-05-20 16:53:33 -04:00
peterjunparkandGitHub 91dc6e26d1 [docs/7.13.0] fix(pytorch): single-line pip install cmd for windows (#6283)
* [docs/7.13.0] fix(pytorch): single-line pip install cmd for windows

* update
2026-05-20 13:39:47 -04:00
peterjunparkandGitHub daa7505bc3 [docs/7.13.0] update ROCm ontology diagram 2026-05-19 17:08:19 -04:00
peterjunparkandGitHub 4d2c114a57 [docs/7.13] add device-all to multi-arch pip install (#6277) 2026-05-19 12:04:03 -04:00
peterjunparkandGitHub f7be0c193f [docs/7.13] fix vLLM whl URLs (#6275) 2026-05-19 09:38:19 -04:00
peterjunparkandGitHub a079d03f73 [docs/7.13] Show OEM kernel note for ubuntu 24.04 only (#6274) 2026-05-19 09:11:57 -04:00
pmoutsias-amdandGitHub c048ac762b [docs/7.13.0] Fix broken links and remove glossary from TOC (#6265) 2026-05-19 09:04:09 -04:00
peterjunparkandGitHub 4b9c7a7e76 [docs/7.13.0] fix missing install cmd for gfx1030 rhel and fix conf.py (#6268)
* fix missing install for gfx1030 rhel

* fix header links
2026-05-15 21:17:56 -04:00
yugang-amdandGitHub ed3b02bc61 Remove xref links to legacy portal (#6267) 2026-05-15 20:47:46 -04:00
yugang-amdandGitHub d0a6559915 Fix RVS link (#6266) 2026-05-15 20:25:11 -04:00
Pratik BasyalandGitHub 21d41b8f98 7.13.0 Ryzen AI 9 PRO HX 475, 470 branding updated (#6264)
* Ryzen AI 9 PRO HX 475, 470 updated

* Selector options updated
2026-05-15 19:58:09 -04:00
pmoutsias-amdandGitHub fe6e8e25cf Merge pull request #6263 from peterjunpark/docs/7.13.0
[docs/7.13.0] Temporarily remove unavailable links and update README link
2026-05-15 19:40:44 -04:00
Peter Park 047c32b40b one more 2026-05-15 19:37:23 -04:00
Peter Park 8a615acc76 comment out unavailable links 2026-05-15 19:33:22 -04:00
Peter Park f6a2652ac7 update release note link 2026-05-15 19:33:22 -04:00
pmoutsias-amdandGitHub bed8ad9815 Merge pull request #6262 from peterjunpark/docs/7.13.0
[docs/7.13.0] Fix component GH links
2026-05-15 19:28:27 -04:00
Peter Park d0e0d07817 docs: fix compo github urls 2026-05-15 19:21:26 -04:00
pmoutsias-amdandGitHub ad14fd8081 Merge pull request #6261 from peterjunpark/docs/7.13.0
[docs/7.13]: Fix component version mismatch
2026-05-15 19:08:01 -04:00
Peter Park 8f0e2cfc94 fix components list 2026-05-15 18:59:38 -04:00
Peter Park 7529766d33 [docs/7.13.0] Update docs for 7.13
[docs/7.13.0] document component packages (#736)

* add package list

* add selector

* remove support col

* update package list

[docs/7.13.0] Allow selector options to set multiple values (#737)

* allow selector options to set multiple values

mostly for gfx

* rename "when" to "cond"

* update wordlist

* simplify vllm and ai-ecosystem

* simplify pages using gfx option

update configs

add .editorconfig

add yaml indent_size

add more options

add glossary and gpu hardware specs pages

add components to toc and remove rocm-packages

conf: remove "preview"

install: add amdgpu-lib to pkgman method

fix extra code block && make tabs sync

fix selector option resolution

add blurb explaining graphics vs headless

install: add multi-arch

install: apply feedback (multi-arch)

fix selector

fix

install: fix OEM kernel dropdown for multi-arch selection

install: add intro explaining install method

fix

fix selector js

Reorg and fix compat page

reorg compat

feat(js): add support for dropdown input

chore: reorg inference and dlf pages

chore: update TOC

fix

update toc

update

conf: improve substitutions

conf: add back datatemplates plugin

chore: move RELEASE to about/release-notes.md and don't copy

reorg toc

put contribute under about

install: reorg files and add (gfxXYZ)

reorg selectors

update

use dropdown-input for large list of GPUs and rm selector-info icon

install: remove centos and azl

remove centos and azl from release notes

fix "os-version" tags

feat(selector.py): make dropdown input a separate directive

docs: org deep learning frameworks

yep

fix(selector.js): fix url state flickering

docs: clean up comfyui page

clean up

colocate images

docs: add back env vars page

add remote-content extension

add AI Playbooks to toc and header nav

Restore fine-tuning pages and do some clean-up and fmt

docs: clean up comfyui page

clean up

consistency

Add Docker reference doc

Add Docker reference doc

make anchor text consistent

asdf

install: reorg files and add (gfxXYZ) and add missing Ryzen AI 400
Series (GPT)

reorg selectors

update

update compat

add gfx1030 to selector (radeon pro)

add gfx1152 to selector [Ryzen AI (PRO) 300 Series]

docs: add sglang page stub

fix

chore: update and fmt configs

update release

feat: add metadata to define selector TOC2 heading and icon

restore xdit page

delete xtra xdit data file

bump TomSelect to 2.6.0

organize sphinx extensions and enable legacy selector

add selector metadata

restore HIP programming guide page

Add back setting-cus.rst under env vars page

add back hpc page and organize under resources/

restore gpu-isolation reference page

delete super stale & redudant pages

delete more garb

restore inf/fine-tuning optimization guides and reorg images

update toc and fix other stuff

restore system optimization guides

move up

restore gpu-arch pages

wip

add graph-safe support and hardware atomics support to toc

also update docker doc

use old table in gpu-specs

add sglang doc

make docker run args consistently ordered

organzie gpu-arch files

org

add sphinx_substitution_extensions

update

update me

add release highlights

remove centos and azl from gpu-selector

fix link

update highlights

remove hover effects on diagrams and fix toc

fix toc

fix

remove cursor: pointer

remove miopen release highlight

update

release: add partitioning and virtu support

add component index pages

fix

remove pytorch docker for rn

update ai ecosystem in release notes

add compos to release notes

fix

fix

ROCM-21212 7.13.0 Known issue added (#757)

* add compos to release notes

fix

* ROCM-21212 Known issue added

---------

Co-authored-by: Peter Park <peter.park@amd.com>

ROCM-23565 known issue added (#758)

css/js: add flash on content change

js only

fix links and templating

restructure wip

fixes

fix links

update adrenalin to 26.5.1

update GIM to 9.0.0

update LICENSE copyright year

use includes for release notes

update stuff

update jax, vllm, sglang

add driver prereqs

fix selector in vllm

fix

update hw support table in release notes

fix

Fix compat info and combine Radeon PRO w/ Radeon

Add hipBLASLt, rocSOLVER, and rocSPARSE release highlights

test canonical url

test

Duplication removed

Add RCCL multi-node performance optimization highlight

Remove non-public internal API details from ROCprofiler-SDK changelog

Apply suggestion from @anisha-amd

Replace Strix codename with Ryzen in SQTT decoder highlight

Add Composable Kernel FP8 quantization and SageAttention v2 highlights

Apply suggestions from @amd-jnovotny

Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
Co-authored-by: yugang-amd <yugang.wang@amd.com>

Apply suggestions from @amd-jnovotny

Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>

Update docs/about/includes/core-sdk-components-aggregated-changelog.md

Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>

Update docs/about/includes/core-sdk-components-aggregated-changelog.md

Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>

Confirm version numbers for hipCUB, hipRAND, hipSOLVER, rocPRIM,
rocRAND, rocSOLVER, rocSPARSE, rocThrust, and rocWMMA

install: remove runfile from multi-arch

Add version numbers for rocminfo and ROCprofiler-SDK; fix env var name
in Systems Profiler highlight

Shorten rocSPARSE highlight heading

FFT changelog additions from blocker PR 7089

Remove Ryzen OEM kernel prereq from runfile

7130 known issue batch2 (#763)

* Known issue for ROCM-21815 added

* ROCM-21824 Known issue added

* Systems profiler release highlight updated

* ROCm Systems Profiler link added

remove rocr runtime from diagram

Update SQTT decoder heading and add ROCprofiler-SDK links; minor wording
fix

Remove version numbers from highlight body text for Composable Kernel,
rocSOLVER, and rocSPARSE

Remove ROCprofiler-SDK links from trace decoder highlight

Clarify hipBLASLt General Batched GEMM description

Docs 7.13 structure update (#759)

* Update reference section structure and add MI350 series to GPU arch

* Update docs/reference/gpu-specs.rst

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Update glossary title

* Update docs/components/runtimes-and-compilers.rst

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Split up glossary

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Update docs/sphinx/_toc.yml.in

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Correct CDNA mentions

* Fix atomic add support page

---------

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

Fix typo: GMM -> GEMM in hipBLASLt highlight

fix typo: support --> supported

Update docs/about/release-notes.md

Co-authored-by: pmoutsias-amd <peter.moutsias@amd.com>

Re-add ROCprofiler-SDK links to trace decoder highlight

add links in rocm core sdk diagram

add rocdecode and rocjpeg to stack diagram

remove rocsolver LANGE and GECON changes from release notes

Apply suggestion from @lpaoletti

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

Update docs/about/release-notes.md

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

Update docs/about/release-notes.md

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

fix

Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

Apply PR review suggestions to release notes

- Update APU branding: add "AI Max PRO 300 series" alongside Ryzen AI
Max 300 series
- Broaden CK heading from FP8-specific to general quantization and
capabilities
- Remove premature Roofline limitation note for RDNA 3.5

Add hipBLASLt and rocSOLVER doc links per PR review feedback

Add re-attach to profiled process section to ROCm Systems Profiler
highlights

Apply ROCprofiler-SDK 1.3.0 changelog edits from PR review

remove graphics from "all"

Add resolved issues for RPM install, vLLM, and PyTorch DDP

Add resolved issue for vLLM tensor parallelism launch failure

Update docs/about/includes/core-sdk-components-aggregated-changelog.md

Co-authored-by: spolifroni-amd <Sandra.Polifroni@amd.com>

Fix AMD SMI version to 26.4.0 in aggregated changelog

add component versions to components table

remove gfx1153

Add ROCm 7.13 component versions to components table

Update component table links to therock-7.13

Fix ROCdbgapi version to 0.80.0 in components table

Group release highlights by category for better navigation

Organize the 16 flat H3 highlights into 4 scannable category groups
(Platform and hardware support, AI inference and frameworks, Developer
tools and profiling, Libraries) following the pattern used in ROCm
7.0.0.
Flatten Compute Profiler and Systems Profiler sub-items into bullet
lists to avoid H5 depth.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

add links to 7.9, 7.10, 7.11, and 7.12 release notes

fix gpu lists in release notes and gpu-selector

Add rocSPARSE documentation link to sparse factorization section

fmt html

fix missing version for LLVM and hipinfo

add trademark symbols

Minor fixes (#764)

* Minor fixes

* Minor fixes

archive stale pages

Fix RCCL branding, changelog versions, and doc links

- Use "AMD Ryzen AI Max 300 series" in RCCL heading and body
- Fix ROCdbgapi version in changelog (0.80 -> 0.80.0) to match
components table
- Fix rocSHMEM version in changelog (3.3.0 -> 3.4.0) to match components
table
- Change Systems Profiler doc links from /en/develop/ to /en/latest/

Fix typos, broken links, and style issues in release notes

- Fix 7.11.0 preview link pointing to 7.10.0-preview URL
- Fix re-attach doc link pointing to GitHub instead of rocm.docs.amd.com
- Add missing commas in version list (7.9.0, 7.10.0, 7.11.0)
- Fix "These issue" → "These issues" (grammar)
- Fix cross-reference from "Detailed component changes" to "ROCm
component changelogs"
- Fix "may" → "might" per style guide
- Fix "statisitcs" typo, "MI 350" → "MI350", "API's" → "APIs"
- Replace "Strix Halo" codename with "AMD Ryzen AI Max 300 series"
- Fix "masks values" → "mask values", "Fix" → "Fixed" for consistency
- Merge duplicate Changed sections in AMD SMI changelog
- Add missing blank line before rocSHMEM Changed heading
- Capitalize "cpu" → "CPU"

clean up changelogs formatting

fix missed directory renaming

fix

update install os selectors

update ryzen OSes and kernel versions

update firmware versions

fix table formatting

fix malformed selector

update component support table in release notes

fix include path for windows version selector

remove "graphics and mixed compute" from uninstall

add amdgpu-install and graphics/mixed use case

install: simplify "add additional package repositories"

remove amdgpu-lib

fix os selector (graphics)

remove graphics gpg keys and repos

add all options to multi-arch installation

update runfile installer url

fix compat selector

update install instructions for 7.13

fix rhel ver selector

fix os selector for ryzen

add python 3.14 to windows pip prereq

Fix punctuation, tense, branding, and style in component changelogs

- Add terminal periods to bullet points missing them
- Fix tense consistency (Improves → Improved, fails → failed)
- Add backticks to amd_dbgapi_process_get_info() function name
- Replace Strix Halo codename with AMD Ryzen AI Max 300 series
- Add commas after e.g. per style guide
- Remove spurious commas in restrictive clauses
- Soften future fix commitment (will be fixed → planned to be addressed)
- Add missing article (Fixed an issue)
- Merge upstream branding (AMD Instinct MI350P) with punctuation fixes
- Add colons before sub-lists, periods on sub-list items

Editorial pass on release notes highlights and profiler sections

- Add summary paragraph under Release highlights
- Wrap long lines for maintainability
- Add backticks to code identifiers in HIP section
- Fix RDNA3 → RDNA 3 spacing in CK heading
- Tighten Systems Profiler bullets: remove repetitive product name,
  use consistent declarative tense
- Restructure hipBLASLt and Compute Profiler sections for clarity
- Minor grammar fixes (variants are available, via → through)

Add rocWMMA HIP RTC resolved issue entry for 7.13

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

Remove comp pages (#766)

* Remove comp pages

* Add hw list on indox

* Apply suggestions from code review

Co-authored-by: Istvan Kiss <neon60@gmail.com>

* Apply suggestions from code review

Co-authored-by: Istvan Kiss <neon60@gmail.com>

fix pytorch page selector

fix jax page selector

update 'mixed graphics and compute' wording

update jax install

Remove duplicate trademark symbols after first occurrence

First use of Instinct™, Radeon™, and Ryzen™ is in the release
highlights intro paragraph. All subsequent uses are now plain text.

Apply suggestions from code review

Adding Leo's comments

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

Add rocDecode and rocJPEG highlight and rocWMMA resolved issue

update jax and vllm

removed redundant text in resolved issues

Add rocDecode and rocJPEG highlight and rocWMMA resolved issue

Update docs/about/release-notes.md

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

Fix component pages (#767)

* Fix component pages

fix release notes inbox kernel driver for ryzen

fix virtualization hl

redundant "fixed" in every bullet

Change TheRock to ROCm in rocDecode/rocJPEG highlight heading

update `sudo dnf update` to version specific for OL

rocm -> rocm core sdk

clarify jax env var step

fix typo

Remove redundant TheRock delivery note from rocDecode changelog

Remove MI350 content (#768)

* Remove MI350 content

* Remove AMD GPU driver link

* Update component pages

add LD_LIBRARY_PATH note for jax

fix oracle linux typo

remove install sys libs from jax page

specify jax version on page

fix headings in fw pages

amdgpu-install: remove graphics for ryzen

Add AMD SMI feature highlights section

add links in compo table to changelogs

Remove GPU targets not listed in supported hardware table

Removed from component changelogs:
- gfx1250 (unannounced — MI400/CDNA 5, not in 7.13 hardware table)
- gfx90c (Ryzen APU, not in 7.13 hardware table)
- gfx1153 (not in 7.13 hardware table)
- MI350P (not in 7.13 hardware table)

To be restored if confirmed by hardware PM.

Restore MI350P reference in ROCm Compute Profiler changelog

Link AMD SMI highlights to component changelogs section

Update ROCprofiler-SDK changelog with consolidated 7.13 entries

Port changelog updates from rocm-systems PR #5825:
- Add KFD event tracing, multi-pass counter collection, PC sampling
- Add Removed and Resolved Issues sections
- Fix redundancies, casing, punctuation

move ryzen ai 9 365 to gfx1150

add Ryzen AI 9 HX PRO 375

Apply suggestion from @yugang-amd

Co-authored-by: yugang-amd <yugang.wang@amd.com>

Update rocWMMA version to 2.2.1

New version confirmed by Yugang.

remove selector from rocm packages page

Restore gfx1250 and gfx90c entries in hipBLAS and rocBLAS changelogs

Reverts partial removal from ec26df2. gfx1250 and gfx90c confirmed
for inclusion; gfx1153 and MI350P remain excluded pending hardware PM
confirmation.

Add ROCm Runfile Installer updates to release highlights

Added Installation section with Runfile Installer 7.13 feature summary,
sourced from Jeffrey Novotny. Placed after Platform and hardware
support.

fix

reorganize some images closer to source doc

remove underlines in diagram

Redesign components table support column and update RDC version

- Rename "Support" column to "Supported platforms"
- Replace verbose OS:/GPU: bold labels with compact inline format
  using middot separators and slash grouping
- Drop redundant "only" from support descriptors
- Update RDC version to 1.3.0
- All original support data preserved; no information removed

add versions to components in compat

remove ai playbooks

7130 known issue batch3 (#765)

* Review feedback added

* ROCM-21706 Known issue added

* ASAN issue added

* Known issue added

* PyTorch issue added

* Known issues updated

* Known issue for ASAN added

* Known issue added

* Known issue added

* Minor change

* Leo's review feedback incorporared

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

---------

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

fix selectors

clean up selectors

fix compat mat

update vllm version to 0.19.1 in highlights

update selector padding and colors

add hover effect to page dropdown

fix component versions, move installer section, and style cleanup

- Correct versions: hipSOLVER 3.4.0, MIOpen 3.5.1, rocSOLVER 3.34.0, CK
1.2.0
- Move Runfile Installer updates out of highlights into own section
- Collapse double-spaced lists, fix punctuation, hyphenate closed-source
- Fix rocminfo and ROCprofiler-SDK changelog heading format

editorial sweep: highlights, changelog, and link fixes

- Merge two Composable Kernel highlight sections into one
- Fix rocDecode/rocJPEG platform support to include Ryzen AI
- Update external doc links from /en/latest/ to /en/docs-7.13/
- Remove duplicate API list in AMD SMI changelog
- Normalize changelog headings (Resolved issues, Optimized)
- Fix typos and formatting (ROCprofiler-SDK casing, hyphenation,
fragments)
- Align vLLM/SGLang wording

clean up archived vllm pages

update wordlist

add mi350p

add ryzen pro 200 series

add ryzen ai PRO 400

fix

add ryzen ai PRO 7 / 5 (gfx1152)

rm gpus not listed in go/no-go

Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

update wordlist and fix typos

improve toc

address linting issues

address markdown linting

fix md linting issues

fix rest of linting issues

conf: reenable intersphinx fetching

update vllm and sglang toc text

fix vllm selector

editorial: fix version string and ASAN sentence in release notes

- "ROCm 7.13" → "ROCm 7.13.0" (two occurrences)
- Reword ASAN intro sentence for directness

docs: add AMD SMI (BM) 26.4.0 changelog entry and fix anchor link

docs: add RDC 1.3.0 changelog entry and anchor link

docs: add RDC release highlight to 7.13 release notes

Add ROCm Data Center Tool (RDC) entry under the Libraries section,
documenting its addition to the ROCm Core SDK for Linux with
AMD Instinct GPUs.

Ref: ROCM-21708, AIROCDOC-3688

7.13.0 Known Issues Batch 4 (#775)

* LLVM issue added

* Minor spacing issue

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Space fixed

* Linting error fixed

---------

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

docs: remove re-attach highlight from release notes

docs: remove consolidated changelog note from release notes

The consolidated changelog reference is not applicable to the
ROCm Core SDK release notes.

docs: add RCCL GDA alltoall highlight and fix version references

- Add RCCL GDA-based alltoall via rocSHMEM integration highlight
(ROCM-2288)
- Fix ROCm 7.12 → 7.12.0 version references for consistency

docs: update Composable Kernel version to 1.3.0

install: update runfile quick start cmd

add pytorch install commands for gfx1030 and and gfx1152

Fix RDNA3.5 mention and update ROCm Programming page (#774)

* Fix RDNA3.5 mention

* Update HIP Programming Guide

* Update HIP Programming Guide

* Update HIP Programming Guide

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Reorg reference section

* Update HIP Programming

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* WIP

* WIP

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* WIP

* Update RDNA2 system optimization page

* Update RDNA3.5

* WIP

* WIP

* WIP

* WIP

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Update HIP Programming

* WIP

* Update HIP Programming

* Update RDNA2 system optimization page

* Update RDNA2 system optimization page

* Update HIP Programming

* Update ROCm programming title

---------

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

docs: add transition guide to TOC and docs

docs: move transition guide from conceptual to about

docs: fill empty cells in summary table

docs: fill remaining empty cells in summary table

Update transition-guide-TheRock.md

Update transition-guide-TheRock.md

Update transition-guide-TheRock.md

Update transition-guide-TheRock.md

Update transition-guide-TheRock.md

Update transition-guide-TheRock.md

Update transition-guide-TheRock.md

Update transition-guide-TheRock.md

Update transition-guide-TheRock.md

Update transition-guide-TheRock.md

docs: convert remaining tables to HTML for consistent styling

docs: fix inline code styling in HTML tables

js: fix query param resolution

js: fix selected content resolution

docs: update ROCm Compute Profiler highlights to match AIPROFCOMP-493

- Correct product branding (AMD Ryzen AI Max 300 series processors)
- Narrow scope from RDNA to RDNA 3.5 devices
- Add roofline limitation note
- Add roofline.csv detail for profile mode

docs: clarify Core SDK package annotations in transition guide

- Add legacy package names for rocDecode, rocJPEG, and RDC
- Change "new" to "newly included" for clarity

docs: scope Compute Profiler RDNA 3.5 support to Ryzen AI Max 300

Per dev feedback, support is gfx1151 only — not all RDNA 3.5 devices.

docs: remove analysis mode dependency mention per dev feedback

docs: restore analysis mode dependency note per dev clarification

docs: consolidate planned components under future releases

Per dev feedback, hipfort, rocALUTION, rocPyDecode, rocAL, and MIVisionX
will release with ROCm-Extras, not separately.

docs(jax): add jax version selector

docs(jax): use `--index-url` instead of `--extra-index-url`

Update 200-install.rst

Update 200-install.rst

docs(jax): update install snippet formatting

docs(pyt): add pytorch version selector

docs: remove rocm packages page

Revert "Update 200-install.rst"

This reverts commit 291c1fe3747227372f0ac183447ff43561fb6ed6.

Revert "Update 200-install.rst"

This reverts commit babcaeb48582e4e8168073b7fcb70b127814eace.

7.13.0 Known issues updated Batch 5 (#776)

* Known issues updated

* Review feedback added

* Known Issue added

* Feedback incoporated

* Known issue added

* hipDF known issues removed

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Review feedback added

* PLDM updated

---------

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

docs: add associated packages column to meta packages table

- Add 4th column listing associated meta packages and packages
- Remove amdrocm-opencl7.13 row (not a meta package)

docs: add cross-reference link to transition guide from meta packages
table

docs(jax): fix env var order for clarity

docs: fix meta packages table formatting with line blocks

docs: add consistent package/meta package labels to all table rows

docs: remove deleted rocm-packages reference from conf.py

docs: fix spacing in meta packages table for consistent formatting

docs: move transition guide cross-reference into core-dev table row

docs: fix transition guide cross-reference to target specific section

docs(vllm): update with addition skus and selectors

docs: move cross-reference link back to after meta packages table

css: add hover highlight to dropdown selector

docs: fix cross-reference to transition guide using Sphinx ref label

docs(jax/pyt): show amdgpu prereq for instinct/radeon only

docs: merge use case and packages columns in meta packages table

docs(vllm): fix vllm version key

docs: clean up and fmt

docs(vllm): make amdgpu prereq conditional

fix

docs: add Windows package availability note to transition guide

docs: wrap Windows package note in admonition block

docs: fix subproject doc links from docs-7.13 to docs-7.13.0

docs: fix incorrect component names in transition guide packages table

- rocm-systems → rocprofiler-systems
- rocm-compute → rocprofiler-compute
- tracer → roctracer

fix mi350p fields in compat page

Clarify ROCm Core SDK 7.13.0 as a preview release

Updated the description of ROCm Core SDK 7.13.0 to reflect its status as
a preview release.

update extras urls in toc

Update ROCm documentation theme options

Update conf.py

rocThrust environment variable changelog removed (#778)

clean up wording

docs(install): remove uninstall step 3

docs(jax): remove LD_LIBRARY_PATH and AMD_COMGR_NAMESPACE workarounds

docs(sglang): remove and keep in exclude dir/

docs(vllm): fix missing gfx1150

docs: add technology preview messaging to release notes intro

Aligns 7.13 release notes with the preview release stream framing
used in prior releases (7.9-7.12). Updates the admonition from note
to important and consolidates prior version links into a single
release history link.

remove SGLang from release notes, compat

docs: fix missing article in transition guide intro

docs: editorial fixes in transition guide

Standardize note format, fix column placement for newly included
annotation, replace em dashes with double dashes, and minor copy edits.

docs: editorial fixes in transition guide

Standardize note format, fix line wrap, clarify parentheticals in
package tables, add deprecation link for ROCm SMI, and minor copy edits.

remove tb and rbt

update docker pull tag to 0.19.1 vllm

7.13.0 Known Issues Batch 6 (#779)

* dev/devel issue added

* Feedback incorporated

* Review feedback added

* Update docs/about/release-notes.md

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

---------

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

docs(vllm): update tags and whl urls

docs(vllm): add miising pytorch install for gfx110X

docs(vllm): indentation fix

docs: clean up links

docs: get rid of system-setup for inference for now

docs: update GA date to 5-15

docs: convert component pages to table format with working hyperlinks

Replace dead :doc: cross-references with versioned URLs to component
documentation. Convert category pages from bullet lists to list-table
format for scannability.

docs: switch component tables from list-table to matrix directive

docs: update runfile installer to 7.13.0-2

docs(vllm): fix torch version

docs: remove contributing docs for now

docs: fix ai developer hub link

docs: reorder resolved issues

docs: fix links and sphinx warnings

update GA date in versions list

docs: fix the rest of the sphinx warnings

docs(xdit): update to 26.5

docs: update .wordlist.txt

docs(jax): update 0.8.2 jaxlib instructions

docs(jax): rmove LD_LIBRARY_PATH note

docs(install): update multi-arch repo path for pkgman

fix install

docs(install): fix debian multi-arch

docs: revert meta packages table to original 3-column layout

docs(install): update amdgpu-install url

PLDM table updated (#781)

Revert "docs: switch component tables from list-table to matrix
directive"

This reverts commit 774dcf59d97c4e59c67db13899aed4f6efbd210d.

Revert "docs: convert component pages to table format with working
hyperlinks"

This reverts commit 04b3aa9378c57549c44fa2bee43a5c288fd2dd6b.

docs: update toc for extras

update mi350p ifwi in compat matrix

fix rowspan

reorder toc

docs(post-install): update amd-smi and rocminfo sample outputs

docs: update wordlist

docs: remove Windows package note from transition guide

docs(amdgpu-install): remove mention of amdgpu-install doc

docs: remove unsupported arch rows and fix extras path and library name
in transition guide

docs(amdgpu-install): add uninstall step to rm repos

docs: add ROCdbgapi and ROCr Debug Agent to amdrocm-debugger package
contents

docs(install): fix missing zypper install for multi-arch

PLDM version, link updated (#782)

docs(install): fix missing gfx1030 gfx1152 tarball

docs(install): fix

fix

fix

docs: add missing Radeon RX 9070 GRE link in hardware support table

fix unreliable javascript

docs(install): OEM kernel ub2404 only

docs(isntall): fix incorrect windows nesting

fix extra backslash

add mi350p to vllm

fix

make values explicit and update 31.30.0 driver url to preview

add new vllm docker tags

make oem kernel prereq 24.04 only

remove amdgpu-install

remove graphics note for radeon/ryzen

fix
2026-05-15 18:31:19 -04:00
Peter Park f45bd5a570 [docs/7.12.0] Update docs for 7.12 preview release
build: update RTD env config

docs(release): update virtualization support

docs: reorganize custom extensions

add tom-select lib

[docs/7.12.0] Update documentation for TheRock preview release 7.12.0
(#6072)

docs(release): add gpus and xrefs

docs(release): update release notes for 7.12

wip: selector-toc2 dropdown input

remove selector tiles from install sections

docs(compat): virtualization support

docs(uninstall): remove unneeded headings

js(selector toc2): improve semantic html in toc2 selector dropdowns

fix(js/css): TomSelect

fix(toc2.js): heading query

fix(css): maximize dropdown input widths to prevent stuttering

fix and reorg

bump rocm-docs-core to 1.32.1

[docs/7.12.0] Update documentation for TheRock 7.12.0

Includes related enhancements:
- Improve secondary sidebar display
- Fix install instruction issues
- Add JAX and vLLM
- Reorganize site structure
- Tweak CSS/JS/Py

update versions list

wording

update driver docs links

docs: fix jax instructions

docs: fix vllm instructions

docs: reorganize some files

reorg

[docs/7.12.0] update vllm docker pull tags

fix

[docs/7.12.0] fix install docs

[docs/7.12.0] fix install instructions and conditional sections

[docs/7.12.0] Fix repo.amd.com url for gfx110X

[docs/7.12.0] Add JAX known issue

[docs/7.12.0] Update JAX known issue description

[docs/7.12.0] document `CK_AMD_GPU_GFX*` known issue

[docs/7.12.0] document more known issues

[docs/7.12.0] Remove GPU partitioning support note (#6084)

Revert "[docs/7.12.0] Remove GPU partitioning support note (#6084)"
(#6085)

This reverts commit 6ac90bdf9e.

[docs/7.12.0] Add ROCm Optiq (Beta) release highlight (#6086)

[docs/7.12.0] docs(install): Fix pip install url `gfx120x-all` -->
`gfx120X-all` (#6087)

X needs to be capitalized

[docs/7.12.0] Update install instructions (#6104)

* Link to i=runfile query param in release note

* fix debian version selector for mi300a, mi250x, mi250

* remove sles prereq

* fix bashrc tarball post-install

add uninstall

* fix relative url

* remove `sudo` from user-local env setup

[docs/7.12.0] Add `bash` to docker run cmds (#6107)

[docs/7.12.0] Document workarounds for vLLM installation via pip (#6108)

* [docs/7.12.0] Document workarounds for vLLM

* [docs/7.12.0] Update `amd-smi version` sample output

* update text

update formatting

note that PyT 2.9.1 is required

update

* fix

fix links

rm extra gfx103x sections

[docs/7.12.0] mention nightlies for gpus w/o "official" support (#6147)

fix markup

wording

words

consistency

cleanup

[docs/7.12.0] Update post-install and uninstall with LD_LIBRARY_PATH
workaround + update runfile installer to 7.12.0-2 (#6151)

* docs(install): Update ROCm post-install and uninstall with
LD_LIBRARY_PATH workarounds

* docs(install): Update runfile installer to 7.12.0-2.run
2026-05-08 17:19:52 -04:00
Peter Park 3375a3a668 [docs/7.11.0] Update docs for 7.11 preview release
update ROCM_VERSION in conf

7.11 known issues added

Minor change

docs(RELEASE): supported OSes and hw

docs(release.md): update hw support and os support

js(selector-toc): remove unused code

docs(index): update rocm ontology diagram

add TM symbols

docs(RELEASE): complete virtualization support tbl

docs(toc): add rocm-examples to toc

chore: bump rocm-docs-core to 1.31.3

chore: update version histor page

docs(RELEASE): add oses

docs(RELEASE): add virtu sup

fix

docs: update RELEASE.md

docs(RELEASE): clean up

py(selector): allow percentage widths

docs(compat): update system-instinct table

docs: finalize components lists

docs(compat): update system-radeon-pro

docs(compat): update system-radeon

docs: compat

docs(RELEASE): clean up

docs(RELEASE): update tables

docs(compat): add virtu sup and fix stuff

docs(RELEASE): update AMD GPU Driver vers

dcos(compat): add missing mi2xx options

docs(RELEASE): update firmware for instinct

docs(RELEASE): fix xref and fmt

docs(rocm-ontology diagram): fix sideways tm

fix

oops, fix

docs: remove ROCgdb

docs(compat): add rhel 8.10 to mi35x

docs: clean up some wording

docs(install): update selector

docs(install/compat): fix selector data

docs: fix

js: fix reconcile selections

wip: install

wip: install

js: add URLSearchParams

wip: install

js: inline page-specific js

docs(compat): add missing rocky linux to mi300x

docs(install): windows adrenalin prereq

wip: install

docs(compat): virtu sup link

docs: windows tar

docs(compat/install): add missing gfx120x, gfx103x, and gfx110x gpus

docs(RELEASE): add missing gfx120x, gfx103x, and gfx110x gpus

docs(install): fix selected content

docs(install): windows tar

rm jira notes

docs(comfyui): update

docs(comfyui): update path

docs: remove hipDNN

chore(conf.py): exclude `**/includes/**` to build only when included

docs(index): fix diagram colors

Update RELEASE.md

remove fn

docs: remove gfx1030 cards

docs: remove windows for gfx12

js: add localStorage to selector for persistence

docs(RELEASE): fix gpu list in tables

js(primary toc install headings): don't reload page if already on the
install page

docs(img): add oses to ontology diagram

docs(rocgdb): add it back + document known issue

docs(comfyui): update selector

docs(RELEASE): fix

docs(compat): fix - add rocky linux for mi300a

docs(install): sles - remove --gpg-auto-import-keys

py(selector): add static assets only if exists

js(selector): clear URLSearchParams if page doesn't have selector

docs: note GIM driver 8.7.0K for virtualization on instinct

add link to gim docs

docs: reorder list of components

docs(install): add oem kernel prereq

docs: clean up diagrams

fix link

docs(RELEASE): update firmware explanation

docs: update

chore: clean up extra files

chore: linting errors

docs(compat): add iGPU to ryzen cards in compat matrix

docs(install): add install libatomic1

docs(install): clean up

fix

docs(install): add meta packages table

docs(compat): remove amdgpu 7.0.3 and 6.4.2

docs(install): uninstall add meta package note

blurb

docs(RELEASE): apex known issue

docs: update components lists in RELEASE and compat matrix

docs(install): clean up meta packages sections

docs(compat): update fw version for mi300x

update known issues

docs: remove minor version from OL and Rocky

docs: fix windows whl urls

docs(install): update rocminfo and amd-smi version example outputs

docs(RELEASE): add known issues

docs: fix linux components list

docs(compat): make igpu heading consistent

Add rocm-examples known issue

Co-authored-by: Istvan Kiss <istvan.kiss@amd.com>

docs(conf): update release date

docs(install): update windows install instructions

update

docs(RELEASE): update known issues

hipblaslt matmul known issue

update known issues

add hipify-clang known issue

js(selector): fix syncStateToURL

update amdgpu driver versions and known issues

update known issues

fix

fix

fix install urls

update RELEASE.md

fix

update for consistency

add known issue

add rccl known issue

update components table

fix amdgpu versions

update prereqs intro

docs(conf): turn of toc_exclude_missing (#5957)

[docs/7.11.0] Fix some release notes documentation and remove unneeded
SLES packages (#5960)

* fix release date and known issue

* add llama.cpp known issue and fix link to amdgpu 31.10.0

* docs: minor fixes

* fx

* clean up known issues

* clean up

[docs/7.11.0] Add minor corrections (#5961)

* fix comfyui linux/windows options

* rm minor version from oracle linux ver

update LD_LIBRARY_PATH config (#5964)

[docs/7.11.0] Add llama.cpp known issue (#5962)

[docs/7.11.0] fix `cd C:\TheRock` (#5965)

* [docs/7.11.0] fix `cd C:\TheRock`

* fix description

[docs/7.11.0] Add minor corrections (#5976)

* add hipinfo to release notes

* fix typo

* fix rccl links

* fix warning: pygments lexer `cmd` unknown

* fix table row alignment

[docs/7.11.0] Update known issues #5975

update ROCM_VERSION

wip: add compat/install data
2026-05-08 17:19:52 -04:00
Peter Park f6c4835da7 [docs/7.10.0] Update docs for 7.10 preview release
fix conf.py (#5762)

[docs/7.10.0] selector.py: Make ids even more unique for selected
content (#5765)

[docs/7.10.0] Fix typo in venv command and installation prerequisites
(#5766)

* fmt

* Fix installation prerequisites and source venv typo

clean up ./.venv/... --> .venv/... (#5767)

[docs/7.10.0] selector: responsive css for narrow viewports (#5769)

Fix element overflow on narrow/mobile screens.

Fix incorrect warning in matrix.py extension

[docs/7.10.0] Radeon cards: suggest amdgpu over RSL driver (#5770)

[docs/7.10] post-install - improve exports (#5783)
2026-05-08 17:19:52 -04:00
Peter Park be0b93b8c6 [docs/7.9.0] Add docs for 7.9.0 preview release
Add release notes

Add install instructions

Add PyTorch + ComfyUI instructions

Add custom selector directives

Add JS and CSS for selector

Add custom icon directive and utils

Clean up conf.py

update PLDM bundle version for MI355

Add "preview" to headings (#5545)

Add custom version history (#5551)

[docs/7.9.0] Fix GPU marketing names in 7.9.0 release.md and
compatibility matrix / Add SD3.5 ComfyUI example (#5552)

* Add SD3.5 example to comfyui doc

* Fix Ryzen AI Max (PRO) SKU names

* Add names in multi line format

[docs/7.9.0] Add build from source overview page / Point to
`therocm-7.9.0` in components list (#5546)

* Update links to components to point the `therock-7.9.0` ref

* Add build from source page

* lint: fix caps and update .wordlist.txt

* add link to "development manuals" list

* add links to TheRock's development guide and fix step 4

* wording and fmt

* fix spacing

* fix fmt

* Fix documentation linting errors

* fix spacing

[docs/7.9.0] Note rocprofiler-sdk is Instinct only. Reorg some files to
match `docs/7.0.x`. (#5563)

* move versions.md and compat-matrix to match prod; note rocprofiler-sdk
is instinct only

* update href in versions.md

[docs/7.9.0] Use "generic" rocm-docs-core theme (#5568)

* use "generic" rocm-docs-core theme to tweak header

* restore "nav_secondary_items"

[docs/7.9.0] Fix xref in Ubuntu prerequsites and RST heading overline
(#5569)

update rocm-docs-core to 1.29.0

(cherry picked from commit 39de859bd1)

update rocm-docs-core to show preview banner

add preview announcement

[docs/7.9.0] Add xDiT diffusion inference doc (#5676)

[docs/7.9.0] Fix rocm-cmake github link due to non-existent tag (#5684)

[docs/7.9.0] Update banner msg (#5704)

update
2026-05-08 17:19:43 -04:00
410 changed files with 20354 additions and 7039 deletions
-1
View File
@@ -17,6 +17,5 @@ __pycache__/
# avoid duplicating contributing.md due to conf.py
docs/contribute/index.md
docs/about/release-notes.md
docs/release/changelog.md
.claude/settings.local.json
+10 -12
View File
@@ -4,23 +4,21 @@
version: 2
sphinx:
configuration: docs/conf.py
configuration: docs/conf.py
formats: [htmlzip]
formats: []
python:
install:
- requirements: docs/sphinx/requirements.txt
install:
- requirements: docs/sphinx/requirements.txt
build:
os: ubuntu-22.04
tools:
python: "3.10"
apt_packages:
- "doxygen"
- "gfortran" # For pre-processing fortran sources
- "graphviz" # For dot graphs in doxygen
os: ubuntu-24.04
tools:
python: "3.10"
search:
ignore:
- "**/previous-versions/**"
- "**/archive/**"
- "**/include/**"
- "**/redirect/**"
+320 -139
View File
File diff suppressed because it is too large Load Diff
+1 -1
View File
@@ -1,6 +1,6 @@
MIT License
Copyright (c) 2023 - 2025 Advanced Micro Devices, Inc. All rights reserved.
Copyright (c) 2023 - 2026 Advanced Micro Devices, Inc. All rights reserved.
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
+1 -1
View File
@@ -136,7 +136,7 @@ For a complete list of ROCm components and version information, see the
## Release notes
- [Latest version of ROCm](https://rocm.docs.amd.com/en/latest/about/release-notes.html) - production
- [ROCm 7.12.0](https://rocm.docs.amd.com/en/7.12.0-preview/about/release-notes.html) preview stream
- [ROCm 7.13.0](https://rocm.docs.amd.com/en/7.13.0-preview/about/release-notes.html) preview stream
---
-607
View File
@@ -1,607 +0,0 @@
<!-- Do not edit this file! -->
<!-- This file is autogenerated with -->
<!-- tools/autotag/tag_script.py -->
<!-- Disable lints since this is an auto-generated file. -->
<!-- markdownlint-disable blanks-around-headers -->
<!-- markdownlint-disable no-duplicate-header -->
<!-- markdownlint-disable no-blanks-blockquote -->
<!-- markdownlint-disable ul-indent -->
<!-- markdownlint-disable no-trailing-spaces -->
<!-- markdownlint-disable reference-links-images -->
<!-- markdownlint-disable no-missing-space-atx -->
<!-- spellcheck-disable -->
# ROCm 7.2.3 release notes
The release notes provide a summary of notable changes since the previous ROCm release.
- [Release highlights](#release-highlights)
- [Supported hardware, operating system, and virtualization changes](#supported-hardware-operating-system-and-virtualization-changes)
- [User space, driver, and firmware dependent changes](#user-space-driver-and-firmware-dependent-changes)
- [ROCm components versioning](#rocm-components)
- [Detailed component changes](#detailed-component-changes)
- [ROCm known issues](#rocm-known-issues)
- [ROCm upcoming changes](#rocm-upcoming-changes)
```{note}
If youre using AMD Radeon™ GPUs or Ryzen™ for graphics workloads, see the [Use ROCm on Radeon and Ryzen](https://rocm.docs.amd.com/projects/radeon-ryzen/en/latest/index.html) documentation to verify compatibility and system requirements.
```
## Release highlights
The following are notable new features and improvements in ROCm 7.2.3. For changes to individual components, see
[Detailed component changes](#detailed-component-changes).
### Supported hardware, operating system, and virtualization changes
Hardware, operating system, and virtualization support remains unchanged in this release.
For more information about:
* AMD hardware, see [Supported GPUs (Linux)](https://rocm.docs.amd.com/projects/install-on-linux/en/docs-7.2.3/reference/system-requirements.html#supported-gpus).
* Operating systems, see [Supported operating systems](https://rocm.docs.amd.com/projects/install-on-linux/en/docs-7.2.3/reference/system-requirements.html#supported-operating-systems) and [ROCm installation for Linux](https://rocm.docs.amd.com/projects/install-on-linux/en/docs-7.2.3/).
* Virtualization support, see [Virtualization support](https://rocm.docs.amd.com/projects/install-on-linux/en/docs-7.2.3/reference/system-requirements.html#virtualization-support).
### User space, driver, and firmware dependent changes
The software for AMD Data Center GPU products requires maintaining a hardware
and software stack with interdependencies among the GPU and baseboard
firmware, AMD GPU drivers, and the ROCm user space software. While AMD publishes drivers and ROCm user space components, your server or infrastructure provider publishes the GPU and baseboard firmware by bundling AMDs firmware releases via the AMD Platform Level Data Model (PLDM) bundle, which includes the Integrated Firmware Image (IFWI).
GPU and baseboard firmware versioning might differ across GPU families.
<div class="pst-scrollable-table-container">
<table class="table table--middle-left">
<thead>
<tr>
<th class="head">
<p>ROCm Version</p>
</th>
<th class="head">
<p>GPU</p>
</th>
<th class="head">
<p>PLDM Bundle (Firmware)</p>
</th>
<th class="head">
<p>AMD GPU Driver (amdgpu)</p>
</th>
<th class="head">
<p>AMD GPU <br>
Virtualization Driver (GIM)</p>
</th>
</tr>
</thead>
<style>
tbody#virtualization-support-instinct tr:last-child {
border-bottom: 2px solid var(--pst-color-primary);
}
</style>
<tr>
<td rowspan="9" style="vertical-align: middle;">ROCm 7.2.3</td>
<td>MI355X</td>
<td>
01.26.00.02<br>
01.25.17.07<br>
01.25.16.03
</td>
<td>
30.30.x where x (0-3)<br>
30.20.x where x (0-1)<br>
30.10.x where x (0-2)
</td>
<td rowspan="3" style="vertical-align: middle;">8.7.1.K</td>
</tr>
<tr>
<td>MI350X</td>
<td>
01.26.00.02<br>
01.25.17.07<br>
01.25.16.03
</td>
<td>
30.30.x where x (0-2)<br>
30.20.x where x (0-1)<br>
30.10.x where x (0-2)
</td>
</tr>
<tr>
<td>MI325X<a href="#footnote1"><sup>[1]</sup></a></td>
<td>
01.25.06.08<br>
01.25.04.02
</td>
<td>30.30.x where x (0-2)<br>
30.20.x where x (0-1)<a href="#footnote1"><sup>[1]</sup></a><br>
30.10.x where x (0-2)<br>
6.4.z where z (0-3)<br>
6.3.3
</td>
</tr>
<tr>
<td>MI300X<a href="#footnote2"><sup>[2]</sup></a></td>
<td>01.25.06.04<br>
01.25.03.12<br>
01.25.02.04</td>
<td rowspan="6" style="vertical-align: middle;">
30.30.x where x (0-2)<br>
30.20.x where x (0-1)<br>
30.10.x where x (0-2)<br>
6.4.z where z (03)<br>
6.3.3
</td>
<td>8.7.1.K</td>
</tr>
<tr>
<td>MI300A</td>
<td>BKC 26.1</td>
<td rowspan="3" style="vertical-align: middle;">Not Applicable</td>
</tr>
<tr>
<td>MI250X</td>
<td>IFWI 47 (or later)</td>
</tr>
<tr>
<td>MI250</td>
<td>MU5 w/ IFWI 75 (or later)</td>
</tr>
<tr>
<td>MI210</td>
<td>MU5 w/ IFWI 75 (or later)</td>
<td>8.7.1.K</td>
</tr>
<tr>
<td>MI100</td>
<td>VBIOS D3430401-037</td>
<td>Not Applicable</td>
</tr>
</table>
</div>
<p id="footnote1">[1]: For AMD Instinct MI325X KVM SR-IOV users, don't use AMD GPU driver (amdgpu) 30.20.0.</p>
<p id="footnote2">[2]: AMD Instinct MI300X KVM SR-IOV with Multi-VF (8 VF) support requires a compatible firmware BKC bundle, which will be released in the coming months.</p>
### Improved profiling accuracy for vLLM workloads
ROCm 7.2.3 improves profiling stability for vLLM workloads traced with PyTorch `torch.profiler`. The large, sporadic idle gaps that previously appeared between GPU kernels in the trace have been substantially reduced in common configurations, and the traces now more accurately reflect actual runtime behavior. Coverage may vary depending on model and parallelism settings; additional improvements are in progress.
### MIGraphX update
[MIGraphX](https://rocm.docs.amd.com/projects/AMDMIGraphX/en/docs-7.2.3/index.html) has the following enhancements:
#### Improved performance of the Gather operator
Performance for embeddingheavy inference workloads is improved by merging multiple independent gather operations from similar embedding tables into a single batched operation. Multigather workloads now run more efficiently with fewer kernel launches and reduced memory traffic by adding horizontal fusion for cross-embedding gather operators. These gather operators have been updated to use `transpose`/`reshape`/`broadcast`/`slice`, enabling better optimization across different backends and data layouts.
#### ONNX Runtime reliability improvement
ONNX Runtime workloads accelerated with MIGraphX now provide a more reliable experience through external stream support in the MIGraphX Execution Provider, with improved memory allocation and deallocation for multi-stream inference.
### ROCm documentation updates
ROCm documentation has been updated with ROCm XIO documentation. ROCm XIO provides an API for Accelerator-Initiated IO (XIO) for an AMD GPU `__device__` code. It enables AMD GPUs to perform direct IO operations to hardware devices without CPU intervention. ROCm XIO was initially released in April 2026 as an early-access software technology preview. Running production workloads is not recommended.
For more information, see the [ROCm XIO documentation](https://rocm.docs.amd.com/projects/rocm-xio/en/beta-0.1.0/index.html) and {fab}`github` [ROCm/rocm-xio](https://github.com/ROCm/rocm-xio) GitHub repository.
## ROCm components
The following table lists the versions of ROCm components for ROCm 7.2.3, including any version
changes from 7.2.2/7.2.1 to 7.2.3. Click the component's updated version to go to a list of its changes.
Click {fab}`github` to go to the component's source code on GitHub.
<div class="pst-scrollable-table-container">
<table id="rocm-rn-components" class="table">
<thead>
<tr>
<th>Category</th>
<th>Group</th>
<th>Name</th>
<th>Version</th>
<th></th>
</tr>
</thead>
<colgroup>
<col span="1">
<col span="1">
</colgroup>
<tbody class="rocm-components-libs rocm-components-ml">
<tr>
<th rowspan="9">Libraries</th>
<th rowspan="9">Machine learning and computer vision</th>
<td><a href="https://rocm.docs.amd.com/projects/composable_kernel/en/docs-7.2.3/index.html">Composable Kernel</a></td>
<td>1.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/composablekernel"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/AMDMIGraphX/en/docs-7.2.3/index.html">MIGraphX</a></td>
<td>2.15.0&nbsp;&Rightarrow;&nbsp;<a href="#migraphx-2-15-0">2.15.0</a></td>
<td><a href="https://github.com/ROCm/AMDMIGraphX"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/MIOpen/en/docs-7.2.3/index.html">MIOpen</a></td>
<td>3.5.1</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/miopen"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/MIVisionX/en/docs-7.2.3/index.html">MIVisionX</a></td>
<td>3.5.0</a></td>
<td><a href="https://github.com/ROCm/MIVisionX"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocAL/en/docs-7.2.3/index.html">rocAL</a></td>
<td>2.5.0</a></td>
<td><a href="https://github.com/ROCm/rocAL"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocDecode/en/docs-7.2.3/index.html">rocDecode</a></td>
<td>1.7.0</a></td>
<td><a href="https://github.com/ROCm/rocDecode"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocJPEG/en/docs-7.2.3/index.html">rocJPEG</a></td>
<td>1.4.0</a></td>
<td><a href="https://github.com/ROCm/rocJPEG"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocPyDecode/en/docs-7.2.3/index.html">rocPyDecode</a></td>
<td>0.8.0</a></td>
<td><a href="https://github.com/ROCm/rocPyDecode"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rpp/en/docs-7.2.3/index.html">RPP</a></td>
<td>2.2.1</a></td>
<td><a href="https://github.com/ROCm/rpp"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
</tbody>
<tbody class="rocm-components-libs rocm-components-communication tbody-reverse-zebra">
<tr>
<th rowspan="2"></th>
<th rowspan="2">Communication</th>
<td><a href="https://rocm.docs.amd.com/projects/rccl/en/docs-7.2.3/index.html">RCCL</a></td>
<td>2.27.7</a></td>
<td><a href="https://github.com/ROCm/rccl"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocSHMEM/en/docs-7.1.0/index.html">rocSHMEM</a></td>
<td>3.2.0</a></td>
<td><a href="https://github.com/ROCm/rocSHMEM"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
</tbody>
<tbody class="rocm-components-libs rocm-components-math tbody-reverse-zebra">
<tr>
<th rowspan="16"></th>
<th rowspan="16">Math</th>
<td><a href="https://rocm.docs.amd.com/projects/hipBLAS/en/docs-7.2.3/index.html">hipBLAS</a></td>
<td>3.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipblas"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/hipBLASLt/en/docs-7.2.3/index.html">hipBLASLt</a></td>
<td>1.2.2</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipblaslt"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/hipFFT/en/docs-7.2.3/index.html">hipFFT</a></td>
<td>1.0.22</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipfft"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/hipfort/en/docs-7.2.3/index.html">hipfort</a></td>
<td>0.7.1</a></td>
<td><a href="https://github.com/ROCm/hipfort"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/hipRAND/en/docs-7.2.3/index.html">hipRAND</a></td>
<td>3.1.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hiprand"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/hipSOLVER/en/docs-7.2.3/index.html">hipSOLVER</a></td>
<td>3.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipsolver"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/hipSPARSE/en/docs-7.2.3/index.html">hipSPARSE</a></td>
<td>4.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipsparse"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/hipSPARSELt/en/docs-7.2.3/index.html">hipSPARSELt</a></td>
<td>0.2.6</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipsparselt"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocALUTION/en/docs-7.2.3/index.html">rocALUTION</a></td>
<td>4.1.0</a></td>
<td><a href="https://github.com/ROCm/rocALUTION"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocBLAS/en/docs-7.2.3/index.html">rocBLAS</a></td>
<td>5.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocblas"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocFFT/en/docs-7.2.3/index.html">rocFFT</a></td>
<td>1.0.36</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocfft"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocRAND/en/docs-7.2.3/index.html">rocRAND</a></td>
<td>4.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocrand"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocSOLVER/en/docs-7.2.3/index.html">rocSOLVER</a></td>
<td>3.32.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocsolver"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocSPARSE/en/docs-7.2.3/index.html">rocSPARSE</a></td>
<td>4.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocsparse"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocWMMA/en/docs-7.2.3/index.html">rocWMMA</a></td>
<td>2.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocwmma"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/Tensile/en/docs-7.2.3/src/index.html">Tensile</a></td>
<td>4.45.0</td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/shared/tensile"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
</tbody>
<tbody class="rocm-components-libs rocm-components-primitives tbody-reverse-zebra">
<tr>
<th rowspan="4"></th>
<th rowspan="4">Primitives</th>
<td><a href="https://rocm.docs.amd.com/projects/hipCUB/en/docs-7.2.3/index.html">hipCUB</a></td>
<td>4.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipcub"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/hipTensor/en/docs-7.2.3/index.html">hipTensor</a></td>
<td>2.2.0</td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hiptensor"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocPRIM/en/docs-7.2.3/index.html">rocPRIM</a></td>
<td>4.2.0</td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocprim"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocThrust/en/docs-7.2.3/index.html">rocThrust</a></td>
<td>4.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocthrust"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
</tbody>
<tbody class="rocm-components-tools rocm-components-system tbody-reverse-zebra">
<tr>
<th rowspan="7">Tools</th>
<th rowspan="7">System management</th>
<td><a href="https://rocm.docs.amd.com/projects/amdsmi/en/docs-7.2.3/index.html">AMD SMI</a></td>
<td>26.2.2</a></td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/amdsmi"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rdc/en/docs-7.2.3/index.html">ROCm Data Center Tool</a></td>
<td>1.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rdc/"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocminfo/en/docs-7.2.3/index.html">rocminfo</a></td>
<td>1.0.0</td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocminfo/"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocm_smi_lib/en/docs-7.2.3/index.html">ROCm SMI</a></td>
<td>7.8.0</td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocm-smi-lib/"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/ROCmValidationSuite/en/docs-7.2.3/index.html">ROCm Validation Suite</a></td>
<td>1.3.0</a></td>
<td><a href="https://github.com/ROCm/ROCmValidationSuite"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
</tbody>
<tbody class="rocm-components-tools rocm-components-perf">
<tr>
<th rowspan="6"></th>
<th rowspan="6">Performance</th>
<td><a href="https://rocm.docs.amd.com/projects/rocm_bandwidth_test/en/docs-7.2.3/index.html">ROCm Bandwidth
Test</a></td>
<td>2.6.0</a></td>
<td><a href="https://github.com/ROCm/rocm_bandwidth_test/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocprofiler-compute/en/docs-7.2.3/index.html">ROCm Compute Profiler</a></td>
<td>3.4.0</a></td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocprofiler-compute"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocprofiler-systems/en/docs-7.2.3/index.html">ROCm Systems Profiler</a></td>
<td>1.3.0</a></td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocprofiler-systems/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocprofiler/en/docs-7.2.3/index.html">ROCProfiler</a></td>
<td>2.0.0</a></td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocprofiler/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/docs-7.2.3/index.html">ROCprofiler-SDK</a></td>
<td>1.1.0</a></td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocprofiler-sdk/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr >
<td><a href="https://rocm.docs.amd.com/projects/roctracer/en/docs-7.2.3/index.html">ROCTracer</a></td>
<td>4.1.0</td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/roctracer/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
</tbody>
<tbody class="rocm-components-tools rocm-components-dev">
<tr>
<th rowspan="5"></th>
<th rowspan="5">Development</th>
<td><a href="https://rocm.docs.amd.com/projects/HIPIFY/en/docs-7.2.3/index.html">HIPIFY</a></td>
<td>22.0.0</td>
<td><a href="https://github.com/ROCm/HIPIFY/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/ROCdbgapi/en/docs-7.2.3/index.html">ROCdbgapi</a></td>
<td>0.77.4</a></td>
<td><a href="https://github.com/ROCm/ROCdbgapi/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/ROCmCMakeBuildTools/en/docs-7.2.3/index.html">ROCm CMake</a></td>
<td>0.14.0</td>
<td><a href="https://github.com/ROCm/rocm-cmake/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/ROCgdb/en/docs-7.2.3/index.html">ROCm Debugger (ROCgdb)</a>
</td>
<td>16.3</a></td>
<td><a href="https://github.com/ROCm/ROCgdb/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocr_debug_agent/en/docs-7.2.3/index.html">ROCr Debug Agent</a>
</td>
<td>2.1.0</td>
<td><a href="https://github.com/ROCm/rocr_debug_agent/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
</tbody>
<tbody class="rocm-components-compilers tbody-reverse-zebra">
<tr>
<th rowspan="2" colspan="2">Compilers</th>
<td><a href="https://rocm.docs.amd.com/projects/HIPCC/en/docs-7.2.3/index.html">HIPCC</a></td>
<td>1.1.1</td>
<td><a href="https://github.com/ROCm/llvm-project/tree/amd-staging/amd/hipcc"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/llvm-project/en/docs-7.2.3/index.html">llvm-project</a></td>
<td>22.0.0</a></td>
<td><a href="https://github.com/ROCm/llvm-project/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
</tbody>
<tbody class="rocm-components-runtimes tbody-reverse-zebra">
<tr>
<th rowspan="2" colspan="2">Runtimes</th>
<td><a href="https://rocm.docs.amd.com/projects/HIP/en/docs-7.2.3/index.html">HIP</a></td>
<td>7.2.1</a></td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/hip"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/ROCR-Runtime/en/docs-7.2.3/index.html">ROCr Runtime</a></td>
<td>1.18.0</td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocr-runtime"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
</tbody>
</table>
</div>
## Detailed component changes
The following sections describe key changes to ROCm components.
```{note}
For a historical overview of ROCm component updates, see the {doc}`ROCm consolidated changelog </release/changelog>`.
```
### **MIGraphX** (2.15.0)
#### Added
* External stream support to the MIGraphX context, allowing external HIP streams to be used during execution.
* Ability to return a vector for output alias, supporting operators like `make_tuple`.
#### Changed
* Refactored `move_output_instructions_after` into the module class.
* Updated rocMLIR to fix `bert_squad` and `bert_tf` regressions.
#### Optimized
* Rewrote the `gather` operator to use `transpose`/`reshape`/`broadcast`/`slice` for improved performance.
* Horizontally fuse cross-embedding `gather` operators.
* Improved tuning for Split-K.
* Removed extra assignments and inserts in `find_nop_reshapes` to reduce overhead.
#### Resolved issues
The following issues have been fixed:
* `int` to `bf16`/`fp16` conversion errors.
* Comparison logic in `find_concat_op` to match the correct I/O.
* `shape_transform_descriptor::rebase` when flattening a broadcasted dimension.
* An error with `rewrite_reshapes`.
* A gather rewrite crash by validating the strided view element count.
* A bug in gather rewrite with NHWC shapes.
* A crash in rocMLIR with Inception v3 on RDNA3 architecture-based Radeon GPUs.
* Filter zero-argument operators during ONNX parsing to prevent errors.
* Conflict for missing `no_broadcast` parameter on ROCm 7.2.x.
## ROCm known issues
ROCm known issues are noted on {fab}`github` [GitHub](https://github.com/ROCm/ROCm/labels/Verified%20Issue). For known
issues related to individual components, review the [Detailed component changes](#detailed-component-changes).
### Minor performance regression for MIGraphX with int8-quantized models
You might observe a slight performance regression when running int8-quantized models with MIGraphX. This impact is generally minimal and does not affect correctness. However, workloads sensitive to peak throughput might have reduced performance when compared to non-quantized or alternative execution paths. This issue is currently under investigation and will be fixed in a future ROCm release. See [GitHub issue #6195](https://github.com/ROCm/ROCm/issues/6195).
## ROCm upcoming changes
The following changes to the ROCm software stack are anticipated for future releases.
### ROCTracer, ROCProfiler, rocprof, and rocprofv2 deprecation
ROCTracer, ROCProfiler, `rocprof`, and `rocprofv2` are deprecated. It's strongly recommended to upgrade to the latest version of the [ROCprofiler-SDK](https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/latest/) library and the (`rocprofv3`) tool to ensure continued support and access to new features.
To learn about key feature improvements and benefits of ROCprofiler-SDK over the deprecated ROCProfiler and ROCTracer, see [Comparing ROCprofiler-SDK to legacy ROCm profiling tools](https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/latest/conceptual/comparing-with-legacy-tools.html).
It's anticipated that ROCTracer, ROCProfiler, `rocprof`, and `rocprofv2` will reach end of support (EoS) by the end of 2026 Q2.
### ROCm SMI deprecation
[ROCm SMI](https://github.com/ROCm/rocm_smi_lib) will be phased out in an
upcoming ROCm release and will enter maintenance mode. After this transition,
only critical bug fixes will be addressed and no further feature development
will take place.
It's strongly recommended to transition your projects to [AMD
SMI](https://github.com/ROCm/rocm-systems/tree/develop/projects/amdsmi), the successor to ROCm SMI. AMD SMI
includes all the features of the ROCm SMI and will continue to receive regular
updates, new functionality, and ongoing support. For more information on AMD
SMI, see the [AMD SMI documentation](https://rocm.docs.amd.com/projects/amdsmi/en/latest/).
### Changes to ROCm Object Tooling
ROCm Object Tooling tools ``roc-obj-ls``, ``roc-obj-extract``, and ``roc-obj`` were
deprecated in ROCm 6.4, and will be removed in a future release. Functionality
has been added to the ``llvm-objdump --offloading`` tool option to extract all
clang-offload-bundles into individual code objects found within the objects
or executables passed as input. The ``llvm-objdump --offloading`` tool option also
supports the ``--arch-name`` option, and only extracts code objects found with
the specified target architecture. See [llvm-objdump](https://llvm.org/docs/CommandGuide/llvm-objdump.html)
for more information.
+15
View File
@@ -0,0 +1,15 @@
root = true
[*]
charset = utf-8
end_of_line = lf
indent_style = space
indent_size = 4
insert_final_newline = true
trim_trailing_whitespace = true
[*.rst]
indent_size = 3
[*.{html,md,yaml}]
indent_size = 2
@@ -0,0 +1,71 @@
<table class="rocm-docs-table table">
<thead>
<tr>
<th class="head">
<p>Framework</p>
</th>
<th class="head">
<p>Supported versions</p>
</th>
<th class="head">
<p>Supported OS</p>
</th>
<th class="head">
<p>Supported Python versions</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="2" style="vertical-align: middle;">
<p>PyTorch</p>
</td>
<td style="vertical-align: middle;">
<p>2.11.0, 2.10.0, 2.9.1</p>
</td>
<td>
<p>Linux</p>
</td>
<td rowspan="2">
<p>3.14, 3.13, 3.12, 3.11</p>
</td>
</tr>
<tr>
<td style="vertical-align: middle;">
<p>2.11.0</p>
</td>
<td style="vertical-align: middle;">
<p>Windows</p>
</td>
</tr>
<tr>
<td style="vertical-align: middle;">
<p>JAX</p>
</td>
<td style="vertical-align: middle;">
<p>0.9.1, 0.8.2</p>
</td>
<td>
<p>Linux</p>
</td>
<td>
<p>3.14, 3.13, 3.12, 3.11</p>
</td>
</tr>
<tr>
<td style="vertical-align: middle;">
<p>vLLM<br>(<a href="#release-supported-hw">gfx950, gfx942, gfx1200,<br>gfx1201, gfx1100,
gfx1101,<br>gfx1102, gfx1151 GPUs only</a>)</p>
</td>
<td>
<p>0.19.1<br>(requires PyTorch 2.10.0)</p>
</td>
<td>
<p>Linux</p>
</td>
<td>
<p>3.13</p>
</td>
</tr>
</tbody>
<table>
@@ -0,0 +1,778 @@
#### **AMD SMI (BM)** (26.4.0)
##### Added
* **Added APU metrics support (table versions 2.4 and 3.0)**.
* New `amdsmi_apu_metrics_t` struct accessible via `amdsmi_gpu_metrics_t.apu_metrics` pointer (non-null when APU-specific metrics are available).
* **v2.4 metrics**:
* `temperature_gfx`, `temperature_soc`, `temperature_core[8]`, `temperature_l3[2]`
* `average_gfx_activity`, `average_mm_activity`
* `average_socket_power`, `average_cpu_power`, `average_soc_power`, `average_gfx_power`, `average_core_power[8]`
* Average clocks: `gfxclk`, `socclk`, `uclk`, `fclk`, `vclk`, `dclk`
* Current clocks: `gfxclk`, `socclk`, `uclk`, `fclk`, `vclk`, `dclk`, `coreclk[8]`, `l3clk[2]`
* `average_temperature_gfx`, `average_temperature_soc`, `average_temperature_core[8]`, `average_temperature_l3[2]`
* `average_cpu_voltage`, `average_soc_voltage`, `average_gfx_voltage`, `average_cpu_current`, `average_soc_current`, `average_gfx_current`
* `throttle_status`, `indep_throttle_status`
* `fan_pwm`
* **v3.0 metrics**:
* `temperature_core[16]`, `temperature_skin`
* `average_vcn_activity`, `average_ipu_activity[8]`, `average_core_c0_activity[16]`
* `average_dram_reads`, `average_dram_writes`, `average_ipu_reads`, `average_ipu_writes`
* `average_apu_power`, `average_dgpu_power`, `average_all_core_power`, `average_ipu_power`, `average_sys_power`
* `stapm_power_limit`, `current_stapm_power_limit`
* `average_core_power[16]`, `current_coreclk[16]`
* `current_core_maxfreq`, `current_gfx_maxfreq`
* `average_vpeclk_frequency`, `average_ipuclk_frequency`, `average_mpipu_frequency`
* `throttle_residency_prochot`, `throttle_residency_spl`, `throttle_residency_fppt`, `throttle_residency_sppt`, `throttle_residency_thm_core`, `throttle_residency_thm_gfx`, `throttle_residency_thm_soc`
* `time_filter_alphavalue`
* Fields not applicable to the current version are set to sentinel values: `0xFFFF` for `uint16_t`, `0xFFFFFFFF` for `uint32_t`, and `UINT64_MAX` for `uint64_t` fields.
* Python bindings updated with `AmdSmiApuMetrics` ctypes structure.
* **Added `oam_id` to `amdsmi_enumeration_info_t`**.
* `amd-smi list -e` now displays `OAM_ID` (Physical XGMI ID / OAM ID).
* Added `--enumeration` as a long-form alias for `-e` in `amd-smi list`.
* **Added support for GPU metrics v1.9 new fields**.
* Added new temperature fields to `amdsmi_gpu_metrics_t`:
* `temperature_hbm_stacks` — per-stack HBM temperatures (°C)
* `temperature_mid` — per-MID temperatures (°C)
* `temperature_aid` — per-AID temperatures (°C)
* `temperature_xcd` — per-XCC compute die temperatures (°C)
* Added new per-die clock fields to `amdsmi_gpu_metrics_t`:
* `current_uclk_aid` — per-AID uclk (MHz)
* `current_socclks_mid` — per-MID SOC clock (MHz)
* Added new constants:
* `AMDSMI_MAX_NUM_HBM_STACKS` (12)
* `AMDSMI_MAX_NUM_AID` (2)
* `AMDSMI_MAX_NUM_MID` (2)
* `AMDSMI_MAX_NUM_CLKS_PER_AID` (2)
* `AMDSMI_MAX_NUM_CLKS_PER_MID` (2)
* **Added VRAM and GTT tuning interface**.
* New `amd-smi static --mem-carveout` to view VRAM carveout options.
* New `amd-smi set --mem-carveout` to change the VRAM carveout (APU).
* New `amd-smi set --gtt` and `amd-smi reset --gtt` for system-wide GTT size tuning.
* New APIs: `amdsmi_get_gpu_uma_carveout_info()`, `amdsmi_set_gpu_uma_carveout()`, `amdsmi_get_ttm_info()`, `amdsmi_set_ttm_pages_limit()`, `amdsmi_reset_ttm_pages_limit()`.
* **Added UBB power and power_limit fields to `amdsmi_power_info_t` and `amdsmi_npm_info_t`**.
* `amd-smi metric --power` now displays `ubb_power` when available.
* `amd-smi node -p` now displays UBB power threshold when available.
* **Added CPU support for family 1A Models 50h-57h**.
* New APIs: `amdsmi_get_cpu_xgmi_pstate_range()`, `amdsmi_get_cpu_core_ccd_power()`, `amdsmi_get_cpu_tdelta()`, `amdsmi_get_cpu_dimm_sb_reg()`, `amdsmi_get_cpu_svi3_vr_controller_temp()`, `amdsmi_get_cpu_pc6_enable()`, `amdsmi_get_cpu_cc6_enable()`, `amdsmi_get_cpu_sdps_limit()`, `amdsmi_get_cpu_core_floor_freq_limit()`, `amdsmi_get_cpu_core_eff_floor_freq_limit()`, and corresponding set APIs.
* **Note**: `amdsmi_get_dfc_ctrl()` renamed to `amdsmi_get_cpu_dfc_ctrl()` and `amdsmi_set_dfc_ctrl()` renamed to `amdsmi_set_cpu_dfc_ctrl()` for naming consistency.
* **Updated memory API documentation**
Added note that the sum of per-process memory usage is not expected to equal total usage.
##### Changed
* **Renamed `processor_type_t` enum typedef to `amdsmi_processor_type_t`**.
* The unprefixed typedef name did not follow the `amdsmi_*_t` convention used throughout `amdsmi.h` and was easy to collide with identifiers defined by other system-management libraries. New code should use `amdsmi_processor_type_t`. The old name is preserved as a backward-compatibility typedef alias, so existing callers continue to compile unchanged.
* **Package install no longer modifies the system-wide `logrotate` timer or cron schedule**.
* Previously, installing `amd-smi-lib` overwrote `/lib/systemd/system/logrotate.timer` (or moved `/etc/cron.daily/logrotate` to `/etc/cron.hourly/`) to force hourly rotation, which affected every other package using `logrotate`.
* The package now only ships `/etc/logrotate.d/amd_smi.conf`, which sets its own `hourly` + `size 1M` cadence. AMD-SMI logs still rotate at the same frequency; system-wide settings stay as the distribution configured them.
##### Optimized
* **Optimized `rsmi_dev_device_identifiers_get()` in the ROCm-SMI device layer**.
* Removed unnecessary iteration by directly indexing the device list.
* Added bounds checking for `device_id`, with clearer error handling/logging.
* Improves performance for device identifier queries.
##### Resolved issues
* **Fixed `amd-smi metric` crashing with `TypeError` on MI300A when no CPU flags are specified**.
* When no CPU arguments are passed, `metric_cpu()` sets all boolean CPU args to `True` to display all available data. `--cpu-svi3-vr-controller-temp` takes a TYPE argument (and optional RAIL_INDEX) rather than a boolean flag — setting it to `True` caused a `TypeError` crash when the code tried to subscript it with `[0][0]`. Added `cpu_svi3_vr_controller_temp` to the show-all exclusion list, following the existing pattern for `cpu_lclk_dpm_level`, `cpu_io_bandwidth`, `cpu_dimm_sb_reg`, and similar argument-taking flags.
* **Fixed `amdsmi_get_gpu_accelerator_partition_profile()` returning incorrect `num_partitions` when `num_partition` is unavailable from GPU metrics**.
* GPU metrics no longer always provides `num_partition`. The function now derives the partition count from the active partition type when `num_partition` is not available:
* SPX → 1, DPX → 2, TPX → 3, QPX → 4
* CPX → derived from the XCD counter via `amdsmi_get_gpu_xcd_counter()`
* **Fixed `amdsmi_topo_get_p2p_status()` returning a raw `ctypes.c_uint32` object instead of an integer for the `type` field**.
* The `'type'` key in the returned dictionary now correctly returns `type_32.value` (an `int`) rather than the unwrapped ctypes object, consistent with the pattern used in `amdsmi_topo_get_link_type()`.
* **Adjusted KFD process caching to be more responsive**.
* Updated process caching to allow cache duration adjustment via the `AMDSMI_PROCESS_INFO_CACHE_MS` environment variable for workflows with rapid metric polling.
* **Fixed CLI exit codes to use absolute values**.
* Invalid GPU parameters now return positive error codes as documented.
* **Fixed CLI breakage when `amdgpu` driver is not present**.
* Improved init to better catch driver loading issues.
* **Aligned `amdsmi_get_gpu_device_uuid()` with HIP/rocminfo UUID format**.
* Modified `amdsmi_asic_info_t.asic_serial` to report per-socket serial using KFD's `unique_id`.
* **Fixed multiple bugs in NIC/switch code and `amdsmi_init()` NIC handling**.
* Fixed `sizeof` operator precedence, `hw_mon` reset, NUMA=65535 handling, and several CLI function call errors.
* Fixed `amdsmi_init()` to succeed when no NIC hardware is present.
* **Fixed shared mutex and self-heal**.
* Improved self-heal logic to correctly identify and recover from corrupted or uninitialized mutex state.
* **Fixed `cu_occupancy` displaying `0%` instead of `N/A` when file is unavailable**.
* Process `cu_occupancy` is now initialized to `INVALID` instead of zero, so `amd-smi process` displays `N/A` rather than a misleading `0%` when the sysfs file is not accessible.
* **Fixed CLI set commands silently succeeding on invalid input values**.
* `amd-smi set --profile <INVALID>` now returns a non-zero exit code and lists available profiles in the error message; invalid profile names are rejected at parse time.
* `amd-smi set --clk-level <CLK_TYPE>` (missing performance level indices) now returns a non-zero exit code with a usage hint instead of silently succeeding.
* `amd-smi set --power-cap <OUT_OF_RANGE>` now returns a non-zero exit code.
* `amd-smi set --fan <INVALID>%` no longer prompts the out-of-spec warning before validating the percentage range; invalid values are rejected immediately.
* **Fixed `amd-smi set --profile` help text omitting `BOOTUP_DEFAULT`**.
* `BOOTUP_DEFAULT` was always accepted at runtime but was missing from the `--help` profile list. Auditing invalid-input handling exposed this gap. `amd-smi reset --profile` can also be used to return to the bootup default power profile.
* **Fixed `amd-smi monitor --brcm_nic` and `--brcm_switch` flags being registered on non-BRCM systems**.
* These flags are now only registered when BRCM hardware is present, preventing spurious failures on AMD GPU-only systems.
* **Fixed `amd-smi` default command alignment**.
* Updated default `amd-smi` output to align values to the left for improved readability.
Several items were misaligned in the default output, and this change ensures a consistent left-aligned format across all fields.
* *This change is purely cosmetic and does not affect any functionality.*
* **Renamed `lc_perf_other_end_recovery` to `lc_perf_other_end_recovery_count` in `amd-smi metric` CLI output for unification**.
* **Removed references to deprecated `amd-smi reset -r`**.
* CLI help text and memory partition change warnings no longer reference `amd-smi reset -r` for driver reloading.
* Users are now directed to use `sudo modprobe -r amdgpu && sudo modprobe amdgpu` to reload the driver after partition changes.
* **Changed CPU power APIs to return values in milliwatts (mW) for higher precision**.
* Removed lossy integer rounding (`(mW + 500) / 1000`) from 6 CPU power get APIs. Values are now
returned in milliwatts directly from the ESMI library, preserving sub-watt precision.
* **C API**: Output parameter type remains `uint32_t*`, but the unit changed from watts to milliwatts (mW).
* `amdsmi_get_cpu_socket_power`
* `amdsmi_get_cpu_socket_power_cap`
* `amdsmi_get_cpu_socket_power_cap_max`
* `amdsmi_get_cpu_pwr_efficiency_mode` (ppt_limit field)
* `amdsmi_get_cpu_core_ccd_power`
* `amdsmi_get_cpu_sdps_limit`
* **Python API (breaking)**: These functions now return `int` (milliwatts) instead of `str` (e.g., `"240 Watts"`).
Callers that parsed the string output must update to handle the numeric return value.
* **CLI output**: Power values now display with milliwatt precision (e.g., `240.500 Watts`).
* Added missing null-pointer validation for output parameters in `amdsmi_get_cpu_socket_power_cap`
and `amdsmi_get_cpu_socket_power_cap_max`.
* Updated header documentation to specify milliwatt units for all affected get and set API parameters.
* **Changed power APIs to have consistent output parameter types**.
* Modified 6 CPU power APIs to have consistent output power types. All set and get APIs have `uint32_t` output values.
* Modified get and set APIs that had double output types to have `uint32_t` output types in milliwatts (mW).
* `amdsmi_get_cpu_socket_power(amdsmi_processor_handle processor_handle, uint32_t* ppower)`
* `amdsmi_get_cpu_socket_power_cap(amdsmi_processor_handle processor_handle, uint32_t* pcap)`
* `amdsmi_get_cpu_socket_power_cap_max(amdsmi_processor_handle processor_handle, uint32_t* pmax)`
* `amdsmi_get_cpu_pwr_efficiency_mode(amdsmi_processor_handle processor_handle, uint32_t* power_efficiency_mode, uint32_t* utilization, uint32_t* ppt_limit)`
* `amdsmi_get_cpu_core_ccd_power(amdsmi_processor_handle processor_handle, uint32_t* power)`
* `amdsmi_get_cpu_sdps_limit(amdsmi_processor_handle processor_handle, uint32_t* sdps_limit)`
#### **Composable Kernel** (1.3.0)
##### Added
* Added overload of `load_tile_transpose` that takes reference to output tensor as output parameter.
* Use data type from LDS tensor view when determining tile distribution for transpose in the GEMM pipeline.
* Added `eightwarps` support for abquant mode in blockscale GEMM.
* Added `preshuffleB` support for abquant mode in blockscale GEMM.
* Added support for explicit GEMM in `CK_TILE` grouped convolution forward and backward weight.
* Added TF32 convolution support on gfx942 and gfx950 in CK. It can be enabled or disabled via `DTYPES` of `tf32`.
* Added `streamingllm` sink support for FMHA FWD, include `qr_ks_vs`, `qr_async` and `splitkv` pipelines.
* Added support for microscaling (MX) FP8/FP4 mixed data types to Flatmm pipeline.
* Added support for fp8 dynamic tensor-wise quantization of FP8 fmha fwd kernel.
* Added FP8 KV cache support for FMHA batch prefill.
* Added FMHA batch prefill kernel support for several KV cache layouts, flexible page sizes, and different lookup table configurations.
* Added gpt-oss sink support for FMHA FWD, include `qr_ks_vs`, `qr_async`, `qr_async_trload` and `splitkv` pipelines.
* Added persistent async input scheduler for CK Tile universal GEMM kernels to support asynchronous input streaming.
* Added FP8 block scale quantization for FMHA forward kernel.
* Added gfx11xx support for FMHA.
* Added microscaling (MX) FP8/FP4 support on gfx950 for FMHA forward kernel (`qr` pipeline only).
* Added FP8 per-tensor quantization support for FMHA forward V3 pipeline on gfx950.
#### **HIP** (7.13)
##### Added
* New HIP APIs
* `cooperative_groups::reduce()` allows calling reduce operators on `thread_block_tile` and `coalesced_threads`. The implementation is based on the `__reduce_*_sync` operations, so the macro `HIP_ENABLE_EXTRA_WARP_SYNC_TYPES` might be needed to unlock some optimizations.
* New device attribute `hipDeviceAttributeGPUDirectRDMAWithHipVMMSupported`, indicating support for GPU Direct RDMA when using HIP VMM. This attribute corresponds to the CUDA `CU_DEVICE_ATTRIBUTE_GPU_DIRECT_RDMA_WITH_CUDA_VMM_SUPPORTED`.
##### Resolved issues
* A segmentation fault that occurred in child graphs during the graphlaunch phase. The issue originated from the entire graph being launched solely according to the parent graphs scheduling logic. The HIP runtime now introduces a pergraph segmentscheduling control flag and propagates the parent graphs scheduling mode to its child graphs, ensuring consistent scheduling behavior (classic vs. segment) and preventing failures when the parent falls back to classic scheduling.
* A segmentation fault caused by passing a null pointer to the hipMemGetAddressRange API. The function now handles null pointers correctly, matching the behavior of the corresponding CUDA API.
##### Changed
* `__reduce_and_sync()`, `__reduce_or_sync()` and `__reduce_xor_sync()` now provide a consistent behavior for all mask values and with CUDA. Previously, some masks were translated into bitwise operations, but others were not (such as those containing "holes"). Now, all masks cause bitwise instructions to be emitted. This is a change in behavior compared to previous versions.
##### Optimized
* Improved HIP runtime error logging when an application's fat binary does not include a compatible code object for the detected GPU architecture, offering clearer guidance to rebuild with the appropriate `--offload-arch=gfxXXXX` option.
* Enables inmemory and backgroundthread asynchronous logging in the HIP runtime by default to improve overall logging capability. This behavior can be disabled by setting the environment variable `AMD_LOG_ASYNC=0`.
#### **hipBLAS** (3.4.0)
##### Added
* gfx1250 and gfx90c support to clients.
* Version and other properties to Windows `hipblas.dll`.
* Support for `OpenBLAS` ILP64-based API usage in clients.
##### Resolved issues
* Restored the fallback of using the deprecated rocBLAS API `rocblas_set_device_memory_size` if allocations are failing.
#### **hipBLASLt** (1.3.0)
##### Added
* General Batched GEMM support.
##### Changed
* Replaced `install.sh` with an invoke-based task runner (`tasks.py`) to support cross-platform builds including Windows (ROCm 7.0+).
* `gtest` and `msgpack-cxx` are now fetched automatically using CMake FetchContent if not found on the system.
#### **hipCUB** (4.4.0)
##### Optimized
* Reduced build times for unit tests.
##### Resolved issues
* Fixed more memory leak issues with some unit tests.
#### **hipFFT** (1.0.23)
##### Added
* hipFFTW plan creation functions for advanced and general plans:
* `fftw_plan_many_dft`
* `fftwf_plan_many_dft`
* `fftw_plan_many_dft_r2c`
* `fftwf_plan_many_dft_r2c`
* `fftw_plan_many_dft_c2r`
* `fftwf_plan_many_dft_c2r`
* `fftw_plan_guru_dft`
* `fftwf_plan_guru_dft`
* `fftw_plan_guru_dft_r2c`
* `fftwf_plan_guru_dft_r2c`
* `fftw_plan_guru_dft_c2r`
* `fftwf_plan_guru_dft_c2r`
* `fftw_plan_guru64_dft`
* `fftwf_plan_guru64_dft`
* `fftw_plan_guru64_dft_r2c`
* `fftwf_plan_guru64_dft_r2c`
* `fftw_plan_guru64_dft_c2r`
* `fftwf_plan_guru64_dft_c2r`
* Support for gfx1150 architecture.
##### Changed
* Moved library to C++20 standard.
* Removed Boost as a dependency for clients and samples.
* Callback functions will be deprecated in a future release.
##### Resolved issues
* Fixed potential launch failure of data generation kernels in test and benchmark programs.
#### **hipRAND** (3.3.0)
##### Added
* `hiprand.dll` now contains embedded file version metadata.
#### **hipSOLVER** (3.4.0)
##### Added
* Compatibility-only functions:
* `geev`
* `hipsolverDnXgeev_bufferSize`
* `hipsolverDnXgeev`
* `syevBatched`
* `hipsolverDnXsyevBatched_bufferSize`
* `hipsolverDnXsyevBatched`
* `syevd`
* `hipsolverDnXsyevd_bufferSize`
* `hipsolverDnXsyevd`
* `sytrs`
* `hipsolverDnXsytrs_bufferSize`
* `hipsolverDnXsytrs`
#### **hipSPARSELt** (0.2.8)
##### Added
* CTest and test categories support (`--smoke`, `--pre_checkin`, and `--nightly`).
##### Optimized
* Provided more kernels for the `FP16`, `BF16`, and `Int8` datatypes.
* Improved the performance of the `HIPSPARSELT_PRUNE_SPMMA_TILE` function.
##### Resolved issues
* Fixed incorrect behavior when retrieving the PCI chip ID.
* Fixed LDS out-of-bounds read in `prune_tile_kernel`.
* Fixed out-of-bounds access for compress function test cases.
* Fixed missing null terminator in the return value of `hipsparseLtGetArchName()`.
* Fixed incorrect CPU result when `bias_type` is `BF16` for spmm test cases.
* Fixed double-free issue in the example code `example_prune_strip`.
* Fixed symbol interposition in the hipSPARSELt library.
#### **MIOpen** (3.5.1)
##### Added
* Added `MIOPEN_LOG_BUFFER_SIZE` option: when set to non-zero, dumps recent MIOpen logs to file on error.
* [Conv] Added `ConvDepthwiseFwd3D` solver for optimizing specific 3D depthwise convolutions.
* [Conv] Added NHWC layout support for Winograd convolution solvers.
* [Conv] Added regular GEMM solver support for Conv3D forward and backward-data with 1x1x1 filters.
* [Conv] Added configurable problem size threshold (`MIOPEN_CONV_DIRECT_MAX_SIZE`) for direct solver.
* [Softmax] Added tuning support via Generic Search.
##### Changed
* [Conv] Improved default kernel selection for Composable Kernel (CK) convolution solvers with ranked shortlists.
* [Conv] Split CK grouped convolution kernels into per-architecture runtime-loaded dynamic libraries.
##### Optimized
* Optimized transpose operations with tiled and vectorized variants for NCHW/NHWC conversions.
* [BatchNorm] Optimized batchnorm reduction using warp shuffle intrinsics.
* [Conv] Added heuristic filtering of slow GEMM solver configurations during tuning.
##### Deprecated
* [Conv] Deprecated CK non-grouped convolution forward and backward solvers.
* Deprecated `miopenConvolutionBackwardBias`: the underlying OpenCL kernel (`MIOpenConvBwdBias.cl`) has been removed. The function now returns `miopenStatusNotImplemented` and will be removed in a future release.
##### Removed
* Removed GraphAPI experimental feature and related code.
##### Resolved issues
* [Conv] Fixed Winograd Fury grouped convolution correctness on gfx12xx when G > 1.
* [Conv] Fixed bf16 WrW convolution precision loss in inter-batch accumulation.
* [Conv] Fixed GPU memory fault in Winograd v3.0 WrW solver for large tensor shapes.
* Fixed BF16 `abs` function precision error caused by unnecessary cast through FP16.
* Fixed pooling kernel runtime compilation failure.
* Fixed gfx1151 inline assembly compilation errors in batchnorm kernels.
* Fixed use-after-free in HIPOCProgram binary loading.
#### **ROCm Data Center Tool (RDC)** (1.3.0)
##### Resolved issues
* **Fixed broken partition metrics**.
* Regardless of whether the GPU was partitioned, RDC only saw the GPU index and no instances due to upstream gpu_metrics changes.
#### **rocBLAS** (5.4.0)
##### Added
* gfx1250 and gfx90c enabled.
* Trace logging using `ROCBLAS_LAYER=1` for `rocblas_gemm_ex_get_solutions`, `rocblas_gemm_batched_ex_get_solutions`, `rocblas_gemm_ex_get_solutions_by_type`, and `rocblas_gemm_batched_ex_get_solutions_by_type`.
* Version and other properties to Windows `rocblas.dll`.
* Support for `OpenBLAS` ILP64 API for host reference in clients.
* Dockerfiles in the `docker` directory to assist in setting up development.
##### Optimized
* Improved the performance of Level 3 `geam` for pure transpose scale use cases.
* Improved the performance of Level 2 `tpsv`.
##### Resolved issues
* Fix for querying solutions when using the `hipBLASLt` backend with `rocblas_gemm_batched_ex_get_solutions` if using null data pointers.
#### **ROCdbgapi** (0.80.0)
##### Added
* `amd_dbgapi_process_get_info()` adds a new query to get a mask spanning
over all the bits used by all the address spaces. The query is called
`AMD_DBGAPI_PROCESS_INFO_SIGNIFICANT_ADDRESS_BITS`.
#### **rocDecode** (1.8.0)
##### Added
* Logging improvement: Added function entry and exit logs (at Info log level).
* Logging improvement: Added duration to function exit logs and optimized log message formatting to reduce runtime overhead.
* Logging improvement: Merged all logger instances into one global instance.
* Logging improvement: Unified logging format in utility classes with core library logging format.
* Logging improvement: Moved debug logging from a compile-time switch to the runtime logger level controlled by `ROCDEC_LOG_LEVEL` (debug = 4).
* Added support for user-set output surface format.
##### Changed
* Removed CPack packaging (DEB/RPM/NSIS/TGZ/ZIP generation and all related CPACK variables).
* Removed `rocDecode-setup.py` dependency installer script.
* Removed Docker files.
* Removed package install documentation; updated all documentation to reference TheRock for installation.
* Simplified libva version check (single `>= 1.22` requirement).
* Cleaned up CMake error messages.
#### **rocFFT** (1.0.37)
##### Optimized
* Allow plans to share hipModules if they use the same kernels. This reduces time spent and memory used when
creating plans that exist concurrently.
* Improved performance of unit-strided, interleaved, complex-to-complex and real-to-complex FFTs on gfx1201, gfx90a, gfx942, and gfx950.
Single-precision lengths:
* (160,72,72)
* (160,80,72)
* (160,80,80)
* (72,72,72)
* (80,80,80)
* (84,84,72)
* (96,96,96)
* (108,108,80)
Double-precision lengths:
* (72,72,52)
* (60,60,60)
* (64,64,52)
* (64,64,64)
##### Changed
* Moved library to C++20 standard.
* Removed Boost as a dependency for clients and samples.
* Split the precompiled kernel cache file (`rocfft_kernel_cache.db`) into per-architecture files (`rocfft_kernel_cache_gfx950.db`, `rocfft_kernel_cache_gfx1201.db`, etc).
* `rocfft_plan_create` returns `rocfft_status_invalid_offset` for any usage of non-zero offsets in plan descriptions. The feature is not supported yet.
* Callback functions will be deprecated in a future release.
##### Resolved issues
* Potential issue with data generation for multi-dimensional transforms in rocfft-tests and rocfft-bench.
* An issue that sometimes blocked complex-to-complex FFT plan creation when using noncontiguous strides in multiple dimensions.
* An issue that sometimes blocked complex-to-real FFT plan creation when using noncontiguous strides in multiple dimensions.
* An issue that sometimes blocked complex-to-real FFT plan creation when using noncontiguous strides with small lengths on the two fastest dimensions.
* Potential launch failure of data generation kernels in test and benchmark programs.
* Incorrect results on some strided real-complex FFTs on gfx90a.
* Incorrect results on some even-length real FFTs that have odd-length strides on higher dimensions.
* Callbacks on MPI transforms when not all ranks have the same number of data bricks.
* Functional issues for multi-device, in-place real transforms.
* Functional issues for multi-dimensional, multi-device transforms involving some unit length(s).
* Functional issues for multi-device transforms involving data divisions along the slowest-varying axis (only) for some bricks but not all.
* Functional issues for multi-device transforms setting no field on input or output.
* Automatic allocation of work memory at plan execution time, when work memory is required on multiple devices.
#### **rocJPEG** (1.5.0)
##### Changed
* rocJPEG is now delivered as part of [TheRock](https://github.com/ROCm/TheRock). All core dependencies are provided by the TheRock build.
* Removed CPack packaging (DEB/RPM/NSIS/TGZ/ZIP generation and all related CPACK variables).
* Removed `rocJPEG-setup.py` dependency installer script.
* Removed Docker files.
* Removed package install documentation; updated all documentation to reference TheRock for installation.
* Simplified libva version check (single `>= 1.22` requirement).
* Cleaned up CMake error messages.
#### **ROCm Compute Profiler** (3.6.0)
##### Added
* Added L2 memory bandwidth derived metrics under `--membw-analysis` to allow L2 memory bandwidth specific profiling and analysis metric block 30.
* Added AMD Ryzen AI Max 300 series (gfx1151) support.
* New memory hierarchy visualization for RDNA 3.5 (gfx115X) in analyze CLI mode.
* Introduced support for AMD Instinct MI350P GPU.
* ``--view table`` option in analyze mode to force all TTY output to plain tables and ignore ``cli_style`` from YAML config (for example, mem_chart, Roofline charts render as tables). The ``--view`` argument is reserved for future TTY views (for example, other chart styles).
* Added EA memory bandwidth derived metrics under `--membw-analysis` to allow EA memory bandwidth specific profiling and analysis metric block 30.
##### Changed
* Standalone roofline (`--roof-only` option) in profile mode now creates `roofline.csv` only. HTML roofline charts are generated via `rocprof-compute analyze`. The `calc_ai_profile()` function has been removed; `calc_ai_analyze()` is the single source of truth for arithmetic intensity calculation.
* Roofline visualization options (`--sort`, `--mem-level`, `--roofline-data-type`) have moved from profile mode to analyze mode.
* Standardized unit naming in analysis configs and Python utilities: `pct`/`Pct``Percent`, `instr``Instructions`.
* Profile mode output format:
* Profile mode now creates separate counter collection files for each application replay (pmc_perf_*.csv or results_*.csv).
* Analyze mode automatically merges these files into a unified pmc_perf.csv containing information from all application replays during pre-processing.
* ROCm Compute Profiler now builds and runs profile mode with vanilla Python without requiring any Python dependencies to be installed via `pip`.
* Note that analysis mode will still require Python dependencies and will report any missing packages.
##### Removed
* Removed HIP API tracing since it's out-of-scope for ROCm Compute Profiler and the trace files were not being analyzed.
##### Optimized
* Filtering for block 21 (`-b 21`) in profile mode now only performs pc sampling and skips unnecessary counter collection.
* Filtering for block 21 in analysis mode now skips metrics calculations and only shows kernel/dispatch/system statistics and pc sampling table.
##### Resolved issues
* Fixed roofline benchmark MFMA FP16/BF16/INT8 peaks for MI350.
* Fixed an issue where pc sampling profiling failed with multi-argument commands and live process attachment.
##### Upcoming changes
* `--path` and `--subpath` options are deprecated and will be removed in a future release.
* Intermediate CSV generation (`results_*.csv`) from rocpd databases during profiling is deprecated and will be removed in a future release. The analyze step will read `.db` files directly.
* `--retain-rocpd-output` is deprecated and will be removed in a future release. `.db` files will be retained by default.
##### Known issues
* For AMD Ryzen AI Max 300 series, the roofline metrics table will have N/A values for "peak" field.
* This is planned to be addressed by adding empirical benchmark support for AMD Ryzen AI Max 300 series in a future release.
#### **ROCm Systems Profiler** (1.6.0)
##### Added
* Kernel Fusion Driver (KFD) event tracing support to capture page faults, page migrations, queue evictions, GPU unmap events, and dropped events. Requires ROCprofiler-SDK 1.2.1 or later. Enable with `ROCPROFSYS_ROCM_DOMAINS=kfd_events`.
* Support for pause and resume of profiling via `roctxProfilerPause` and `roctxProfilerResume`.
* Support for selective region tracing via the `ROCPROFSYS_SELECTED_REGIONS` environment variable, limiting tracing to specified regions.
* `--selected-regions` CLI argument to `rocprof-sys-sample`, `rocprof-sys-run`, and `rocprof-sys-instrument` for specifying selective region tracing from the command line.
* Support for re-attaching to a previously profiled process. After detaching, `rocprof-sys-attach` can re-attach to the same PID for a new profiling session.
* MPI-rank-based file output filtering feature controlled with two new CLI arguments: `--rank-filter-output` and `--rank-filter-id`.
* JSON-based configurable preset system with `--preset=<name>` flag, replacing the old `--<preset-name>` flags. Presets are now loaded from JSON files in `source/bin/common/presets/`, making them extensible and exportable. Use `--list-presets` to see available presets and `--explain=<name>` for detailed preset information.
* Domain flags for composable configuration: `--gpu[=metrics]`, `--rocm[=domains]`, `--cpu[=hz]`, `--parallel[=runtimes]`. Domain flags can be combined with presets to customize profiling without editing configuration files.
* Configuration export via `--export-config[=file]` to save resolved settings as reusable JSON configuration files. Exported configs can be loaded back with `--preset=./config.json`.
* Topic-based help system: `--help` now shows a compact summary with essential options and a list of help topics. Use `--help=<topic>` (e.g., `--help=sampling`, `--help=gpu`, `--help=tracing`) to see only relevant options. Use `--help=all` for the full option listing.
* Post-run output summary during library finalization showing result file locations.
* JSON schema file (`share/rocprofiler-systems/presets/schema.json`) for preset validation.
* Documentation (`docs/how-to/instrumenting-rewriting-binary-application.rst`) describing what to do when Dyninst reports a "Failed to transform trace" error during instrumentation.
##### Changed
* `rocprof-sys-avail` no longer queries GPU devices or hardware counters unless `--hw-counters` or `--all` is requested, reducing startup time and allowing settings/component queries in environments without GPU/ROCm.
* `rocprof-sys-instrument` diagnostic file dumps (available, instrumented, excluded, coverage, overlapping) are now gated behind the `--dump-info` flag instead of being generated unconditionally.
* Preset flags changed from `--balanced` to `--preset=balanced` syntax. The old `--<preset-name>` flags are still supported and handled within `preset_registry`.
* Removed the `ROCPROFSYS_USE_ROCM` CMake option. ROCm is now required for building the ROCm Systems Profiler.
##### Resolved issues
* Fixed an issue where the `--rocm-domains` CLI option for `rocprof-sys-run` was not recognized.
#### **rocminfo** (1.0.0)
##### Resolved issues
* Fixed BDF (Bus:Device.Function) ID truncation issue that caused incorrect display of PCI device identifiers. The `bdf_id` field was incorrectly declared as `uint16_t` instead of `uint32_t`, causing silent truncation when HSA runtime returned the full 32-bit BDF ID value. This has been corrected to properly display complete BDF information for all GPU agents.
#### **rocPRIM** (4.4.0)
##### Added
* Added type trait definitions for `__hip_bfloat16`. This should resolve issues where this type did not work with radix-based algorithms.
* Unit tests for config_types.
##### Optimized
* Reduced build times for unit tests.
* Reduced memory usage in unit tests.
##### Resolved issues
* Fixed a silent overflow in `rocprim::device_segmented_reduce` where it could exceed the maximum number of HIP threads, resulting in missing output.
* Certain large unit tests now properly detect if insufficient system memory is present and skip the test case accordingly.
* Fixed out-of-bounds memory access in block run length decode.
* Fixed memory leak in unit tests.
#### **ROCprofiler-SDK** (1.3.0)
##### Added
**API:**
* Late-start profiling support: Enables profiling when `rocprofiler-sdk` is loaded after HSA/HIP runtimes have already initialized.
* `rocprofiler_force_configure()` now automatically detects and profiles runtimes initialized before the SDK loads.
* Integrates with `rocprofiler-register` to retrieve the registered API tables.
* Supports all runtime types (HSA, HIP, ROCTX, RCCL, rocDecode, rocJPEG, and more) automatically.
* No explicit late-start API calls required; works transparently.
* KFD (Kernel Fusion Driver) event tracing support:
* Buffer service configurations for each KFD buffer tracing type.
* New type `tool_buffer_tracing_kfd_record_t` using `std::variant` to wrap 8 different KFD buffer tracing types.
* Each KFD event generates `rocpd_info_pmc`, `rocpd_event`, `rocpd_region`, and `rocpd_pmc_event` rows.
* Fixed handling for special SVM location in KFD prefetch location reporting.
* Fixed parsing for queue restore events to handle both correct format (character '0') and broken driver format (NULL character '\0').
**rocprofv3 (CLI):**
* Multi-pass counter collection support: Support for multiple `--pmc` flags to define separate counter groups for different profiling passes.
* Ability to combine command-line `--pmc` flags with input file counter groups.
* Each pass generates output in a separate `pass_n` subdirectory.
* Example: `rocprofv3 --pmc SQ_WAVES --pmc GRBM_COUNT -- <app>` creates two profiling passes.
* KFD (Kernel Fusion Driver) event tracing support:
* KFD record dumping to `rocpd` with support for 8 main KFD event types.
* Support for `rocpd` to Perfetto conversion for KFD events.
* `--kfd-trace` flag to enable KFD event tracing.
* ROCTx support for ATT: Added ROCtx support to device thread trace when using `--att --selected-regions`.
* Allows `roctxProfilerPause` and `roctxProfilerResume` to explicitly control when ATT data collection starts and stops.
* Enables more precise, region-focused ATT tracing with reduced overhead and noise.
* Supports multiple resume/pause cycles, each producing separate trace output files.
* Incompatible with `--att-consecutive-kernels`.
* PC sampling support for dynamic attach: Allows users to attach to a running application and collect PC samples without restarting the workload.
* Enables profiling long-running or production-style jobs at the point of interest.
* Results integrate with the existing PC sampling analysis flow.
**Documentation:**
* Added marker-controlled thread tracing section to the thread trace how-to guide.
* Added cross-reference from ROCTx documentation to ATT with `selected-regions`.
##### Changed
**Implementation:**
* Late-start architecture redesign: Removed direct runtime symbol access in favor of proper rocprofiler-register integration.
* Replaced ~600 lines of `dlopen`/`dlsym` bypass logic with ~80 lines by using `rocprofiler_register_invoke_all_registrations()`.
* Late-start now works by requesting `rocprofiler-register` to re-propagate stored API tables.
* Extensible design. Automatically supports new runtimes without SDK code changes.
* Provides a proper separation of concerns. `rocprofiler-register` manages the table storage while SDK manages the table wrapping.
* Counter dimension encoding changed from fixed-width to variable-width allocation per dimension type.
* Dimension selection and reduction logic now uses explicit dimension masks and single-index selection.
* HSA queue interception extended to handle AMD extended kernel dispatch packets.
##### Removed
* Counter collection support for plain text (`.txt`) input files. Only structured file formats (JSON and YAML) with schema validation are now supported.
##### Resolved issues
* Fixed rocpd OTF2 output to add `ACCELERATOR_DEVICE` as system tree node domain for AMD devices.
* Fixed `rocprofv3` input file parsing where comment lines containing `pmc:` were incorrectly processed as valid counter collection directives, causing unintended profiling passes.
#### **rocRAND** (4.4.0)
##### Added
* gfx1150 and gfx1152 support.
* rocrand.dll now contains embedded file version metadata.
##### Resolved issues
* Fixed memory leak in unit tests.
#### **rocSHMEM** (3.4.0)
##### Added
* Added new APIs:
* `rocshmem_quiet_on_stream`
* `rocshmem_sync_all_on_stream`
* `rocshmem_TYPENAME_alltoall_wg`
* `rocshmem_TYPENAME_alltoallv_wg`
* `rocshmem_team_my_pe`
* `rocshmem_team_n_pes`
* `rocshmem_barrier`
* `rocshmem_barrier_wave`
* `rocshmem_barrier_wg`
* `rocshmem_buffer_register`
* `rocshmem_buffer_unregister`
* `rocshmem_info_get_version`
* `rocshmem_info_get_name`
* `rocshmem_vendor_get_version_info`
* Added library constants: `ROCSHMEM_MAJOR_VERSION`, `ROCSHMEM_MINOR_VERSION`,
`ROCSHMEM_MAX_NAME_LEN`, `ROCSHMEM_VENDOR_STRING`, `ROCSHMEM_VERSION`,
`ROCSHMEM_VENDOR_MAJOR_VERSION`, `ROCSHMEM_VENDOR_MINOR_VERSION`,
`ROCSHMEM_VENDOR_PATCH_VERSION`.
* Added vendor string and backend metadata to the `rocshmem_info` output.
* Added `ROCSHMEM_TEAM_WORLD` for device code.
* Added `ROCSHMEM_TEAM_SHARED` predefined team for PEs sharing a common memory domain (same node).
* Added new environment variables:
* `ROCSHMEM_GDA_OVERRIDE_NIC_FIRMWARE_CHECK`
* `ROCSHMEM_GDA_NUM_QPS_PER_PE_DEFAULT_CTX`
* `ROCSHMEM_GDA_NUM_QPS_PER_PE_USR_CTX`
* Added VMM POSIX memory allocator (`USE_HEAP_DEVICE_VMM_POSIX`):
* Uses HIP Virtual Memory Management (VMM) APIs for fine-grained memory control.
* Requires ROCm 7.0+ and Linux kernel 5.6+.
* Not compatible with MPI-based initialization (use `ROCSHMEM_INIT_WITH_UNIQUEID` instead).
##### Changed
* Use CQ collapsing for the Mellanox MLX5 GDA conduit.
#### **rocSOLVER** (3.34.0)
##### Added
* Computation of solution for LU factorization without pivoting:
* GETRS_NPVT (with batched and strided\_batched versions)
* GETRS_NPVT_64 (with batched and strided\_batched versions)
* Linear solver routines for symmetric matrices:
* SYTRS (with batched and strided\_batched versions)
* SYTRS_64 (with batched and strided\_batched versions)
##### Optimized
* Improved the performance of POTF2 and downstream functions such as POTRF.
##### Resolved issues
* Fixed a memory access error in SYTRF and synchronization issues in LASYF and SYTF2.
#### **rocSPARSE** (4.6.0)
##### Added
* `rocsparse_create_const_bsr_descr` routine for creating a const sparse BSR matrix descriptor.
* `rocsparse_spic0` and `rocsparse_spilu0` routines for incomplete factorizations, with strided batched computations enabled.
* `rocsparse_sptrsv_descr_create` and `rocsparse_sptrsv_descr_destroy` routines.
* `rocsparse_singularity` enumeration.
* `rocsparse_sptrsv_output_singularity` and `rocsparse_sptrsv_output_singularity_position` in `rocsparse_sptrsv_output`.
* Strided batched computations for `rocsparse_sptrsv`.
##### Optimized
* Significant performance improvement for `rocsparse_Xgtsv_no_pivot_strided_batch`.
* Significant performance improvement for `rocsparse_Xgtsv_no_pivot`.
##### Resolved issues
* Fixed incorrect usage of `__syncthreads` in `bsrmm`, `csrmm` (row_split), and `csritilu0x`.
* Fixed incorrect usage of `__syncthreads` in `csx2dense`, `dense2csx`, `prune_dense2csr`, `csrcolor`, and `csrmm` (`nnz_split`).
* Fixed `rocsparse_[s|d|c|z]csric0` where `rocsparse_status_invalid_value` was being returned when the maximum number of non-zeros in any row is between 513 and 1024.
* Fixed compilation when using `--rocsparse_ILP64`.
* Fixed off-by-one heap-buffer-overflow in temporary buffer allocation for `rocsparse_csrsort`, `rocsparse_check_matrix_csr`, and `rocsparse_check_matrix_gebsr` (and their delegating routines `rocsparse_cscsort`, `rocsparse_coosort`, `rocsparse_check_matrix_csc`, and `rocsparse_check_matrix_gebsc`) where the `shift_offsets_kernel` temp buffer was sized for `m` elements instead of `m+1`.
##### Removed
* The deprecated C++14 support, which is no longer supported by the rocPRIM dependency.
#### **rocThrust** (4.4.0)
##### Resolved issues
* Fixed memory leak in unit test.
* Fixed unit test compatibility with ASAN.
#### **rocWMMA** (2.2.1)
##### Added
* Added the following community samples for external contributions, with build support and documentation:
* `simple_gemm_silu`: demonstrates a GEMM + SiLU fused operator using the rocWMMA API.
* `simple_gemm_fusion`: demonstrates block-tile-level dual-GEMM fusion using the rocWMMA API.
* `simple_gemm_swiglu`: demonstrates a SwiGLU fused dual-GEMM kernel (LLaMA/Mistral FFN gate layer) using the rocWMMA API.
##### Changed
* Updated the `find_package` search for OpenMP to prefer the `openmp-config.cmake` provided by ROCm, with a fallback to module search mode.
* Updated `INSTALL_RPATH` and added `BUILD_RPATH` for OpenMP.
##### Resolved issues
* Improved HIP RTC regression test portability when deployed outside the default path.
@@ -0,0 +1,193 @@
<table class="rocm-docs-table table">
<thead>
<tr>
<th class="head"><p>Component group</p></th>
<th class="head"><p>Component name</p></th>
<th class="head"><p>Version</p></th>
<th class="head"><p>Supported platforms</p></th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="18" style="vertical-align: middle;">
<p>Math and compute libraries</p>
</td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipblas">hipBLAS</a></td>
<td><a href="#hipblas-3-4-0">3.4.0</a></td>
<td rowspan="16" style="vertical-align: middle;">Linux/Windows · Instinct/Radeon/Ryzen</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipblaslt">hipBLASLt</a></td>
<td><a href="#hipblaslt-1-3-0">1.3.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipcub">hipCUB</a></td>
<td><a href="#hipcub-4-4-0">4.4.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipfft">hipFFT</a></td>
<td><a href="#hipfft-1-0-23">1.0.23</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hiprand">hipRAND</a></td>
<td><a href="#hiprand-3-3-0">3.3.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipsolver">hipSOLVER</a></td>
<td><a href="#hipsolver-3-4-0">3.4.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipsparse">hipSPARSE</a></td>
<td>4.5.0</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/miopen">MIOpen</a></td>
<td><a href="#miopen-3-5-1">3.5.1</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocblas">rocBLAS</a></td>
<td><a href="#rocblas-5-4-0">5.4.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocfft">rocFFT</a></td>
<td><a href="#rocfft-1-0-37">1.0.37</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocrand">rocRAND</a></td>
<td><a href="#rocrand-4-4-0">4.4.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocsolver">rocSOLVER</a></td>
<td><a href="#rocsolver-3-34-0">3.34.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocsparse">rocSPARSE</a></td>
<td><a href="#rocsparse-4-6-0">4.6.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocprim">rocPRIM</a></td>
<td><a href="#rocprim-4-4-0">4.4.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocthrust">rocThrust</a></td>
<td><a href="#rocthrust-4-4-0">4.4.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocwmma">rocWMMA</a></td>
<td><a href="#rocwmma-2-2-1">2.2.1</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/composablekernel">Composable
Kernel</a></td>
<td><a href="#composable-kernel-1-3-0">1.3.0</a></td>
<td>Linux/Windows · Instinct/Radeon</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipsparselt">hipSPARSELt</a></td>
<td><a href="#hipsparselt-0-2-8">0.2.8</a></td>
<td>Linux/Windows · Instinct (gfx950/gfx942)</td>
</tr>
<tr>
<td rowspan="2" style="vertical-align: middle;">
<p>Communication libraries</p>
</td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rccl">RCCL</a></td>
<td>2.28.3</td>
<td>Linux · Instinct/Radeon/Ryzen</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocshmem">rocSHMEM</a></td>
<td><a href="#rocshmem-3-4-0">3.4.0</a></td>
<td>Linux · Instinct (gfx950/gfx942/gfx90a) · Radeon (gfx1201/gfx1200/gfx1100/gfx1101/gfx1102)</td>
</tr>
<tr>
<td rowspan="2" style="vertical-align: middle;">
<p>Media libraries</p>
</td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocdecode">rocDecode</a></td>
<td><a href="#rocdecode-1-8-0">1.8.0</a></td>
<td rowspan="2" style="vertical-align: middle;">Linux · Instinct/Radeon · Ryzen (gfx1150/gfx1151/gfx1152)</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocjpeg">rocJPEG</a></td>
<td><a href="#rocjpeg-1-5-0">1.5.0</a></td>
</tr>
<tr>
<td rowspan="5" style="vertical-align: middle;">
<p>Runtimes and compilers</p>
</td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/hip">HIP</a></td>
<td><a href="#hip-7-13">7.13</a></td>
<td rowspan="4" style="vertical-align: middle;">Linux/Windows · Instinct/Radeon/Ryzen</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/HIPIFY/tree/therock-7.13">HIPIFY</a></td>
<td>7.13</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/llvm-project/tree/therock-7.13">LLVM</a></td>
<td>23.0.0</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/SPIRV-LLVM-Translator/tree/therock-7.13">SPIRV-LLVM-Translator</a></td>
<td>23.0.0</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocr-runtime">ROCr Runtime</a></td>
<td>1.21.0</td>
<td>Linux · Instinct/Radeon/Ryzen</td>
</tr>
<tr>
<td rowspan="6" style="vertical-align: middle;">
<p>Profiling and debugging tools</p>
</td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocprofiler-compute">ROCm
Compute Profiler (rocprofiler-compute)</a></td>
<td><a href="#rocm-compute-profiler-3-6-0">3.6.0</a></td>
<td rowspan="2" style="vertical-align: middle;">Linux · Instinct · Ryzen (gfx1150/gfx1151/gfx1152)</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocprofiler-systems">ROCm
Systems Profiler (rocprofiler-systems)</a></td>
<td><a href="#rocm-systems-profiler-1-6-0">1.6.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocprofiler-sdk">ROCprofiler-SDK</a></td>
<td><a href="#rocprofiler-sdk-1-3-0">1.3.0</a></td>
<td>Linux · Instinct/Radeon · Ryzen (gfx1150/gfx1151/gfx1152)</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocdbgapi">ROCdbgapi</a></td>
<td><a href="#rocdbgapi-0-80-0">0.80.0</a></td>
<td rowspan="3" style="vertical-align: middle;">Linux · Instinct/Radeon</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/ROCgdb/tree/therock-7.13">ROCm Debugger (ROCgdb)</a></td>
<td>16.3</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocr-debug-agent">ROCr Debug
Agent</a></td>
<td>2.1.0</td>
</tr>
<tr>
<td rowspan="3" style="vertical-align: middle;">
<p>Control and monitoring tools</p>
</td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/amdsmi">AMD SMI (BM)</a></td>
<td><a href="#amd-smi-bm-26-4-0">26.4.0</a></td>
<td>Linux · Instinct/Radeon</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocminfo">rocminfo</a></td>
<td><a href="#rocminfo-1-0-0">1.0.0</a></td>
<td>Linux · Instinct/Radeon/Ryzen</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rdc">ROCm Data Center Tool
(RDC)</a></td>
<td><a href="#rocm-data-center-tool-rdc-1-3-0">1.3.0</a></td>
<td>Linux · Instinct</td>
</tr>
</tbody>
</table>
@@ -0,0 +1,268 @@
::::{tab-set}
:::{tab-item} Instinct
:sync: instinct
<table class="rocm-docs-table table">
<colgroup style="width: 25%;">
<thead>
<tr>
<th class="head">
<p>AMD device</p>
</th>
<th class="head">
<p>Firmware</p>
</th>
<th class="head">
<p>Linux driver</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<td>
<p>Instinct MI355X</p>
</td>
<td rowspan="2" style="vertical-align: middle">
<p>PLDM bundle 01.26.00.02</p>
</td>
<td rowspan="10" style="vertical-align: middle">
<p>
<strong>AMD GPU Driver (amdgpu)</strong><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/31.30.0-preview/documentation/release-notes.html"
target="_blank"
>31.30.0</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/31.20.0-preview/documentation/release-notes.html"
target="_blank"
>31.20.0</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/31.10.0-preview/documentation/release-notes.html"
target="_blank"
>31.10.0</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.3/documentation/release-notes.html"
target="_blank"
>30.30.3</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.2/documentation/release-notes.html"
target="_blank"
>30.30.2</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.1/documentation/release-notes.html"
target="_blank"
>30.30.1</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.0/documentation/release-notes.html"
target="_blank"
>30.30.0</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.20.1/documentation/release-notes.html"
target="_blank"
>30.20.1</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.20.0/documentation/release-notes.html"
target="_blank"
>30.20.0</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.10.2/documentation/release-notes.html"
target="_blank"
>30.10.2</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.10.1/documentation/release-notes.html"
target="_blank"
>30.10.1</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.10/documentation/release-notes.html"
target="_blank"
>30.10.0</a><br>
</p>
</td>
</tr>
<tr>
<td>
<p>Instinct MI350X</p>
</td>
</tr>
<tr>
<td>
<p>Instinct MI350P</p>
</td>
<td style="vertical-align: middle">
<p>IFWI 00185129</p>
</td>
</tr>
<tr>
<td>
<p>Instinct MI325X</p>
</td>
<td style="vertical-align: middle">
<p>PLDM bundle 01.25.04.02</p>
</td>
</tr>
<tr>
<td>
<p>Instinct MI300X</p>
</td>
<td>
<p>PLDM bundle 01.26.00.02</p>
</td>
</tr>
<tr>
<td>
<p>Instinct MI300A</p>
</td>
<td>
<p>BKC 26.1</p>
</td>
</tr>
<tr>
<td>
<p>Instinct MI250X</p>
</td>
<td>
<p>IFWI 75 (or later)</p>
</td>
</tr>
<tr>
<td>
<p>Instinct MI250</p>
</td>
<td rowspan="2">
<p>Maintenance update (MU) 5 with IFWI 75 (or later)</p>
</td>
</tr>
<tr>
<td>
<p>Instinct MI210</p>
</td>
</tr>
<tr>
<td>
<p>Instinct MI100</p>
</td>
<td>
<p>VBIOS D3430401-037</p>
</td>
</tr>
</tbody>
</table>
:::
:::{tab-item} Radeon
:sync: radeon
<table class="rocm-docs-table table">
<colgroup style="width: 50%;">
<thead>
<tr>
<th class="head">
<p>Linux driver</p>
</th>
<th class="head">
<p>Windows driver</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<td style="vertical-align: middle">
<p>
<strong>AMD GPU Driver (amdgpu)</strong><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/31.30.0-preview/documentation/release-notes.html"
target="_blank"
>31.30.0</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/31.20.0-preview/documentation/release-notes.html"
target="_blank"
>31.20.0</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/31.10.0-preview/documentation/release-notes.html"
target="_blank"
>31.10.0</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.3/documentation/release-notes.html"
target="_blank"
>30.30.3</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.2/documentation/release-notes.html"
target="_blank"
>30.30.2</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.1/documentation/release-notes.html"
target="_blank"
>30.30.1</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.0/documentation/release-notes.html"
target="_blank"
>30.30.0</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.20.1/documentation/release-notes.html"
target="_blank"
>30.20.1</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.20.0/documentation/release-notes.html"
target="_blank"
>30.20.0</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.10.2/documentation/release-notes.html"
target="_blank"
>30.10.2</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.10.1/documentation/release-notes.html"
target="_blank"
>30.10.1</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.10/documentation/release-notes.html"
target="_blank"
>30.10.0</a><br>
</p>
</td>
<td style="vertical-align: middle">
<p>
<strong>AMD Software: Adrenalin Edition</strong>
<a
href="https://www.amd.com/en/resources/support-articles/release-notes/RN-RAD-WIN-26-5-1.html"
target="_blank"
>26.5.1</a>
</td>
</tr>
</tbody>
</table>
:::
:::{tab-item} Ryzen
:sync: ryzen
<table class="rocm-docs-table table">
<colgroup style="width: 50%;">
<thead>
<tr>
<th class="head">
<p>Linux driver</p>
</th>
<th class="head">
<p>Windows driver</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<td style="vertical-align: middle">
<p>Inbox kernel driver in Ubuntu 26.04 or 24.04.4</p>
</td>
<td rowspan="30" style="vertical-align: middle">
<p>
<strong>AMD Software: Adrenalin Edition</strong>
<a
href="https://www.amd.com/en/resources/support-articles/release-notes/RN-RAD-WIN-26-5-1.html"
target="_blank"
>26.5.1</a>
</p>
</td>
</tr>
</tbody>
</table>
:::
::::
@@ -0,0 +1,538 @@
::::{tab-set}
:::{tab-item} Instinct
:sync: instinct
<table class="rocm-docs-table table">
<thead>
<colgroup style="width: 33%;">
<colgroup style="width: 32%;">
<tr>
<th class="head">
<p>Device series</p>
</th>
<th class="head">
<p>Device</p>
</th>
<th class="head">
<p>LLVM target</p>
</th>
<th class="head">
<p>Architecture</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<td class="stub">
<a href="https://www.amd.com/en/products/accelerators/instinct/mi350.html" target="_blank">AMD Instinct MI350
Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi350/mi355x.html" target="_blank">Instinct
MI355X</a></p>
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi350/mi350x.html" target="_blank">Instinct
MI350X</a></p>
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi350/mi350p.html" target="_blank">Instinct
MI350P</a></p>
</td>
<td>
<p>gfx950</p>
</td>
<td>
<a href="https://www.amd.com/en/technologies/cdna.html#cdna4" target="_blank">CDNA 4</a>
</td>
</tr>
<tr>
<td class="stub">
<a href="https://www.amd.com/en/products/accelerators/instinct/mi300.html" target="_blank">AMD Instinct MI300
Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi300/mi325x.html" target="_blank">Instinct
MI325X</a></p>
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi300/mi300x.html" target="_blank">Instinct
MI300X</a></p>
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi300/mi300a.html" target="_blank">Instinct
MI300A</a></p>
</td>
<td>
<p>gfx942</p>
</td>
<td>
<a href="https://www.amd.com/en/technologies/cdna.html#cdna3" target="_blank">CDNA 3</a>
</td>
</tr>
<tr>
<td class="stub">
<a href="https://www.amd.com/en/products/accelerators/instinct/mi200.html" target="_blank">AMD Instinct MI200
Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi200/mi250x.html" target="_blank">Instinct
MI250X</a></p>
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi200/mi250.html" target="_blank">Instinct
MI250</a></p>
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi200/mi210.html" target="_blank">Instinct
MI210</a></p>
</td>
<td>
<p>gfx90a</p>
</td>
<td>
<a href="https://www.amd.com/en/technologies/cdna.html#cdna2" target="_blank">CDNA 2</a>
</td>
</tr>
<tr>
<td class="stub">
<a href="https://www.amd.com/en/products/accelerators/instinct/mi100.html" target="_blank">AMD Instinct MI100
Series</a>
</td>
<td>
<a href="https://www.amd.com/en/products/accelerators/instinct/mi100.html" target="_blank">Instinct MI100</a>
</td>
<td>
<p>gfx908</p>
</td>
<td>
<a href="https://www.amd.com/en/technologies/cdna.html#cdna" target="_blank">CDNA</a>
</td>
</tr>
</tbody>
</table>
:::
:::{tab-item} Radeon
:sync: radeon
<table class="rocm-docs-table table">
<thead>
<colgroup style="width: 33%;">
<colgroup style="width: 32%;">
<tr>
<th class="head">
<p>Device series</p>
</th>
<th class="head">
<p>Device</p>
</th>
<th class="head">
<p>LLVM target</p>
</th>
<th class="head">
<p>Architecture</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<td>
<a href="https://www.amd.com/en/products/graphics/workstations/radeon-ai-pro.html#tabs-95fa144b96-item-b95ec9e1ca-tab"
target="_blank">AMD Radeon AI PRO R9000 Series</a>
</td>
<td>
<p><a
href="https://www.amd.com/en/products/graphics/workstations/radeon-ai-pro/ai-9000-series/amd-radeon-ai-pro-r9700s.html"
target="_blank">Radeon AI PRO R9700S</a></p>
<p><a
href="https://www.amd.com/en/products/graphics/workstations/radeon-ai-pro/ai-9000-series/amd-radeon-ai-pro-r9700.html"
target="_blank">Radeon AI PRO R9700</a></p>
<p><a
href="https://www.amd.com/en/products/graphics/workstations/radeon-ai-pro/ai-9000-series/amd-radeon-ai-pro-r9600d.html"
target="_blank">Radeon AI PRO R9600D</a></p>
</td>
<td>
<p>gfx1201</p>
</td>
<td rowspan="3">
<a href="https://www.amd.com/en/technologies/rdna.html#tabs-1fabb91c39-item-330ee548f0-tab" target="_blank">RDNA
4</p>
</td>
</tr>
<tr>
<td rowspan="2" class="stub">
<a href="https://www.amd.com/en/products/graphics/desktops/radeon.html#tabs-ff9c5c3863-item-37fb38a236-tab"
target="_blank">AMD Radeon RX 9000 Series</p>
</td>
<td>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9070xt.html"
target="_blank">Radeon RX 9070 XT</a></p>
<p><a href="https://www.amd.com/en/support/downloads/drivers.html/graphics/radeon-rx/radeon-rx-9000-series/amd-radeon-rx-9070-gre.html"
target="_blank">Radeon RX 9070 GRE</a></p>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9070.html"
target="_blank">Radeon RX 9070</a></p>
</td>
<td>
<p>gfx1201</p>
</td>
</tr>
<tr>
<td>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9060xt-lp.html"
target="_blank">Radeon RX 9060 XT LP</a></p>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9060xt.html"
target="_blank">Radeon RX 9060 XT</a></p>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9060.html"
target="_blank">Radeon RX 9060</a></p>
</td>
<td>
<p>gfx1200</p>
</td>
</tr>
<tr>
<td rowspan="2" class="stub">
<a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro.html#tabs-990fdead92-item-20daa37284-tab"
target="_blank">AMD Radeon PRO W7000 Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w7900-dual-slot.html"
target="_blank">Radeon PRO W7900 Dual Slot</a></p>
<p><a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w7900.html" target="_blank">Radeon
PRO W7900</a></p>
<p><a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w7800-48gb.html"
target="_blank">Radeon PRO W7800 48GB</a></p>
<p><a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w7800.html" target="_blank">Radeon
PRO W7800</a></p>
</td>
<td>
<p>gfx1100</p>
</td>
<td rowspan="6">
<a href="https://www.amd.com/en/technologies/rdna.html#tabs-1fabb91c39-item-05915f6044-tab" target="_blank">RDNA
3</p>
</td>
</tr>
<tr>
<td>
<p><a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w7700.html" target="_blank">Radeon
PRO W7700</a></p>
</td>
<td>
<p>gfx1101</p>
</td>
</tr>
<tr>
<td rowspan="3" class="stub">
<a href="https://www.amd.com/en/products/graphics/desktops/radeon.html#tabs-ff9c5c3863-item-b55a56bf12-tab"
target="_blank">AMD Radeon RX 7000 Series</p>
</td>
<td>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7900xtx.html"
target="_blank">Radeon RX 7900 XTX</a></p>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7900xt.html"
target="_blank">Radeon RX 7900 XT</a></p>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7900-gre.html"
target="_blank">Radeon RX 7900 GRE</a></p>
</td>
<td>
<p>gfx1100</p>
</td>
</tr>
<tr>
<td>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7800-xt.html"
target="_blank">Radeon RX 7800 XT</a></p>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7700-xt.html"
target="_blank">Radeon RX 7700 XT</a></p>
<p>Radeon RX 7700 XE</p>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7700.html"
target="_blank">Radeon RX 7700</a></p>
</td>
<td>
<p>gfx1101</p>
</td>
</tr>
<tr>
<td>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7600.html"
target="_blank">Radeon RX 7600</a></p>
</td>
<td>
<p>gfx1102</p>
</td>
</tr>
<tr>
<td rowspan="2">
<a href="https://www.amd.com/en/products/accelerators/radeon-pro.html" target="_blank">AMD Radeon PRO V
Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/accelerators/radeon-pro/amd-radeon-pro-v710.html"
target="_blank">Radeon PRO V710</a></p>
</td>
<td>
<p>gfx1101</p>
</td>
</tr>
<tr>
<td>
<p><a href="https://www.amd.com/en/products/accelerators/radeon-pro/amd-radeon-pro-v620.html"
target="_blank">Radeon PRO V620</a></p>
</td>
<td>
<p>gfx1030</p>
</td>
<td rowspan="2">
<a href="https://www.amd.com/en/technologies/rdna.html#tabs-1fabb91c39-item-9ed969eddf-tab" target="_blank">RDNA
2</p>
</td>
</tr>
<tr>
<td class="stub">
<a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w6800.html" target="_blank">AMD Radeon
PRO W6000 Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w6800.html" target="_blank">Radeon
PRO W6800</a></p>
</td>
<td>
<p>gfx1030</p>
</td>
</tr>
</tbody>
</table>
:::
:::{tab-item} Ryzen
:sync: ryzen
<table class="rocm-docs-table table">
<thead>
<colgroup style="width: 26%;">
<colgroup style="width: 40%;">
<tr>
<th class="head">
<p>Device series</p>
</th>
<th class="head">
<p>Device</p>
</th>
<th class="head">
<p>LLVM target</p>
</th>
<th class="head">
<p>Architecture</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<td class="stub">
<a href="https://www.amd.com/en/products/processors/workstations/mobile.html#tabs-7f0c432fb2-item-5116ab7a74-tab"
target="_blank">AMD Ryzen AI Max PRO<br>300 Series</a>
</td>
<td>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-max-pro-300-series/amd-ryzen-ai-max-plus-pro-395.html"
target="_blank">Ryzen AI Max+ PRO 395</a> (Radeon 8060S)</p>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-max-pro-300-series/amd-ryzen-ai-max-pro-390.html"
target="_blank">Ryzen AI Max PRO 390</a> (Radeon 8050S)</p>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-max-pro-300-series/amd-ryzen-ai-max-pro-385.html"
target="_blank">Ryzen AI Max PRO 385</a> (Radeon 8050S)</p>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-max-pro-300-series/amd-ryzen-ai-max-pro-380.html"
target="_blank">Ryzen AI Max PRO 380</a> (Radeon 8040S)</p>
</td>
<td>
<p>gfx1151</p>
</td>
<td rowspan="10">
<p>RDNA 3.5</p>
</td>
</tr>
<tr>
<td class="stub">
<a href="https://www.amd.com/en/products/processors/laptop/ryzen.html#tabs-1181ea0b44-item-6ccfea5f65-tab"
target="_blank">AMD Ryzen AI Max<br>300 Series</a>
</td>
<td>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-plus-395.html"
target="_blank">Ryzen AI Max+ 395</a> (Radeon 8060S)</p>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-plus-392.html"
target="_blank">Ryzen AI Max+ 392</a> (Radeon 8060S)</p>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-plus-388.html"
target="_blank">Ryzen AI Max+ 388</a> (Radeon 8060S)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-390.html"
target="_blank">Ryzen AI Max 390</a> (Radeon 8050S)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-385.html"
target="_blank">Ryzen AI Max 385</a> (Radeon 8050S)</p>
</td>
<td>
<p>gfx1151</p>
</td>
</tr>
<tr>
<td rowspan="2" class="stub">
<a href="https://www.amd.com/en/products/processors/laptop/ryzen-for-business.html#tabs-0d174caf43-item-87690677fc-tab"
target="_blank">AMD Ryzen AI PRO<br>400 Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-400-series/amd-ryzen-ai-9-hx-pro-475.html"
target="_blank">Ryzen AI 9 HX PRO 475</a> (Radeon 890M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-400-series/amd-ryzen-ai-9-hx-pro-470.html"
target="_blank">Ryzen AI 9 HX PRO 470</a> (Radeon 890M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-400-series/amd-ryzen-ai-9-pro-465.html"
target="_blank">Ryzen AI 9 PRO 465</a> (Radeon 880M)</p>
</td>
<td>
<p>gfx1150</p>
</td>
</tr>
<tr>
<td>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-400-series/amd-ryzen-ai-7-pro-450.html"
target="_blank">Ryzen AI 7 PRO 450</a> (Radeon 860M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-400-series/amd-ryzen-ai-5-pro-440.html"
target="_blank">Ryzen AI 5 PRO 440</a> (Radeon 840M)</p>
</td>
<td>
<p>gfx1152</p>
</td>
</tr>
<tr>
<td rowspan="2" class="stub">
<a href="https://www.amd.com/en/products/processors/consumer/ryzen-ai.html#tabs-f556098628-item-808b56dca3-tab"
target="_blank">AMD Ryzen AI<br>400 Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-400-series/amd-ryzen-ai-9-hx-475.html"
target="_blank">Ryzen AI 9 HX 475</a> (Radeon 890M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-400-series/amd-ryzen-ai-9-hx-470.html"
target="_blank">Ryzen AI 9 HX 470</a> (Radeon 890M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-400-series/amd-ryzen-ai-9-465.html"
target="_blank">Ryzen AI 9 465</a> (Radeon 880M)</p>
</td>
<td>
<p>gfx1150</p>
</td>
</tr>
<tr>
<td>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-400-series/amd-ryzen-ai-7-450.html"
target="_blank">Ryzen AI 7 450</a> (Radeon 860M)</p>
</td>
<td>
<p>gfx1152</p>
</td>
</tr>
<tr>
<td rowspan="2" class="stub">
<a href="https://www.amd.com/en/products/processors/workstations/mobile.html#tabs-7f0c432fb2-item-387526c6cc-tab"
target="_blank">AMD Ryzen AI PRO<br>300 Series</a>
</td>
<td>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-300-series/amd-ryzen-ai-9-hx-pro-375.html"
target="_blank">Ryzen AI 9 HX PRO 375</a> (Radeon 890M)</p>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-300-series/amd-ryzen-ai-9-hx-pro-370.html"
target="_blank">Ryzen AI 9 HX PRO 370</a> (Radeon 890M)</p>
</td>
<td>
<p>gfx1150</p>
</td>
</tr>
<tr>
<td>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-300-series/amd-ryzen-ai-7-pro-350.html"
target="_blank">Ryzen AI 7 PRO 350</a> (Radeon 860M)</p>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-300-series/amd-ryzen-ai-5-pro-340.html"
target="_blank">Ryzen AI 5 PRO 340</a> (Radeon 840M)</p>
</td>
<td>
<p>gfx1152</p>
</td>
</tr>
<tr>
<td rowspan="2" class="stub">
<a href="https://www.amd.com/en/products/processors/consumer/ryzen-ai.html#tabs-f556098628-item-54e149d850-tab"
target="_blank">AMD Ryzen AI<br>300 Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-9-hx-375.html"
target="_blank">Ryzen AI 9 HX 375</a> (Radeon 890M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-9-hx-370.html"
target="_blank">Ryzen AI 9 HX 370</a> (Radeon 890M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-9-365.html"
target="_blank">Ryzen AI 9 365</a> (Radeon 880M)</p>
</td>
<td>
<p>gfx1150</p>
</td>
</tr>
<tr>
<td>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-7-350.html"
target="_blank">Ryzen AI 7 350</a> (Radeon 860M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-7-345.html"
target="_blank">Ryzen AI 7 345</a> (Radeon 840M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-5-340.html"
target="_blank">Ryzen AI 5 340</a> (Radeon 840M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-5-330.html"
target="_blank">Ryzen AI 5 330</a> (Radeon 820M)</p>
</td>
<td>
<p>gfx1152</p>
</td>
</tr>
<tr>
<td class="stub">
<a href="https://www.amd.com/en/products/processors/laptop/ryzen-for-business.html#tabs-0d174caf43-item-a8ec88d07e-tab"
target="_blank">AMD Ryzen PRO<br>200 Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/200-series/amd-ryzen-7-pro-250.html"
target="_blank">Ryzen 7 PRO 250</a> (Radeon 780M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/200-series/amd-ryzen-5-pro-230.html"
target="_blank">Ryzen 5 PRO 230</a> (Radeon 760M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/200-series/amd-ryzen-5-pro-220.html"
target="_blank">Ryzen 5 PRO 220</a> (Radeon 740M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/200-series/amd-ryzen-5-pro-215.html"
target="_blank">Ryzen 5 PRO 215</a> (Radeon 740M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/200-series/amd-ryzen-3-pro-210.html"
target="_blank">Ryzen 3 PRO 210</a> (Radeon 740M)</p>
</td>
<td>
<p>gfx1103</p>
</td>
<td rowspan="2">
<a href="https://www.amd.com/en/technologies/rdna.html#tabs-1fabb91c39-item-05915f6044-tab" target="_blank">RDNA
3</a>
</td>
</tr>
<tr>
<td class="stub">
<a href="https://www.amd.com/en/products/processors/laptop/ryzen.html#tabs-1181ea0b44-item-895d56feed-tab"
target="_blank">AMD Ryzen<br>200 Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-9-270.html"
target="_blank">Ryzen 9 270</a> (Radeon 780M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-7-260.html"
target="_blank">Ryzen 7 260</a> (Radeon 780M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-7-250.html"
target="_blank">Ryzen 7 250</a> (Radeon 780M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-5-240.html"
target="_blank">Ryzen 5 240</a> (Radeon 760M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-5-230.html"
target="_blank">Ryzen 5 230</a> (Radeon 760M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-5-220.html"
target="_blank">Ryzen 5 220</a> (Radeon 740M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-3-210.html"
target="_blank">Ryzen 3 210</a> (Radeon 740M)</p>
</td>
<td>
<p>gfx1103</p>
</td>
</tr>
</tbody>
</table>
:::
::::
+311
View File
@@ -0,0 +1,311 @@
::::{tab-set}
:::{tab-item} Instinct
:sync: instinct
<table class="rocm-docs-table table">
<thead>
<colgroup style="width: 33%;">
<tr>
<th class="head">
<p>Linux distribution</p>
</th>
<th class="head">
<p>Supported versions</p>
</th>
<th class="head">
<p>Linux kernel version</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<th rowspan="3" class="stub" style="vertical-align: middle">
<p>Ubuntu</p>
</th>
<td>
<p>26.04</p>
</td>
<td>
<p>GA 7.0</p>
</td>
</tr>
<tr>
<td>
<p>24.04.4</p>
</td>
<td>
<p>GA 6.8</p>
</td>
</tr>
<tr>
<td>
<p>22.04.5</p>
</td>
<td>
<p>GA 5.15</p>
</td>
</tr>
<tr>
<th rowspan="2" class="stub" style="vertical-align: middle">
<p>Debian</p>
</th>
<td>
<p>13</p>
</td>
<td>
<p>6.12</p>
</td>
</tr>
<tr>
<td>
<p>12</p>
</td>
<td>
<p>6.1.0</p>
</td>
</tr>
<tr>
<th rowspan="6" class="stub" style="vertical-align: middle">
<p>Red Hat Enterprise Linux (RHEL)</p>
</th>
<td>
<p>10.1</p>
</td>
<td>
<p>6.12.0-124</p>
</td>
</tr>
<tr>
<td>
<p>10.0</p>
</td>
<td>
<p>6.12.0-55</p>
</td>
</tr>
<tr>
<td>
<p>9.7</p>
</td>
<td>
<p>5.14.0-611</p>
</td>
</tr>
<tr>
<td>
<p>9.6</p>
</td>
<td>
<p>5.14.0-570</p>
</td>
</tr>
<tr>
<td>
<p>9.4</p>
</td>
<td>
<p>5.14.0-427</p>
</td>
</tr>
<tr>
<td>
<p>8.10</p>
</td>
<td>
<p>4.18.0-553</p>
</td>
</tr>
<tr>
<th rowspan="3" class="stub" style="vertical-align: middle">
<p>Oracle Linux</p>
</th>
<td>
<p>10</p>
</td>
<td>
<p>UEK 8.1</p>
</td>
</tr>
<tr>
<td>
<p>9</p>
</td>
<td>
<p>UEK 8</p>
</td>
</tr>
<tr>
<td>
<p>8</p>
</td>
<td>
<p>UEK 7</p>
</td>
</tr>
<tr>
<th class="stub" style="vertical-align: middle">
<p>Rocky Linux</p>
</th>
<td>
<p>9</p>
</td>
<td>
<p>5.14.0-570</p>
</td>
</tr>
<tr>
<th rowspan="2" class="stub" style="vertical-align: middle">
<p>SUSE Linux Enterprise Server (SLES)</p>
</th>
<td>
<p>16.0</p>
</td>
<td>
<p>6.12</p>
</td>
</tr>
<tr>
<td>
<p>15.7</p>
</td>
<td>
<p>6.4.0-150700.51</p>
</td>
</tr>
</tbody>
</table>
:::
:::{tab-item} Radeon
:sync: radeon
<table class="rocm-docs-table table">
<thead>
<colgroup style="width: 33%;">
<tr>
<th class="head">
<p>Operating system</p>
</th>
<th class="head">
<p>Supported versions</p>
</th>
<th class="head">
<p>Linux kernel version</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<th rowspan="3" class="stub" style="vertical-align: middle">
<p>Ubuntu</p>
</th>
<td>
<p>26.04</p>
</td>
<td>
<p>GA 7.0</p>
</td>
</tr>
<tr>
<td>
<p>24.04.4</p>
</td>
<td>
<p>GA 6.8</p>
</td>
</tr>
<tr>
<td>
<p>22.04.5</p>
</td>
<td>
<p>GA 5.15</p>
</td>
</tr>
<tr>
<th rowspan="2" class="stub" style="vertical-align: middle">
<p>Red Hat Enterprise Linux (RHEL)</p>
</th>
<td>
<p>10.1</p>
</td>
<td>
<p>6.12.0-124</p>
</td>
</tr>
<tr>
<td>
<p>9.7</p>
</td>
<td>
<p>5.14.0-611</p>
</td>
</tr>
<tr>
<th class="stub" style="vertical-align: middle">
<p>Windows</p>
</th>
<td>
<p>11 25H2</p>
</td>
<td>
<p style="text-align: center;"> — </p>
</td>
</tr>
</tbody>
</table>
:::
:::{tab-item} Ryzen
:sync: ryzen
<table class="rocm-docs-table table">
<thead>
<colgroup style="width: 33%;">
<tr>
<th class="head">
<p>Operating system</p>
</th>
<th class="head">
<p>Supported versions</p>
</th>
<th class="head">
<p>Linux kernel version</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<th rowspan="2" class="stub" style="vertical-align: middle">
<p>Ubuntu</p>
</th>
<td>
<p>26.04</p>
</td>
<td>
<p>GA 7.0</p>
</td>
</tr>
<tr>
<td>
<p>24.04.4</p>
</td>
<td>
<p>HWE 6.17</p>
</td>
</tr>
<tr>
<th class="stub" style="vertical-align: middle">
<p>Windows</p>
</th>
<td>
<p>11 25H2</p>
</td>
<td>
<p style="text-align: center;"> — </p>
</td>
</tr>
</tbody>
</table>
:::
::::
@@ -0,0 +1,88 @@
<table class="rocm-docs-table table">
<thead>
<tr>
<th class="head">
<p>Device</p>
</th>
<th class="head">
<p>Compute partition mode</p>
</th>
<th class="head">
<p>NPS mode</p>
</th>
<th class="head">
<p>Deployment</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="3" style="vertical-align: middle;">
<p>Instinct MI355X, MI350X</p>
</td>
<td>
<p>CPX</p>
</td>
<td>
<p>NPS2</p>
</td>
<td rowspan="7" style="vertical-align: middle;">
<p>Bare metal</p>
</td>
</tr>
<tr>
<td>
<p>DPX</p>
</td>
<td>
<p>NPS2</p>
</td>
</tr>
<tr>
<td>
<p>QPX</p>
</td>
<td>
<p>NPS2</p>
</td>
</tr>
<tr>
<td rowspan="2" style="vertical-align: middle;">
<p>Instinct MI350P</p>
</td>
<td>
<p>CPX</p>
</td>
<td>
<p>NPS1</p>
</td>
</tr>
<tr>
<td>
<p>SPX</p>
</td>
<td>
<p>NPS1</p>
</td>
</tr>
<tr>
<td rowspan="2" style="vertical-align: middle;">
<p>Instinct MI300X</p>
</td>
<td>
<p>CPX</p>
</td>
<td>
<p>NPS4</p>
</td>
</tr>
<tr>
<td>
<p>DPX</p>
</td>
<td>
<p>NPS2</p>
</td>
</tr>
</tbody>
</table>
@@ -0,0 +1,227 @@
<table class="rocm-docs-table table">
<colgroup style="width: 14%;">
<colgroup style="width: 14%;">
<colgroup style="width: 17%;">
<colgroup style="width: 17%;">
<colgroup style="width: 19%;">
<colgroup style="width: 19%;">
<thead>
<tr>
<th class="head">
<p>AMD GPU</p>
</th>
<th class="head">
<p>Hypervisor</p>
</th>
<th class="head">
<p>Virtualization technology</p>
</th>
<th class="head">
<a>Virtualization driver</a>
</th>
<th class="head">
<p>Host OS</p>
</th>
<th class="head">
<p>Guest OS</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="5" style="vertical-align: middle">
<p>Instinct MI355X</p>
</td>
<td rowspan="4" style="vertical-align: middle">
<p>KVM</p>
</td>
<td style="vertical-align: middle">
<p>Passthrough</p>
</td>
<td>
<p style="text-align: center"></p>
</td>
<td rowspan="4" style="vertical-align: middle">
<p>Ubuntu 24.04</p>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 24.04</p>
</td>
</tr>
<tr>
<td rowspan="3" style="vertical-align: middle">
<p>SR-IOV</p>
</td>
<td rowspan="3" style="vertical-align: middle">
<a href="https://github.com/amd/MxGPU-Virtualization/releases/tag/9.0.0.K" target="_blank">GIM 9.0.0K
</a>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 24.04</p>
</td>
</tr>
<tr>
<td style="vertical-align: middle">
<p>RHEL 10.0</p>
</td>
</tr>
<tr>
<td style="vertical-align: middle">
<p>RHEL 9.6</p>
</td>
</tr>
<tr>
<td style="vertical-align: middle">
<p>ESXi</p>
</td>
<td>
<p style="text-align: center"></p>
</td>
<td>
<p style="text-align: center"></p>
</td>
<td style="vertical-align: middle">
<p>VMware ESXi 9.1</p>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 24.04</p>
</td>
</tr>
<tr>
<td rowspan="3" style="vertical-align: middle">
<p>Instinct MI350X</p>
</td>
<td rowspan="3" style="vertical-align: middle">
<p>KVM</p>
</td>
<td style="vertical-align: middle">
<p>Passthrough</p>
</td>
<td>
<p style="text-align: center"></p>
</td>
<td rowspan="3" style="vertical-align: middle">
<p>Ubuntu 24.04</p>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 24.04</p>
</td>
</tr>
<tr>
<td rowspan="2" style="vertical-align: middle">
<p>SR-IOV</p>
</td>
<td rowspan="2" style="vertical-align: middle">
<a href="https://github.com/amd/MxGPU-Virtualization/releases/tag/9.0.0.K" target="_blank">GIM 9.0.0K
</a>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 24.04</p>
</td>
</tr>
<tr>
<td style="vertical-align: middle">RHEL 9.6</td>
</tr>
<tr>
<td style="vertical-align: middle">
<p>Instinct MI325X</p>
</td>
<td style="vertical-align: middle">
<p>KVM</p>
</td>
<td style="vertical-align: middle">
<p>SR-IOV</p>
</td>
<td style="vertical-align: middle">
<a href="https://github.com/amd/MxGPU-Virtualization/releases/tag/9.0.0.K" target="_blank">GIM 9.0.0K
</a>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 22.04</p>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 22.04</p>
</td>
</tr>
<tr>
<td rowspan="3" style="vertical-align: middle">
<p>Instinct MI300X</p>
</td>
<td rowspan="3" style="vertical-align: middle">
<p>KVM</p>
</td>
<td style="vertical-align: middle">
<p>Passthrough</p>
</td>
<td>
<p style="text-align: center"></p>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 22.04</p>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 22.04</p>
</td>
</tr>
<tr>
<td rowspan="2" style="vertical-align: middle">
<p>SR-IOV</p>
</td>
<td rowspan="2" style="vertical-align: middle">
<a href="https://github.com/amd/MxGPU-Virtualization/releases/tag/9.0.0.K" target="_blank">GIM 9.0.0K
</a>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 24.04</p>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 24.04</p>
</td>
</tr>
<tr>
<td style="vertical-align: middle">
<p>Ubuntu 22.04</p>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 22.04</p>
</td>
</tr>
<tr>
<td rowspan="3" style="vertical-align: middle">
<p>Instinct MI210</p>
</td>
<td rowspan="3" style="vertical-align: middle">
<p>KVM</p>
</td>
<td style="vertical-align: middle">
<p>Passthrough</p>
</td>
<td>
<p style="text-align: center"></p>
</td>
<td rowspan="3" style="vertical-align: middle">
<p>RHEL 9.4</p>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 22.04</p>
</td>
</tr>
<tr>
<td rowspan="2" style="vertical-align: middle">
<p>SR-IOV</p>
</td>
<td rowspan="2" style="vertical-align: middle">
<a href="https://github.com/amd/MxGPU-Virtualization/releases/tag/9.0.0.K" target="_blank">GIM 9.0.0K
</a>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 22.04</p>
</td>
</tr>
<tr>
<td style="vertical-align: middle">
<p>RHEL 9.4</p>
</td>
</tr>
</tbody>
</table>
+1 -1
View File
@@ -4,7 +4,7 @@
<meta name="keywords" content="license, licensing terms">
</head>
# ROCm license
# ROCm licenses
```{include} ../../LICENSE
```
+616
View File
@@ -0,0 +1,616 @@
# ROCm Core SDK {{ ROCM_VERSION }} release notes
ROCm Core SDK {{ ROCM_VERSION }} continues the technology preview release stream
that began with ROCm 7.9.0, advancing the transition to the new
[TheRock](https://github.com/rocm/therock) build and release system. To learn
more, see the [transition guide](/about/transition-guide-TheRock).
(preview-stream-note)=
:::{important}
ROCm {{ ROCM_VERSION }} follows the
<a href="https://rocm.docs.amd.com/en/7.9.0-preview/about/release-notes.html#preview-stream-note"
target="_blank">versioning discontinuity that began with the 7.9.0 preview release</a>
and remains separate from the 7.0 to 7.2 production releases. For the latest
production stream release, see the
<a href="https://rocm.docs.amd.com/en/latest/">ROCm documentation</a>.
Maintaining parallel release streams -- preview and production -- gives
users ample time to evaluate and adopt the new build system and dependency
changes. The technology preview stream is planned to continue through
mid-2026, after which it will replace the current production stream.
For previous preview releases, see the
<a target="_blank" href="https://rocm.docs.amd.com/en/7.12.0-preview/release/versions.html">release history</a>.
:::
## Release highlights
ROCm Core SDK {{ ROCM_VERSION }} with TheRock builds upon the [7.12.0 preview
release](https://rocm.docs.amd.com/en/7.12.0-preview/about/release-notes.html).
This release expands support for AI inference, distributed workloads, and
profiling workflows across AMD Instinct™, Radeon™, and Ryzen™ AI platforms.
ROCm 7.13.0 adds inference-ready vLLM containers, expands GPU virtualization
and partitioning support, introduces new profiling and tracing capabilities,
and improves AI kernel, sparse math, and communication libraries.
### Platform and hardware support
This release expands GPU, operating system, virtualization, and partitioning support.
#### Expanded AMD GPU support
ROCm 7.13.0 adds support for the following AMD GPUs and APUs:
* AMD Instinct MI350P (gfx950)
* AMD Radeon AI PRO R9700S (gfx1201)
* AMD Radeon PRO W6800 (gfx1030)
* AMD Radeon PRO V620 (gfx1030)
* AMD Ryzen AI 7 PRO 360 (gfx1152)
* AMD Ryzen AI 7 PRO 350 (gfx1152)
* AMD Ryzen AI 5 PRO 340 (gfx1152)
* AMD Ryzen AI 7 350 (gfx1152)
* AMD Ryzen AI 7 345 (gfx1152)
* AMD Ryzen AI 5 340 (gfx1152)
* AMD Ryzen AI 5 330 (gfx1152)
For the complete list of supported AMD hardware, see [AMD hardware support](#amd-hardware-support).
#### Expanded Ubuntu support
ROCm 7.13.0 adds support for Ubuntu 26.04 on Instinct, Radeon, and Ryzen
devices.
24.04.4 is now the validated Ubuntu 24 release instead of Ubuntu 24.04.3.
For the full list of supported Linux distributions, see [Operating system support](#operating-system-support).
#### Expanded GPU virtualization support for Instinct GPUs
ROCm 7.13.0 adds support for the following virtualization configurations on AMD Instinct GPUs.
* On MI355X: VMware ESXi 9.1 with Ubuntu 24.04 guest OS.
* On MI300X: KVM SR-IOV with Ubuntu 24.04 host OS and Ubuntu 24.04 guest OS.
* On MI210:
* KVM passthrough with RHEL 9.4 host OS and Ubuntu 22.04 guest OS.
* KVM SR-IOV with RHEL 9.4 host OS and Ubuntu 22.04 guest OS.
* KVM SR-IOV with RHEL 9.4 host OS and RHEL 9.4 guest OS.
Supported SR-IOV configurations require the [GIM Driver
9.0.0K](https://github.com/amd/MxGPU-Virtualization/releases/tag/9.0.0.K). For
details, see [GPU virtualization support](#gpu-virtualization-support).
#### Expanded Instinct GPU partitioning support
ROCm 7.13.0 enables the following GPU partitioning configurations in bare metal deployments:
* On MI355X and MI350X: QPX compute partition mode with NPS2 memory partitioning.
* On MI350P:
* CPX compute partition mode with NPS1 memory partition.
* SPX compute partition mode with NPS1 memory partition.
For details, see [GPU partitioning support](#gpu-partitioning-support).
### AI inference and frameworks
This release adds inference-ready container images and improves multi-node communication for distributed workloads.
#### vLLM 0.19.1 Docker images and pip packages
With ROCm 7.13.0, Docker images for running vLLM inference workloads are
available. Images include vLLM 0.19.1, PyTorch 2.10, and Python 3.13 on Ubuntu 24.04.
Architecture-specific images are available for:
* AMD Instinct GPUs: gfx942 (MI325X, MI300X, MI300A) and gfx950 (MI355X, MI350X, MI350P)
* AMD Radeon GPUs: gfx1100, gfx1101, gfx1102, gfx1200, gfx1201
* AMD Ryzen AI APUs: gfx1150, gfx1151, gfx1152
See [](../ai-inference/vllm) to get started.
#### RCCL multi-node optimization for AMD Ryzen AI Max 300 series
RCCL improves multi-node clustering performance on systems with AMD Ryzen AI
Max 300 series connected over Ethernet. Building on the initial
multi-node enablement in ROCm 7.12.0, this release optimizes collective
communication for distributed AI inference workloads using tensor parallelism
(TP) and expert parallelism (EP) across up to 4 Ethernet-connected nodes.
#### RCCL GDA-based alltoall via rocSHMEM integration (experimental)
RCCL adds experimental support for GPU Direct Async (GDA)-based alltoall and
alltoallv collective operations through rocSHMEM integration. When enabled,
RCCL invokes rocSHMEM operations that use GDA to reduce latency for small
message alltoall patterns.
This feature requires building RCCL with the `--rocshmem` flag and setting
`RCCL_ROCSHMEM_ENABLE=1` at runtime. GDA support currently requires Broadcom
NICs with GDA capability.
### Developer tools and profiling
This release adds new profiling capabilities, introduces the open-source ROCprof Trace Decoder, and extends HIP programming APIs.
#### ROCprof Trace Decoder open source release
ROCprof Trace Decoder, previously delivered as a closed-source
component within ROCprofiler-SDK, is now available as the open-source
rocprof-trace-decoder library. The decoder converts raw SQTT data from AMD GPUs
into structured execution traces for performance analysis and debugging. It
supports a wide range of AMD GPUs spanning Instinct, Radeon, and Ryzen
architectures, with unit and integration tests across all supported hardware.
See [AMD hardware support](#amd-hardware-support) for the complete list.
<!-- For more information, see [ROCprof Trace Decoder and thread trace APIs](https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/docs-7.13.0/api-reference/thread_trace.html) and [Using thread trace](https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/docs-7.13.0/how-to/using-thread-trace.html) in the ROCprofiler-SDK documentation. -->
#### HIP cooperative groups reduce operations
HIP adds `cooperative_groups::reduce()` for performing reduction operations
across `thread_block_tile` and `coalesced_threads` groups. The implementation
is based on `__reduce_*_sync` operations, and the
`HIP_ENABLE_EXTRA_WARP_SYNC_TYPES` macro might be required to enable some
optimizations.
Additionally, `__reduce_and_sync()`, `__reduce_or_sync()`, and
`__reduce_xor_sync()` now provide consistent behavior for all mask values. All
masks now emit bitwise instructions, aligning behavior with NVIDIA CUDA. This
is a change from previous versions, where some masks were translated to bitwise
operations, and others were not.
#### ROCm Compute Profiler feature highlights
The following are notable enhancements to the ROCm Compute Profiler
(rocprofiler-compute).
* **RDNA 3.5 support:** ROCm Compute Profiler now supports GPU performance
profiling and analysis on AMD Ryzen AI Max 300 series processors.
<!-- An [RDNA 3 -->
<!-- section](https://rocm.docs.amd.com/projects/rocprofiler-compute/en/docs-7.13.0/conceptual/rdna/rdna-performance-model.html) -->
<!-- has been added to the performance model documentation explaining the supported -->
<!-- performance metrics for AMD Ryzen AI Max 300 series processors. A new memory -->
<!-- chart visualization accommodates the architectural differences between -->
<!-- RDNA 3.5 and CDNA GPUs. Roofline is not yet supported for AMD Ryzen AI -->
<!-- Max 300 series processors. -->
* **Removed dependency requirements for profiling:** Building ROCm Compute
Profiler and using profile mode no longer requires installing Python
dependencies from the `requirements.txt` file. Analysis mode still requires
Python dependencies.
This change moves several operations from profile mode to analysis mode,
including roofline HTML generation, roofline-related options
(`--sort`, `--mem-level`, `--roofline-data-type`), and creation of the
combined `pmc_perf.csv` file. Profile mode now only runs the roofline
empirical benchmark, creates a `roofline.csv` file, and creates per-replay
CSV files without merging them.
#### ROCm Systems Profiler feature highlights
The following are notable enhancements to the ROCm Systems Profiler
(rocprofiler-systems).
* **Pause and resume profiling:** ROCm Systems Profiler now supports pausing
and resuming profiling at runtime through the `roctxProfilerPause` and
`roctxProfilerResume` APIs. This allows you to capture profiling data only
during specific execution phases, reducing overhead and minimizing output size
for long-running workloads.
<!-- For more information, see [Configuring runtime -->
<!-- options](https://rocm.docs.amd.com/projects/rocprofiler-systems/en/docs-7.13.0/how-to/configuring-runtime-options.html) -->
<!-- in the ROCm Systems Profiler documentation. -->
* **Selective region tracing:** You can now restrict tracing to defined regions
of interest using the `ROCPROFSYS_SELECTED_REGIONS` environment variable,
reducing noise and limiting data collection to relevant workload segments.
<!-- For more -->
<!-- information, see -->
<!-- [ROCPROFSYS_SELECTED_REGIONS](https://rocm.docs.amd.com/projects/rocprofiler-systems/en/docs-7.13.0/how-to/configuring-runtime-options.html#rocprofsys-selected-regions) -->
<!-- in the ROCm Systems Profiler documentation. -->
* **KFD event tracing:** Kernel Fusion Driver (KFD) event tracing is now
available for GPU memory management analysis, including page faults, page
migrations, queue evictions, GPU unmap events, and dropped events. Requires
an XNACK-capable GPU and ROCprofiler-SDK 1.2.1 or later.
<!-- For more -->
<!-- information, see [Configuring runtime -->
<!-- options](https://rocm.docs.amd.com/projects/rocprofiler-systems/en/docs-7.13.0/how-to/configuring-runtime-options.html#exploring-gpu-metrics) -->
<!-- in the ROCm Systems Profiler documentation. -->
* **MPI file-output filtering:** You can now filter profiler output files based
on MPI rank using the `--rank-filter-output` CLI option or the
`ROCPROFSYS_RANK_FILTER_OUTPUT` configuration setting, suppressing output
from all other ranks. An optional `--rank-filter-id` option
(`ROCPROFSYS_RANK_FILTER_ID`) allows specifying a custom environment variable
for rank identification.
<!-- For more information, see [Selective rank -->
<!-- profiling](https://rocm.docs.amd.com/projects/rocprofiler-systems/en/docs-7.13.0/how-to/communication-runtime-profiling.html#selective-rank-profiling) -->
<!-- in the ROCm Systems Profiler documentation. -->
* **JSON-based profiling presets and domain flags:** You can now configure
common profiling workflows using JSON-based presets and a single
`--preset=<name>` flag instead of manually setting multiple `ROCPROFSYS_*`
environment variables. Eleven built-in presets cover common profiling scenarios, including GPU
tracing, HPC workloads, and API-level analysis. Composable domain flags
(`--gpu`, `--rocm`, `--cpu`, `--parallel`) and a topic-based
`--help=<topic>` system further simplify configuration and discoverability.
<!-- For more information, see [Using preset profiles and domain -->
<!-- flags](https://rocm.docs.amd.com/projects/rocprofiler-systems/en/docs-7.13.0/how-to/using-preset-profiles.html) -->
<!-- in the ROCm Systems Profiler documentation. -->
#### AMD SMI feature highlights
* **APU metrics and memory tuning**: New APU telemetry provides per-core
temperature, power, clock, voltage, current, and throttle monitoring, with
additional support for IPU activity and DRAM bandwidth metrics. New VRAM
carveout and GTT tuning controls enable configurable memory allocation on
supported APU platforms.
* **Per-component GPU temperature and clock monitoring**: GPU metrics table
version 1.9 adds HBM stack temperatures, per-die temperature monitoring, and
per-die memory and SOC clock reporting for data center deployments.
* **CPU power APIs report in milliwatts (breaking change)**: CPU power APIs now
return values in milliwatts (mW) instead of watts. Python bindings now return
numeric integer values instead of formatted strings. Existing applications
that parse previous string-based outputs must be updated.
For more information, see the AMD SMI section in the [ROCm component changelogs](#rocm-component-changelogs).
### Libraries
This release adds new routines, data type support, and performance improvements across ROCm math and AI libraries.
#### Composable Kernel adds quantization and attention kernel capabilities
Composable Kernel adds several capabilities for AI and large language model
workloads:
* **Microscaling (MX) FP8/FP4 support:** Mixed data type support for MX FP8 and
FP4 in GEMM and Flash Multi-Head Attention (FMHA) forward kernels on AMD
Instinct MI350 Series GPUs.
* **FP8 quantization for FMHA:** FMHA forward kernels now support multiple FP8
quantization modes, including dynamic tensor-wise quantization, block scale
quantization, per-tensor quantization, and FP8 KV cache support for batch
prefill.
* **StreamingLLM and long-context inference:** Sink token support for FMHA
forward enables StreamingLLM-style long-context inference.
* **Batch prefill enhancements:** FMHA batch prefill kernels now support
multiple KV cache layouts, flexible page sizes, and configurable lookup table
configurations.
* **RDNA 3 FMHA support:** Flash Attention kernels are now available on RDNA 3
architectures.
* **SageAttention v2 forward kernel:** Multi-granularity quantization for Q, K,
and V tensors with FP8, INT8, and INT4 data types and per-tensor, per-block,
per-warp, and per-thread scale granularities on AMD Instinct MI300 Series and
MI350 Series GPUs.
#### General Batched GEMM support in hipBLASLt
hipBLASLt adds native support for General Batched GEMM, where all matrices in
a batch share the same problem dimensions but can have independent leading
dimensions and strides. This replaces the previous implementation through the
`hipblaslt_ext` Grouped GEMM APIs, which had known limitations.
The new implementation includes support for Global Split-U (GSU) to improve
performance at large problem sizes. General Batched GEMM is important for
inference workloads that dispatch batches of same-shape GEMM operations.
<!-- For more information, see the [hipBLASLt -->
<!-- documentation](https://rocm.docs.amd.com/projects/hipBLASLt/en/docs-7.13.0/index.html). -->
#### rocSOLVER adds new solver routines and matrix analysis functions
rocSOLVER adds the following new routines, all with 64-bit index support:
* **GETRS_NPVT:** Solution of linear systems using LU factorization without
pivoting. Batched and strided-batched variants are available.
* **SYTRS:** Solution of linear systems for symmetric matrices. Batched and
strided-batched variants are available.
Additionally, POTF2 and downstream POTRF Cholesky factorization performance
have been improved.
<!-- For more information, see the [rocSOLVER -->
<!-- documentation](https://rocm.docs.amd.com/projects/rocSOLVER/en/docs-7.13.0/index.html). -->
#### rocSPARSE adds sparse factorization routines
rocSPARSE adds new generic API routines for sparse incomplete factorization and
triangular solve:
* `rocsparse_spic0` and `rocsparse_spilu0`: Generic incomplete Cholesky (IC0)
and incomplete LU (ILU0) factorization routines with strided-batched
computation support.
* `rocsparse_sptrsv`: Extended with strided-batched computation support and
singularity detection through the new `rocsparse_singularity` enumeration.
Performance of tridiagonal solvers `rocsparse_Xgtsv_no_pivot` and
`rocsparse_Xgtsv_no_pivot_strided_batch` has been improved.
<!-- For more -->
<!-- information, see the [rocSPARSE -->
<!-- documentation](https://rocm.docs.amd.com/projects/rocSPARSE/en/docs-7.13.0/index.html). -->
#### Added rocDecode and rocJPEG libraries to the ROCm Core SDK
rocDecode provides hardware-accelerated video decoding for H.264, H.265/HEVC,
AV1, and VP9 codecs, while rocJPEG provides hardware-accelerated JPEG decoding
on AMD GPUs. Together, they enable
efficient GPU-based media processing pipelines for data-intensive workloads
such as AI training.
Both libraries are supported on Linux on AMD Instinct, Radeon, and Ryzen AI. See
the projects in [ROCm/rocm-systems](https://github.com/ROCm/rocm-systems) for
more information.
#### Added ROCm Data Center Tool to the ROCm Core SDK
ROCm Data Center Tool (RDC) provides telemetry collection, health monitoring,
and job-level GPU statistics for data center deployments with AMD Instinct
accelerators. RDC enables system administrators and cluster managers to monitor
GPU health, collect telemetry data, and track per-job GPU usage across
multi-node environments.
RDC is supported on Linux with AMD Instinct GPUs.
<!-- See the -->
<!-- [RDC documentation](https://rocm.docs.amd.com/projects/rdc/en/docs-7.13.0/index.html) -->
<!-- for more information. -->
(release-supported-hw)=
## AMD hardware support
The following table lists supported AMD Instinct GPUs, Radeon GPUs, and Ryzen
APUs. Each supported device is listed with its corresponding GPU
microarchitecture and LLVM target.
:::{note}
If your GPU is not listed, it might be community-enabled through TheRock
nightly builds. For more information, see [TheRock supported
GPUs](https://github.com/ROCm/TheRock/blob/main/SUPPORTED_GPUS.md). For
installation guidance, see [TheRock
releases](https://github.com/ROCm/TheRock/blob/main/RELEASES.md).
:::
```{include} ./include/hardware-support-table.md
:parser: myst
```
(release-supported-os)=
## Operating system support
ROCm supports the following Linux distribution and Microsoft Windows versions.
If you're running ROCm on Linux, ensure your system is using a supported kernel
version.
:::{important}
The following table is a general overview of supported OSes. Actual support
might vary by AMD GPU or APU. Use the {doc}`Compatibility matrix
</compatibility/compatibility-matrix>` to verify support for your specific
setup before installation.
:::
```{include} ./include/os-support-table.md
:parser: myst
```
## Installation updates
ROCm 7.13.0 introduces several improvements to the Runfile Installer:
* Performance improvements for installing and uninstalling gfx architectures.
* ROCm component tests are now included.
* Support for prerequisite OEM kernel installation as part of the dependency install on Ryzen systems. You no longer need to install it manually.
* Auto-detection of the GPU when using the GUI or when the `gfx=` argument is not provided on the command line. If the installer cannot detect the GPU, you must specify the gfx architecture using the GUI or the `gfx=` argument.
(release-supported-fw)=
## Kernel driver and firmware bundle support
ROCm requires a coordinated stack of compatible firmware, driver, and user
space components. Maintaining version alignment between these layers ensures
correct GPU operation and performance, especially for AMD data center products.
While AMD publishes the AMD GPU driver and ROCm user space components, your
server OEM (original equipment manufacturer) or infrastructure provider
distributes the firmware packages. AMD supplies those firmware images (PLDM
bundles), which the OEM integrates and distributes.
```{include} ./include/driver-firmware-support-table.md
:parser: myst
```
(release-virtualization-support)=
## GPU virtualization support
AMD Instinct data center GPUs support virtualization in the following
configurations. Supported SR-IOV configurations require the AMD GPU
Virtualization Driver (GIM) 9.0.0K -- see the [AMD Instinct Virtualization
Driver
documentation](https://instinct.docs.amd.com/projects/virt-drv/en/mainline-9.0.0.k/)
for more information.
```{include} ./include/virtualization-support-table.html
:parser: myst
```
(release-gpu-partitioning-support)=
## GPU partitioning support
The following compute partition and NUMA-per-socket (NPS) configurations are
available on AMD Instinct GPUs in bare metal deployments.
```{include} ./include/partitioning-support-table.html
:parser: myst
```
See the [AMD GPU partitioning](https://instinct.docs.amd.com/projects/amdgpu-docs/en/latest/gpu-partitioning/index.html)
topic in the AMD GPU Driver documentation to learn more.
(release-ai-ecosystem)=
## AI ecosystem support
ROCm 7.13.0 provides optimized support for popular deep learning frameworks and
AI inference engines. The following table lists supported frameworks and
libraries, their compatible operating systems, and validated versions.
```{include} ./include/ai-ecosystem-support-table.html
:parser: myst
```
(release-components)=
## ROCm Core SDK components
The following table lists core tools and libraries included in the ROCm 7.13.0
release.
:::{important}
The following table is a general overview of ROCm Core SDK components. Actual
support for these libraries and tools can vary by GPU and OS. Use the
{doc}`Compatibility matrix </compatibility/compatibility-matrix>` to verify
support for your specific setup.
:::
```{include} ./include/core-sdk-components-table.html
:parser: myst
```
### ROCm component changelogs
The following sections describe key changes to ROCm Core SDK components.
```{include} ./include/core-sdk-components-aggregated-changelog.md
:parser: myst
```
## ROCm known issues
ROCm known issues are noted on {fab}`github` [GitHub](https://github.com/ROCm/ROCm/labels/Verified%20Issue). These issues will be fixed in a future ROCm release. For known issues related to individual components, review the [ROCm component changelogs](#rocm-component-changelogs).
### ROCm Compute Profiler might fail when profiling bash script or command
Running a bash script or command as a target for ROCm Compute Profiler might fail because bash overwrites the required environment variables. As a workaround, use `--no-native-tool` option in the profile mode. Note that this will disable iteration multiplexing.
### hipFFT and rocFFT callback examples fail to build on Windows
The hipFFT and rocFFT callback examples in [rocm-examples](https://github.com/rocm/rocm-examples) fail to build on a Windows operating system due to a linker error. CMake configuration and HIP object compilation will complete successfully, but the final link step fails with `clang: error: invalid linker name in argument '-fuse-ld=lld-link'` This issue affects all Windows configurations using Relocatable Device Code (RDC) mode. Linux builds are not affected. As a workaround, skip the hipFFT and rocFFT callback examples on Windows, and refer to the Linux builds or [callback](https://github.com/ROCm/rocm-examples/tree/amd-staging/Libraries/rocFFT/callback/) functionality documentation.
### QMCPACK might become unresponsive during DMC simulation on AMD Instinct MI300A GPUs
QMCPACK might become unresponsive when running Diffusion Monte Carlo (DMC) simulations with certain inputs on AMD Instinct MI300A GPUs. The application stops making progress after initialization and must be terminated manually.
### Resource-intensive workloads might result in GPU memory faults
Applications that pass large, complex data structures between device functions using scratch memory, and particularly rely on compiler optimization to minimize the number of copy operations, might encounter GPU memory access faults and become unresponsive.
### Increased binary size for multi-target GPU builds
Applications targeting multiple AMD GPU architectures might observe significantly larger binary sizes. Multi-target builds can produce binaries up to 54 percent larger. Single-target builds add approximately 8 MB of additional size per GPU target. As a workaround, reduce the number of GPU targets in multi-target builds, or strip the resource-usage symbols from release binaries.
### HIP cooperative groups might fail when compiled using the SPIR-V path
HIP applications that use cooperative groups might fail at kernel launch when compiled with `--offload-arch=amdgcnspirv`. The application fails at runtime with `LLVM ERROR: Cannot select: intrinsic %llvm.amdgcn.s.wait.asynccnt` error message. This
affects all GPU architectures when using the SPIR-V compilation path. As a workaround, compile using a direct GPU architecture target (for example, `--offload-arch=gfx942`) instead of `--offload-arch=amdgcnspirv`.
### Illegal memory address error when using placement new with device function returns
HIP kernels that use the placement new operators to construct objects in the `hipMalloc` device memory might crash with `hipErrorIllegalAddress` error message when you pass a `__device__` function return value as the constructor argument. This only affects non-trivially-copyable types (for example, types with user-defined or deleted copy/move constructors). Trivially-copyable types are not affected. As a workaround, assign the device function return value to a local variable before passing it to placement new.
### LLVM-based compilers might fail when compiling half-precision vector operations
LLVM-based compilers might fail, returning `Failed to find subregs!` error message in `SIInstrInfo::copyPhysReg`, when compiling half-precision vector operations with optimization enabled. The issue was observed at optimization levels `-O1` to `-O3`.
### hipBLAS test suites failure on Windows
When using hipBLAS on Windows, the test suites might return non-zero exit codes, even when all mathematical correctness tests pass. This issue can affect CI/CD pipeline validation and block automated testing workflows on Windows systems, because the test framework might fail to detect successful test completion.
### ROCm Systems Profiler overwrites ROCPD output after process re-attachment
When you use `rocprof-sys-attach` to re-attach to a previously profiled process, the `ROCPD` output database files (.db) are written to the initial session's output directory instead of a new timestamped directory. This makes it difficult to distinguish profiling data between sessions. Perfetto trace files are not affected. As a workaround, back up your output directory before re-attaching to a previously profiled process.
### Missing dependencies when installing ROCm Core SDK
Installing the ROCm Core SDK using `amdrocm-core-sdk` or `amdrocm-core-dev/devel` might succeed, but some dependencies from the dev/devel meta packages might not be installed. As a workaround, install the dev packages manually:
```bash
sudo apt install amdrocm-*
```
### Issues related to AddressSanitizer
Multiple issues associated with AddressSanitizer (ASAN) `-fsanitize=address` being enabled have been observed including:
#### ASAN reports false errors for GPU kernels using shared memory
When you compile GPU kernels with ASAN enabled, kernels that use `__shared__` memory might produce false heap-buffer-overflow errors or GPU memory faults. As a workaround, disable ASAN by removing `-fsanitize=address` setting for affected kernels.
#### GPU kernels fail to launch in ASAN builds with large thread counts
When you build GPU libraries with ASAN enabled, kernels configured with large thread counts might fail to launch with `HSA_STATUS_ERROR_INVALID_ISA` error. As a workaround, reduce the thread block sizes to 256 threads or fewer for ASAN builds. The issue is currently under investigation.
#### ASAN breaks multi-architecture HIP binary builds
HIP applications built with ASAN enabled, targeting multiple GPU architectures, might fail to launch with `RuntimeError: .hipFatBinSegment size N is not a multiple of wrapper size (24)` and `RuntimeError: Unexpected magic 0x00000000 at wrapper i` error messages. Single-architecture builds are not affected. As a workaround, build single-architecture binaries using `--offload-arch` targeting only one GPU architecture, or disable ASAN by removing `-fsanitize=address` for HIP compilation.
#### ASAN produces incorrect results with ternary operators on struct kernel arguments
When you compile GPU kernels with ASAN enabled, ternary operators with struct kernel arguments might produce incorrect results. This can mask real bugs and produce false-positive results during memory-safety validation. The issue doesn't occur when the kernel arguments are first copied to local variables, or when compiled without ASAN. As a workaround, copy kernel arguments to local variables before using them in ternary expressions:
```cpp
auto local_arg = kernel_arg;
result = condition ? local_arg : other_arg;
```
Alternatively, disable ASAN by removing `-fsanitize=address` when compiling GPU kernels.
## ROCm resolved issues
The following notable issues have been fixed in ROCm 7.13.0.
### Multi-ROCm installation failed on RPM-based distributions
Previously, installing multiple ROCm versions side by side on RPM-based distributions (RHEL and SLES) failed due to `.build-id` file conflicts between versioned packages.
### vLLM server failed to launch in ROCm Docker images
Previously, the vLLM server failed to start in ROCm 7.12.0 Docker images with an `ImportError` for `librocm_smi64.so.1` due to missing library path configuration.
### vLLM server failed to launch with tensor parallelism
Previously, the vLLM server failed to start with an invalid device pointer error when launching models with tensor parallelism set to 8 on AMD Instinct MI300 and MI355X GPUs.
### PyTorch DDP Gloo backend test failed on AMD GPUs
Previously, the PyTorch Distributed Data Parallel (DDP) test `test_ddp_apply_optim_in_backward_grad_as_bucket_view_false` failed when using the Gloo backend.
### rocWMMA header produced unknown type errors in HIP RTC
Previously, HIP RTC programs that included the `rocwmma/rocwmma.hpp` header failed to compile with unknown type name errors.
## ROCm upcoming changes
Future releases will add support for:
* Additional ROCm Core SDK components
* Domain-specific expansion toolkits (data science, life science, finance,
simulation, and other HPC domains)
* More AMD hardware support
+338
View File
@@ -0,0 +1,338 @@
# Transition guide from legacy ROCm release stream
[ROCm Core SDK 7.13.0](https://rocm.docs.amd.com/en/7.13.0-preview/index.html#rocm-core-sdk) marks a step change from the ROCm legacy release stream. It is a preview release built on our new build system, TheRock.
## Major changes
<table class="rocm-docs-table table">
<thead>
<tr>
<th class="head">Feature</th>
<th class="head">ROCm Core SDK</th>
<th class="head">ROCm Legacy</th>
<th class="head">Description</th>
</tr>
</thead>
<tbody>
<tr>
<td>Installation directory</td>
<td><code class="docutils literal notranslate"><span class="pre">/opt/rocm/core</span></code></td>
<td><code class="docutils literal notranslate"><span class="pre">/opt/rocm/</span></code></td>
<td>To support additional release streams downstream of the ROCm Core SDK</td>
</tr>
<tr>
<td>Package names</td>
<td><code class="docutils literal notranslate"><span class="pre">amdrocm-[$component]</span></code></td>
<td><code class="docutils literal notranslate"><span class="pre">rocm-[$component]</span></code> or <code class="docutils literal notranslate"><span class="pre">roc[$component]</span></code> or <code class="docutils literal notranslate"><span class="pre">hip[$component]</span></code></td>
<td>Unique package prefix to avoid conflicts with upstream packages</td>
</tr>
<tr>
<td>Extras directory</td>
<td><code class="docutils literal notranslate"><span class="pre">/opt/rocm/extras-7/</span></code></td>
<td>N/A</td>
<td>Shared install prefix scoped to each ROCm major version for projects built on the ROCm Core SDK</td>
</tr>
</tbody>
</table>
## Paths and linking
ROCm Core SDK 7.13.0 maintains ABI and API compatibility with the ROCm 7.2
legacy releases, so recompilation is not required. For installations using your
Linux distribution's package manager, the `amdrocm` meta package configures
`update-alternatives` and provides backward-compatible symlinks for
`/opt/rocm/bin`, `/opt/rocm/lib`, and other `/opt/rocm/` directories. For
tarball installs, update `PATH`, `LD_LIBRARY_PATH`, `ROCM_PATH`, or other
environment variables to reflect the new installation path (`/opt/rocm/core`).
## Software packages
ROCm Core SDK packages are more consolidated than the legacy ROCm release
stream. For example, hipBLAS and rocBLAS are now combined into one package,
`amdrocm-blas`. The table below lists new packages, their contents, and the
corresponding legacy packages.
> **Note:** ASAN packages are not available in 7.13.0 and are planned for a future release.
(linux-packages-available-in-rocm-7-13-0)=
### Linux packages available in ROCm 7.13.0
<table class="rocm-docs-table table">
<thead>
<tr>
<th class="head">ROCm Core SDK Package</th>
<th class="head">Package Contents</th>
<th class="head">ROCm Legacy Package</th>
</tr>
</thead>
<tbody>
<tr>
<td>amdrocm-amdsmi</td>
<td>amd-smi</td>
<td>amd-smi-lib, rocm-smi-lib</td>
</tr>
<tr>
<td>amdrocm-llvm</td>
<td>amdclang++, hipcc, flang</td>
<td>rocm-llvm, rocm-llvm-dev, Fortran compiler (included in rocm-llvm OpenMP runtime)</td>
</tr>
<tr>
<td>amdrocm-runtime</td>
<td>HIP, ROCR, runtime compilation</td>
<td>hip-runtime-amd, rocm-hip-runtime, rocm-language-runtime, hsa-rocr, comgr</td>
</tr>
<tr>
<td>amdrocm-fft</td>
<td>rocFFT, hipFFT, hipFFTW</td>
<td>rocfft, hipfft</td>
</tr>
<tr>
<td>amdrocm-blas</td>
<td>rocBLAS, hipBLAS, hipBLASLt, hipSPARSELt</td>
<td>rocblas, hipblas, hipblaslt, hipsparselt</td>
</tr>
<tr>
<td>amdrocm-sparse</td>
<td>rocSPARSE, hipSPARSE</td>
<td>rocsparse, hipsparse</td>
</tr>
<tr>
<td>amdrocm-solver</td>
<td>rocSOLVER, hipSOLVER</td>
<td>rocsolver, hipsolver, rocalution</td>
</tr>
<tr>
<td>amdrocm-dnn</td>
<td>hipDNN, MIOpen</td>
<td>miopen-hip</td>
</tr>
<tr>
<td>amdrocm-rand</td>
<td>rocRAND, hipRAND</td>
<td>rocrand, hiprand</td>
</tr>
<tr>
<td>amdrocm-ccl</td>
<td>rocPRIM, rocThrust, hipCUB</td>
<td>rocprim, rocthrust, hipcub, rocwmma</td>
</tr>
<tr>
<td>amdrocm-profiler</td>
<td>rocprofiler-systems, rocprofiler-compute, rocprofiler-sdk, roctracer</td>
<td>rocprofiler, rocprofiler-compute, rocprofiler-systems, rocprofiler-sdk, roctracer</td>
</tr>
<tr>
<td>amdrocm-profiler-base</td>
<td>rocprofiler-sdk, roctracer</td>
<td>rocprofiler-register, roctracer, hsa-amd-aqlprofile</td>
</tr>
<tr>
<td>amdrocm-base</td>
<td>rocminfo, rocm-core</td>
<td>rocm-core, rocminfo, rocm-cmake, half</td>
</tr>
<tr>
<td>amdrocm-ck</td>
<td>Composable Kernel</td>
<td>composablekernel</td>
</tr>
<tr>
<td>amdrocm-debugger</td>
<td>rocgdb, ROCdbgapi, ROCr Debug Agent</td>
<td>rocm-gdb, rocm-dbgapi, rocm-debug-agent</td>
</tr>
<tr>
<td>amdrocm-hipify</td>
<td>HIPIFY</td>
<td>hipify-clang</td>
</tr>
<tr>
<td>amdrocm-opencl</td>
<td>OpenCL runtime and ICD loader</td>
<td>rocm-opencl-runtime, rocm-opencl, hip-opencl</td>
</tr>
<tr>
<td>amdrocm-decode</td>
<td>rocDecode (newly included in the ROCm Core SDK)</td>
<td>rocdecode</td>
</tr>
<tr>
<td>amdrocm-jpeg</td>
<td>rocJPEG (newly included in the ROCm Core SDK)</td>
<td>rocjpeg</td>
</tr>
<tr>
<td>amdrocm-rccl</td>
<td>rccl</td>
<td>rccl</td>
</tr>
<tr>
<td>amdrocm-rocshmem</td>
<td>rocSHMEM</td>
<td>rocshmem</td>
</tr>
<tr>
<td>amdrocm-rdc</td>
<td>ROCm Data Center Tool (newly included in the ROCm Core SDK)</td>
<td>rdc</td>
</tr>
<tr>
<td>amdrocm-sysdeps</td>
<td>Bundled third-party dependencies (libdrm, libelf, numa, libVA)</td>
<td>System dependencies</td>
</tr>
</tbody>
</table>
Packages are offered in the following variants:
- **For all supported GPUs** -- works across all GPUs supported by ROCm (for example, `apt install amdrocm-core-sdk7.13`).
- **For a specific GPU architecture** -- smaller install size, but requires you to know the GPU installed in your system (for example, `apt install amdrocm-core-sdk7.13-gfx110x`).
Installing all GPU architectures is not required. You can install packages for a specific architecture, multiple architectures side by side, or all supported GPU architectures.
When redistributing software built on the ROCm Core SDK (for example, via containers), we recommend the all GPU package variant for broad hardware support. If disk footprint is a concern, you can use a single GPU architecture package variant instead.
### Architecture-specific packages available in ROCm 7.13.0
<table class="rocm-docs-table table">
<thead>
<tr>
<th class="head">Architecture Family</th>
<th class="head">Package Suffix</th>
<th class="head">Product Name (Not Exhaustive)</th>
</tr>
</thead>
<tbody>
<tr>
<td>CDNA4</td>
<td>-gfx950</td>
<td>AMD Instinct MI355X / MI350X</td>
</tr>
<tr>
<td>CDNA3</td>
<td>-gfx94x</td>
<td>AMD Instinct MI325X / MI300X / MI300A</td>
</tr>
<tr>
<td>CDNA2</td>
<td>-gfx90a</td>
<td>AMD Instinct MI250X / MI250 / MI210</td>
</tr>
<tr>
<td>CDNA</td>
<td>-gfx908</td>
<td>AMD Instinct MI100</td>
</tr>
<tr>
<td>RDNA4</td>
<td>-gfx120x</td>
<td>AMD Radeon RX 9070 / AMD Radeon RX 9060 / AMD Radeon RX 9070 XT / AMD Radeon RX 9060 XT / AMD Radeon RX 9070 GRE / AMD Radeon AI PRO R9700S / AMD Radeon AI PRO R9700 / AMD Radeon AI PRO R9600D / AMD Radeon RX 9060 XT LP</td>
</tr>
<tr>
<td>RDNA3.5</td>
<td>-gfx1150<br>-gfx1151<br>-gfx1152</td>
<td>AMD Ryzen AI 9 465 / AMD Ryzen AI 9 365 / AMD Ryzen AI 9 HX 475 / AMD Ryzen AI 9 HX 470 / AMD Ryzen AI 9 HX 375 / AMD Ryzen AI 9 HX 370 / AMD Ryzen AI 9 PRO 465 / AMD Ryzen AI 9 PRO HX 475 / AMD Ryzen AI 9 PRO HX 470 / AMD Ryzen AI 9 HX PRO 375 / AMD Ryzen AI 9 HX PRO 370 / AMD Ryzen AI Max 390 / AMD Ryzen AI Max 385 / AMD Ryzen AI Max+ 395 / AMD Ryzen AI Max+ 392 / AMD Ryzen AI Max+ 388 / AMD Ryzen AI Max PRO 390 / AMD Ryzen AI Max PRO 385 / AMD Ryzen AI Max PRO 380 / AMD Ryzen AI Max+ PRO 395 / AMD Ryzen AI 7 450 / AMD Ryzen AI 7 350 / AMD Ryzen AI 7 345 / AMD Ryzen AI 5 340 / AMD Ryzen AI 5 330 / AMD Ryzen AI 7 PRO 450 / AMD Ryzen AI 5 PRO 440 / AMD Ryzen AI 7 PRO 350 / AMD Ryzen AI 5 PRO 340</td>
</tr>
<tr>
<td>RDNA3</td>
<td>-gfx110x</td>
<td>AMD Radeon RX 7700 / AMD Radeon RX 7600 / AMD Radeon PRO V710 / AMD Radeon PRO W7900 / AMD Radeon PRO W7800 / AMD Radeon PRO W7700 / AMD Radeon RX 7900 XT / AMD Radeon RX 7800 XT / AMD Radeon RX 7700 XT / AMD Radeon RX 7700 XE / AMD Radeon RX 7900 XTX / AMD Radeon RX 7900 GRE / AMD Radeon PRO W7800 48GB / AMD Radeon PRO W7900 Dual Slot</td>
</tr>
<tr>
<td>RDNA2</td>
<td>-gfx1030</td>
<td>AMD Radeon PRO V620 / AMD Radeon PRO W6800</td>
</tr>
</tbody>
</table>
## ROCm Core SDK component changes (moved or removed)
### Planned for future releases
- ROCm Core SDK: RPP
- ROCm-Extras: hipfort, rocALUTION, rocPyDecode, rocAL, MIVisionX
### Moved to ROCm-Extras
- ROCm Validation Suite
- ROCm Bandwidth Test
- TransferBench
- MIGraphX
### Moved to Standalone/ONNX
- ONNX runtime
### Removed
- [ROCm SMI](https://rocm.docs.amd.com/en/latest/about/release-notes.html#rocm-smi-deprecation) (replaced by AMD SMI)
## Notable package relocations
- rocMLIR (now included in MIGraphX)
- HIPCC (now included in `amdrocm-llvm`)
- FLANG (now included in `amdrocm-llvm`)
- ROCm CMake (now in `amdrocm-base`)
- ROCTracer (now in `amdrocm-profiler-base`)
- ROCProfiler (functionality in `amdrocm-profiler`)
## Components available in the ROCm Core SDK, ROCm-Extras, and Standalone/ONNX
<table class="rocm-docs-table table">
<thead>
<tr>
<th class="head"></th>
<th class="head">Category</th>
<th class="head">Present</th>
<th class="head">Absent/Moved</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="6" class="stub" style="vertical-align: middle"><strong>ROCm Core SDK</strong></td>
<td>Math and compute libraries</td>
<td>CK, hipBLAS, hipBLASLt, hipCUB, hipFFT, hipRAND, hipSOLVER, hipSPARSE/SPARSELt, MIOpen, rocBLAS, rocFFT, rocRAND, rocSOLVER, rocSPARSE, rocPRIM, rocThrust, rocWMMA</td>
<td>hipfort, rocALUTION</td>
</tr>
<tr>
<td>Communication libraries</td>
<td>RCCL, rocSHMEM</td>
<td>—</td>
</tr>
<tr>
<td>Media libraries</td>
<td>rocDecode, rocJPEG, ROCm Performance Primitives (RPP planned for a future release)</td>
<td>rocPyDecode, rocAL, MIVisionX, MIGraphX, CK (moved to math and compute)</td>
</tr>
<tr>
<td>Runtime, compilers, build tools</td>
<td>HIP, HIPIFY, LLVM</td>
<td>HIPCC, FLANG, ROCm CMake</td>
</tr>
<tr>
<td>Profiling and debugging tools</td>
<td>ROCm Compute Profiler, ROCm Systems Profiler, ROCprofiler-SDK, ROCdbgapi, ROCm Debugger, ROCr Debug Agent</td>
<td>ROCTracer, ROCProfiler</td>
</tr>
<tr>
<td>Control and monitoring tools</td>
<td>AMD SMI, ROCm Data Center Tool, rocminfo, hipinfo</td>
<td>ROCm SMI (removed), ROCm Validation Suite, ROCm Bandwidth Test</td>
</tr>
<tr>
<td style="vertical-align: middle"><strong>ROCm-Extras</strong></td>
<td>—</td>
<td>ROCm Validation Suite, ROCm Bandwidth Test, TransferBench, MIGraphX</td>
<td>—</td>
</tr>
<tr>
<td style="vertical-align: middle"><strong>Standalone/ONNX</strong></td>
<td>—</td>
<td>rocMLIR, ONNX runtime</td>
<td>—</td>
</tr>
</tbody>
</table>
@@ -8,7 +8,7 @@ docker:
- "Parallel VAE decode support for Wan models"
- "Batch inference and data parallel support"
components:
TheRock:
TheRock:
version: 9b611c6
url: https://github.com/ROCm/TheRock
rocm-libraries:
@@ -75,7 +75,7 @@ docker:
- '--guidance_scale 6.0 \'
- '--use_torch_compile \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- model: Hunyuan Video 1.5
model_repo: hunyuanvideo-community/HunyuanVideo-1.5-Diffusers-720p_t2v
url: https://huggingface.co/hunyuanvideo-community/HunyuanVideo-1.5-Diffusers-720p_t2v
@@ -97,7 +97,7 @@ docker:
- '--enable_tiling --enable_slicing \'
- '--use_torch_compile \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- group: Wan-AI
js_tag: wan
models:
@@ -123,7 +123,7 @@ docker:
- '--num_inference_steps 40 \'
- '--use_torch_compile \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- model: Wan2.2
model_repo: Wan-AI/Wan2.2-I2V-A14B-Diffusers
url: https://huggingface.co/Wan-AI/Wan2.2-I2V-A14B-Diffusers
@@ -146,7 +146,7 @@ docker:
- '--num_inference_steps 40 \'
- '--use_torch_compile \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- group: FLUX
js_tag: flux
models:
@@ -172,7 +172,7 @@ docker:
- '--guidance_scale 0.0 \'
- '--num_iterations 50 \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- model: FLUX.1 Kontext
model_repo: black-forest-labs/FLUX.1-Kontext-dev
url: https://huggingface.co/black-forest-labs/FLUX.1-Kontext-dev
@@ -196,7 +196,7 @@ docker:
- '--guidance_scale 2.5 \'
- '--num_iterations 25 \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- model: FLUX.2
model_repo: black-forest-labs/FLUX.2-dev
url: https://huggingface.co/black-forest-labs/FLUX.2-dev
@@ -220,7 +220,7 @@ docker:
- '--guidance_scale 4.0 \'
- '--num_iterations 25 \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- model: FLUX.2 Klein
model_repo: black-forest-labs/FLUX.2-klein-9B
url: https://huggingface.co/black-forest-labs/FLUX.2-klein-9B
@@ -242,7 +242,7 @@ docker:
- '--guidance_scale 1.0 \'
- '--num_iterations 25 \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- group: StableDiffusion
js_tag: stablediffusion
models:
@@ -263,7 +263,7 @@ docker:
- '--use_cfg_parallel \'
- '--use_torch_compile \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- group: Z-Image
js_tag: z_image
models:
@@ -289,7 +289,7 @@ docker:
- '--guidance_scale 4.0 \'
- '--num_iterations 25 \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- group: LTX
js_tag: ltx
models:
@@ -313,7 +313,7 @@ docker:
- '--guidance_scale 4.0 \'
- '--num_iterations 1 \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- group: Qwen-Image
js_tag: qwen_image
models:
@@ -336,7 +336,7 @@ docker:
- '--use_torch_compile \'
- '--num_iterations 1 \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- model: Qwen-Image-Edit
model_repo: Qwen/Qwen-Image-Edit
url: https://huggingface.co/Qwen/Qwen-Image-Edit
@@ -358,4 +358,4 @@ docker:
- '--use_torch_compile \'
- '--num_iterations 1 \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
@@ -17,7 +17,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker-812:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.0_20250812-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.0-20250812.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -64,7 +64,7 @@ The following is summary of notable changes since the :doc:`previous ROCm/vLLM D
Supported models
================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.0_20250812-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.0-20250812.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -123,8 +123,6 @@ Supported models
vLLM is a toolkit and library for LLM inference and serving. AMD implements
high-performance custom kernels and modules in vLLM to enhance performance.
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
more information.
.. _vllm-benchmark-performance-measurements-812:
@@ -157,7 +155,7 @@ To test for optimal performance, consult the recommended :ref:`System health ben
<rocm-for-ai-system-health-bench>`. This suite of tests will help you verify and fine-tune your
system's configuration.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.0_20250812-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.0-20250812.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -296,21 +294,21 @@ system's configuration.
* -
- ``configs/extended.csv``
-
-
* -
- ``configs/performance.csv``
-
-
* - ``--benchmark``
- ``throughput``
- Measure offline end-to-end throughput.
* -
* -
- ``serving``
- Measure online serving performance.
* -
* -
- ``all``
- Measure both throughput and serving.
@@ -433,9 +431,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -16,7 +16,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker-909:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20250909-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.1-20250909.yaml
{% set docker = data.dockers[0] %}
@@ -57,7 +57,7 @@ The following is summary of notable changes since the :doc:`previous ROCm/vLLM D
Supported models
================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20250909-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.1-20250909.yaml
{% set docker = data.dockers[0] %}
{% set model_groups = data.model_groups %}
@@ -146,7 +146,7 @@ To test for optimal performance, consult the recommended :ref:`System health ben
<rocm-for-ai-system-health-bench>`. This suite of tests will help you verify and fine-tune your
system's configuration.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20250909-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.1-20250909.yaml
{% set docker = data.dockers[0] %}
{% set model_groups = data.model_groups %}
@@ -433,9 +433,6 @@ Further reading
- To learn more about system settings and management practices to configure your system for
AMD Instinct MI300X Series accelerators, see `AMD Instinct MI300X system optimization <https://instinct.docs.amd.com/projects/amdgpu-docs/en/latest/system-optimization/mi300x.html>`_.
- See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
a brief introduction to vLLM and optimization strategies.
- For application performance optimization strategies for HPC and AI workloads,
including inference with vLLM, see :doc:`/how-to/rocm-for-ai/inference-optimization/workload`.
@@ -16,7 +16,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker-930:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20251006-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.1-20251006.yaml
{% set docker = data.dockers[0] %}
@@ -75,7 +75,7 @@ The following is summary of notable changes since the :doc:`previous ROCm/vLLM D
Supported models
================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20251006-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.1-20251006.yaml
{% set docker = data.dockers[0] %}
{% set model_groups = data.model_groups %}
@@ -178,7 +178,7 @@ system's configuration.
Pull the Docker image
=====================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20251006-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.1-20251006.yaml
{% set docker = data.dockers[0] %}
@@ -192,7 +192,7 @@ Pull the Docker image
Benchmarking
============
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20251006-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.1-20251006.yaml
{% set docker = data.dockers[0] %}
{% set model_groups = data.model_groups %}
@@ -441,7 +441,7 @@ To reproduce this ROCm-enabled vLLM Docker image release, follow these steps:
2. Use the following command to build the image directly from the specified commit.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20251006-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.1-20251006.yaml
{% set docker = data.dockers[0] %}
.. code-block:: shell
@@ -467,9 +467,6 @@ Further reading
- To learn more about system settings and management practices to configure your system for
AMD Instinct MI300X Series GPUs, see `AMD Instinct MI300X system optimization <https://instinct.docs.amd.com/projects/amdgpu-docs/en/latest/system-optimization/mi300x.html>`_.
- See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
a brief introduction to vLLM and optimization strategies.
- For application performance optimization strategies for HPC and AI workloads,
including inference with vLLM, see :doc:`/how-to/rocm-for-ai/inference-optimization/workload`.
@@ -16,7 +16,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker-1103:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.11.1_20251103-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.11.1-20251103.yaml
{% set docker = data.dockers[0] %}
@@ -61,7 +61,7 @@ The following is summary of notable changes since the :doc:`previous ROCm/vLLM D
Supported models
================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.11.1_20251103-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.11.1-20251103.yaml
{% set docker = data.dockers[0] %}
{% set model_groups = data.model_groups %}
@@ -164,7 +164,7 @@ system's configuration.
Pull the Docker image
=====================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.11.1_20251103-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.11.1-20251103.yaml
{% set docker = data.dockers[0] %}
@@ -178,7 +178,7 @@ Pull the Docker image
Benchmarking
============
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.11.1_20251103-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.11.1-20251103.yaml
{% set docker = data.dockers[0] %}
{% set model_groups = data.model_groups %}
@@ -431,7 +431,7 @@ To reproduce this ROCm-enabled vLLM Docker image release, follow these steps:
2. Use the following command to build the image directly from the specified commit.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.11.1_20251103-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.11.1-20251103.yaml
{% set docker = data.dockers[0] %}
.. code-block:: shell
@@ -457,9 +457,6 @@ Further reading
- To learn more about system settings and management practices to configure your system for
AMD Instinct MI300X Series GPUs, see `AMD Instinct MI300X system optimization <https://instinct.docs.amd.com/projects/amdgpu-docs/en/latest/system-optimization/mi300x.html>`_.
- See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
a brief introduction to vLLM and optimization strategies.
- For application performance optimization strategies for HPC and AI workloads,
including inference with vLLM, see :doc:`/how-to/rocm-for-ai/inference-optimization/workload`.
@@ -44,9 +44,7 @@ optimizing performance with popular AI models.
consumption and increases throughput by leveraging dynamic key and value
allocation in GPU memory. vLLM also incorporates many LLM acceleration
and quantization algorithms. In addition, AMD implements high-performance
custom kernels and modules in vLLM to enhance performance further. See
:ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for more
information.
custom kernels and modules in vLLM to enhance performance further.
Getting started
===============
@@ -277,7 +275,7 @@ options and their descriptions.
Latency benchmark example
^^^^^^^^^^^^^^^^^^^^^^^^^
Use this command to benchmark the latency of the Llama 3.1 8B model on one GPU with the ``float16`` data type.
.. code-block::
@@ -334,9 +332,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -68,8 +68,6 @@ optimizing performance with popular AI models.
vLLM is a toolkit and library for LLM inference and serving. AMD implements
high-performance custom kernels and modules in vLLM to enhance performance.
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
more information.
Getting started
===============
@@ -342,7 +340,7 @@ options and their descriptions.
Example 1: latency benchmark
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
Use this command to benchmark the latency of the Llama 3.1 8B model on one GPU with the ``float16`` and ``float8`` data types.
.. code-block::
@@ -404,9 +402,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -41,8 +41,6 @@ and :ref:`standalone benchmarking <vllm-benchmark-standalone-v066-options>`.
vLLM is a toolkit and library for LLM inference and serving. AMD implements
high-performance custom kernels and modules in vLLM to enhance performance.
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
more information.
Getting started
===============
@@ -387,7 +385,7 @@ options and their descriptions.
Example 1: latency benchmark
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
Use this command to benchmark the latency of the Llama 3.1 70B model on eight GPUs with the ``float16`` and ``float8`` data types.
.. code-block::
@@ -449,9 +447,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -17,7 +17,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.7.3_20250325-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.7.3-20250325.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -93,8 +93,6 @@ vLLM inference performance testing
vLLM is a toolkit and library for LLM inference and serving. AMD implements
high-performance custom kernels and modules in vLLM to enhance performance.
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
more information.
.. _vllm-benchmark-performance-measurements-v073:
@@ -317,9 +315,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -12,7 +12,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.8.3_20250415-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.8.3-20250415.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -88,8 +88,6 @@ vLLM inference performance testing
vLLM is a toolkit and library for LLM inference and serving. AMD implements
high-performance custom kernels and modules in vLLM to enhance performance.
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
more information.
.. _vllm-benchmark-performance-measurements-v083:
@@ -333,9 +331,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -17,7 +17,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.8.5_20250513-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.8.5-20250513.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -97,8 +97,6 @@ vLLM inference performance testing
vLLM is a toolkit and library for LLM inference and serving. AMD implements
high-performance custom kernels and modules in vLLM to enhance performance.
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
more information.
.. _vllm-benchmark-performance-measurements-v085-20250513:
@@ -342,9 +340,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -17,7 +17,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.8.5_20250521-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.8.5-20250521.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -97,8 +97,6 @@ vLLM inference performance testing
vLLM is a toolkit and library for LLM inference and serving. AMD implements
high-performance custom kernels and modules in vLLM to enhance performance.
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
more information.
.. _vllm-benchmark-performance-measurements-v085-20250521:
@@ -342,9 +340,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -17,7 +17,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.9.0.1_20250605-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.9.0.1-20250605.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -97,8 +97,6 @@ vLLM inference performance testing
vLLM is a toolkit and library for LLM inference and serving. AMD implements
high-performance custom kernels and modules in vLLM to enhance performance.
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
more information.
.. _vllm-benchmark-performance-measurements-v0901-20250605:
@@ -341,9 +339,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -17,7 +17,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker-702:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.9.1_20250702-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.9.1-20250702.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -97,8 +97,6 @@ vLLM inference performance testing
vLLM is a toolkit and library for LLM inference and serving. AMD implements
high-performance custom kernels and modules in vLLM to enhance performance.
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
more information.
.. _vllm-benchmark-performance-measurements-20250702:
@@ -341,9 +339,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -17,7 +17,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker-715:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.9.1_20250715-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.9.1-20250715.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -70,7 +70,7 @@ The following is summary of notable changes since the :doc:`previous ROCm/vLLM D
Supported models
================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.9.1_20250715-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.9.1-20250715.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -129,8 +129,6 @@ Supported models
vLLM is a toolkit and library for LLM inference and serving. AMD implements
high-performance custom kernels and modules in vLLM to enhance performance.
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
more information.
.. _vllm-benchmark-performance-measurements-715:
@@ -163,7 +161,7 @@ To test for optimal performance, consult the recommended :ref:`System health ben
<rocm-for-ai-system-health-bench>`. This suite of tests will help you verify and fine-tune your
system's configuration.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.9.1_20250715-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.9.1-20250715.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -438,9 +436,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -6,13 +6,19 @@
prebuilt and optimized docker images.
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
************************
xDiT diffusion inference
************************
******************************
xDiT diffusion inference 25.10
******************************
.. caution::
This documentation does not reflect the latest version of the xDiT diffusion
inference performance documentation. See
:doc:`/ai-inference/archive/xdit-history` for the latest version.
.. _xdit-video-diffusion-2510:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.10-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.10.yaml
{% set docker = data.xdit_diffusion_inference.docker %}
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
@@ -59,7 +65,7 @@ The following models are supported for inference performance benchmarking.
Some instructions, commands, and recommendations in this documentation might
vary by model -- select one to get started.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.10-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.10.yaml
{% set docker = data.xdit_diffusion_inference.docker %}
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
@@ -122,7 +128,7 @@ guide to properly configure your system settings before starting.
Pull the Docker image
=====================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.10-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.10.yaml
{% set docker = data.xdit_diffusion_inference.docker %}
@@ -139,7 +145,7 @@ Validate and benchmark
Once the image has been downloaded you can follow these steps to
run benchmarks and generate outputs.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.10-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.10.yaml
{% set model_groups = data.xdit_diffusion_inference.model_groups %}
{% for model_group in model_groups %}
@@ -166,7 +172,7 @@ Prepare the model
You can either use an existing Hugging Face cache or download the model fresh inside the container.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.10-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.10.yaml
{% set docker = data.xdit_diffusion_inference.docker %}
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
@@ -264,7 +270,7 @@ Run inference
You can benchmark models through `MAD <https://github.com/ROCm/MAD>`__-integrated automation or standalone
torchrun commands.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.10-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.10.yaml
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
{% for model_group in model_groups %}
@@ -295,7 +301,7 @@ torchrun commands.
--tags {{model.mad_tag}} \
--keep-model-dir \
--live-output
MAD launches a Docker container with the name
``container_ci-{{model.mad_tag}}``. The throughput and serving reports of the
model are collected in the following paths: ``{{ model.mad_tag }}_throughput.csv``
@@ -395,5 +401,5 @@ Further reading
Previous versions
=================
See :doc:`xdit-history` to find documentation for previous releases
of xDiT diffusion inference performance testing.
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
releases of xDiT diffusion inference performance testing.
@@ -6,20 +6,19 @@
prebuilt and optimized docker images.
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
************************
xDiT diffusion inference
************************
******************************
xDiT diffusion inference 25.11
******************************
.. caution::
This documentation does not reflect the latest version of ROCm vLLM
This documentation does not reflect the latest version of the xDiT diffusion
inference performance documentation. See
:doc:`/how-to/rocm-for-ai/inference/xdit-diffusion-inference` for the latest
version.
:doc:`/ai-inference/archive/xdit-history` for the latest version.
.. _xdit-video-diffusion-2511:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
{% set docker = data.xdit_diffusion_inference.docker | selectattr("version", "equalto", "v25-11") | first %}
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
@@ -48,7 +47,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
What's new
==========
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
{% set docker = data.xdit_diffusion_inference.docker | selectattr("version", "equalto", "v25-11") | first %}
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
@@ -66,7 +65,7 @@ The following models are supported for inference performance benchmarking.
Some instructions, commands, and recommendations in this documentation might
vary by model -- select one to get started.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
{% set docker = data.xdit_diffusion_inference.docker | selectattr("version", "equalto", "v25-11") | first %}
{% set model_groups = data.xdit_diffusion_inference.model_groups %}
@@ -145,7 +144,7 @@ system's configuration.
Pull the Docker image
=====================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
{% set docker = data.xdit_diffusion_inference.docker | selectattr("version", "equalto", "v25-11") | first %}
@@ -162,7 +161,7 @@ Validate and benchmark
Once the image has been downloaded you can follow these steps to
run benchmarks and generate outputs.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
{% for model_group in model_groups %}
{% for model in model_group.models %}
@@ -180,7 +179,7 @@ Choose your setup method
You can either use an existing Hugging Face cache or download the model fresh inside the container.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
{% set docker = data.xdit_diffusion_inference.docker | selectattr("version", "equalto", "v25-11") | first %}
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
@@ -270,7 +269,7 @@ You can either use an existing Hugging Face cache or download the model fresh in
Run inference
=============
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
{% for model_group in model_groups %}
@@ -384,7 +383,5 @@ Run inference
Previous versions
=================
See
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
to find documentation for previous releases of xDiT diffusion inference
performance testing.
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
releases of xDiT diffusion inference performance testing.
@@ -6,20 +6,19 @@
prebuilt and optimized docker images.
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
************************
xDiT diffusion inference
************************
******************************
xDiT diffusion inference 25.12
******************************
.. caution::
This documentation does not reflect the latest version of xDiT diffusion
This documentation does not reflect the latest version of the xDiT diffusion
inference performance documentation. See
:doc:`/how-to/rocm-for-ai/inference/xdit-diffusion-inference` for the latest
version.
:doc:`/ai-inference/archive/xdit-history` for the latest version.
.. _xdit-video-diffusion-2512:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
{% set docker = data.docker %}
@@ -51,7 +50,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
What's new
==========
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
{% set docker = data.docker %}
@@ -68,7 +67,7 @@ The following models are supported for inference performance benchmarking.
Some instructions, commands, and recommendations in this documentation might
vary by model -- select one to get started.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
{% set docker = data.docker %}
@@ -133,7 +132,7 @@ system's configuration.
Pull the Docker image
=====================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
{% set docker = data.docker %}
@@ -147,7 +146,7 @@ Pull the Docker image
Validate and benchmark
======================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
{% set docker = data.docker %}
@@ -170,7 +169,7 @@ Choose your setup method
You can either use an existing Hugging Face cache or download the model fresh inside the container.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
{% set docker = data.docker %}
@@ -262,7 +261,7 @@ You can either use an existing Hugging Face cache or download the model fresh in
Run inference
=============
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
{% set docker = data.docker %}
@@ -406,7 +405,5 @@ Run inference
Previous versions
=================
See
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
to find documentation for previous releases of xDiT diffusion inference
performance testing.
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
releases of xDiT diffusion inference performance testing.
@@ -6,20 +6,19 @@
prebuilt and optimized docker images.
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
************************
xDiT diffusion inference
************************
******************************
xDiT diffusion inference 25.13
******************************
.. caution::
This documentation does not reflect the latest version of the xDiT diffusion
inference performance documentation. See
:doc:`/how-to/rocm-for-ai/inference/xdit-diffusion-inference` for the latest
version.
:doc:`/ai-inference/archive/xdit-history` for the latest version.
.. _xdit-video-diffusion-2513:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
{% set docker = data.docker %}
@@ -52,7 +51,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
What's new
==========
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
{% set docker = data.docker %}
@@ -69,7 +68,7 @@ The following models are supported for inference performance benchmarking.
Some instructions, commands, and recommendations in this documentation might
vary by model -- select one to get started.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
{% set docker = data.docker %}
@@ -150,7 +149,7 @@ system's configuration.
Pull the Docker image
=====================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
{% set docker = data.docker %}
@@ -164,7 +163,7 @@ Pull the Docker image
Validate and benchmark
======================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
{% set docker = data.docker %}
@@ -187,7 +186,7 @@ Choose your setup method
You can either use an existing Hugging Face cache or download the model fresh inside the container.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
{% set docker = data.docker %}
@@ -279,7 +278,7 @@ You can either use an existing Hugging Face cache or download the model fresh in
Run inference
=============
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
{% set docker = data.docker %}
@@ -469,7 +468,5 @@ Run inference
Previous versions
=================
See
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
to find documentation for previous releases of xDiT diffusion inference
performance testing.
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
releases of xDiT diffusion inference performance testing.
@@ -6,20 +6,19 @@
prebuilt and optimized docker images.
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
************************
xDiT diffusion inference
************************
*****************************
xDiT diffusion inference 26.1
*****************************
.. caution::
This documentation does not reflect the latest version of the xDiT diffusion
inference performance documentation. See
:doc:`/how-to/rocm-for-ai/inference/xdit-diffusion-inference` for the latest
version.
:doc:`/ai-inference/archive/xdit-history` for the latest version.
.. _xdit-video-diffusion-v261-v261:
.. _xdit-video-diffusion-v261:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
{% set docker = data.docker %}
@@ -47,7 +46,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
What's new
==========
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
{% set docker = data.docker %}
@@ -64,7 +63,7 @@ The following models are supported for inference performance benchmarking.
Some instructions, commands, and recommendations in this documentation might
vary by model -- select one to get started.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
{% set docker = data.docker %}
@@ -129,7 +128,7 @@ system's configuration.
Pull the Docker image
=====================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
{% set docker = data.docker %}
@@ -143,7 +142,7 @@ Pull the Docker image
Validate and benchmark
======================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
{% set docker = data.docker %}
@@ -166,7 +165,7 @@ Choose your setup method
You can either use an existing Hugging Face cache or download the model fresh inside the container.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
{% set docker = data.docker %}
@@ -258,7 +257,7 @@ You can either use an existing Hugging Face cache or download the model fresh in
Run inference
=============
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
{% set docker = data.docker %}
@@ -317,7 +316,5 @@ Run inference
Previous versions
=================
See
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
to find documentation for previous releases of xDiT diffusion inference
performance testing.
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
releases of xDiT diffusion inference performance testing.
@@ -14,12 +14,11 @@ xDiT diffusion inference
This documentation does not reflect the latest version of the xDiT diffusion
inference performance documentation. See
:doc:`/how-to/rocm-for-ai/inference/xdit-diffusion-inference` for the latest
version.
:doc:`/ai-inference/archive/xdit-history` for the latest version.
.. _xdit-video-diffusion-262:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
{% set docker = data.docker %}
@@ -47,7 +46,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
What's new
==========
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
{% set docker = data.docker %}
@@ -64,7 +63,7 @@ The following models are supported for inference performance benchmarking.
Some instructions, commands, and recommendations in this documentation might
vary by model -- select one to get started.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
{% set docker = data.docker %}
@@ -129,7 +128,7 @@ system's configuration.
Pull the Docker image
=====================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
{% set docker = data.docker %}
@@ -143,7 +142,7 @@ Pull the Docker image
Validate and benchmark
======================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
{% set docker = data.docker %}
@@ -166,7 +165,7 @@ Choose your setup method
You can either use an existing Hugging Face cache or download the model fresh inside the container.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
{% set docker = data.docker %}
@@ -258,7 +257,7 @@ You can either use an existing Hugging Face cache or download the model fresh in
Run inference
=============
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
{% set docker = data.docker %}
@@ -315,7 +314,5 @@ Run inference
Previous versions
=================
See
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
to find documentation for previous releases of xDiT diffusion inference
performance testing.
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
releases of xDiT diffusion inference performance testing.
@@ -6,20 +6,19 @@
prebuilt and optimized docker images.
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
************************
xDiT diffusion inference
************************
*****************************
xDiT diffusion inference 26.3
*****************************
.. caution::
This documentation does not reflect the latest version of the xDiT diffusion
inference performance documentation. See
:doc:`/how-to/rocm-for-ai/inference/xdit-diffusion-inference` for the latest
version.
:doc:`/ai-inference/archive/xdit-history` for the latest version.
.. _xdit-video-diffusion-263:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
{% set docker = data.docker %}
@@ -47,7 +46,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
What's new
==========
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
{% set docker = data.docker %}
@@ -64,7 +63,7 @@ The following models are supported for inference performance benchmarking.
Some instructions, commands, and recommendations in this documentation might
vary by model -- select one to get started.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
{% set docker = data.docker %}
@@ -129,7 +128,7 @@ system's configuration.
Pull the Docker image
=====================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
{% set docker = data.docker %}
@@ -143,7 +142,7 @@ Pull the Docker image
Validate and benchmark
======================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
{% set docker = data.docker %}
@@ -166,7 +165,7 @@ Choose your setup method
You can either use an existing Hugging Face cache or download the model fresh inside the container.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
{% set docker = data.docker %}
@@ -258,7 +257,7 @@ You can either use an existing Hugging Face cache or download the model fresh in
Run inference
=============
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
{% set docker = data.docker %}
@@ -315,7 +314,5 @@ Run inference
Previous versions
=================
See
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
to find documentation for previous releases of xDiT diffusion inference
performance testing.
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
releases of xDiT diffusion inference performance testing.
+314
View File
@@ -0,0 +1,314 @@
:orphan:
:no-search:
:selector-toc2: Model
:selector-toc2-icon: fa-solid fa-robot
.. meta::
:description: Learn to validate diffusion model video generation on MI300X, MI350X and MI355X accelerators using
prebuilt and optimized docker images.
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
*****************************
xDiT diffusion inference 26.4
*****************************
.. caution::
This documentation does not reflect the latest version of the xDiT diffusion
inference performance documentation. See
:doc:`/ai-inference/archive/xdit-history` for the latest version.
.. _xdit-video-diffusion-264:
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
{% set docker = data.docker %}
The `rocm/pytorch-xdit <{{ docker.docker_hub_url }}>`_ Docker image offers a prebuilt, optimized environment based on `xDiT <https://github.com/xdit-project/xDiT>`_ for
benchmarking diffusion model video and image generation on gfx942 and gfx950 series (AMD Instinct™ MI300X, MI325X, MI350X, and MI355X) GPUs.
The image runs `ROCm {{docker.ROCm}} (preview) <https://rocm.docs.amd.com/en/7.12.0-preview/about/release-notes.html>`__ based on `TheRock <https://github.com/ROCm/TheRock>`_
and includes the following components:
.. dropdown:: Software components - {{ docker.pull_tag.split('-')|last }}
.. list-table::
:header-rows: 1
* - Software component
- Version
{% for component_name, component_data in docker.components.items() %}
* - `{{ component_name }} <{{ component_data.url }}>`_
- {{ component_data.version }}
{% endfor %}
Follow this guide to pull the required image, spin up a container, download the model, and run a benchmark.
For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.docker.com/r/amdsiloai/pytorch-xdit>`_.
What's new
==========
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
{% set docker = data.docker %}
{% for item in docker.whats_new %}
* {{ item }}
{% endfor %}
.. _xdit-video-diffusion-supported-models-264:
Supported models
================
The following models are supported for inference performance benchmarking.
Some instructions, commands, and recommendations in this documentation might
vary by model -- select one to get started.
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
{% set docker = data.docker %}
.. selector:: Model
:key: model-group
{% for model_group in docker.supported_models %}
.. selector-option:: {{ model_group.group }}
:value: {{ model_group.js_tag }}
:width: 25%
{% endfor %}
{% for model_group in docker.supported_models %}
.. selector:: Variant
:key: model
:show-cond: model-group={{ model_group.js_tag }}
{% set models = model_group.models %}
{% for model in models %}
.. selector-option:: {{ model.model }}
:value: {{ model.js_tag }}
{% endfor %}
{% endfor %}
{% for model_group in docker.supported_models %}
{% for model in model_group.models %}
.. selected:: model={{ model.js_tag }}
.. note::
To learn more about your specific model see the `{{ model.model }} model card on Hugging Face <{{ model.url }}>`_
or visit the `GitHub page <{{ model.github }}>`__. Note that some models require access authorization before use via an
external license agreement through a third party.
{% endfor %}
{% endfor %}
System validation
=================
Before running AI workloads, it's important to validate that your AMD hardware is configured
correctly and performing optimally.
If you have already validated your system settings, including aspects like NUMA auto-balancing, you
can skip this step. Otherwise, complete the procedures in the :ref:`System validation and
optimization <rocm-for-ai-system-optimization>` guide to properly configure your system settings
before starting.
To test for optimal performance, consult the recommended :ref:`System health benchmarks
<rocm-for-ai-system-health-bench>`. This suite of tests will help you verify and fine-tune your
system's configuration.
Pull the Docker image
=====================
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
{% set docker = data.docker %}
For this tutorial, it's recommended to use the latest ``{{ docker.pull_tag }}`` Docker image.
Pull the image using the following command:
.. code-block:: shell
docker pull {{ docker.pull_tag }}
Validate and benchmark
======================
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
{% set docker = data.docker %}
Once the image has been downloaded you can follow these steps to
run benchmarks and generate outputs.
{% for model_group in docker.supported_models %}
{% for model in model_group.models %}
.. selected:: model={{ model.js_tag }}
The following commands are written for {{ model.model }}.
See :ref:`xdit-video-diffusion-supported-models-264` to switch to another available model.
{% endfor %}
{% endfor %}
Choose your setup method
------------------------
You can either use an existing Hugging Face cache or download the model fresh inside the container.
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
{% set docker = data.docker %}
{% for model_group in docker.supported_models %}
{% for model in model_group.models %}
.. selected:: model={{model.js_tag}}
.. tab-set::
.. tab-item:: Option 1: Use existing Hugging Face cache
If you already have models downloaded on your host system, you can mount your existing cache.
1. Set your Hugging Face cache location.
.. code-block:: shell
export HF_HOME=/your/hf_cache/location
2. Download the model (if not already cached).
.. code-block:: shell
hf download {{ model.model_repo }} {% if model.revision %} --revision {{ model.revision }} {% endif %}
3. Launch the container with mounted cache.
.. code-block:: shell
docker run \
-it --rm \
--cap-add=SYS_PTRACE \
--security-opt seccomp=unconfined \
--user root \
--device=/dev/kfd \
--device=/dev/dri \
--group-add video \
--ipc=host \
--network host \
--privileged \
--shm-size 128G \
--name pytorch-xdit \
-e HSA_NO_SCRATCH_RECLAIM=1 \
-e OMP_NUM_THREADS=16 \
-e CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
-e HF_HOME=/app/huggingface_models \
-v $HF_HOME:/app/huggingface_models \
{{ docker.pull_tag }}
.. tab-item:: Option 2: Download inside container
If you prefer to keep the container self-contained or don't have an existing cache.
1. Launch the container
.. code-block:: shell
docker run \
-it --rm \
--cap-add=SYS_PTRACE \
--security-opt seccomp=unconfined \
--user root \
--device=/dev/kfd \
--device=/dev/dri \
--group-add video \
--ipc=host \
--network host \
--privileged \
--shm-size 128G \
--name pytorch-xdit \
-e HSA_NO_SCRATCH_RECLAIM=1 \
-e OMP_NUM_THREADS=16 \
-e CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
{{ docker.pull_tag }}
2. Inside the container, set the Hugging Face cache location and download the model.
.. code-block:: shell
export HF_HOME=/app/huggingface_models
hf download {{ model.model_repo }} {% if model.revision %} --revision {{ model.revision }} {% endif %}
.. warning::
Models will be downloaded to the container's filesystem and will be lost when the container is removed unless you persist the data with a volume.
{% endfor %}
{% endfor %}
Run inference
=============
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
{% set docker = data.docker %}
{% for model_group in docker.supported_models %}
{% for model in model_group.models %}
.. selected:: model={{ model.js_tag }}
.. tab-set::
.. tab-item:: MAD-integrated benchmarking
1. Clone the ROCm Model Automation and Dashboarding (`<https://github.com/ROCm/MAD>`__) repository to a local
directory and install the required packages on the host machine.
.. code-block:: shell
git clone https://github.com/ROCm/MAD
cd MAD
pip install -r requirements.txt
2. On the host machine, use this command to run the performance benchmark test on
the `{{model.model}} <{{ model.url }}>`_ model using one node.
.. code-block:: shell
export MAD_SECRETS_HFTOKEN="your personal Hugging Face token to access gated models"
madengine run \
--tags {{model.mad_tag}} \
--keep-model-dir \
--live-output
MAD launches a Docker container with the name
``container_ci-{{model.mad_tag}}``. The throughput and serving reports of the
model are collected in the following paths: ``{{ model.mad_tag }}_throughput.csv``
and ``{{ model.mad_tag }}_serving.csv``.
.. tab-item:: Standalone benchmarking
To run the benchmarks for {{ model.model }}, use the following command:
.. code-block:: shell
{{ model.benchmark_command
| map('replace', '{model_repo}', model.model_repo)
| map('trim')
| join('\n ') }}
The generated content and timing information will be stored under the results directory.
{% endfor %}
{% endfor %}
Previous versions
=================
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
releases of xDiT diffusion inference performance testing.
@@ -1,8 +1,8 @@
:orphan:
************************************************************
xDiT diffusion inference performance testing version history
************************************************************
****************************************
xDiT diffusion inference version history
****************************************
This table lists previous versions of the ROCm xDiT diffusion inference performance
testing environment. For detailed information about available models for
@@ -15,12 +15,20 @@ benchmarking, see the version-specific documentation.
- Components
- Resources
* - ``rocm/pytorch-xdit:v26.4`` (latest)
* - ``rocm/pytorch-xdit:v26.5`` (latest)
-
* ROCm 7.13.0
* TheRock cbff3d1
-
* :doc:`Documentation <../xdit>`
* `Docker Hub <https://hub.docker.com/layers/rocm/pytorch-xdit/v26.5/images/sha256-b8ad9fd4b41bc116ac2aff07c1066bf369cf7fc110b1a323f6302191985a51fd>`__
* - ``rocm/pytorch-xdit:v26.4``
-
* `ROCm 7.12.0 preview <https://rocm.docs.amd.com/en/7.12.0-preview/about/release-notes.html>`__
* TheRock 9b611c6
-
* :doc:`Documentation </how-to/rocm-for-ai/inference/xdit-diffusion-inference>`
* :doc:`Documentation <xdit-26.4>`
* `Docker Hub <https://hub.docker.com/layers/rocm/pytorch-xdit/v26.4/images/sha256-b4296a638eb8dc7ebcafc808e180b78a3c44177580c21986082ec9539496067c>`__
* - ``rocm/pytorch-xdit:v26.3``
+124
View File
@@ -0,0 +1,124 @@
********************************
ComfyUI image generation on ROCm
********************************
`ComfyUI <https://github.com/comfyanonymous/ComfyUI>`__ is an open-source,
node-based interface for building and running image generation workflows with
diffusion models such as Stable Diffusion. Its modular graph-based design lets
you construct, customize, and share complex pipelines without writing code. This
page walks through installing and running ComfyUI on AMD GPUs.
Prerequisites
=============
Ensure your working environment is running ROCm-enabled PyTorch on
a :ref:`supported system <compat-matrix>`. See :ref:`pytorch-install` for
instructions.
.. important::
On Windows, ComfyUI might not start if Smart App Control is enabled in your
Windows security settings.
Installation
============
After installing ROCm and PyTorch in your Python environment, follow these
steps to install ComfyUI.
1. Clone the ComfyUI repository.
.. code-block:: shell
git clone https://github.com/comfyanonymous/ComfyUI.git
2. Activate your Python virtual environment and install dependencies.
.. tab-set::
.. tab-item:: Linux
:sync: linux
.. code-block:: bash
pip install -r ComfyUI/requirements.txt
.. tab-item:: Windows
:sync: windows
.. code-block:: bat
pip install -r ComfyUI\requirements.txt
Run ComfyUI
===========
Use the following steps for a simple example of running ComfyUI.
1. Start the ComfyUI server from the command line.
.. tab-set::
.. tab-item:: Linux
:sync: linux
.. code-block:: bash
python ComfyUI/main.py
.. tab-item:: Windows
:sync: windows
.. code-block:: bat
python ComfyUI\main.py
This starts the server, displaying a prompt like:
.. code-block:: text
To see the GUI go to: http://127.0.0.1:8188
2. Go to ``http://127.0.0.1:8188`` in your web browser. You might need to
replace ``8188`` with the appropriate port.
.. image:: ./images/comfyui/comfyui-main.png
:align: center
3. Search for one of the following templates and download any missing
models.
.. tab-set::
.. tab-item:: SD3.5 Simple
Select **Template****Model Filter****SD3.5****SD3.5 Simple**
.. image:: ./images/comfyui/sd3_5-simple-card.png
:align: center
Download required models, if missing.
.. image:: ./images/comfyui/sd3_5-missing-models.png
:align: center
.. tab-item:: Chroma1 Radiance text to image
Select **Template****Model Filter****Chroma****Chroma1 Radiance text to image**
.. image:: ./images/comfyui/chroma1-radiance-tti-card.png
:align: center
Download required models, if missing.
.. image:: ./images/comfyui/chroma1-radiance-tti-missing-models.png
:align: center
4. Click the **Run** button.
The application will use your AMD GPU to convert the prompted text to an image.
.. seealso::
To learn more about the ComfyUI interface and workflows, see the `ComfyUI
documentation <https://docs.comfy.org/development/core-concepts/workflow>`__.

Before

Width:  |  Height:  |  Size: 44 KiB

After

Width:  |  Height:  |  Size: 44 KiB

Before

Width:  |  Height:  |  Size: 28 KiB

After

Width:  |  Height:  |  Size: 28 KiB

Before

Width:  |  Height:  |  Size: 112 KiB

After

Width:  |  Height:  |  Size: 112 KiB

Before

Width:  |  Height:  |  Size: 188 KiB

After

Width:  |  Height:  |  Size: 188 KiB

Before

Width:  |  Height:  |  Size: 129 KiB

After

Width:  |  Height:  |  Size: 129 KiB

Before

Width:  |  Height:  |  Size: 80 KiB

After

Width:  |  Height:  |  Size: 80 KiB

Before

Width:  |  Height:  |  Size: 153 KiB

After

Width:  |  Height:  |  Size: 153 KiB

Before

Width:  |  Height:  |  Size: 219 KiB

After

Width:  |  Height:  |  Size: 219 KiB

Before

Width:  |  Height:  |  Size: 310 KiB

After

Width:  |  Height:  |  Size: 310 KiB

Before

Width:  |  Height:  |  Size: 342 KiB

After

Width:  |  Height:  |  Size: 342 KiB

@@ -2,35 +2,35 @@
:description: How to Use ROCm for AI inference optimization
:keywords: ROCm, LLM, AI inference, Optimization, GPUs, usage, tutorial
*******************************************
Use ROCm for AI inference optimization
*******************************************
**********************
Inference optimization
**********************
AI inference optimization is the process of improving the performance of machine learning models and speeding up the inference process. It includes:
- **Quantization**: This involves reducing the precision of model weights and activations while maintaining acceptable accuracy levels. Reduced precision improves inference efficiency because lower precision data requires less storage and better utilizes the hardware's computation power.
- **Quantization**: This involves reducing the precision of model weights and activations while maintaining acceptable accuracy levels. Reduced precision improves inference efficiency because lower precision data requires less storage and better utilizes the hardware's computation power.
- **Kernel optimization**: This technique involves optimizing computation kernels to exploit the underlying hardware capabilities. For example, the kernels can be optimized to use multiple GPU cores or utilize specialized hardware like tensor cores to accelerate the computations.
- **Kernel optimization**: This technique involves optimizing computation kernels to exploit the underlying hardware capabilities. For example, the kernels can be optimized to use multiple GPU cores or utilize specialized hardware like tensor cores to accelerate the computations.
- **Libraries**: Libraries such as Flash Attention, xFormers, and PyTorch TunableOp are used to accelerate deep learning models and improve the performance of inference workloads.
- **Libraries**: Libraries such as Flash Attention, xFormers, and PyTorch TunableOp are used to accelerate deep learning models and improve the performance of inference workloads.
- **Hardware acceleration**: Hardware acceleration techniques, like GPUs for AI inference, can significantly improve performance due to their parallel processing capabilities.
- **Hardware acceleration**: Hardware acceleration techniques, like GPUs for AI inference, can significantly improve performance due to their parallel processing capabilities.
- **Pruning**: This involves removing unnecessary connections, layers, or weights from a pre-trained model while maintaining acceptable accuracy levels, resulting in a smaller model that requires fewer computational resources to run inference.
- **Pruning**: This involves removing unnecessary connections, layers, or weights from a pre-trained model while maintaining acceptable accuracy levels, resulting in a smaller model that requires fewer computational resources to run inference.
Utilizing these optimization techniques with the ROCm™ software platform can significantly reduce inference time, improve performance, and reduce the cost of your AI applications.
Utilizing these optimization techniques with the ROCm™ software platform can significantly reduce inference time, improve performance, and reduce the cost of your AI applications.
Throughout the following topics, this guide discusses optimization techniques for inference workloads.
- :doc:`Model quantization <model-quantization>`
- :doc:`Model acceleration libraries <model-acceleration-libraries>`
- :doc:`Model acceleration libraries <model-acceleration-libs>`
- :doc:`Optimizing with Composable Kernel <optimizing-with-composable-kernel>`
- :doc:`Optimizing with Composable Kernel <optimize-with-composable-kernel>`
- :doc:`Optimizing Triton kernels <optimizing-triton-kernel>`
- :doc:`Optimizing Triton kernels <optimize-triton-kernels>`
- :doc:`Profiling and debugging <profiling-and-debugging>`
- :doc:`Workload tuning <workload-optimization>`
- :doc:`Workload tuning <workload>`
- :ref:`Profiling and debugging <mi300x-profiling-tools>`
@@ -7,9 +7,7 @@ LLM inference frameworks
************************
This section discusses how to implement `vLLM <https://docs.vllm.ai/en/latest>`_ and `Hugging Face TGI
<https://huggingface.co/docs/text-generation-inference/en/index>`_ using
:doc:`single-accelerator <../fine-tuning/single-gpu-fine-tuning-and-inference>` and
:doc:`multi-accelerator <../fine-tuning/multi-gpu-fine-tuning-and-inference>` systems.
<https://huggingface.co/docs/text-generation-inference/en/index>`_.
.. _fine-tuning-llms-vllm:
@@ -68,7 +66,7 @@ Installing vLLM
The following log message is displayed in your command line indicates that the server is listening for requests.
.. image:: ../../../data/how-to/llm-fine-tuning-optimization/vllm-single-gpu-log.png
.. image:: ./images/llm-inference-frameworks/vllm-single-gpu-log.png
:alt: vLLM API server log message
:align: center
@@ -141,7 +139,7 @@ Installing vLLM
ROCm provides a prebuilt optimized Docker image for validating the performance of LLM inference with vLLM
on the MI300X GPU. The Docker image includes ROCm, vLLM, and PyTorch.
For more information, see :doc:`/how-to/rocm-for-ai/inference/benchmark-docker/vllm`.
For more information, see :doc:`/ai-inference/vllm`.
.. _fine-tuning-llms-tgi:
@@ -20,7 +20,7 @@ Attention (GQA), and Multi-Query Attention (MQA). This reduction in memory movem
time-to-first-token (TTFT) latency for large batch sizes and long prompt sequences, thereby enhancing overall
performance.
.. image:: ../../../data/how-to/llm-fine-tuning-optimization/attention-module.png
.. image:: ./images/model-acceleration-libs/attention-module.png
:alt: Attention module of a large language module utilizing tiling
:align: center
@@ -36,7 +36,7 @@ These can be installed by following the official
`PyTorch installation guide <https://pytorch.org/get-started/locally/>`_. Alternatively, for a simpler setup, you can use a preconfigured
:ref:`ROCm PyTorch Docker image <using-docker-with-pytorch-pre-installed>`, which already includes the required libraries.
Installing Flash Attention 2
Installing Flash Attention 2
----------------------------
`Flash Attention <https://github.com/Dao-AILab/flash-attention>`_ supports two backend implementations on AMD GPUs.
@@ -61,29 +61,29 @@ To install Flash Attention 2, use the following commands:
pip install ninja
# To install the CK backend flash attention
python setup.py install
python setup.py install
# To install the Triton backend flash attention
FLASH_ATTENTION_TRITON_AMD_ENABLE="TRUE" python setup.py install
FLASH_ATTENTION_TRITON_AMD_ENABLE="TRUE" python setup.py install
# To install both CK and Triton backend flash attention
FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE && FLASH_ATTENTION_SKIP_CK_BUILD=FALSE python setup.py install
For detailed installation instructions, see `Flash Attention <https://github.com/Dao-AILab/flash-attention>`_.
Benchmarking Flash Attention 2
Benchmarking Flash Attention 2
------------------------------
Benchmark scripts to evaluate the performance of Flash Attention 2 are stored in the ``flash-attention/benchmarks/`` directory.
To benchmark the CK backend
To benchmark the CK backend
.. code-block:: shell
cd flash-attention/benchmarks
pip install transformers einops ninja
python3 benchmark_flash_attention.py
python3 benchmark_flash_attention.py
To benchmark the Triton backend
@@ -91,7 +91,7 @@ To benchmark the Triton backend
FLASH_ATTENTION_TRITON_AMD_ENABLE="TRUE" python3 benchmark_flash_attention.py
Using Flash Attention 2
Using Flash Attention 2
-----------------------
.. code-block:: python
@@ -128,13 +128,13 @@ xFormers also improves the performance of attention modules. Although xFormers a
similarly to Flash Attention 2 due to its tiling behavior of query, key, and value, its widely used for LLMs and
Stable Diffusion models with the Hugging Face Diffusers library.
Installing CK xFormers
Installing CK xFormers
----------------------
Use the following commands to install CK xFormers.
.. code-block:: shell
# Install from source
git clone https://github.com/ROCm/xformers.git
cd xformers/
@@ -175,20 +175,20 @@ of the PyTorch compilation.
os.environ["TOKENIZERS_PARALLELISM"] = "false"
model_name = "NousResearch/Meta-Llama-3-8B"
prompts = []
for b in range(1):
prompts.append("New york city is where "
)
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.float16).to(device).eval()
inputs = tokenizer(prompts, return_tensors="pt").to(model.device)
def decode_one_tokens(model, cur_token, input_pos, cache_position):
logits = model(cur_token, position_ids=input_pos, cache_position=cache_position, return_dict=False, use_cache=True)[0]
new_token = torch.argmax(logits[:, -1], dim=-1)[:, None]
return new_token
batch_size, seq_length = inputs["input_ids"].shape
# Static key-value cache
@@ -198,16 +198,16 @@ of the PyTorch compilation.
cache_position = torch.arange(seq_length, device=device)
generated_ids = torch.zeros(batch_size, seq_length + max_new_tokens + 1, dtype=torch.int, device=device)
generated_ids[:, cache_position] = inputs["input_ids"].to(device).to(torch.int)
logits = model(**inputs, cache_position=cache_position, return_dict=False, use_cache=True)[0]
next_token = torch.argmax(logits[:, -1], dim=-1)[:, None]
# torch compilation
decode_one_tokens = torch.compile(decode_one_tokens, mode="max-autotune-no-cudagraphs",fullgraph=True)
generated_ids[:, seq_length] = next_token[:, 0]
cache_position = torch.tensor([seq_length + 1], device=device)
with torch.no_grad():
for _ in range(1, max_new_tokens):
with torch.backends.cuda.sdp_kernel(enable_flash=False, enable_mem_efficient=False, enable_math=True):
@@ -235,7 +235,7 @@ page describes the options.
# To turn on TunableOp, simply set this environment variable
export PYTORCH_TUNABLEOP_ENABLED=1
# Python
import torch
import torch.nn as nn
@@ -244,7 +244,7 @@ page describes the options.
W = torch.rand(200, 20, device="cuda")
Out = F.linear(A, W)
print(Out.size())
# tunableop_results0.csv
Validator,PT_VERSION,2.4.0
Validator,ROCM_VERSION,6.1.0.0-82-5fabb4c
@@ -253,7 +253,7 @@ page describes the options.
Validator,ROCBLAS_VERSION,4.1.0-cefa4a9b-dirty
GemmTunableOp_float_TN,tn_200_100_20,Gemm_Rocblas_32323,0.00669595
.. image:: ../../../data/how-to/llm-fine-tuning-optimization/tunableop.png
.. image:: ./images/model-acceleration-libs/tunableop.png
:alt: GEMM and TunableOp
:align: center
@@ -270,7 +270,7 @@ and as a back end for PyTorch quantized operators. FBGEMM offers optimized on-CP
strong performance on native tensor formats, and the ability to generate
high-performance shape- and size-specific kernels at runtime.
FBGEMM_GPU collects several high-performance PyTorch GPU operator libraries
FBGEMM_GPU collects several high-performance PyTorch GPU operator libraries
for use in training and inference. It provides efficient table-batched embedding functionality,
data layout transformation, and quantization support.
@@ -288,7 +288,7 @@ Installing FBGEMM_GPU consists of the following steps:
* Install ROCm using Docker or the :doc:`package manager <rocm-install-on-linux:install/install-methods/package-manager-index>`
* Install the nightly `PyTorch <https://pytorch.org/>`_ build
* Complete the pre-build and build tasks
.. note::
FBGEMM_GPU doesn't require the installation of FBGEMM. To optionally install
@@ -375,7 +375,7 @@ and run the ROCm Docker image, use this command:
You can also install ROCm using the package manager. FBGEMM_GPU requires the installation of the full ROCm package.
For more information, see :doc:`the ROCm installation guide <rocm-install-on-linux:install/detailed-install>`.
The ROCm package also requires the :doc:`MIOpen <miopen:index>` component as a dependency.
The ROCm package also requires the :doc:`MIOpen <miopen:index>` component as a dependency.
To install MIOpen, use the ``apt install`` command.
.. code-block:: shell
@@ -407,7 +407,7 @@ Install `PyTorch <https://pytorch.org/>`_ using ``pip`` for the most reliable an
Perform the prebuild and build
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
#. Clone the FBGEMM repository and the relevant submodules. Use ``pip`` to install the
#. Clone the FBGEMM repository and the relevant submodules. Use ``pip`` to install the
components in ``requirements.txt``. Run the following commands inside the Miniconda environment.
.. code-block:: shell
@@ -447,7 +447,7 @@ Perform the prebuild and build
# Set the Python platform name for the Linux case
export python_plat_name="manylinux2014_${ARCH}"
#. Build FBGEMM_GPU for the ROCm platform. Set ``ROCM_PATH`` to the path to your ROCm installation.
#. Build FBGEMM_GPU for the ROCm platform. Set ``ROCM_PATH`` to the path to your ROCm installation.
Run these commands from the ``fbgemm_gpu/`` directory inside the Miniconda environment.
.. code-block:: shell
@@ -474,7 +474,7 @@ Perform the prebuild and build
--package_variant=rocm \
-DHIP_ROOT_DIR="${ROCM_PATH}" \
-DCMAKE_C_FLAGS="-DTORCH_USE_HIP_DSA" \
-DCMAKE_CXX_FLAGS="-DTORCH_USE_HIP_DSA"
-DCMAKE_CXX_FLAGS="-DTORCH_USE_HIP_DSA"
Post-build validation
----------------------
@@ -533,8 +533,8 @@ follow these instructions:
# Run the test
python -m pytest -v -rsx -s -W ignore::pytest.PytestCollectionWarning split_table_batched_embeddings_test.py
To run the FBGEMM_GPU ``uvm`` test, use these commands. These tests only support the AMD MI210 and
more recent GPUs.
To run the FBGEMM_GPU ``uvm`` test, use these commands. These tests only support the AMD MI210 and
more recent GPUs.
.. code-block:: shell
@@ -6,14 +6,14 @@
Optimizing Triton kernels
*************************
This section introduces the general steps for
This section introduces the general steps for
`Triton <https://openai.com/index/triton/>`_ kernel optimization. Broadly,
Triton kernel optimization is similar to :doc:`HIP <hip:how-to/performance_guidelines>`
and CUDA kernel optimization.
Refer to the
:ref:`Triton kernel performance optimization <mi300x-triton-kernel-performance-optimization>`
section of the :doc:`workload` guide
section of the :doc:`workload-optimization` guide
for detailed information.
Triton kernel performance optimization includes the following topics.
@@ -29,11 +29,11 @@ The template parameters of the instance are grouped into four parameter types:
- [Parameters for determining extra operations on matrix elements](matrix-element-operation)
- [Performance-oriented tunable parameters](tunable-parameters)
<!--
<!--
================
### Figure 2
================ -->
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-template_parameters.jpg
```{figure} ./images/optimize-with-composable-kernel/ck-template_parameters.jpg
The template parameters of the selected GEMM kernel are classified into four groups. These template parameter groups should be defined properly before running the instance.
```
@@ -100,7 +100,7 @@ struct AddRelu
(tunable-parameters)=
#### Tunable parameters
#### Tunable parameters
The CK instance includes a series of tunable template parameters to control the parallel granularity of the workload to achieve load balancing on different hardware platforms.
@@ -123,11 +123,11 @@ After determining the template parameters, we instantiate the kernel with actual
The row and column, and stride information of input matrices are also passed to the instance. For batched GEMM, you must pass in additional batch count and batch stride values. The extra operations for pre and post-processing are also passed with an actual argument; for example, α and β for GEMM scaling operations. Afterward, the instantiated kernel is launched by the invoker, as illustrated in Figure 3.
<!--
<!--
================
### Figure 3
================ -->
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-kernel_launch.jpg
```{figure} ./images/optimize-with-composable-kernel/ck-kernel_launch.jpg
Templated kernel launching consists of kernel instantiation, making arguments by passing in actual application parameters, creating an invoker, and running the instance through the invoker.
```
@@ -152,11 +152,11 @@ The following section discusses the analysis of the operation flow of `Linear_Re
The first operation in the process is to perform the multiplication of input matrices A and B. The resulting matrix C is then scaled with α to obtain T1. At the same time, the process performs a scaling operation on D elements to obtain T2. Afterward, the process performs matrix addition between T1 and T2, element activation calculation using ReLU, and element rounding sequentially. The operations to generate E1, E2, and E are encapsulated and completed by a user-defined template function in CK (given in the next sub-section). This template function is integrated into the fundamental instance directly during the compilation phase so that all these steps can be fused in a single GPU kernel.
<!--
<!--
================
### Figure 4
================ -->
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-operation_flow.jpg
```{figure} ./images/optimize-with-composable-kernel/ck-operation_flow.jpg
Operation flow.
```
@@ -168,11 +168,11 @@ Third, consider the platform for implementing CK instances. The instances suffix
Here, we use [DeviceBatchedGemmMultiD_Xdl](https://github.com/ROCm/composable_kernel/tree/develop/example/24_batched_gemm) as the fundamental instance to implement the functionalities in the previous table.
<!--
<!--
================
### Figure 5
================ -->
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-root_instance.jpg
```{figure} ./images/optimize-with-composable-kernel/ck-root_instance.jpg
Use the DeviceBatchedGemmMultiD_Xdl instance as a root.
```
@@ -189,7 +189,7 @@ The inference of SQ quantized models relies on using PyTorch and Transformer lib
In GEMM, the A and B inputs are two-dimensional matrices, and the required input matrices of the selected fundamental CK instance are three-dimensional matrices. Therefore, we must convert the input 2-D tensors to 3-D tensors, by using `tensor`'s `unsqueeze()` method before passing these matrices to the instance. For batched GEMM in the preceding table, ignore this step.
```c++
// Function input and output
// Function input and output
torch::Tensor linear_relu_abde_i8(
torch::Tensor A_,
torch::Tensor B_,
@@ -197,10 +197,10 @@ torch::Tensor linear_relu_abde_i8(
float alpha,
float beta)
{
// Convert torch::Tensor A_ (M, K) to torch::Tensor A (1, M, K)
// Convert torch::Tensor A_ (M, K) to torch::Tensor A (1, M, K)
auto A = A_.unsqueeze(0);
// Convert torch::Tensor B_ (K, N) to torch::Tensor A (1, K, N)
// Convert torch::Tensor B_ (K, N) to torch::Tensor A (1, K, N)
auto B = B_.unsqueeze(0);
...
```
@@ -232,7 +232,7 @@ As shown in the following code block, we obtain M, N, and K values using input t
auto D = D_.view({1,-1}).repeat({M, 1});
// Allocate memory for E
auto E = torch::empty({batch_count, M, N},
auto E = torch::empty({batch_count, M, N},
torch::dtype(torch::kInt8).device(A.device()));
```
@@ -241,7 +241,7 @@ In the following code block, `ADataType`, `BDataType` and `D0DataType` are used
`AccDataType` determines the data precision used to represent the multiply-add results of A and B elements. Generally, a larger range data type is applied to store the multiply-add results of A and B to avoid result overflow; `I32` is applied in this case. The `CShuffleDataType I32` data type indicates that the multiply-add results continue to be stored in LDS as an `I32` data format. All of this is implemented through the following code block.
```c++
// Data precision
// Data precision
using ADataType = I8;
using BDataType = I8;
using AccDataType = I32;
@@ -265,7 +265,7 @@ Following the convention of various linear algebra libraries, row-major and colu
In CK, `PassThrough` is a struct denoting if an operation is applied to the tensor it binds to. To fuse the operations between E1, E2, and E introduced in section [Operation flow analysis](#operation-flow-analysis), we define a custom C++ struct, `ScaleScaleAddRelu`, and bind it to `CDEELementOp`. It determines the operations that will be applied to `CShuffle` (A×B results), tensor D, α, and β.
```c++
// No operations bound to the elements of A and B
// No operations bound to the elements of A and B
using AElementOp = PassThrough;
using BElementOp = PassThrough;
@@ -290,17 +290,17 @@ struct ScaleScaleAddRelu {
// Perform addition operation
F32 temp = c_scale + d_scale;
// Perform RELU operation
temp = temp > 0 ? temp : 0;
// Perform rounding operation
// Perform rounding operation
temp = temp > 127 ? 127 : temp;
// Return to E
e = ck::type_convert<I8>(temp);
}
F32 alpha;
F32 beta;
};
@@ -315,16 +315,16 @@ static constexpr auto GemmDefault = ck::tensor_operation::device::GemmSpecializa
The template parameters of the target fundamental instance are initialized with the above parameters and includes default tunable parameters. For specific tuning methods, see [Tunable parameters](#tunable-parameters).
```c++
using DeviceOpInstance = ck::tensor_operation::device::DeviceBatchedGemmMultiD_Xdl<
using DeviceOpInstance = ck::tensor_operation::device::DeviceBatchedGemmMultiD_Xdl<
// Tensor layout
ALayout, BLayout, DsLayout, ELayout,
ALayout, BLayout, DsLayout, ELayout,
// Tensor data type
ADataType, BDataType, AccDataType, CShuffleDataType, DsDataType, EDataType,
ADataType, BDataType, AccDataType, CShuffleDataType, DsDataType, EDataType,
// Tensor operation
AElementOp, BElementOp, CDEElementOp,
// Padding strategy
AElementOp, BElementOp, CDEElementOp,
// Padding strategy
GemmDefault,
// Tunable parameters
// Tunable parameters
tunable parameters>;
```
@@ -356,13 +356,13 @@ invoker.Run(argument, StreamConfig{nullptr, 0});
The output of the fundamental instance is a calculated batched matrix E (batch, M, N). Before the return, it needs to be converted to a 2-D matrix if a normal GEMM result is required.
```c++
// Convert (1, M, N) to (M, N)
// Convert (1, M, N) to (M, N)
return E.squeeze(0);
```
### Binding to Python
Since these functions are written in C++ and `torch::Tensor`, you can use `pybind11` to bind the functions and import them as Python modules. For the example, the necessary binding code for exposing the functions in the table spans but a few lines.
Since these functions are written in C++ and `torch::Tensor`, you can use `pybind11` to bind the functions and import them as Python modules. For the example, the necessary binding code for exposing the functions in the table spans but a few lines.
```c++
#include <torch/extension.h>
@@ -390,7 +390,7 @@ os.environ["CXX"] = "hipcc"
sources = [
'torch_int/kernels/linear.cpp',
'torch_int/kernels/bmm.cpp',
'torch_int/kernels/pybind.cpp',
'torch_int/kernels/pybind.cpp',
]
include_dirs = ['torch_int/kernels/include']
@@ -418,11 +418,11 @@ setup(
Run `python setup.py install` to build and install the extension. It should look something like Figure 6:
<!--
<!--
================
### Figure 6
================ -->
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-compilation.jpg
```{figure} ./images/optimize-with-composable-kernel/ck-compilation.jpg
Compilation and installation of the INT8 kernels.
```
@@ -430,11 +430,11 @@ Compilation and installation of the INT8 kernels.
The implementation architecture of running SmoothQuant models on MI300X GPUs is illustrated in Figure 7, where (a) shows the decoder layer composition components of the target model, (b) shows the major implementation class for the decoder layer components, and \(c\) denotes the underlying GPU kernels implemented by CK instance.
<!--
<!--
================
### Figure 7
================ -->
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-inference_flow.jpg
```{figure} ./images/optimize-with-composable-kernel/ck-inference_flow.jpg
The implementation architecture of running SmoothQuant models on AMD MI300X GPUs.
```
@@ -456,11 +456,11 @@ Note that since the default values were used for the tunable parameters of the f
Figure 8 shows the performance comparisons between the original FP16 and the SmoothQuant-quantized INT8 models on a single MI300X GPU. The GPU memory footprints of SmoothQuant-quantized models are significantly reduced. It also indicates the per-sample inference latency is significantly reduced for all SmoothQuant-quantized OPT models (illustrated in (b)). Notably, the performance of the CK instance-based INT8 kernel steadily improves with an increase in model size.
<!--
<!--
================
### Figure 8
================ -->
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-comparisons.jpg
```{figure} ./images/optimize-with-composable-kernel/ck-comparisons.jpg
Performance comparisons between the original FP16 and the SmoothQuant-quantized INT8 models on a single MI300X GPU.
```
@@ -31,8 +31,7 @@ The following variables are generally useful for Instinct MI300X/MI355X GPUs and
* ``export HIP_FORCE_DEV_KERNARG=1`` — improves kernel launch performance by
forcing device kernel arguments. This is already set by default in
:doc:`vLLM ROCm Docker images
</how-to/rocm-for-ai/inference/benchmark-docker/vllm>`. Bare-metal users
:doc:`vLLM ROCm Docker images </ai-inference/vllm>`. Bare-metal users
should set this manually.
* ``export TORCH_BLAS_PREFER_HIPBLASLT=1`` — explicitly prefers hipBLASLt
over hipBLAS for GEMM operations. By default, PyTorch uses heuristics to
@@ -128,8 +127,7 @@ Quick start examples:
``VLLM_ROCM_USE_AITER=1`` to enable all optimizations. ROCm provides a
prebuilt optimized Docker image for validating the performance of LLM
inference with vLLM on MI300X Series GPUs. The Docker image includes ROCm,
vLLM, and PyTorch. For more information, see
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/vllm`.
vLLM, and PyTorch. For more information, see :doc:`/ai-inference/vllm`.
.. _vllm-optimization-aiter-moe-requirements:
@@ -373,7 +371,7 @@ vLLM V1 on ROCm provides these attention implementations:
* Automatically selected when ``VLLM_ROCM_USE_AITER=1`` and model is not MLA
5. **vLLM Triton Multi-head Latent Attention (MLA)** (for DeepSeek-V3/R1/V2)
* Automatically selected when ``VLLM_ROCM_USE_AITER=0`` (or unset)
6. **AITER Multi-head Latent Attention (MLA)** (for DeepSeek-V3/R1/V2)
@@ -652,8 +650,8 @@ balancing KV-cache capacity.
.. code-block:: bash
for i in $(seq 0 7); do
CUDA_VISIBLE_DEVICES="$i" vllm bench throughput
-tp 1 --model /path/to/model
CUDA_VISIBLE_DEVICES="$i" vllm bench throughput
-tp 1 --model /path/to/model
--dataset /path/to/ShareGPT_V3_unfiltered_cleaned_split.json &
done
@@ -801,15 +799,15 @@ CUDA graphs reduce kernel launch overhead by capturing and replaying GPU operati
Quantization support
====================
vLLM supports FP4/FP8 (4-bit/8-bit floating point) weight and activation quantization using hardware acceleration on the Instinct MI300X and MI355X.
Quantization of models with FP4/FP8 allows for a **2x-4x** reduction in model memory requirements and up to a **1.6x**
improvement in throughput with minimal impact on accuracy.
vLLM supports FP4/FP8 (4-bit/8-bit floating point) weight and activation quantization using hardware acceleration on the Instinct MI300X and MI355X.
Quantization of models with FP4/FP8 allows for a **2x-4x** reduction in model memory requirements and up to a **1.6x**
improvement in throughput with minimal impact on accuracy.
vLLM ROCm supports a variety of quantization demands:
vLLM ROCm supports a variety of quantization demands:
* On-the-fly quantization
* On-the-fly quantization
* Pre-quantized model through Quark and llm-compressor
* Pre-quantized model through Quark and llm-compressor
Supported quantization methods
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
@@ -982,7 +980,7 @@ Many AWQ-quantized models are available on Hugging Face. Use them directly with
vllm serve hugging-quants/Meta-Llama-3.1-70B-Instruct-AWQ-INT4 \
--quantization awq \
--tensor-parallel-size 1 \
--dtype auto
--dtype auto
**Important Notes:**
@@ -1100,7 +1098,7 @@ Speculative decoding (experimental)
===================================
Recent vLLM versions add support for speculative decoding backends (for example, Eaglev3). Evaluate for your model and latency/throughput goals.
Speculative decoding is a technique to reduce latency when max number of concurrency is low.
Speculative decoding is a technique to reduce latency when max number of concurrency is low.
Depending on the methods, the effective concurrency varies, for example, from 16 to 64.
Example command:
@@ -1128,7 +1126,7 @@ Example command:
It has been observed that more ``num_speculative_tokens`` causes less
acceptance rate of draft model tokens and a decline in throughput. As a
workaround, set ``num_speculative_tokens`` to <= 2.
workaround, set ``num_speculative_tokens`` to <= 2.
Multi-node checklist and troubleshooting
@@ -1142,5 +1140,5 @@ Multi-node checklist and troubleshooting
Further reading
===============
* :doc:`workload`
* :doc:`/how-to/rocm-for-ai/inference/benchmark-docker/vllm`
* :doc:`workload-optimization`
* :doc:`/ai-inference/vllm`
@@ -4,6 +4,8 @@
environment variable, performance, HIP, Triton, PyTorch TunableOp, vLLM, RCCL,
MIOpen, GPU, resource utilization
.. _mi300x-workload-optimization:
*****************************************
AMD Instinct MI300X workload optimization
*****************************************
@@ -66,7 +68,7 @@ When profiling indicates that GPUs are a performance bottleneck, delve deeper
into kernel-level profiling. Tools such as the
:ref:`ROCr Debug Agent <mi300x-rocr-debug-agent>`,
:ref:`ROCProfiler <mi300x-rocprof>`, and
:ref:`ROCm Compute Profiler <mi300x-rocprof-compute>` offer detailed insights
:ref:`ROCm Compute Profiler (rocprofiler-compute) <mi300x-rocprof-compute>` offer detailed insights
into GPU kernel execution. These tools can help isolate problematic GPU
operations and provide data needed for targeted optimizations.
@@ -99,7 +101,7 @@ execution.
.. seealso::
See :doc:`vllm-optimization` to learn more about vLLM performance
See :doc:`vllm-v1-optimization` to learn more about vLLM performance
optimization techniques.
.. _mi300x-auto-tune:
@@ -151,29 +153,30 @@ address any new bottlenecks that may emerge.
ROCm provides a prebuilt optimized Docker image that has everything required to implement
the LLM inference tips in this section. It includes ROCm, PyTorch, and vLLM.
For more information, see :doc:`/how-to/rocm-for-ai/inference/benchmark-docker/vllm`.
For more information, see :doc:`/ai-inference/vllm`.
.. _mi300x-profiling-tools:
Profiling tools
===============
AMD profiling tools provide valuable insights into how efficiently your
application utilizes hardware and help diagnose potential bottlenecks that
contribute to poor performance. Developers targeting AMD GPUs have multiple
tools available depending on their specific profiling needs.
AMD profiling tools help you understand how efficiently an application uses
hardware resources and help identify bottlenecks that contribute to poor
performance. Developers targeting AMD GPUs can choose from multiple tools,
depending on the type of analysis they need.
* ROCProfiler tool collects kernel execution performance
metrics. For more information, see the
:doc:`ROCProfiler <rocprofiler:index>`
documentation.
* :doc:`ROCprofiler-SDK <rocprofiler-sdk:index>` provides low-level APIs and
CLI utilities for collecting GPU hardware performance counters, runtime
traces, and user annotations.
* ROCm Compute Profiler builds upon ROCProfiler but provides more guided analysis.
For more information, see
:doc:`ROCm Compute Profiler documentation <rocprofiler-compute:index>`.
* :doc:`ROCm Compute Profiler <rocprofiler-compute:index>` builds on
ROCprofiler-SDK to provide guided, kernel-centric analysis of GPU
performance data.
Refer to :doc:`profiling-and-debugging`
to explore commonly used profiling tools and their usage patterns.
* :doc:`ROCm Systems Profiler <rocprofiler-systems:index>` provides
application- and system-level profiling for CPU and GPU workloads. For ROCm
GPU tracing and profiling data, it uses ROCprofiler-SDK and combines that
data with CPU and system metrics collected from additional sources.
Once performance bottlenecks are identified, you can implement an informed workload
tuning strategy. If kernels are the bottleneck, consider:
@@ -204,6 +207,12 @@ collect CPU and GPU performance metrics while the script is running. See the `Py
You can then visualize and view these metrics using an open-source profile visualization tool like
`Perfetto UI <https://ui.perfetto.dev>`_.
.. note::
In PyTorch on ROCm, the profiler and device APIs continue to use the
``torch.cuda`` namespace for compatibility. As a result,
``ProfilerActivity.CUDA`` is expected on ROCm systems.
#. Use the following snippet to invoke PyTorch Profiler in your code.
.. code-block:: python
@@ -211,6 +220,7 @@ You can then visualize and view these metrics using an open-source profile visua
import torch
import torchvision.models as models
from torch.profiler import profile, record_function, ProfilerActivity
model = models.resnet18().cuda()
inputs = torch.randn(2000, 3, 224, 224).cuda()
@@ -222,9 +232,9 @@ You can then visualize and view these metrics using an open-source profile visua
#. Profile results in ``resnet18_profile.json`` can be viewed by the Perfetto visualization tool. Go to
`<https://ui.perfetto.dev>`__ and import the file. In your Perfetto visualization, you'll see that the upper section
shows transactions denoting the CPU activities that launch GPU kernels while the lower section shows the actual GPU
activities where it processes the ``resnet18`` inferences layer by layer.
activities where it processes the ``resnet18`` inferences layer by layer.
.. figure:: ../../../data/how-to/tuning-guides/perfetto-trace.svg
.. figure:: ./images/workload-optimization/perfetto-trace.svg
:width: 800
Perfetto trace visualization example.
@@ -232,88 +242,98 @@ You can then visualize and view these metrics using an open-source profile visua
ROCm profiling tools
--------------------
Heterogenous systems, where programs run on both CPUs and GPUs, introduce additional complexities. Understanding the
critical path and kernel execution is all the more important. So, performance tuning is a necessary component in the
benchmarking process.
Heterogeneous systems, where programs execute on both CPUs and GPUs,
introduce additional complexity. Understanding the critical path, runtime
behavior, and kernel execution is essential for performance tuning and
benchmarking.
With AMD's profiling tools, developers are able to gain important insight into how efficiently their application is
using hardware resources and effectively diagnose potential bottlenecks contributing to poor performance. Developers
working with AMD Instinct GPUs have multiple tools depending on their specific profiling needs; these include:
AMD provides multiple profiling tools for AMD Instinct GPU workloads,
including:
* :ref:`ROCProfiler <mi300x-rocprof>`
* :ref:`ROCm Compute Profiler <mi300x-rocprof-compute>`
* :ref:`ROCm Compute Profiler (rocprofiler-compute) <mi300x-rocprof-compute>`
* :ref:`ROCm Systems Profiler <mi300x-rocprof-systems>`
* :ref:`ROCm Systems Profiler (rocprofiler-systems) <mi300x-rocprof-systems>`
.. _mi300x-rocprof:
.. _mi300x-rocprofiler-sdk:
ROCProfiler
^^^^^^^^^^^
ROCprofiler-SDK
^^^^^^^^^^^^^^^
:doc:`ROCProfiler <rocprofiler:index>` is primarily a low-level API for accessing and extracting GPU hardware performance
metrics, commonly called *performance counters*. These counters quantify the performance of the underlying architecture
showcasing which pieces of the computational pipeline and memory hierarchy are being utilized.
:doc:`ROCprofiler-SDK <rocprofiler-sdk:index>` is the low-level profiling and
tracing toolkit for AMD GPUs. It provides APIs and CLI utilities for
accessing GPU hardware performance metrics, commonly called *performance
counters*, as well as runtime traces and user annotations.
Your ROCm installation contains a script or executable command called ``rocprof`` which provides the ability to list all
available hardware counters for your specific GPU, and run applications while collecting counters during
their execution.
This ``rocprof`` utility also depends on the :doc:`ROCTracer and ROC-TX libraries <roctracer:index>`, giving it the
ability to collect timeline traces of the GPU software stack as well as user-annotated code regions.
Depending on your ROCm version, your installation may include the
ROCprofiler-SDK-based ``rocprofv3`` CLI.
.. note::
``rocprof`` is a CLI-only utility where inputs and outputs take the form of text and CSV files. These
formats provide a raw view of the data and puts the onus on the user to parse and analyze. ``rocprof``
gives the user full access and control of raw performance profiling data, but requires extra effort to analyze the
collected data.
ROCprofiler-SDK tooling is intentionally low level. Depending on the
collection mode, it can emit raw text, CSV, or trace files. These formats
provide detailed profiling data, but typically require additional parsing,
visualization, or analysis by the user.
.. _mi300x-rocprof-compute:
ROCm Compute Profiler
^^^^^^^^^^^^^^^^^^^^^
:doc:`ROCm Compute Profiler <rocprofiler-compute:index>` is a system performance profiler for high-performance computing (HPC) and
machine learning (ML) workloads using Instinct GPUs. Under the hood, ROCm Compute Profiler uses
:ref:`ROCProfiler <mi300x-rocprof>` to collect hardware performance counters. The ROCm Compute Profiler tool performs
system profiling based on all approved hardware counters for Instinct
GPU architectures. It provides high level performance analysis features including System Speed-of-Light, IP
block Speed-of-Light, Memory Chart Analysis, Roofline Analysis, Baseline Comparisons, and more.
:doc:`ROCm Compute Profiler <rocprofiler-compute:index>` is a GPU performance
analysis tool for high-performance computing (HPC) and machine learning (ML)
workloads running on AMD Instinct GPUs. Under the hood, ROCm Compute Profiler
uses :ref:`ROCprofiler-SDK <mi300x-rocprofiler-sdk>` to collect hardware
performance counters.
ROCm Compute Profiler takes the guesswork out of profiling by removing the need to provide text input files with lists of counters
to collect and analyze raw CSV output files as is the case with ROCProfiler. Instead, ROCm Compute Profiler automates the collection
of all available hardware counters in one command and provides graphical interfaces to help users understand and
analyze bottlenecks and stressors for their computational workloads on AMD Instinct GPUs.
ROCm Compute Profiler performs guided kernel- and GPU-level analysis using
approved hardware counters for AMD Instinct GPU architectures. It provides
high-level analysis features including System Speed-of-Light, IP block
Speed-of-Light, Memory Chart Analysis, Roofline Analysis, Baseline
Comparisons, and more.
ROCm Compute Profiler removes much of the manual work required when using
ROCprofiler-SDK directly. Instead of requiring users to specify counter lists
and analyze raw output files, it automates multi-pass counter collection and
provides graphical and command-line interfaces to help identify bottlenecks in
computational workloads.
.. note::
ROCm Compute Profiler collects hardware counters in multiple passes, and will therefore re-run the application during each pass
to collect different sets of metrics.
ROCm Compute Profiler collects hardware counters in multiple passes and
therefore re-runs the application during each pass to collect different
sets of metrics.
.. figure:: ../../../data/how-to/tuning-guides/rocprof-compute-analysis.png
.. figure:: ./images/workload-optimization/rocprof-compute-analysis.png
:width: 800
ROCm Compute Profiler memory chart analysis panel.
In brief, ROCm Compute Profiler provides details about hardware activity for a particular GPU kernel. It also supports both
a web-based GUI or command-line analyzer, depending on your preference.
In brief, ROCm Compute Profiler provides detailed insight into the hardware
behavior of individual GPU kernels. It supports both a web-based GUI and a
command-line analyzer.
.. _mi300x-rocprof-systems:
ROCm Systems Profiler
^^^^^^^^^^^^^^^^^^^^^
:doc:`ROCm Systems Profiler <rocprofiler-systems:index>` is a comprehensive profiling and tracing tool for parallel applications,
including HPC and ML packages, written in C, C++, Fortran, HIP, OpenCL, and Python which execute on the CPU or CPU and
GPU. It is capable of gathering the performance information of functions through any combination of binary
instrumentation, call-stack sampling, user-defined regions, and Python interpreter hooks.
:doc:`ROCm Systems Profiler <rocprofiler-systems:index>` is a comprehensive
profiling and tracing tool for parallel applications written in C, C++,
Fortran, HIP, OpenCL, and Python which execute on the CPU or CPU and GPU. It is
capable of gathering the performance information of functions through any
combination of binary instrumentation, call-stack sampling, user-defined
regions, and Python interpreter hooks.
ROCm Systems Profiler supports interactive visualization of comprehensive traces in the web browser in addition to high-level
ROCm Systems Profiler supports interactive visualization of comprehensive
traces in `Perfetto <https://perfetto.dev/>`__ in addition to high-level
summary profiles with ``mean/min/max/stddev`` statistics. Beyond runtime
information, ROCm Systems Profiler supports the collection of system-level metrics such as CPU frequency, GPU temperature, and GPU
utilization. Process and thread level metrics such as memory usage, page faults, context switches, and numerous other
hardware counters are also included.
information, ROCm Systems Profiler supports the collection of system-level
metrics such as CPU frequency, GPU temperature, and GPU utilization. Process
and thread level metrics such as memory usage, page faults, context switches,
and numerous other hardware counters are also included.
.. tip::
@@ -322,7 +342,7 @@ hardware counters are also included.
have the greatest impact on the end-to-end execution of the application and to discover what else is happening on the
system during a performance bottleneck.
.. figure:: ../../../data/how-to/tuning-guides/rocprof-systems-timeline.png
.. figure:: ./images/workload-optimization/rocprof-systems-timeline.png
:width: 800
ROCm Systems Profiler timeline trace example.
@@ -332,7 +352,7 @@ vLLM performance optimization
vLLM is a high-throughput and memory efficient inference and serving engine for
large language models that has gained traction in the AI community for its
performance and ease of use. See :doc:`vllm-optimization`, where you'll learn
performance and ease of use. See :doc:`vllm-v1-optimization`, where you'll learn
how to:
* Enable AITER (AI Tensor Engine for ROCm) to speed up on LLM models.
@@ -356,9 +376,8 @@ In short, it will try up to thousands of matrix multiply algorithms that are ava
A caveat is that as the math libraries improve over time, there is a less benefit to using TunableOp,
and there is also no guarantee that the workload being tuned will be able to outperform the default GEMM algorithm in hipBLASLt.
Some additional references for PyTorch TunableOp include `ROCm blog <https://rocm.blogs.amd.com/artificial-intelligence/pytorch-tunableop/README.html>`__,
TunableOp `README <https://github.com/pytorch/pytorch/blob/main/aten/src/ATen/cuda/tunable/README.md>`__, and
`llm tuning <https://rocm.docs.amd.com/en/latest/how-to/llm-fine-tuning-optimization/model-acceleration-libraries.html#fine-tuning-llms-pytorch-tunableop>`__.
Some additional references for PyTorch TunableOp include `ROCm blog <https://rocm.blogs.amd.com/artificial-intelligence/pytorch-tunableop/README.html>`__ and
TunableOp `README <https://github.com/pytorch/pytorch/blob/main/aten/src/ATen/cuda/tunable/README.md>`__.
The three most important environment variables for controlling TunableOp are:
@@ -383,7 +402,7 @@ Use these environment variables to enable TunableOp for any applications or libr
The first step is the tuning pass:
1. Enable TunableOp and tuning. Optionally enable verbose mode:
1. Enable TunableOp and tuning. Optionally enable verbose mode:
.. code-block:: shell
@@ -466,8 +485,8 @@ shape. See the configurations in PyTorch source code:
* `matmul configurations for "max-autotune" <https://github.com/pytorch/pytorch/blob/a1d02b423c6b4ccacd25ebe86de43f650463bbc6/torch/_inductor/kernel/mm_common.py#L118>`_
This tuning will select the best Triton ``gemm`` configurations according to tile-size
``(BLOCK_M, BLOCK_N, BLOCK_K), num_stages, num_warps`` and ``mfma`` instruction size ( ``matrix_instr_nonkdim`` )
This tuning will select the best Triton ``gemm`` configurations according to tile-size
``(BLOCK_M, BLOCK_N, BLOCK_K), num_stages, num_warps`` and ``mfma`` instruction size ( ``matrix_instr_nonkdim`` )
(see "Triton kernel optimization" section for more details).
* Set ``torch._inductor.config.max_autotune = True`` or ``TORCHINDUCTOR_MAX_AUTOTUNE=1``.
@@ -482,7 +501,7 @@ This tuning will select the best Triton ``gemm`` configurations according to til
``torch._inductor.max_autotune_gemm_backends`` or ``TORCHINDUCTOR_MAX_AUTOTUNE_GEMM_BACKENDS``
Selects the candidate backends for ``mm`` auto-tuning. Defaults to
``TRITON,ATEN``.
``TRITON,ATEN``.
Limiting this to ``TRITON`` might improve performance by
enabling more fused ``mm`` kernels instead of going to rocBLAS.
@@ -496,10 +515,10 @@ This tuning will select the best Triton ``gemm`` configurations according to til
``torch._inductor.config.cpp_wrapper=True`` or ``TORCHINDUCTOR_CPP_WRAPPER=1``
* Convolution workloads might see a performance benefit by specifying
* Convolution workloads might see a performance benefit by specifying
``torch._inductor.config.layout_optimization=True`` or ``TORCHINDUCTOR_LAYOUT_OPTIMIZATION=1``.
This can help performance by enforcing ``channel_last`` memory format on the
convolution in TorchInductor, avoiding any unnecessary transpose operations.
convolution in TorchInductor, avoiding any unnecessary transpose operations.
Note that ``PYTORCH_MIOPEN_SUGGEST_NHWC=1`` is recommended if using this.
* To extract the Triton kernels generated by ``inductor``, set the environment variable
@@ -557,7 +576,7 @@ ROCm library tuning involves optimizing the performance of routine computational
operations (such as ``GEMM``) provided by ROCm libraries like
:ref:`hipBLASLt <mi300x-hipblaslt>`, :ref:`Composable Kernel <mi300x-ck>`,
:ref:`MIOpen <mi300x-miopen>`, and :ref:`RCCL <mi300x-rccl>`. This tuning aims
to maximize efficiency and throughput on Instinct MI300X GPUs to gain
to maximize efficiency and throughput on Instinct MI300X GPUs to gain
improved application performance.
.. _mi300x-library-gemm:
@@ -633,7 +652,7 @@ Create a working folder for the auto-tuning tool, for example, ``tuning/``.
1. Set the ``ProblemType``, ``TestConfig``, and ``TuningParameters`` in the YAML file. You can modify the template YAML file in ``hipblaslt/utilities``.
.. figure:: ../../../data/how-to/tuning-guides/hipblaslt_yaml_template.png
.. figure:: ./images/workload-optimization/hipblaslt_yaml_template.png
:align: center
:alt: HipBLASLt auto-tuning yaml file template
@@ -649,12 +668,12 @@ Create a working folder for the auto-tuning tool, for example, ``tuning/``.
Output
''''''
The tool will create two output folders. The first one is the benchmark results,
the second one is the generated equality kernels. If ``SplitK`` is used, the solution's ``GlobalSplitU`` will
also change if the winner is using a different ``SplitK`` from the solution. The YAML files generated inside the
The tool will create two output folders. The first one is the benchmark results,
the second one is the generated equality kernels. If ``SplitK`` is used, the solution's ``GlobalSplitU`` will
also change if the winner is using a different ``SplitK`` from the solution. The YAML files generated inside the
folder ``1_LogicYaml`` are logic ones. These YAML files are just like those generated from TensileLite.
.. figure:: ../../../data/how-to/tuning-guides/hipblaslt_auto_tuning_output_files.png
.. figure:: ./images/workload-optimization/hipblaslt_auto_tuning_output_files.png
:align: center
:alt: HipBLASLt auto-tuning output folder
@@ -698,9 +717,9 @@ The tuning tool is a two-step tool. It first runs the benchmark, then it creates
3. ``AlgoMethod``: We recommended to keep this unchanged because method "all" returns all the available solutions for the problem type.
4. ``ApiMethod``: We have c, mix, and cpp. Doesn't affect the result much.
5. ``RotatingBuffer``: This is a size in the unit of MB. Recommended to set the value equal to the size of the cache of the card to avoid the kernel fetching data from the cache.
* ``TuningParameters``
``SplitK``: Divide ``K`` into ``N`` portions. Not every solution supports ``SplitK``.
``SplitK``: Divide ``K`` into ``N`` portions. Not every solution supports ``SplitK``.
The solution will be skipped if not supported.
* ``CreateLogic``
@@ -783,7 +802,7 @@ of problem sizes.
.. _tensilelite-tuning-flow-fig:
.. figure:: ../../../data/how-to/tuning-guides/tensilelite-tuning-flow.png
.. figure:: ./images/workload-optimization/tensilelite-tuning-flow.png
:align: center
:alt: TensileLite tuning flow
@@ -818,7 +837,7 @@ Step 3: Fork parameters
Rather than continuing to determine globally fastest parameters, which eventually leads
to a single fastest kernel, forking creates many different kernels,
all of which will be considered for use. All forked
parameters are considered determined, i.e., they aren't measured to determine
parameters are considered determined, i.e., they aren't measured to determine
which is fastest. The :ref:`preceding figure <tensilelite-tuning-flow-fig>` shows 7 kernels being forked in Step 3.
Step 4: Benchmark fork parameters
@@ -870,7 +889,7 @@ command:
.. code-block:: shell
merge.py original_dir new_tuned_yaml_dir output_dir
merge.py original_dir new_tuned_yaml_dir output_dir
The following table describes the logic YAML files.
@@ -942,10 +961,10 @@ Additionally, ``bf16`` matrix core execution is noticeably faster than ``f16``.
Distributing workgroups with data sharing on the same XCD can enhance
performance (reduce latency) and improve benchmarking stability.
CK provides a rich set of template parameters for generating flexible accelerated
CK provides a rich set of template parameters for generating flexible accelerated
computing kernels for difference application scenarios.
See :doc:`optimizing-with-composable-kernel`
See :doc:`optimize-with-composable-kernel`
for an overview of Composable Kernel GEMM kernels, information on tunable
parameters, and examples.
@@ -988,7 +1007,7 @@ Tuning in MIOpen
What does :doc:`PerfDb <miopen:conceptual/perfdb>` look like?
.. code-block::
.. code-block::
[
2x128x56xNHWCxF, [
@@ -1021,7 +1040,7 @@ Finding the fastest kernel
What does :doc:`FindDb <miopen:conceptual/finddb>` look like?
.. code-block::
.. code-block::
[
@@ -1089,7 +1108,7 @@ case of 2- or 4-GPU collective operations (generally less than 8 GPUs),
you can only use a fraction of the potential bandwidth on the node.
The following figure shows an
:doc:`MI300X node-level architecture </conceptual/gpu-arch/mi300>` of a
:doc:`MI300X node-level architecture </reference/gpu-arch/mi300>` of a
system with AMD EPYC processors in a dual-socket configuration and eight
AMD Instinct MI300X GPUs. The MI300X OAMs attach to the host system via
PCIe Gen 5 x16 links (yellow lines). The GPUs use seven high-bandwidth,
@@ -1098,7 +1117,7 @@ low-latency AMD Infinity Fabric™ links (red lines) to form a fully connected
.. _mi300x-node-level-arch-fig:
.. figure:: ../../../data/shared/mi300-node-level-arch.png
.. figure:: /images/shared/mi300-node-level-arch.png
MI300 Series node-level architecture showing 8 fully interconnected MI300X
OAM modules connected to (optional) PCIe switches via re-timers and HGX
@@ -1301,7 +1320,7 @@ Register access is the fastest yet smallest among the three.
.. _mi300x-cu-fig:
.. figure:: ../../../data/shared/compute-unit.png
.. figure:: /images/shared/compute-unit.png
Schematic representation of a CU in the CDNA2 or CDNA3 architecture.
@@ -1341,7 +1360,7 @@ efficiency and throughput of various computational kernels.
.. _mi300x-occupancy-vgpr-table:
.. figure:: ../../../data/shared/occupancy-vgpr.png
.. figure:: /images/shared/occupancy-vgpr.png
:alt: Occupancy related to VGPR usage in an Instinct MI300X GPU.
:align: center
@@ -1380,12 +1399,12 @@ Overall GPU resource utilization
--------------------------------
As depicted in the following figure, each XCD in
:doc:`MI300X </conceptual/gpu-arch/mi300>` contains 40 compute units (CUs),
:doc:`MI300X </reference/gpu-arch/mi300>` contains 40 compute units (CUs),
with 38 active. Each MI300X contains eight vertical XCDs, and a total of 304
active compute units capable of parallel computation. The first consideration is
the number of CUs a kernel can distribute its task across.
.. figure:: ../../../data/shared/xcd-sys-arch.png
.. figure:: /images/shared/xcd-sys-arch.png
XCD-level system architecture showing 40 compute units,
each with 32 KB L1 cache, a unified compute system with 4 ACE compute
@@ -1725,4 +1744,4 @@ topology during startup, reducing the need for extensive manual tuning.
Further reading
===============
* :doc:`vllm-optimization`
* :doc:`vllm-v1-optimization`
+363
View File
@@ -0,0 +1,363 @@
docker:
pull_tag: rocm/pytorch-xdit:v26.5
docker_hub_url: https://hub.docker.com/layers/rocm/pytorch-xdit/v26.5/images/sha256-b8ad9fd4b41bc116ac2aff07c1066bf369cf7fc110b1a323f6302191985a51fd
ROCm: 7.13.0
whats_new:
- "Hunyuan Video 1.5 sparse attention (SSTA) support"
- "Support fp8 MLA for MI355"
- "Block wise sparsity support for AMD triton FAv3 Sage attention"
components:
TheRock:
version: cbff3d1
url: https://github.com/ROCm/TheRock
rocm-libraries:
version: a668483b
url: https://github.com/ROCm/rocm-libraries
rocm-systems:
version: c76140fa
url: https://github.com/ROCm/rocm-systems
torch:
version: ff65f5b
url: https://github.com/ROCm/pytorch
torchaudio:
version: e3c6ee2
url: https://github.com/pytorch/audio
torchvision:
version: b919bd0
url: https://github.com/pytorch/vision
triton:
version: a272dfa
url: https://github.com/ROCm/triton
accelerate:
version: 46ba481
url: https://github.com/huggingface/accelerate
aiter:
version: bc5ea32c
url: https://github.com/ROCm/aiter
diffusers:
version: 447e571a
url: https://github.com/huggingface/diffusers
distvae:
version: 5a0fcbb
url: https://github.com/xdit-project/DistVAE
xfuser:
version: 051db68f
url: https://github.com/xdit-project/xDiT
yunchang:
version: 631bdfd
url: https://github.com/feifeibear/long-context-attention
supported_models:
- group: Hunyuan Video
js_tag: hunyuan
models:
- model: Hunyuan Video
model_repo: tencent/HunyuanVideo
revision: refs/pr/18
url: https://huggingface.co/tencent/HunyuanVideo
github: https://github.com/Tencent-Hunyuan/HunyuanVideo
mad_tag: pyt_xdit_hunyuanvideo
js_tag: hunyuan_tag
benchmark_command:
- mkdir results
- 'xdit \'
- '--model {model_repo} \'
- '--prompt "In the large cage, two puppies were wagging their tails at each other." \'
- '--batch_size 1 \'
- '--height 720 --width 1280 \'
- '--seed 1168860793 \'
- '--num_frames 129 \'
- '--num_inference_steps 50 \'
- '--warmup_calls 1 \'
- '--num_iterations 1 \'
- '--ulysses_degree 8 \'
- '--enable_tiling --enable_slicing \'
- '--guidance_scale 6.0 \'
- '--use_torch_compile \'
- '--attention_backend aiter \'
- '--output_directory results'
- model: Hunyuan Video 1.5
model_repo: hunyuanvideo-community/HunyuanVideo-1.5-Diffusers-720p_t2v
url: https://huggingface.co/hunyuanvideo-community/HunyuanVideo-1.5-Diffusers-720p_t2v
github: https://github.com/Tencent-Hunyuan/HunyuanVideo-1.5
mad_tag: pyt_xdit_hunyuanvideo_1_5
js_tag: hunyuan_1_5_tag
benchmark_command:
- mkdir results
- 'xdit \'
- '--model {model_repo} \'
- '--prompt "In the large cage, two puppies were wagging their tails at each other." \'
- '--task t2v \'
- '--height 720 --width 1280 \'
- '--seed 1168860793 \'
- '--num_frames 129 \'
- '--num_inference_steps 50 \'
- '--num_iterations 1 \'
- '--ulysses_degree 8 \'
- '--enable_tiling --enable_slicing \'
- '--use_torch_compile \'
- '--attention_backend aiter \'
- '--output_directory results'
- group: Wan-AI
js_tag: wan
models:
- model: Wan2.1
model_repo: Wan-AI/Wan2.1-I2V-14B-720P-Diffusers
url: https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-720P-Diffusers
github: https://github.com/Wan-Video/Wan2.1
mad_tag: pyt_xdit_wan_2_1
js_tag: wan_21_tag
benchmark_command:
- mkdir results
- 'xdit \'
- '--model {model_repo} \'
- '--prompt "Summer beach vacation style, a white cat wearing sunglasses sits on a surfboard. The fluffy-furred feline gazes directly at the camera with a relaxed expression. Blurred beach scenery forms the background featuring crystal-clear waters, distant green hills, and a blue sky dotted with white clouds. The cat assumes a naturally relaxed posture, as if savoring the sea breeze and warm sunlight. A close-up shot highlights the feline''s intricate details and the refreshing atmosphere of the seaside." \'
- '--height 720 \'
- '--width 1280 \'
- '--input_images /app/data/wan_input.jpg \'
- '--num_frames 81 \'
- '--ulysses_degree 8 \'
- '--use_parallel_vae \'
- '--seed 42 \'
- '--guidance_scale 3.0 \'
- '--num_iterations 1 \'
- '--num_inference_steps 40 \'
- '--use_torch_compile \'
- '--attention_backend aiter \'
- '--output_directory results'
- model: Wan2.2
model_repo: Wan-AI/Wan2.2-I2V-A14B-Diffusers
url: https://huggingface.co/Wan-AI/Wan2.2-I2V-A14B-Diffusers
github: https://github.com/Wan-Video/Wan2.2
mad_tag: pyt_xdit_wan_2_2
js_tag: wan_22_tag
benchmark_command:
- mkdir results
- 'xdit \'
- '--model {model_repo} \'
- '--prompt "Summer beach vacation style, a white cat wearing sunglasses sits on a surfboard. The fluffy-furred feline gazes directly at the camera with a relaxed expression. Blurred beach scenery forms the background featuring crystal-clear waters, distant green hills, and a blue sky dotted with white clouds. The cat assumes a naturally relaxed posture, as if savoring the sea breeze and warm sunlight. A close-up shot highlights the feline''s intricate details and the refreshing atmosphere of the seaside." \'
- '--height 720 \'
- '--width 1280 \'
- '--input_images /app/data/wan_input.jpg \'
- '--num_frames 81 \'
- '--ulysses_degree 8 \'
- '--use_parallel_vae \'
- '--seed 42 \'
- '--num_iterations 1 \'
- '--num_inference_steps 40 \'
- '--use_torch_compile \'
- '--attention_backend aiter \'
- '--output_directory results'
- group: FLUX
js_tag: flux
models:
- model: FLUX.1
model_repo: black-forest-labs/FLUX.1-dev
url: https://huggingface.co/black-forest-labs/FLUX.1-dev
github: https://github.com/black-forest-labs/flux
mad_tag: pyt_xdit_flux
js_tag: flux_1_tag
benchmark_command:
- mkdir results
- 'xdit \'
- '--model {model_repo} \'
- '--seed 42 \'
- '--prompt "A small cat" \'
- '--height 1024 \'
- '--width 1024 \'
- '--num_inference_steps 25 \'
- '--max_sequence_length 256 \'
- '--warmup_calls 5 \'
- '--ulysses_degree 8 \'
- '--use_torch_compile \'
- '--guidance_scale 0.0 \'
- '--num_iterations 50 \'
- '--attention_backend aiter \'
- '--output_directory results'
- model: FLUX.1 Kontext
model_repo: black-forest-labs/FLUX.1-Kontext-dev
url: https://huggingface.co/black-forest-labs/FLUX.1-Kontext-dev
github: https://github.com/black-forest-labs/flux
mad_tag: pyt_xdit_flux_kontext
js_tag: flux_1_kontext_tag
benchmark_command:
- mkdir results
- 'xdit \'
- '--model {model_repo} \'
- '--seed 42 \'
- '--prompt "Add a cool hat to the cat" \'
- '--height 1024 \'
- '--width 1024 \'
- '--num_inference_steps 30 \'
- '--max_sequence_length 512 \'
- '--warmup_calls 5 \'
- '--ulysses_degree 8 \'
- '--use_torch_compile \'
- '--input_images /app/data/flux_cat.png \'
- '--guidance_scale 2.5 \'
- '--num_iterations 25 \'
- '--attention_backend aiter \'
- '--output_directory results'
- model: FLUX.2
model_repo: black-forest-labs/FLUX.2-dev
url: https://huggingface.co/black-forest-labs/FLUX.2-dev
github: https://github.com/black-forest-labs/flux2
mad_tag: pyt_xdit_flux_2
js_tag: flux_2_tag
benchmark_command:
- mkdir results
- 'xdit \'
- '--model {model_repo} \'
- '--seed 42 \'
- '--prompt "Add a cool hat to the cat" \'
- '--height 1024 \'
- '--width 1024 \'
- '--num_inference_steps 50 \'
- '--max_sequence_length 512 \'
- '--warmup_calls 5 \'
- '--ulysses_degree 8 \'
- '--use_torch_compile \'
- '--input_images /app/data/flux_cat.png \'
- '--guidance_scale 4.0 \'
- '--num_iterations 25 \'
- '--attention_backend aiter \'
- '--output_directory results'
- model: FLUX.2 Klein
model_repo: black-forest-labs/FLUX.2-klein-9B
url: https://huggingface.co/black-forest-labs/FLUX.2-klein-9B
github: https://github.com/black-forest-labs/flux2
mad_tag: pyt_xdit_flux_2_klein
js_tag: flux_2_klein_tag
benchmark_command:
- mkdir results
- 'xdit \'
- '--model {model_repo} \'
- '--seed 42 \'
- '--prompt "A spectacular sunset over the ocean" \'
- '--height 2048 \'
- '--width 2048 \'
- '--num_inference_steps 4 \'
- '--warmup_calls 5 \'
- '--ulysses_degree 8 \'
- '--use_torch_compile \'
- '--guidance_scale 1.0 \'
- '--num_iterations 25 \'
- '--attention_backend aiter \'
- '--output_directory results'
- group: StableDiffusion
js_tag: stablediffusion
models:
- model: stable-diffusion-3.5-large
model_repo: stabilityai/stable-diffusion-3.5-large
url: https://huggingface.co/stabilityai/stable-diffusion-3.5-large
github: https://github.com/Stability-AI/sd3.5
mad_tag: pyt_xdit_sd_3_5
js_tag: stable_diffusion_3_5_large_tag
benchmark_command:
- mkdir results
- 'xdit \'
- '--model {model_repo} \'
- '--prompt "A capybara holding a sign that reads Hello World" \'
- '--num_iterations 50 \'
- '--num_inference_steps 28 \'
- '--pipefusion_parallel_degree 4 \'
- '--use_cfg_parallel \'
- '--use_torch_compile \'
- '--attention_backend aiter \'
- '--output_directory results'
- group: Z-Image
js_tag: z_image
models:
- model: Z-Image
model_repo: Tongyi-MAI/Z-Image
url: https://huggingface.co/Tongyi-MAI/Z-Image
github: https://github.com/Tongyi-MAI/Z-Image
mad_tag: pyt_xdit_z_image
js_tag: z_image_tag
benchmark_command:
- mkdir results
- 'xdit \'
- '--model {model_repo} \'
- '--seed 42 \'
- '--prompt "A crowded beach" \'
- '--height 1088 \'
- '--width 1920 \'
- '--num_inference_steps 50 \'
- '--ulysses_degree 2 \'
- '--ring_degree 2 \'
- '--use_cfg_parallel \'
- '--use_torch_compile \'
- '--guidance_scale 4.0 \'
- '--num_iterations 25 \'
- '--attention_backend aiter \'
- '--output_directory results'
- group: LTX
js_tag: ltx
models:
- model: LTX-2
model_repo: Lightricks/LTX-2
url: https://huggingface.co/Lightricks/LTX-2
github: https://github.com/Lightricks/LTX-2
mad_tag: pyt_xdit_ltx2
js_tag: ltx2_tag
benchmark_command:
- mkdir results
- 'xdit \'
- '--model {model_repo} \'
- '--seed 42 \'
- '--prompt "Cinematic action packed shot. The man says silently: \"We need to run.\". The camera zooms in on his mouth then immediately screams: \"NOW!\". The camera zooms back out, he turns around and bolts it." \'
- '--height 1088 \'
- '--width 1920 \'
- '--num_inference_steps 40 \'
- '--ulysses_degree 8 \'
- '--use_torch_compile \'
- '--guidance_scale 4.0 \'
- '--num_iterations 1 \'
- '--attention_backend aiter \'
- '--output_directory results'
- group: Qwen-Image
js_tag: qwen_image
models:
- model: Qwen-Image
model_repo: Qwen/Qwen-Image-2512
url: https://huggingface.co/Qwen/Qwen-Image-2512
github: https://github.com/QwenLM/Qwen-Image
mad_tag: pyt_xdit_qwen_image
js_tag: qwen_image_tag
benchmark_command:
- mkdir results
- 'xdit \'
- '--model {model_repo} \'
- '--seed 42 \'
- '--prompt "A cat holding a sign that says hello world" \'
- '--height 2048 \'
- '--width 2048 \'
- '--num_inference_steps 50 \'
- '--ulysses_degree 8 \'
- '--use_torch_compile \'
- '--guidance_scale 0.0 \'
- '--num_iterations 25 \'
- '--attention_backend aiter \'
- '--output_directory results'
- model: Qwen-Image-Edit
model_repo: Qwen/Qwen-Image-Edit
url: https://huggingface.co/Qwen/Qwen-Image-Edit
github: https://github.com/QwenLM/Qwen-Image
mad_tag: pyt_xdit_qwen_image_edit
js_tag: qwen_image_edit_tag
benchmark_command:
- mkdir results
- 'xdit \'
- '--model {model_repo} \'
- '--seed 42 \'
- '--prompt "Add a cool hat to the cat." \'
- '--negative_prompt "" \'
- '--input_images /app/data/flux_cat.png \'
- '--height 2048 \'
- '--width 2048 \'
- '--num_inference_steps 50 \'
- '--ulysses_degree 8 \'
- '--use_torch_compile \'
- '--guidance_scale 4.0 \'
- '--num_iterations 25 \'
- '--attention_backend aiter \'
- '--output_directory results'
Binary file not shown.

After

Width:  |  Height:  |  Size: 167 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 47 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 778 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 39 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 171 KiB

+694
View File
@@ -0,0 +1,694 @@
:selector-toc2: Installation environment
:selector-toc2-icon: fa-solid fa-computer
.. meta::
:description: Learn how to validate LLM inference performance on MI300X GPUs using AMD MAD and the ROCm vLLM Docker image.
:keywords: model, MAD, automation, dashboarding, validate
.. |VLLM_VERSION| replace:: 0.19.1
.. |VLLM_DOCKER_TAG_GFX950| replace:: rocm/vllm:rocm7.13.0_gfx950-dcgpu_ubuntu24.04_py3.13_pytorch_2.10.0_vllm_0.19.1
.. |VLLM_DOCKER_TAG_GFX94X| replace:: rocm/vllm:rocm7.13.0_gfx94X-dcgpu_ubuntu24.04_py3.13_pytorch_2.10.0_vllm_0.19.1
.. |VLLM_DOCKER_TAG_GFX120X-ALL| replace:: rocm/vllm:rocm7.13.0_gfx120X-all_ubuntu24.04_py3.13_pytorch_2.10.0_vllm_0.19.1
.. |VLLM_DOCKER_TAG_GFX110X-ALL| replace:: rocm/vllm:rocm7.13.0_gfx110X-all_ubuntu24.04_py3.13_pytorch_2.10.0_vllm_0.19.1
.. |VLLM_DOCKER_TAG_GFX1151| replace:: rocm/vllm:rocm7.13.0_gfx1151_ubuntu24.04_py3.13_pytorch_2.10.0_vllm_0.19.1
.. |VLLM_DOCKER_TAG_GFX1150| replace:: rocm/vllm:rocm7.13.0_gfx1150_ubuntu24.04_py3.13_pytorch_2.10.0_vllm_0.19.1
.. |VLLM_DOCKER_TAG_GFX1152| replace:: rocm/vllm:rocm7.13.0_gfx1152_ubuntu24.04_py3.13_pytorch_2.10.0_vllm_0.19.1
.. |VLLM_DOC| replace:: `vLLM <https://docs.vllm.ai/en/v0.19.1/>`__
.. |VLLM_USAGE_DOC| replace:: `Using vLLM <https://docs.vllm.ai/en/v0.19.1/usage/>`__
.. |VLLM_DOCKER_INSTALL_DOC| replace:: `Set up using Docker (vLLM docs) <https://docs.vllm.ai/en/v0.19.1/getting_started/installation/gpu/#amd-rocm_5>`__
.. |VLLM_PIP_INSTALL_DOC| replace:: `Set up using Python (vLLM docs) <https://docs.vllm.ai/en/v0.19.1/getting_started/installation/gpu/#amd-rocm_3>`__
**********************************
vLLM inference and serving on ROCm
**********************************
|VLLM_DOC| is an open-source library for fast, memory-efficient LLM inference
and serving. This page describes how to set up and run vLLM on AMD GPUs and
APUs using either a prebuilt Docker image (recommended) or pip. It applies to
:ref:`supported AMD GPUs and platforms <release-ai-ecosystem>`.
.. selector:: Device family
:key: fam
.. selector-option:: AMD Instinct™
:value: instinct w=compute
:width: 4
:toc-label: AMD Instinct
.. selector-option:: AMD Radeon™
:value: radeon w=compute
:width: 4
:toc-label: AMD Radeon
.. selector-option:: AMD Ryzen™
:value: ryzen w=compute
:width: 4
:toc-label: AMD Ryzen
.. ================================================================ GPU / APU ==
.. selected:: fam=instinct fam=radeon fam=ryzen
.. selector-dropdown:: Instinct GPU
:key: gpu
:show-cond: fam=instinct
.. selector-option:: AMD Instinct MI355X (gfx950)
:value: mi355x gfx=gfx950
.. selector-option:: AMD Instinct MI350X (gfx950)
:value: mi350x gfx=gfx950
.. selector-option:: AMD Instinct MI350P (gfx950)
:value: mi350p gfx=gfx950
.. selector-option:: AMD Instinct MI325X (gfx942)
:value: mi325x gfx=gfx942
.. selector-option:: AMD Instinct MI300X (gfx942)
:value: mi300x gfx=gfx942
.. selector-option:: AMD Instinct MI300A (gfx942)
:value: mi300a gfx=gfx942
.. selector-dropdown:: Radeon GPU
:key: gpu
:show-cond: fam=radeon
.. selector-option:: AMD Radeon AI PRO R9700 (gfx1201)
:value: ai-r9700 gfx=gfx1201
.. selector-option:: AMD Radeon AI PRO R9600D (gfx1201)
:value: ai-r9600d gfx=gfx1201
.. selector-option:: AMD Radeon RX 9070 XT (gfx1201)
:value: rx-9070-xt gfx=gfx1201
.. selector-option:: AMD Radeon RX 9070 GRE (gfx1201)
:value: rx-9070-gre gfx=gfx1201
.. selector-option:: AMD Radeon RX 9070 (gfx1201)
:value: rx-9070 gfx=gfx1201
.. selector-option:: AMD Radeon RX 9060 XT LP (gfx1200)
:value: rx-9060-xt-lp gfx=gfx1200
.. selector-option:: AMD Radeon RX 9060 XT (gfx1200)
:value: rx-9060-xt gfx=gfx1200
.. selector-option:: AMD Radeon RX 9060 (gfx1200)
:value: rx-9060 gfx=gfx1200
.. selector-option:: AMD Radeon PRO W7900 Dual Slot (gfx1100)
:value: w7900-dual-slot gfx=gfx1100
.. selector-option:: AMD Radeon PRO W7900 (gfx1100)
:value: w7900 gfx=gfx1100
.. selector-option:: AMD Radeon PRO W7800 48GB (gfx1100)
:value: w7800-48gb gfx=gfx1100
.. selector-option:: AMD Radeon PRO W7800 (gfx1100)
:value: w7800 gfx=gfx1100
.. selector-option:: AMD Radeon RX 7900 XTX (gfx1100)
:value: rx-7900-xtx gfx=gfx1100
.. selector-option:: AMD Radeon RX 7900 XT (gfx1100)
:value: rx-7900-xt gfx=gfx1100
.. selector-option:: AMD Radeon RX 7900 GRE (gfx1100)
:value: rx-7900-gre gfx=gfx1100
.. selector-option:: AMD Radeon PRO W7700 (gfx1101)
:value: w7700 gfx=gfx1101
.. selector-option:: AMD Radeon RX 7800 XT (gfx1101)
:value: rx-7800-xt gfx=gfx1101
.. selector-option:: AMD Radeon RX 7700 XT (gfx1101)
:value: rx-7700-xt gfx=gfx1101
.. selector-option:: AMD Radeon RX 7700 XE (gfx1101)
:value: rx-7700-xe gfx=gfx1101
.. selector-option:: AMD Radeon RX 7700 (gfx1101)
:value: rx-7700 gfx=gfx1101
.. selector-option:: AMD Radeon PRO V710 (gfx1101)
:value: v710 gfx=gfx1101
.. selector-option:: AMD Radeon RX 7600 (gfx1102)
:value: rx-7600 gfx=gfx1102
.. selector-dropdown:: Ryzen APU
:key: gpu
:show-cond: fam=ryzen
.. selector-option:: AMD Ryzen AI Max+ PRO 395 (gfx1151)
:value: max-pro-395 gfx=gfx1151
.. selector-option:: AMD Ryzen AI Max PRO 390 (gfx1151)
:value: max-pro-390 gfx=gfx1151
.. selector-option:: AMD Ryzen AI Max PRO 385 (gfx1151)
:value: max-pro-385 gfx=gfx1151
.. selector-option:: AMD Ryzen AI Max PRO 380 (gfx1151)
:value: max-pro-380 gfx=gfx1151
.. selector-option:: AMD Ryzen AI Max+ 395 (gfx1151)
:value: max-395 gfx=gfx1151
.. selector-option:: AMD Ryzen AI Max+ 392 (gfx1151)
:value: max-392 gfx=gfx1151
.. selector-option:: AMD Ryzen AI Max+ 388 (gfx1151)
:value: max-388 gfx=gfx1151
.. selector-option:: AMD Ryzen AI Max 390 (gfx1151)
:value: max-390 gfx=gfx1151
.. selector-option:: AMD Ryzen AI Max 385 (gfx1151)
:value: max-385 gfx=gfx1151
.. selector-option:: AMD Ryzen AI 9 PRO HX 475 (gfx1150)
:value: ai-9-pro-hx-475 gfx=gfx1150
.. selector-option:: AMD Ryzen AI 9 PRO HX 470 (gfx1150)
:value: ai-9-pro-hx-470 gfx=gfx1150
.. selector-option:: AMD Ryzen AI 9 PRO 465 (gfx1150)
:value: ai-9-pro-465 gfx=gfx1150
.. selector-option:: AMD Ryzen AI 7 PRO 450 (gfx1152)
:value: ai-7-pro-450 gfx=gfx1152
.. selector-option:: AMD Ryzen AI 5 PRO 440 (gfx1152)
:value: ai-5-pro-440 gfx=gfx1152
.. selector-option:: AMD Ryzen AI 9 HX 475 (gfx1150)
:value: ai-9-hx-475 gfx=gfx1150
.. selector-option:: AMD Ryzen AI 9 HX 470 (gfx1150)
:value: ai-9-hx-470 gfx=gfx1150
.. selector-option:: AMD Ryzen AI 9 465 (gfx1150)
:value: ai-9-465 gfx=gfx1150
.. selector-option:: AMD Ryzen AI 7 450 (gfx1152)
:value: ai-7-450 gfx=gfx1152
.. selector-option:: AMD Ryzen AI 9 HX PRO 375 (gfx1150)
:value: 9-hx-pro-375 gfx=gfx1150
.. selector-option:: AMD Ryzen AI 9 HX PRO 370 (gfx1150)
:value: 9-hx-pro-370 gfx=gfx1150
.. selector-option:: AMD Ryzen AI 7 PRO 350 (gfx1152)
:value: ai-7-pro-350 gfx=gfx1152
.. selector-option:: AMD Ryzen AI 5 PRO 340 (gfx1152)
:value: ai-5-pro-340 gfx=gfx1152
.. selector-option:: AMD Ryzen AI 9 HX 375 (gfx1150)
:value: 9-hx-375 gfx=gfx1150
.. selector-option:: AMD Ryzen AI 9 HX 370 (gfx1150)
:value: 9-hx-370 gfx=gfx1150
.. selector-option:: AMD Ryzen AI 9 365 (gfx1150)
:value: 9-365 gfx=gfx1150
.. selector-option:: AMD Ryzen AI 7 350 (gfx1152)
:value: ai-7-350 gfx=gfx1152
.. selector-option:: AMD Ryzen AI 7 345 (gfx1152)
:value: ai-7-345 gfx=gfx1152
.. selector-option:: AMD Ryzen AI 5 340 (gfx1152)
:value: ai-5-340 gfx=gfx1152
.. selector-option:: AMD Ryzen AI 5 330 (gfx1152)
:value: ai-5-330 gfx=gfx1152
.. selector:: vLLM version
:key: vllm-ver
.. selector-option:: 0.19.1
:value: 0.19.1
:width: 12
.. selector:: Installation method
:key: i
.. selector-option:: Docker
:value: docker
:width: 6
.. selector-option:: pip
:value: pip
:width: 6
Prerequisites
=============
.. selected:: i=docker
.. selected:: fam=instinct fam=radeon
- Ensure your system has the AMD GPU Driver (amdgpu) installed. See the
:ref:`compat-matrix` for driver support information. For installation
instructions, see the `AMD GPU Driver documentation
<https://instinct.docs.amd.com/projects/amdgpu-docs/en/31.30.0-preview/index.html>`__.
- Ensure the host system has `Docker Engine
<https://docs.docker.com/engine/install/>`__ and the AMD GPU Driver
(amdgpu) installed.
.. selected:: fam=ryzen
Ensure the host system has `Docker Engine
<https://docs.docker.com/engine/install/>`__ installed.
.. selected:: i=pip
.. selected:: fam=instinct fam=radeon
- Ensure your system has the AMD GPU Driver (amdgpu) installed. See the
:ref:`compat-matrix` for driver support information. For installation
instructions, see the `AMD GPU Driver documentation
<https://instinct.docs.amd.com/projects/amdgpu-docs/en/31.30.0-preview/index.html>`__.
- Ensure your system has :ref:`Python 3.13 <rocm-compat-python>` installed and
accessible. Review the :ref:`compat-matrix` for more support details.
- Install `uv <https://docs.astral.sh/uv/getting-started/installation/>`__,
a drop-in replacement for pip that handles custom package indexes more
predictably.
.. note::
It's recommended to use `uv <https://docs.astral.sh/uv/pip/>`__ to install
the vLLM wheel. Installing from custom package indexes with pip can be
cumbersome because pip resolves packages from both ``--extra-index-url`` and
the default index, then selects the highest available version. This makes it
difficult to install a wheel from a custom index when all dependency
versions are pinned exactly.
.. selected:: i=docker
:heading: Get started
.. selected:: gfx=gfx950
1. Pull the ROCm vLLM |VLLM_VERSION| Docker image.
.. code-block:: bash
:substitutions:
docker pull |VLLM_DOCKER_TAG_GFX950|
2. Start the Docker container.
.. code-block:: bash
:substitutions:
docker run -it --rm \
--device /dev/kfd \
--device /dev/dri \
--network=host \
--ipc=host \
--group-add=video \
--cap-add=SYS_PTRACE \
--security-opt seccomp=unconfined \
-v <path/to/your/models>:/app/models \
-e HF_HOME="/app/models" \
|VLLM_DOCKER_TAG_GFX950| \
bash
.. selected:: gfx=gfx942
1. Pull the ROCm vLLM |VLLM_VERSION| Docker image.
.. code-block:: bash
:substitutions:
docker pull |VLLM_DOCKER_TAG_GFX94X|
2. Start the Docker container.
.. code-block:: bash
:substitutions:
docker run -it --rm \
--device /dev/kfd \
--device /dev/dri \
--network=host \
--ipc=host \
--group-add=video \
--cap-add=SYS_PTRACE \
--security-opt seccomp=unconfined \
-v <path/to/your/models>:/app/models \
-e HF_HOME="/app/models" \
|VLLM_DOCKER_TAG_GFX94X| \
bash
.. selected:: gfx=gfx1201 gfx=gfx1200
1. Pull the ROCm vLLM Docker image.
.. code-block:: bash
:substitutions:
docker pull |VLLM_DOCKER_TAG_GFX120X-ALL|
2. Start the Docker container.
.. code-block:: bash
:substitutions:
docker run -it --rm \
--device /dev/kfd \
--device /dev/dri \
--network=host \
--ipc=host \
--group-add=video \
--cap-add=SYS_PTRACE \
--security-opt seccomp=unconfined \
-v <path/to/your/models>:/app/models \
-e HF_HOME="/app/models" \
|VLLM_DOCKER_TAG_GFX120X-ALL| \
bash
.. selected:: gfx=gfx1100 gfx=gfx1101 gfx=gfx1102
1. Pull the ROCm vLLM Docker image.
.. code-block:: bash
:substitutions:
docker pull |VLLM_DOCKER_TAG_GFX110X-ALL|
2. Start the Docker container.
.. code-block:: bash
:substitutions:
docker run -it --rm \
--device /dev/kfd \
--device /dev/dri \
--network=host \
--ipc=host \
--group-add=video \
--cap-add=SYS_PTRACE \
--security-opt seccomp=unconfined \
-v <path/to/your/models>:/app/models \
-e HF_HOME="/app/models" \
|VLLM_DOCKER_TAG_GFX110X-ALL| \
bash
.. selected:: gfx=gfx1151
1. Pull the ROCm vLLM Docker image.
.. code-block:: bash
:substitutions:
docker pull |VLLM_DOCKER_TAG_GFX1151|
2. Start the Docker container.
.. code-block:: bash
:substitutions:
docker run -it --rm \
--device /dev/kfd \
--device /dev/dri \
--network=host \
--ipc=host \
--group-add=video \
--cap-add=SYS_PTRACE \
--security-opt seccomp=unconfined \
-v <path/to/your/models>:/app/models \
-e HF_HOME="/app/models" \
|VLLM_DOCKER_TAG_GFX1151| \
bash
.. selected:: gfx=gfx1150
1. Pull the ROCm vLLM Docker image.
.. code-block:: bash
:substitutions:
docker pull |VLLM_DOCKER_TAG_GFX1150|
2. Start the Docker container.
.. code-block:: bash
:substitutions:
docker run -it --rm \
--device /dev/kfd \
--device /dev/dri \
--network=host \
--ipc=host \
--group-add=video \
--cap-add=SYS_PTRACE \
--security-opt seccomp=unconfined \
-v <path/to/your/models>:/app/models \
-e HF_HOME="/app/models" \
|VLLM_DOCKER_TAG_GFX1150| \
bash
.. selected:: gfx=gfx1152
1. Pull the ROCm vLLM Docker image.
.. code-block:: bash
:substitutions:
docker pull |VLLM_DOCKER_TAG_GFX1152|
2. Start the Docker container.
.. code-block:: bash
:substitutions:
docker run -it --rm \
--device /dev/kfd \
--device /dev/dri \
--network=host \
--ipc=host \
--group-add=video \
--cap-add=SYS_PTRACE \
--security-opt seccomp=unconfined \
-v <path/to/your/models>:/app/models \
-e HF_HOME="/app/models" \
|VLLM_DOCKER_TAG_GFX1152| \
bash
.. seealso::
|VLLM_DOCKER_INSTALL_DOC|
3. After setting up your environment, follow the vLLM |VLLM_VERSION| usage
documentation to get started: |VLLM_USAGE_DOC|.
.. selected:: i=pip
:heading: Install vLLM using pip
1. Set up your Python virtual environment.
.. code-block:: shell
python -m venv .venv
2. Activate your Python virtual environment.
.. code-block:: shell
source .venv/bin/activate
3. Install PyTorch 2.10 in your virtual environment. This should also
install the ROCm core libraries as a dependency.
.. note::
``torchvision`` 0.25 must be installed alongside PyTorch — vLLM
requires it and will fail without it.
.. selected:: gfx=gfx950
.. code-block:: bash
python -m pip install --index-url https://repo.amd.com/rocm/whl/gfx950-dcgpu/ \
"torch==2.10.0+rocm7.13.0" \
"torchvision==0.25.0+rocm7.13.0" \
"torchaudio==2.10.0+rocm7.13.0"
.. selected:: gfx=gfx942
.. code-block:: bash
python -m pip install --index-url https://repo.amd.com/rocm/whl/gfx94X-dcgpu/ \
"torch==2.10.0+rocm7.13.0" \
"torchvision==0.25.0+rocm7.13.0" \
"torchaudio==2.10.0+rocm7.13.0"
.. selected:: gfx=gfx1201 gfx=gfx1200
.. code-block:: bash
python -m pip install --index-url https://repo.amd.com/rocm/whl/gfx120X-all/ \
"torch==2.10.0+rocm7.13.0" \
"torchvision==0.25.0+rocm7.13.0" \
"torchaudio==2.10.0+rocm7.13.0"
.. selected:: gfx=gfx1100 gfx=gfx1101 gfx=gfx1102 gfx=gfx1103
.. code-block:: bash
python -m pip install --index-url https://repo.amd.com/rocm/whl/gfx110X-all/ \
"torch==2.10.0+rocm7.13.0" \
"torchvision==0.25.0+rocm7.13.0" \
"torchaudio==2.10.0+rocm7.13.0"
.. selected:: gfx=gfx1151
.. code-block:: bash
python -m pip install --index-url https://repo.amd.com/rocm/whl/gfx1151/ \
"torch==2.10.0+rocm7.13.0" \
"torchvision==0.25.0+rocm7.13.0" \
"torchaudio==2.10.0+rocm7.13.0"
.. selected:: gfx=gfx1150
.. code-block:: bash
python -m pip install --index-url https://repo.amd.com/rocm/whl/gfx1150/ \
"torch==2.10.0+rocm7.13.0" \
"torchvision==0.25.0+rocm7.13.0" \
"torchaudio==2.10.0+rocm7.13.0"
.. selected:: gfx=gfx1152
.. code-block:: bash
python -m pip install --index-url https://repo.amd.com/rocm/whl/gfx1152/ \
"torch==2.10.0+rocm7.13.0" \
"torchvision==0.25.0+rocm7.13.0" \
"torchaudio==2.10.0+rocm7.13.0"
4. Install Flash Attention.
.. selected:: gfx=gfx950
.. code-block:: bash
python -m pip install https://rocm.frameworks.amd.com/whl/gfx950-dcgpu/flash_attn-2.8.3-cp313-cp313-linux_x86_64.whl
.. selected:: gfx=gfx942
.. code-block:: bash
python -m pip install https://rocm.frameworks.amd.com/whl/gfx94X-dcgpu/flash_attn-2.8.3-cp313-cp313-linux_x86_64.whl
.. selected:: gfx=gfx1201 gfx=gfx1200
.. code-block:: bash
python -m pip install https://rocm.frameworks.amd.com/whl/gfx120X-all/flash_attn-2.8.3-py3-none-any.whl
.. selected:: gfx=gfx1100 gfx=gfx1101 gfx=gfx1102
.. code-block:: bash
python -m pip install https://rocm.frameworks.amd.com/whl/gfx110X-all/flash_attn-2.8.3-py3-none-any.whl
.. selected:: gfx=gfx1151
.. code-block:: bash
python -m pip install https://rocm.frameworks.amd.com/whl/gfx1151/flash_attn-2.8.3-py3-none-any.whl
.. selected:: gfx=gfx1150
.. code-block:: bash
python -m pip install https://rocm.frameworks.amd.com/whl/gfx1150/flash_attn-2.8.3-py3-none-any.whl
.. selected:: gfx=gfx1152
.. code-block:: bash
python -m pip install https://rocm.frameworks.amd.com/whl/gfx1152/flash_attn-2.8.3-py3-none-any.whl
5. Install the vLLM |VLLM_VERSION| wheel using ``uv pip``.
.. selected:: gfx=gfx950
.. code-block:: bash
uv pip install https://rocm.frameworks.amd.com/whl/gfx950-dcgpu/vllm-0.19.1.dev3%2Brocm7.13.0.g72ed2b398.d20260513-cp313-cp313-linux_x86_64.whl
.. selected:: gfx=gfx942
.. code-block:: bash
uv pip install https://rocm.frameworks.amd.com/whl/gfx94X-dcgpu/vllm-0.19.1.dev3%2Brocm7.13.0.g72ed2b398.d20260513-cp313-cp313-linux_x86_64.whl
.. selected:: gfx=gfx1201 gfx=gfx1200
.. code-block:: bash
uv pip install https://rocm.frameworks.amd.com/whl/gfx120X-all/vllm-0.19.1.dev3%2Brocm7.13.0.g72ed2b398.d20260513-cp313-cp313-linux_x86_64.whl
.. selected:: gfx=gfx1100 gfx=gfx1101 gfx=gfx1102
.. code-block:: bash
uv pip install https://rocm.frameworks.amd.com/whl/gfx110X-all/vllm-0.19.1.dev3%2Brocm7.13.0.g72ed2b398.d20260513-cp313-cp313-linux_x86_64.whl
.. selected:: gfx=gfx1151
.. code-block:: bash
uv pip install https://rocm.frameworks.amd.com/whl/gfx1151/vllm-0.19.1.dev3%2Brocm7.13.0.g72ed2b398.d20260513-cp313-cp313-linux_x86_64.whl
.. selected:: gfx=gfx1150
.. code-block:: bash
uv pip install https://rocm.frameworks.amd.com/whl/gfx1150/vllm-0.19.1.dev3%2Brocm7.13.0.g72ed2b398.d20260513-cp313-cp313-linux_x86_64.whl
.. selected:: gfx=gfx1152
.. code-block:: bash
uv pip install https://rocm.frameworks.amd.com/whl/gfx1152/vllm-0.19.1.dev3%2Brocm7.13.0.g72ed2b398.d20260513-cp313-cp313-linux_x86_64.whl
6. Set the following environment variables to prevent errors related to ROCm platform and Flash Attention availability when running vLLM.
.. code-block:: bash
export PYTHONPATH=$VIRTUAL_ENV/lib/python3.13/site-packages/_rocm_sdk_core/share/amd_smi
export FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE
7. Check your installation.
.. code-block:: bash
echo "=== vLLM ===" && python -c "import vllm; print('vLLM version:', vllm.__version__)"
echo "=== PyTorch ===" && python -c "import torch; print('PyTorch:', torch.__version__); print('HIP available:', torch.cuda.is_available()); print('HIP built:', torch.backends.hip.is_built() if hasattr(torch.backends, 'hip') else 'N/A')"
echo "=== flash-attn ===" && python -c "import flash_attn; print('flash-attn:', flash_attn.__version__)"
.. seealso::
|VLLM_PIP_INSTALL_DOC|
8. After setting up your environment, follow the vLLM |VLLM_VERSION| usage
documentation to get started: |VLLM_USAGE_DOC|.
@@ -1,3 +1,6 @@
:selector-toc2: Model
:selector-toc2-icon: fa-solid fa-robot
.. meta::
:description: Learn to validate diffusion model video generation on MI300X, MI350X and MI355X accelerators using
prebuilt and optimized docker images.
@@ -9,13 +12,13 @@ xDiT diffusion inference
.. _xdit-video-diffusion:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/xdit-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit.yaml
{% set docker = data.docker %}
The `rocm/pytorch-xdit <{{ docker.docker_hub_url }}>`_ Docker image offers a prebuilt, optimized environment based on `xDiT <https://github.com/xdit-project/xDiT>`_ for
benchmarking diffusion model video and image generation on gfx942 and gfx950 series (AMD Instinct™ MI300X, MI325X, MI350X, and MI355X) GPUs.
The image runs `ROCm {{docker.ROCm}} (preview) <https://rocm.docs.amd.com/en/7.12.0-preview/about/release-notes.html>`__ based on `TheRock <https://github.com/ROCm/TheRock>`_
The image runs `ROCm {{docker.ROCm}} <https://rocm.docs.amd.com/en/7.13.0-preview/about/release-notes.html>`__ based on `TheRock <https://github.com/ROCm/TheRock>`_
and includes the following components:
.. dropdown:: Software components - {{ docker.pull_tag.split('-')|last }}
@@ -37,7 +40,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
What's new
==========
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/xdit-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit.yaml
{% set docker = data.docker %}
@@ -54,43 +57,37 @@ The following models are supported for inference performance benchmarking.
Some instructions, commands, and recommendations in this documentation might
vary by model -- select one to get started.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/xdit-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit.yaml
{% set docker = data.docker %}
.. raw:: html
.. selector:: Model
:key: model-group
<div id="vllm-benchmark-ud-params-picker" class="container-fluid">
<div class="row gx-0">
<div class="col-2 me-1 px-2 model-param-head">Model</div>
<div class="row col-10 pe-0">
{% for model_group in docker.supported_models %}
<div class="col-6 px-2 model-param" data-param-k="model-group" data-param-v="{{ model_group.js_tag }}" tabindex="0">{{ model_group.group }}</div>
{% endfor %}
</div>
</div>
{% for model_group in docker.supported_models %}
.. selector-option:: {{ model_group.group }}
:value: {{ model_group.js_tag }}
:width: 25%
<div class="row gx-0 pt-1">
<div class="col-2 me-1 px-2 model-param-head">Variant</div>
<div class="row col-10 pe-0">
{% for model_group in docker.supported_models %}
{% set models = model_group.models %}
{% for model in models %}
{% if models|length % 3 == 0 %}
<div class="col-4 px-2 model-param" data-param-k="model" data-param-v="{{ model.js_tag }}" data-param-group="{{ model_group.js_tag }}" tabindex="0">{{ model.model }}</div>
{% else %}
<div class="col-6 px-2 model-param" data-param-k="model" data-param-v="{{ model.js_tag }}" data-param-group="{{ model_group.js_tag }}" tabindex="0">{{ model.model }}</div>
{% endif %}
{% endfor %}
{% endfor %}
</div>
</div>
</div>
{% endfor %}
{% for model_group in docker.supported_models %}
.. selector:: Variant
:key: model
:show-cond: model-group={{ model_group.js_tag }}
{% set models = model_group.models %}
{% for model in models %}
.. selector-option:: {{ model.model }}
:value: {{ model.js_tag }}
{% endfor %}
{% endfor %}
{% for model_group in docker.supported_models %}
{% for model in model_group.models %}
.. container:: model-doc {{ model.js_tag }}
.. selected:: model={{ model.js_tag }}
.. note::
@@ -119,7 +116,7 @@ system's configuration.
Pull the Docker image
=====================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/xdit-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit.yaml
{% set docker = data.docker %}
@@ -133,7 +130,7 @@ Pull the Docker image
Validate and benchmark
======================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/xdit-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit.yaml
{% set docker = data.docker %}
@@ -143,7 +140,7 @@ Validate and benchmark
{% for model_group in docker.supported_models %}
{% for model in model_group.models %}
.. container:: model-doc {{model.js_tag}}
.. selected:: model={{ model.js_tag }}
The following commands are written for {{ model.model }}.
See :ref:`xdit-video-diffusion-supported-models` to switch to another available model.
@@ -156,13 +153,13 @@ Choose your setup method
You can either use an existing Hugging Face cache or download the model fresh inside the container.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/xdit-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit.yaml
{% set docker = data.docker %}
{% for model_group in docker.supported_models %}
{% for model in model_group.models %}
.. container:: model-doc {{model.js_tag}}
.. selected:: model={{model.js_tag}}
.. tab-set::
@@ -248,14 +245,14 @@ You can either use an existing Hugging Face cache or download the model fresh in
Run inference
=============
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/xdit-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit.yaml
{% set docker = data.docker %}
{% for model_group in docker.supported_models %}
{% for model in model_group.models %}
.. container:: model-doc {{ model.js_tag }}
.. selected:: model={{ model.js_tag }}
.. tab-set::
@@ -305,7 +302,5 @@ Run inference
Previous versions
=================
See
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
to find documentation for previous releases of xDiT diffusion inference
performance testing.
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
releases of xDiT diffusion inference performance testing.

Some files were not shown because too many files have changed in this diff Show More