Compare commits

..
Author SHA1 Message Date
yugang-amdandGitHub 243df61a88 Remove xref links to legacy portal 2026-05-15 20:32:51 -04:00
Pratik BasyalandGitHub 21d41b8f98 7.13.0 Ryzen AI 9 PRO HX 475, 470 branding updated (#6264)
* Ryzen AI 9 PRO HX 475, 470 updated

* Selector options updated
2026-05-15 19:58:09 -04:00
pmoutsias-amdandGitHub fe6e8e25cf Merge pull request #6263 from peterjunpark/docs/7.13.0
[docs/7.13.0] Temporarily remove unavailable links and update README link
2026-05-15 19:40:44 -04:00
Peter Park 047c32b40b one more 2026-05-15 19:37:23 -04:00
Peter Park 8a615acc76 comment out unavailable links 2026-05-15 19:33:22 -04:00
Peter Park f6a2652ac7 update release note link 2026-05-15 19:33:22 -04:00
pmoutsias-amdandGitHub bed8ad9815 Merge pull request #6262 from peterjunpark/docs/7.13.0
[docs/7.13.0] Fix component GH links
2026-05-15 19:28:27 -04:00
Peter Park d0e0d07817 docs: fix compo github urls 2026-05-15 19:21:26 -04:00
pmoutsias-amdandGitHub ad14fd8081 Merge pull request #6261 from peterjunpark/docs/7.13.0
[docs/7.13]: Fix component version mismatch
2026-05-15 19:08:01 -04:00
Peter Park 8f0e2cfc94 fix components list 2026-05-15 18:59:38 -04:00
Peter Park 7529766d33 [docs/7.13.0] Update docs for 7.13
[docs/7.13.0] document component packages (#736)

* add package list

* add selector

* remove support col

* update package list

[docs/7.13.0] Allow selector options to set multiple values (#737)

* allow selector options to set multiple values

mostly for gfx

* rename "when" to "cond"

* update wordlist

* simplify vllm and ai-ecosystem

* simplify pages using gfx option

update configs

add .editorconfig

add yaml indent_size

add more options

add glossary and gpu hardware specs pages

add components to toc and remove rocm-packages

conf: remove "preview"

install: add amdgpu-lib to pkgman method

fix extra code block && make tabs sync

fix selector option resolution

add blurb explaining graphics vs headless

install: add multi-arch

install: apply feedback (multi-arch)

fix selector

fix

install: fix OEM kernel dropdown for multi-arch selection

install: add intro explaining install method

fix

fix selector js

Reorg and fix compat page

reorg compat

feat(js): add support for dropdown input

chore: reorg inference and dlf pages

chore: update TOC

fix

update toc

update

conf: improve substitutions

conf: add back datatemplates plugin

chore: move RELEASE to about/release-notes.md and don't copy

reorg toc

put contribute under about

install: reorg files and add (gfxXYZ)

reorg selectors

update

use dropdown-input for large list of GPUs and rm selector-info icon

install: remove centos and azl

remove centos and azl from release notes

fix "os-version" tags

feat(selector.py): make dropdown input a separate directive

docs: org deep learning frameworks

yep

fix(selector.js): fix url state flickering

docs: clean up comfyui page

clean up

colocate images

docs: add back env vars page

add remote-content extension

add AI Playbooks to toc and header nav

Restore fine-tuning pages and do some clean-up and fmt

docs: clean up comfyui page

clean up

consistency

Add Docker reference doc

Add Docker reference doc

make anchor text consistent

asdf

install: reorg files and add (gfxXYZ) and add missing Ryzen AI 400
Series (GPT)

reorg selectors

update

update compat

add gfx1030 to selector (radeon pro)

add gfx1152 to selector [Ryzen AI (PRO) 300 Series]

docs: add sglang page stub

fix

chore: update and fmt configs

update release

feat: add metadata to define selector TOC2 heading and icon

restore xdit page

delete xtra xdit data file

bump TomSelect to 2.6.0

organize sphinx extensions and enable legacy selector

add selector metadata

restore HIP programming guide page

Add back setting-cus.rst under env vars page

add back hpc page and organize under resources/

restore gpu-isolation reference page

delete super stale & redudant pages

delete more garb

restore inf/fine-tuning optimization guides and reorg images

update toc and fix other stuff

restore system optimization guides

move up

restore gpu-arch pages

wip

add graph-safe support and hardware atomics support to toc

also update docker doc

use old table in gpu-specs

add sglang doc

make docker run args consistently ordered

organzie gpu-arch files

org

add sphinx_substitution_extensions

update

update me

add release highlights

remove centos and azl from gpu-selector

fix link

update highlights

remove hover effects on diagrams and fix toc

fix toc

fix

remove cursor: pointer

remove miopen release highlight

update

release: add partitioning and virtu support

add component index pages

fix

remove pytorch docker for rn

update ai ecosystem in release notes

add compos to release notes

fix

fix

ROCM-21212 7.13.0 Known issue added (#757)

* add compos to release notes

fix

* ROCM-21212 Known issue added

---------

Co-authored-by: Peter Park <peter.park@amd.com>

ROCM-23565 known issue added (#758)

css/js: add flash on content change

js only

fix links and templating

restructure wip

fixes

fix links

update adrenalin to 26.5.1

update GIM to 9.0.0

update LICENSE copyright year

use includes for release notes

update stuff

update jax, vllm, sglang

add driver prereqs

fix selector in vllm

fix

update hw support table in release notes

fix

Fix compat info and combine Radeon PRO w/ Radeon

Add hipBLASLt, rocSOLVER, and rocSPARSE release highlights

test canonical url

test

Duplication removed

Add RCCL multi-node performance optimization highlight

Remove non-public internal API details from ROCprofiler-SDK changelog

Apply suggestion from @anisha-amd

Replace Strix codename with Ryzen in SQTT decoder highlight

Add Composable Kernel FP8 quantization and SageAttention v2 highlights

Apply suggestions from @amd-jnovotny

Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>
Co-authored-by: yugang-amd <yugang.wang@amd.com>

Apply suggestions from @amd-jnovotny

Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>

Update docs/about/includes/core-sdk-components-aggregated-changelog.md

Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>

Update docs/about/includes/core-sdk-components-aggregated-changelog.md

Co-authored-by: Jeffrey Novotny <jnovotny@amd.com>

Confirm version numbers for hipCUB, hipRAND, hipSOLVER, rocPRIM,
rocRAND, rocSOLVER, rocSPARSE, rocThrust, and rocWMMA

install: remove runfile from multi-arch

Add version numbers for rocminfo and ROCprofiler-SDK; fix env var name
in Systems Profiler highlight

Shorten rocSPARSE highlight heading

FFT changelog additions from blocker PR 7089

Remove Ryzen OEM kernel prereq from runfile

7130 known issue batch2 (#763)

* Known issue for ROCM-21815 added

* ROCM-21824 Known issue added

* Systems profiler release highlight updated

* ROCm Systems Profiler link added

remove rocr runtime from diagram

Update SQTT decoder heading and add ROCprofiler-SDK links; minor wording
fix

Remove version numbers from highlight body text for Composable Kernel,
rocSOLVER, and rocSPARSE

Remove ROCprofiler-SDK links from trace decoder highlight

Clarify hipBLASLt General Batched GEMM description

Docs 7.13 structure update (#759)

* Update reference section structure and add MI350 series to GPU arch

* Update docs/reference/gpu-specs.rst

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Update glossary title

* Update docs/components/runtimes-and-compilers.rst

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Split up glossary

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Update docs/sphinx/_toc.yml.in

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Correct CDNA mentions

* Fix atomic add support page

---------

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

Fix typo: GMM -> GEMM in hipBLASLt highlight

fix typo: support --> supported

Update docs/about/release-notes.md

Co-authored-by: pmoutsias-amd <peter.moutsias@amd.com>

Re-add ROCprofiler-SDK links to trace decoder highlight

add links in rocm core sdk diagram

add rocdecode and rocjpeg to stack diagram

remove rocsolver LANGE and GECON changes from release notes

Apply suggestion from @lpaoletti

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

Update docs/about/release-notes.md

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

Update docs/about/release-notes.md

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

fix

Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

Apply PR review suggestions to release notes

- Update APU branding: add "AI Max PRO 300 series" alongside Ryzen AI
Max 300 series
- Broaden CK heading from FP8-specific to general quantization and
capabilities
- Remove premature Roofline limitation note for RDNA 3.5

Add hipBLASLt and rocSOLVER doc links per PR review feedback

Add re-attach to profiled process section to ROCm Systems Profiler
highlights

Apply ROCprofiler-SDK 1.3.0 changelog edits from PR review

remove graphics from "all"

Add resolved issues for RPM install, vLLM, and PyTorch DDP

Add resolved issue for vLLM tensor parallelism launch failure

Update docs/about/includes/core-sdk-components-aggregated-changelog.md

Co-authored-by: spolifroni-amd <Sandra.Polifroni@amd.com>

Fix AMD SMI version to 26.4.0 in aggregated changelog

add component versions to components table

remove gfx1153

Add ROCm 7.13 component versions to components table

Update component table links to therock-7.13

Fix ROCdbgapi version to 0.80.0 in components table

Group release highlights by category for better navigation

Organize the 16 flat H3 highlights into 4 scannable category groups
(Platform and hardware support, AI inference and frameworks, Developer
tools and profiling, Libraries) following the pattern used in ROCm
7.0.0.
Flatten Compute Profiler and Systems Profiler sub-items into bullet
lists to avoid H5 depth.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

add links to 7.9, 7.10, 7.11, and 7.12 release notes

fix gpu lists in release notes and gpu-selector

Add rocSPARSE documentation link to sparse factorization section

fmt html

fix missing version for LLVM and hipinfo

add trademark symbols

Minor fixes (#764)

* Minor fixes

* Minor fixes

archive stale pages

Fix RCCL branding, changelog versions, and doc links

- Use "AMD Ryzen AI Max 300 series" in RCCL heading and body
- Fix ROCdbgapi version in changelog (0.80 -> 0.80.0) to match
components table
- Fix rocSHMEM version in changelog (3.3.0 -> 3.4.0) to match components
table
- Change Systems Profiler doc links from /en/develop/ to /en/latest/

Fix typos, broken links, and style issues in release notes

- Fix 7.11.0 preview link pointing to 7.10.0-preview URL
- Fix re-attach doc link pointing to GitHub instead of rocm.docs.amd.com
- Add missing commas in version list (7.9.0, 7.10.0, 7.11.0)
- Fix "These issue" → "These issues" (grammar)
- Fix cross-reference from "Detailed component changes" to "ROCm
component changelogs"
- Fix "may" → "might" per style guide
- Fix "statisitcs" typo, "MI 350" → "MI350", "API's" → "APIs"
- Replace "Strix Halo" codename with "AMD Ryzen AI Max 300 series"
- Fix "masks values" → "mask values", "Fix" → "Fixed" for consistency
- Merge duplicate Changed sections in AMD SMI changelog
- Add missing blank line before rocSHMEM Changed heading
- Capitalize "cpu" → "CPU"

clean up changelogs formatting

fix missed directory renaming

fix

update install os selectors

update ryzen OSes and kernel versions

update firmware versions

fix table formatting

fix malformed selector

update component support table in release notes

fix include path for windows version selector

remove "graphics and mixed compute" from uninstall

add amdgpu-install and graphics/mixed use case

install: simplify "add additional package repositories"

remove amdgpu-lib

fix os selector (graphics)

remove graphics gpg keys and repos

add all options to multi-arch installation

update runfile installer url

fix compat selector

update install instructions for 7.13

fix rhel ver selector

fix os selector for ryzen

add python 3.14 to windows pip prereq

Fix punctuation, tense, branding, and style in component changelogs

- Add terminal periods to bullet points missing them
- Fix tense consistency (Improves → Improved, fails → failed)
- Add backticks to amd_dbgapi_process_get_info() function name
- Replace Strix Halo codename with AMD Ryzen AI Max 300 series
- Add commas after e.g. per style guide
- Remove spurious commas in restrictive clauses
- Soften future fix commitment (will be fixed → planned to be addressed)
- Add missing article (Fixed an issue)
- Merge upstream branding (AMD Instinct MI350P) with punctuation fixes
- Add colons before sub-lists, periods on sub-list items

Editorial pass on release notes highlights and profiler sections

- Add summary paragraph under Release highlights
- Wrap long lines for maintainability
- Add backticks to code identifiers in HIP section
- Fix RDNA3 → RDNA 3 spacing in CK heading
- Tighten Systems Profiler bullets: remove repetitive product name,
  use consistent declarative tense
- Restructure hipBLASLt and Compute Profiler sections for clarity
- Minor grammar fixes (variants are available, via → through)

Add rocWMMA HIP RTC resolved issue entry for 7.13

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

Remove comp pages (#766)

* Remove comp pages

* Add hw list on indox

* Apply suggestions from code review

Co-authored-by: Istvan Kiss <neon60@gmail.com>

* Apply suggestions from code review

Co-authored-by: Istvan Kiss <neon60@gmail.com>

fix pytorch page selector

fix jax page selector

update 'mixed graphics and compute' wording

update jax install

Remove duplicate trademark symbols after first occurrence

First use of Instinct™, Radeon™, and Ryzen™ is in the release
highlights intro paragraph. All subsequent uses are now plain text.

Apply suggestions from code review

Adding Leo's comments

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

Add rocDecode and rocJPEG highlight and rocWMMA resolved issue

update jax and vllm

removed redundant text in resolved issues

Add rocDecode and rocJPEG highlight and rocWMMA resolved issue

Update docs/about/release-notes.md

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

Fix component pages (#767)

* Fix component pages

fix release notes inbox kernel driver for ryzen

fix virtualization hl

redundant "fixed" in every bullet

Change TheRock to ROCm in rocDecode/rocJPEG highlight heading

update `sudo dnf update` to version specific for OL

rocm -> rocm core sdk

clarify jax env var step

fix typo

Remove redundant TheRock delivery note from rocDecode changelog

Remove MI350 content (#768)

* Remove MI350 content

* Remove AMD GPU driver link

* Update component pages

add LD_LIBRARY_PATH note for jax

fix oracle linux typo

remove install sys libs from jax page

specify jax version on page

fix headings in fw pages

amdgpu-install: remove graphics for ryzen

Add AMD SMI feature highlights section

add links in compo table to changelogs

Remove GPU targets not listed in supported hardware table

Removed from component changelogs:
- gfx1250 (unannounced — MI400/CDNA 5, not in 7.13 hardware table)
- gfx90c (Ryzen APU, not in 7.13 hardware table)
- gfx1153 (not in 7.13 hardware table)
- MI350P (not in 7.13 hardware table)

To be restored if confirmed by hardware PM.

Restore MI350P reference in ROCm Compute Profiler changelog

Link AMD SMI highlights to component changelogs section

Update ROCprofiler-SDK changelog with consolidated 7.13 entries

Port changelog updates from rocm-systems PR #5825:
- Add KFD event tracing, multi-pass counter collection, PC sampling
- Add Removed and Resolved Issues sections
- Fix redundancies, casing, punctuation

move ryzen ai 9 365 to gfx1150

add Ryzen AI 9 HX PRO 375

Apply suggestion from @yugang-amd

Co-authored-by: yugang-amd <yugang.wang@amd.com>

Update rocWMMA version to 2.2.1

New version confirmed by Yugang.

remove selector from rocm packages page

Restore gfx1250 and gfx90c entries in hipBLAS and rocBLAS changelogs

Reverts partial removal from ec26df2. gfx1250 and gfx90c confirmed
for inclusion; gfx1153 and MI350P remain excluded pending hardware PM
confirmation.

Add ROCm Runfile Installer updates to release highlights

Added Installation section with Runfile Installer 7.13 feature summary,
sourced from Jeffrey Novotny. Placed after Platform and hardware
support.

fix

reorganize some images closer to source doc

remove underlines in diagram

Redesign components table support column and update RDC version

- Rename "Support" column to "Supported platforms"
- Replace verbose OS:/GPU: bold labels with compact inline format
  using middot separators and slash grouping
- Drop redundant "only" from support descriptors
- Update RDC version to 1.3.0
- All original support data preserved; no information removed

add versions to components in compat

remove ai playbooks

7130 known issue batch3 (#765)

* Review feedback added

* ROCM-21706 Known issue added

* ASAN issue added

* Known issue added

* PyTorch issue added

* Known issues updated

* Known issue for ASAN added

* Known issue added

* Known issue added

* Minor change

* Leo's review feedback incorporared

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

---------

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

fix selectors

clean up selectors

fix compat mat

update vllm version to 0.19.1 in highlights

update selector padding and colors

add hover effect to page dropdown

fix component versions, move installer section, and style cleanup

- Correct versions: hipSOLVER 3.4.0, MIOpen 3.5.1, rocSOLVER 3.34.0, CK
1.2.0
- Move Runfile Installer updates out of highlights into own section
- Collapse double-spaced lists, fix punctuation, hyphenate closed-source
- Fix rocminfo and ROCprofiler-SDK changelog heading format

editorial sweep: highlights, changelog, and link fixes

- Merge two Composable Kernel highlight sections into one
- Fix rocDecode/rocJPEG platform support to include Ryzen AI
- Update external doc links from /en/latest/ to /en/docs-7.13/
- Remove duplicate API list in AMD SMI changelog
- Normalize changelog headings (Resolved issues, Optimized)
- Fix typos and formatting (ROCprofiler-SDK casing, hyphenation,
fragments)
- Align vLLM/SGLang wording

clean up archived vllm pages

update wordlist

add mi350p

add ryzen pro 200 series

add ryzen ai PRO 400

fix

add ryzen ai PRO 7 / 5 (gfx1152)

rm gpus not listed in go/no-go

Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

update wordlist and fix typos

improve toc

address linting issues

address markdown linting

fix md linting issues

fix rest of linting issues

conf: reenable intersphinx fetching

update vllm and sglang toc text

fix vllm selector

editorial: fix version string and ASAN sentence in release notes

- "ROCm 7.13" → "ROCm 7.13.0" (two occurrences)
- Reword ASAN intro sentence for directness

docs: add AMD SMI (BM) 26.4.0 changelog entry and fix anchor link

docs: add RDC 1.3.0 changelog entry and anchor link

docs: add RDC release highlight to 7.13 release notes

Add ROCm Data Center Tool (RDC) entry under the Libraries section,
documenting its addition to the ROCm Core SDK for Linux with
AMD Instinct GPUs.

Ref: ROCM-21708, AIROCDOC-3688

7.13.0 Known Issues Batch 4 (#775)

* LLVM issue added

* Minor spacing issue

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Space fixed

* Linting error fixed

---------

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

docs: remove re-attach highlight from release notes

docs: remove consolidated changelog note from release notes

The consolidated changelog reference is not applicable to the
ROCm Core SDK release notes.

docs: add RCCL GDA alltoall highlight and fix version references

- Add RCCL GDA-based alltoall via rocSHMEM integration highlight
(ROCM-2288)
- Fix ROCm 7.12 → 7.12.0 version references for consistency

docs: update Composable Kernel version to 1.3.0

install: update runfile quick start cmd

add pytorch install commands for gfx1030 and and gfx1152

Fix RDNA3.5 mention and update ROCm Programming page (#774)

* Fix RDNA3.5 mention

* Update HIP Programming Guide

* Update HIP Programming Guide

* Update HIP Programming Guide

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Reorg reference section

* Update HIP Programming

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* WIP

* WIP

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* WIP

* Update RDNA2 system optimization page

* Update RDNA3.5

* WIP

* WIP

* WIP

* WIP

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Update HIP Programming

* WIP

* Update HIP Programming

* Update RDNA2 system optimization page

* Update RDNA2 system optimization page

* Update HIP Programming

* Update ROCm programming title

---------

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

docs: add transition guide to TOC and docs

docs: move transition guide from conceptual to about

docs: fill empty cells in summary table

docs: fill remaining empty cells in summary table

Update transition-guide-TheRock.md

Update transition-guide-TheRock.md

Update transition-guide-TheRock.md

Update transition-guide-TheRock.md

Update transition-guide-TheRock.md

Update transition-guide-TheRock.md

Update transition-guide-TheRock.md

Update transition-guide-TheRock.md

Update transition-guide-TheRock.md

Update transition-guide-TheRock.md

docs: convert remaining tables to HTML for consistent styling

docs: fix inline code styling in HTML tables

js: fix query param resolution

js: fix selected content resolution

docs: update ROCm Compute Profiler highlights to match AIPROFCOMP-493

- Correct product branding (AMD Ryzen AI Max 300 series processors)
- Narrow scope from RDNA to RDNA 3.5 devices
- Add roofline limitation note
- Add roofline.csv detail for profile mode

docs: clarify Core SDK package annotations in transition guide

- Add legacy package names for rocDecode, rocJPEG, and RDC
- Change "new" to "newly included" for clarity

docs: scope Compute Profiler RDNA 3.5 support to Ryzen AI Max 300

Per dev feedback, support is gfx1151 only — not all RDNA 3.5 devices.

docs: remove analysis mode dependency mention per dev feedback

docs: restore analysis mode dependency note per dev clarification

docs: consolidate planned components under future releases

Per dev feedback, hipfort, rocALUTION, rocPyDecode, rocAL, and MIVisionX
will release with ROCm-Extras, not separately.

docs(jax): add jax version selector

docs(jax): use `--index-url` instead of `--extra-index-url`

Update 200-install.rst

Update 200-install.rst

docs(jax): update install snippet formatting

docs(pyt): add pytorch version selector

docs: remove rocm packages page

Revert "Update 200-install.rst"

This reverts commit 291c1fe3747227372f0ac183447ff43561fb6ed6.

Revert "Update 200-install.rst"

This reverts commit babcaeb48582e4e8168073b7fcb70b127814eace.

7.13.0 Known issues updated Batch 5 (#776)

* Known issues updated

* Review feedback added

* Known Issue added

* Feedback incoporated

* Known issue added

* hipDF known issues removed

* Apply suggestions from code review

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

* Review feedback added

* PLDM updated

---------

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

docs: add associated packages column to meta packages table

- Add 4th column listing associated meta packages and packages
- Remove amdrocm-opencl7.13 row (not a meta package)

docs: add cross-reference link to transition guide from meta packages
table

docs(jax): fix env var order for clarity

docs: fix meta packages table formatting with line blocks

docs: add consistent package/meta package labels to all table rows

docs: remove deleted rocm-packages reference from conf.py

docs: fix spacing in meta packages table for consistent formatting

docs: move transition guide cross-reference into core-dev table row

docs: fix transition guide cross-reference to target specific section

docs(vllm): update with addition skus and selectors

docs: move cross-reference link back to after meta packages table

css: add hover highlight to dropdown selector

docs: fix cross-reference to transition guide using Sphinx ref label

docs(jax/pyt): show amdgpu prereq for instinct/radeon only

docs: merge use case and packages columns in meta packages table

docs(vllm): fix vllm version key

docs: clean up and fmt

docs(vllm): make amdgpu prereq conditional

fix

docs: add Windows package availability note to transition guide

docs: wrap Windows package note in admonition block

docs: fix subproject doc links from docs-7.13 to docs-7.13.0

docs: fix incorrect component names in transition guide packages table

- rocm-systems → rocprofiler-systems
- rocm-compute → rocprofiler-compute
- tracer → roctracer

fix mi350p fields in compat page

Clarify ROCm Core SDK 7.13.0 as a preview release

Updated the description of ROCm Core SDK 7.13.0 to reflect its status as
a preview release.

update extras urls in toc

Update ROCm documentation theme options

Update conf.py

rocThrust environment variable changelog removed (#778)

clean up wording

docs(install): remove uninstall step 3

docs(jax): remove LD_LIBRARY_PATH and AMD_COMGR_NAMESPACE workarounds

docs(sglang): remove and keep in exclude dir/

docs(vllm): fix missing gfx1150

docs: add technology preview messaging to release notes intro

Aligns 7.13 release notes with the preview release stream framing
used in prior releases (7.9-7.12). Updates the admonition from note
to important and consolidates prior version links into a single
release history link.

remove SGLang from release notes, compat

docs: fix missing article in transition guide intro

docs: editorial fixes in transition guide

Standardize note format, fix column placement for newly included
annotation, replace em dashes with double dashes, and minor copy edits.

docs: editorial fixes in transition guide

Standardize note format, fix line wrap, clarify parentheticals in
package tables, add deprecation link for ROCm SMI, and minor copy edits.

remove tb and rbt

update docker pull tag to 0.19.1 vllm

7.13.0 Known Issues Batch 6 (#779)

* dev/devel issue added

* Feedback incorporated

* Review feedback added

* Update docs/about/release-notes.md

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

---------

Co-authored-by: Leo Paoletti
<164940351+lpaoletti@users.noreply.github.com>

docs(vllm): update tags and whl urls

docs(vllm): add miising pytorch install for gfx110X

docs(vllm): indentation fix

docs: clean up links

docs: get rid of system-setup for inference for now

docs: update GA date to 5-15

docs: convert component pages to table format with working hyperlinks

Replace dead :doc: cross-references with versioned URLs to component
documentation. Convert category pages from bullet lists to list-table
format for scannability.

docs: switch component tables from list-table to matrix directive

docs: update runfile installer to 7.13.0-2

docs(vllm): fix torch version

docs: remove contributing docs for now

docs: fix ai developer hub link

docs: reorder resolved issues

docs: fix links and sphinx warnings

update GA date in versions list

docs: fix the rest of the sphinx warnings

docs(xdit): update to 26.5

docs: update .wordlist.txt

docs(jax): update 0.8.2 jaxlib instructions

docs(jax): rmove LD_LIBRARY_PATH note

docs(install): update multi-arch repo path for pkgman

fix install

docs(install): fix debian multi-arch

docs: revert meta packages table to original 3-column layout

docs(install): update amdgpu-install url

PLDM table updated (#781)

Revert "docs: switch component tables from list-table to matrix
directive"

This reverts commit 774dcf59d97c4e59c67db13899aed4f6efbd210d.

Revert "docs: convert component pages to table format with working
hyperlinks"

This reverts commit 04b3aa9378c57549c44fa2bee43a5c288fd2dd6b.

docs: update toc for extras

update mi350p ifwi in compat matrix

fix rowspan

reorder toc

docs(post-install): update amd-smi and rocminfo sample outputs

docs: update wordlist

docs: remove Windows package note from transition guide

docs(amdgpu-install): remove mention of amdgpu-install doc

docs: remove unsupported arch rows and fix extras path and library name
in transition guide

docs(amdgpu-install): add uninstall step to rm repos

docs: add ROCdbgapi and ROCr Debug Agent to amdrocm-debugger package
contents

docs(install): fix missing zypper install for multi-arch

PLDM version, link updated (#782)

docs(install): fix missing gfx1030 gfx1152 tarball

docs(install): fix

fix

fix

docs: add missing Radeon RX 9070 GRE link in hardware support table

fix unreliable javascript

docs(install): OEM kernel ub2404 only

docs(isntall): fix incorrect windows nesting

fix extra backslash

add mi350p to vllm

fix

make values explicit and update 31.30.0 driver url to preview

add new vllm docker tags

make oem kernel prereq 24.04 only

remove amdgpu-install

remove graphics note for radeon/ryzen

fix
2026-05-15 18:31:19 -04:00
Peter Park f45bd5a570 [docs/7.12.0] Update docs for 7.12 preview release
build: update RTD env config

docs(release): update virtualization support

docs: reorganize custom extensions

add tom-select lib

[docs/7.12.0] Update documentation for TheRock preview release 7.12.0
(#6072)

docs(release): add gpus and xrefs

docs(release): update release notes for 7.12

wip: selector-toc2 dropdown input

remove selector tiles from install sections

docs(compat): virtualization support

docs(uninstall): remove unneeded headings

js(selector toc2): improve semantic html in toc2 selector dropdowns

fix(js/css): TomSelect

fix(toc2.js): heading query

fix(css): maximize dropdown input widths to prevent stuttering

fix and reorg

bump rocm-docs-core to 1.32.1

[docs/7.12.0] Update documentation for TheRock 7.12.0

Includes related enhancements:
- Improve secondary sidebar display
- Fix install instruction issues
- Add JAX and vLLM
- Reorganize site structure
- Tweak CSS/JS/Py

update versions list

wording

update driver docs links

docs: fix jax instructions

docs: fix vllm instructions

docs: reorganize some files

reorg

[docs/7.12.0] update vllm docker pull tags

fix

[docs/7.12.0] fix install docs

[docs/7.12.0] fix install instructions and conditional sections

[docs/7.12.0] Fix repo.amd.com url for gfx110X

[docs/7.12.0] Add JAX known issue

[docs/7.12.0] Update JAX known issue description

[docs/7.12.0] document `CK_AMD_GPU_GFX*` known issue

[docs/7.12.0] document more known issues

[docs/7.12.0] Remove GPU partitioning support note (#6084)

Revert "[docs/7.12.0] Remove GPU partitioning support note (#6084)"
(#6085)

This reverts commit 6ac90bdf9e.

[docs/7.12.0] Add ROCm Optiq (Beta) release highlight (#6086)

[docs/7.12.0] docs(install): Fix pip install url `gfx120x-all` -->
`gfx120X-all` (#6087)

X needs to be capitalized

[docs/7.12.0] Update install instructions (#6104)

* Link to i=runfile query param in release note

* fix debian version selector for mi300a, mi250x, mi250

* remove sles prereq

* fix bashrc tarball post-install

add uninstall

* fix relative url

* remove `sudo` from user-local env setup

[docs/7.12.0] Add `bash` to docker run cmds (#6107)

[docs/7.12.0] Document workarounds for vLLM installation via pip (#6108)

* [docs/7.12.0] Document workarounds for vLLM

* [docs/7.12.0] Update `amd-smi version` sample output

* update text

update formatting

note that PyT 2.9.1 is required

update

* fix

fix links

rm extra gfx103x sections

[docs/7.12.0] mention nightlies for gpus w/o "official" support (#6147)

fix markup

wording

words

consistency

cleanup

[docs/7.12.0] Update post-install and uninstall with LD_LIBRARY_PATH
workaround + update runfile installer to 7.12.0-2 (#6151)

* docs(install): Update ROCm post-install and uninstall with
LD_LIBRARY_PATH workarounds

* docs(install): Update runfile installer to 7.12.0-2.run
2026-05-08 17:19:52 -04:00
Peter Park 3375a3a668 [docs/7.11.0] Update docs for 7.11 preview release
update ROCM_VERSION in conf

7.11 known issues added

Minor change

docs(RELEASE): supported OSes and hw

docs(release.md): update hw support and os support

js(selector-toc): remove unused code

docs(index): update rocm ontology diagram

add TM symbols

docs(RELEASE): complete virtualization support tbl

docs(toc): add rocm-examples to toc

chore: bump rocm-docs-core to 1.31.3

chore: update version histor page

docs(RELEASE): add oses

docs(RELEASE): add virtu sup

fix

docs: update RELEASE.md

docs(RELEASE): clean up

py(selector): allow percentage widths

docs(compat): update system-instinct table

docs: finalize components lists

docs(compat): update system-radeon-pro

docs(compat): update system-radeon

docs: compat

docs(RELEASE): clean up

docs(RELEASE): update tables

docs(compat): add virtu sup and fix stuff

docs(RELEASE): update AMD GPU Driver vers

dcos(compat): add missing mi2xx options

docs(RELEASE): update firmware for instinct

docs(RELEASE): fix xref and fmt

docs(rocm-ontology diagram): fix sideways tm

fix

oops, fix

docs: remove ROCgdb

docs(compat): add rhel 8.10 to mi35x

docs: clean up some wording

docs(install): update selector

docs(install/compat): fix selector data

docs: fix

js: fix reconcile selections

wip: install

wip: install

js: add URLSearchParams

wip: install

js: inline page-specific js

docs(compat): add missing rocky linux to mi300x

docs(install): windows adrenalin prereq

wip: install

docs(compat): virtu sup link

docs: windows tar

docs(compat/install): add missing gfx120x, gfx103x, and gfx110x gpus

docs(RELEASE): add missing gfx120x, gfx103x, and gfx110x gpus

docs(install): fix selected content

docs(install): windows tar

rm jira notes

docs(comfyui): update

docs(comfyui): update path

docs: remove hipDNN

chore(conf.py): exclude `**/includes/**` to build only when included

docs(index): fix diagram colors

Update RELEASE.md

remove fn

docs: remove gfx1030 cards

docs: remove windows for gfx12

js: add localStorage to selector for persistence

docs(RELEASE): fix gpu list in tables

js(primary toc install headings): don't reload page if already on the
install page

docs(img): add oses to ontology diagram

docs(rocgdb): add it back + document known issue

docs(comfyui): update selector

docs(RELEASE): fix

docs(compat): fix - add rocky linux for mi300a

docs(install): sles - remove --gpg-auto-import-keys

py(selector): add static assets only if exists

js(selector): clear URLSearchParams if page doesn't have selector

docs: note GIM driver 8.7.0K for virtualization on instinct

add link to gim docs

docs: reorder list of components

docs(install): add oem kernel prereq

docs: clean up diagrams

fix link

docs(RELEASE): update firmware explanation

docs: update

chore: clean up extra files

chore: linting errors

docs(compat): add iGPU to ryzen cards in compat matrix

docs(install): add install libatomic1

docs(install): clean up

fix

docs(install): add meta packages table

docs(compat): remove amdgpu 7.0.3 and 6.4.2

docs(install): uninstall add meta package note

blurb

docs(RELEASE): apex known issue

docs: update components lists in RELEASE and compat matrix

docs(install): clean up meta packages sections

docs(compat): update fw version for mi300x

update known issues

docs: remove minor version from OL and Rocky

docs: fix windows whl urls

docs(install): update rocminfo and amd-smi version example outputs

docs(RELEASE): add known issues

docs: fix linux components list

docs(compat): make igpu heading consistent

Add rocm-examples known issue

Co-authored-by: Istvan Kiss <istvan.kiss@amd.com>

docs(conf): update release date

docs(install): update windows install instructions

update

docs(RELEASE): update known issues

hipblaslt matmul known issue

update known issues

add hipify-clang known issue

js(selector): fix syncStateToURL

update amdgpu driver versions and known issues

update known issues

fix

fix

fix install urls

update RELEASE.md

fix

update for consistency

add known issue

add rccl known issue

update components table

fix amdgpu versions

update prereqs intro

docs(conf): turn of toc_exclude_missing (#5957)

[docs/7.11.0] Fix some release notes documentation and remove unneeded
SLES packages (#5960)

* fix release date and known issue

* add llama.cpp known issue and fix link to amdgpu 31.10.0

* docs: minor fixes

* fx

* clean up known issues

* clean up

[docs/7.11.0] Add minor corrections (#5961)

* fix comfyui linux/windows options

* rm minor version from oracle linux ver

update LD_LIBRARY_PATH config (#5964)

[docs/7.11.0] Add llama.cpp known issue (#5962)

[docs/7.11.0] fix `cd C:\TheRock` (#5965)

* [docs/7.11.0] fix `cd C:\TheRock`

* fix description

[docs/7.11.0] Add minor corrections (#5976)

* add hipinfo to release notes

* fix typo

* fix rccl links

* fix warning: pygments lexer `cmd` unknown

* fix table row alignment

[docs/7.11.0] Update known issues #5975

update ROCM_VERSION

wip: add compat/install data
2026-05-08 17:19:52 -04:00
Peter Park f6c4835da7 [docs/7.10.0] Update docs for 7.10 preview release
fix conf.py (#5762)

[docs/7.10.0] selector.py: Make ids even more unique for selected
content (#5765)

[docs/7.10.0] Fix typo in venv command and installation prerequisites
(#5766)

* fmt

* Fix installation prerequisites and source venv typo

clean up ./.venv/... --> .venv/... (#5767)

[docs/7.10.0] selector: responsive css for narrow viewports (#5769)

Fix element overflow on narrow/mobile screens.

Fix incorrect warning in matrix.py extension

[docs/7.10.0] Radeon cards: suggest amdgpu over RSL driver (#5770)

[docs/7.10] post-install - improve exports (#5783)
2026-05-08 17:19:52 -04:00
Peter Park be0b93b8c6 [docs/7.9.0] Add docs for 7.9.0 preview release
Add release notes

Add install instructions

Add PyTorch + ComfyUI instructions

Add custom selector directives

Add JS and CSS for selector

Add custom icon directive and utils

Clean up conf.py

update PLDM bundle version for MI355

Add "preview" to headings (#5545)

Add custom version history (#5551)

[docs/7.9.0] Fix GPU marketing names in 7.9.0 release.md and
compatibility matrix / Add SD3.5 ComfyUI example (#5552)

* Add SD3.5 example to comfyui doc

* Fix Ryzen AI Max (PRO) SKU names

* Add names in multi line format

[docs/7.9.0] Add build from source overview page / Point to
`therocm-7.9.0` in components list (#5546)

* Update links to components to point the `therock-7.9.0` ref

* Add build from source page

* lint: fix caps and update .wordlist.txt

* add link to "development manuals" list

* add links to TheRock's development guide and fix step 4

* wording and fmt

* fix spacing

* fix fmt

* Fix documentation linting errors

* fix spacing

[docs/7.9.0] Note rocprofiler-sdk is Instinct only. Reorg some files to
match `docs/7.0.x`. (#5563)

* move versions.md and compat-matrix to match prod; note rocprofiler-sdk
is instinct only

* update href in versions.md

[docs/7.9.0] Use "generic" rocm-docs-core theme (#5568)

* use "generic" rocm-docs-core theme to tweak header

* restore "nav_secondary_items"

[docs/7.9.0] Fix xref in Ubuntu prerequsites and RST heading overline
(#5569)

update rocm-docs-core to 1.29.0

(cherry picked from commit 39de859bd1)

update rocm-docs-core to show preview banner

add preview announcement

[docs/7.9.0] Add xDiT diffusion inference doc (#5676)

[docs/7.9.0] Fix rocm-cmake github link due to non-existent tag (#5684)

[docs/7.9.0] Update banner msg (#5704)

update
2026-05-08 17:19:43 -04:00
420 changed files with 20029 additions and 11580 deletions
+4 -4
View File
@@ -1,8 +1,8 @@
* @ROCm/rocm-documentation
* @amd-aakash @jlgreathouse @samjwu @yhuiYH @ROCm/rocm-documentation
# Documentation files
docs/ @ROCm/rocm-documentation
*.md @ROCm/rocm-documentation
*.rst @ROCm/rocm-documentation
docs/ @amd-aakash @jlgreathouse @samjwu @yhuiYH @ROCm/rocm-documentation
*.md @amd-aakash @jlgreathouse @samjwu @yhuiYH @ROCm/rocm-documentation
*.rst @amd-aakash @jlgreathouse @samjwu @yhuiYH @ROCm/rocm-documentation
# External CI
/.azuredevops/ @ROCm/external-ci
tools/rocm-build/ @ROCm/rocm-devops
-9
View File
@@ -1,9 +0,0 @@
<svg xmlns="http://www.w3.org/2000/svg" width="1280" height="640" viewBox="0 0 1280 640" role="img" aria-label="ROCm">
<rect width="1280" height="640" fill="#0A0A0A"/>
<svg x="96" y="215" width="210" height="210" viewBox="0 0 67 67"><path d="M22.21 67V44.6369H0V67H22.21Z" fill="#fff"/><path d="M66.7038 22.3184H22.2534L0.0878906 44.6367H44.4634L66.7038 22.3184Z" fill="#fff"/><path d="M22.21 0H0V22.3184H22.21V0Z" fill="#fff"/><path d="M66.7198 0H44.5098V22.3184H66.7198V0Z" fill="#fff"/><path d="M66.7198 67V44.6369H44.5098V67H66.7198Z" fill="#fff"/></svg>
<text x="378" y="276" font-family="Inter,system-ui,-apple-system,sans-serif" font-size="78" font-weight="800" letter-spacing="-2" fill="#ffffff">ROCm</text>
<text x="378" y="322" font-family="Inter,system-ui,sans-serif" font-size="30" fill="#ffffff" opacity=".66">AMD ROCm™ Software - GitHub Home</text>
<rect x="378" y="338" width="806" height="3" rx="1.5" fill="#ffffff" opacity=".9"/>
<text x="378" y="390" font-family="Inter,system-ui,sans-serif" font-size="24" font-weight="600" fill="#ffffff" opacity=".5">github.com/hanzoai</text>
<text x="1184" y="390" text-anchor="end" font-family="Inter,system-ui,sans-serif" font-size="24" font-weight="600" fill="#ffffff" opacity=".5">hanzo.ai</text>
</svg>

Before

Width:  |  Height:  |  Size: 1.2 KiB

-1
View File
@@ -17,6 +17,5 @@ __pycache__/
# avoid duplicating contributing.md due to conf.py
docs/contribute/index.md
docs/about/release-notes.md
docs/release/changelog.md
.claude/settings.local.json
+10 -12
View File
@@ -4,23 +4,21 @@
version: 2
sphinx:
configuration: docs/conf.py
configuration: docs/conf.py
formats: [htmlzip]
formats: []
python:
install:
- requirements: docs/sphinx/requirements.txt
install:
- requirements: docs/sphinx/requirements.txt
build:
os: ubuntu-22.04
tools:
python: "3.10"
apt_packages:
- "doxygen"
- "gfortran" # For pre-processing fortran sources
- "graphviz" # For dot graphs in doxygen
os: ubuntu-24.04
tools:
python: "3.10"
search:
ignore:
- "**/previous-versions/**"
- "**/archive/**"
- "**/include/**"
- "**/redirect/**"
+189 -54
View File
@@ -1,16 +1,9 @@
AAC
ABI
ACE
ACEs
ACS
AITER
ALU
AllReduce
AllToAll
AGPR
AGPRs
AITER
ALU
AMD
AMDGPU
AMDGPUs
@@ -31,6 +24,7 @@ ASICs
ASan
ASm
ATI
ATT
AccVGPR
AccVGPRs
AddressSanitizer
@@ -41,10 +35,13 @@ Arb
Async
Autocast
BARs
BDF
BKC
BLAS
BLASLt
BMC
BNXT
BRCM
BSR
BabelStream
Backported
BatchNorm
@@ -68,11 +65,15 @@ CMakeLists
CMakePackage
CNP
CP
CPACK
CPC
CPF
CPP
CPU
CPUs
CPX
CPack
CQ
CSC
CSDATA
CSE
@@ -91,11 +92,13 @@ ChatGPT
Cholesky
CoRR
Codespaces
ComfyUI
Commitizen
CommonMark
Concretized
Conda
ConnectX
Conv
CountOnes
Cron
CuPy
@@ -110,13 +113,14 @@ DIMM
DKMS
DL
DMA
DMC
DNN
DNNL
DOCA
DOMContentLoaded
DPM
DPX
DRI
DSA
DSCP
DW
DWORD
@@ -137,13 +141,16 @@ Disaggregated
Dockerfile
Dockerized
Doxygen
Dyninst
ELMo
ENDPGM
EP
EPYC
ESMI
ESXi
EoS
FBGEMM
FFN
FFT
FFTs
FFmpeg
@@ -151,9 +158,10 @@ FHS
FIFOs
FIXME
FMA
FNUZ
FMHA
FP
FX
FetchContent
FiLM
Filesystem
FindDb
@@ -161,6 +169,7 @@ Flang
FlashAttention
FlashInfer
FlashInfers
Flatmm
FluxBenchmark
Fortran
Fuyu
@@ -173,12 +182,15 @@ GCD
GCDs
GCN
GCNN
GDA
GDB
GDDR
GDR
GDS
GEMM
GEMMs
GETRS
GETRS
GFLOPS
GFXIP
GFortran
@@ -186,15 +198,12 @@ GGUF
GID
GIM
GL
Glibc
GLM
GIM
GL
GLXT
GMI
GNN
GNNs
GPG
GPGPU
GPR
GPT
GPU
@@ -204,6 +213,7 @@ GPUVM
GPUs
GRBM
GRE
GSU
GTT
Gbps
Gemma
@@ -214,15 +224,18 @@ GitHub
Gitpod
Glibc
Gloo
Gluon
GraphAPI
GraphBolt
GraphSage
Graphbolt
HBM
HCA
HEVC
HGX
HIPCC
HIPExtension
HIPIFY
HIPOCProgram
HIPification
HIPify
HLO
@@ -234,27 +247,30 @@ HSA
HW
HWE
HWS
HX
Haswell
Higgs
Huggingface
Hunyuan
HunyuanVideo
HybridEngine
Hyperparameters
IB
IC
ICD
InternVL
ICT
ICV
IDE
IDEs
IFWI
ILP
ILU
IMDb
IOMMU
IOP
IOPM
IOPS
IOV
IPC
IPU
IPs
IRQ
ISA
@@ -271,7 +287,7 @@ Intersphinx
Intra
Ioffe
JAX's
JAXLIB
JPG
JSON
Jinja
Jupyter
@@ -281,23 +297,21 @@ KMD
KV
KVM
Karpathy's
Kimi
KiB
Kineto
Keras
Khronos
KiB
Kineto
LAPACK
LASYF
LCLK
LDS
LLM
LLMs
LLVM
LLaMA
LM
LPDDR
LRU
LSE
LSAN
LSTMs
LSan
@@ -317,6 +331,7 @@ MIOpenGEMM
MIVisionX
MLA
MLM
MLX
MMA
MMIO
MMIOH
@@ -324,14 +339,16 @@ MMU
MNIST
MPI
MPT
MSI
MSVC
MTP
MTU
MVAPICH
MVFFR
MX
MXFP
Makefile
Makefiles
ManyLinux
Matplotlib
Matrox
MaxText
@@ -342,6 +359,7 @@ Mellanox
Mellanox's
Meta's
MiB
Microscaling
Miniconda
MirroredStrategy
Mixtral
@@ -351,9 +369,6 @@ Mooncake
MosaicML
Mpops
Multicore
Multimodal
multimodal
multihost
Multithreaded
MyEnvironment
MyST
@@ -362,6 +377,7 @@ NBIO
NBIOs
NCCL
NCF
NCHW
NCS
NFS
NIC
@@ -372,6 +388,8 @@ NN
NOP
NPKit
NPS
NPVT
NSIS
NSP
NUMA
NVCC
@@ -379,7 +397,6 @@ NVIDIA
NVLink
NVPTX
NaN
NaNs
Nano
Navi
NoReturn
@@ -397,16 +414,17 @@ OMPI
OMPT
OMPX
ONNX
OSL
OOM
OSS
OSU
OOM
OTF
OpenCL
OpenCV
OpenFabrics
OpenGL
OpenMP
OpenMPI
OpenSHMEM
OpenSSL
OpenVX
OpenXLA
@@ -417,11 +435,15 @@ PCI
PCIe
PEFT
PEQT
PEs
PID
PIL
PILImage
PJRT
PLDM
POR
POSIX
POTF
POTRF
PRNG
PRs
PSID
@@ -435,12 +457,10 @@ Pensando
PerfDb
Perfetto
PipelineParallel
Pipelining
PnP
Pollara
PowerEdge
PowerShell
Preshuffled
Pretrained
Pretraining
Primus
@@ -448,12 +468,15 @@ Profiler's
PyPi
PyTorch
Pytest
QMCPACK
QPS
QPX
Qcycles
QoS
Qwen
RAII
RAS
RBT
RCCL
RDC
RDC's
@@ -478,10 +501,14 @@ ROCm
ROCmCC
ROCmSoftwarePlatform
ROCmValidationSuite
ROCprof
ROCprofiler
ROCr
ROCtx
RPP
RST
RTC
RTC
RW
Radeon
Radix
@@ -495,10 +522,10 @@ RoCE
Runfile
Ryzen
SALU
safetensors
SBIOS
SCA
SDK
SDKs
SDMA
SDPA
SDRAM
@@ -519,24 +546,35 @@ SMEM
SMFMA
SMI
SMT
SPARSELt
SPI
SPIR
SPX
SQTT
SQs
SRAM
SRAMECC
SVD
SVM
SWE
SYTF
SYTRF
SYTRS
SageAttention
ScaledGEMM
SerDes
Shardy
ShareGPT
Shlens
SiLU
Skylake
Slurm
Softmax
Spack
SplitK
StreamingLLM
Strix
Supermicro
SwiGLU
Szegedy
TCA
TCC
@@ -547,16 +585,23 @@ TCP
TCR
TF
TFLOPS
TGZ
THREADGROUPS
TP
TPS
TPU
TPUs
TPX
TSME
TT
TTM
TTY
TUI
TVM
TagRAM
Tagram
Taichi
Taichi's
TensileLite
TensorBoard
TensorFloat
@@ -575,11 +620,12 @@ TorchVision
TransferBench
TrapStatus
UAC
UBB
UC
UCC
UCX
UE
UI
UEK
UIF
UMC
USM
@@ -588,6 +634,7 @@ UTCL
UTCL
UTIL
UTIL
UUID
UX
UltraChat
Uncached
@@ -602,20 +649,25 @@ VM
VMEM
VMID
VMIDs
VMM
VMWare
VMs
VMware
VRAM
VSIX
VSkipped
Vanhoucke
Vulkan
WDAG
WGP
WGPs
WR
WX
WikiText
Winograd
Wojna
Workgroups
WrW
Writebacks
XCD
XCDs
@@ -634,6 +686,7 @@ YML
YModel
ZeRO
ZenDNN
abquant
accuracies
activations
addEventListener
@@ -644,8 +697,15 @@ alloc
allocatable
allocator
allocators
alltoall
alltoallv
amd
amdclang
amdgpu
amdrocm
amdsmi
api
aqlprofile
async
aten
atmi
@@ -654,12 +714,12 @@ atomics
autogenerated
autograd
autotune
autotuning
avx
awk
az
backend
backends
batchnorm
bb
benchmarked
benchmarking
@@ -668,11 +728,13 @@ bilinear
bitcode
bitsandbytes
bitwise
blas
blit
blockscale
bootloader
bootup
boson
bosons
bottlenecked
br
btn
buildable
@@ -681,25 +743,31 @@ bzip
cTDP
cacheable
carveout
ccl
cd
centos
centric
cgroups
changelog
changelogs
checkpointing
chiplet
cholesky
classList
cmake
cmd
coalescable
codecs
codename
codenamed
collater
comfyui
comgr
compat
completers
composable
composablekernel
concretization
cond
config
configs
conformant
@@ -709,7 +777,10 @@ convolutional
convolves
copyable
cpp
cron
csn
csv
ctypes
cuBLAS
cuDNN
cuFFT
@@ -721,10 +792,10 @@ cuda
cudnn
customizable
customizations
cxx
dGPU
dGPUs
da
dataflows
dataset
datasets
dataspace
@@ -742,12 +813,14 @@ denoise
denoised
denoises
denormalize
depthwise
dequantization
dequantized
dequantizes
deserializers
detections
dev
devel
devicelibs
devsel
dgl
@@ -756,9 +829,13 @@ disagg
disaggregated
disaggregation
disambiguates
discoverability
distro
distros
dkms
dll
dnf
dnn
dropless
dtype
eb
@@ -779,17 +856,22 @@ eth
ethernet
exascale
executables
factorizations
fam
fam
fas
ffmpeg
fft
filesystem
flashinfer
flang
fmha
forEach
foreach
fortran
fp
framebuffer
gRPC
galb
gb
gcc
gdb
gemm
@@ -800,21 +882,27 @@ githooks
github
globals
gnupg
gpt
gpu
granularities
grayscale
gre
gtest
gx
gz
gzip
hardcoded
heterogenous
hipBLAS
hipBLASLt
hipBLASLt's
hipCUB
hipDNN
hipDataType
hipFFT
hipFFTW
hipFORT
hipLIB
hipMemGetAddressRange
hipModules
hipRAND
hipSOLVER
hipSPARSE
@@ -829,8 +917,11 @@ hipfft
hipfort
hipification
hipify
hipinfo
hiprand
hipsolver
hipsparse
hipsparselt
hlist
hostname
hotspotting
@@ -839,6 +930,7 @@ hpp
href
hsa
hsakmt
hx
hyperparameter
hyperparameters
iDRAC
@@ -862,6 +954,7 @@ invariants
invocating
ipo
jax
jpeg
js
json
kdb
@@ -870,9 +963,13 @@ kv
lang
latencies
len
libVA
libdrm
libelf
libfabric
libjpeg
libs
libva
linalg
linearized
linter
@@ -885,15 +982,20 @@ logits
logsumexp
loopback
lossy
lp
lstsq
mW
macOS
matchers
maxtext
megablocks
megatron
microarchitecture
microscaling
migraphx
migratable
milliwatt
milliwatts
miopen
miopengemm
mivisionx
@@ -904,24 +1006,26 @@ mlirmiopen
mtypes
mul
multihost
multimodal
mutex
mvffr
mx
namespace
namespaces
nanoGPT
netplan
noncontiguous
num
numa
numref
ocl
ol
openai
opencl
opencv
openmp
openssl
optimizers
os
oss
oversubscription
pageable
pallas
@@ -931,15 +1035,16 @@ param
parameterization
params
passthrough
pb
pc
pe
perf
perfcounter
performant
perl
piecewise
pipelined
pipelining
pkgman
pmc
ppt
pragma
pre
prebuild
@@ -977,6 +1082,7 @@ querySelectorAll
queueing
qwen
radeon
radix
rc
rccl
rdc
@@ -985,10 +1091,12 @@ reStructuredText
reachability
recommender
recommenders
redhat
redirections
refactorization
reformats
reinforcememt
relocations
repo
repos
representativeness
@@ -1012,6 +1120,7 @@ rocMLIR
rocPRIM
rocPyDecode
rocRAND
rocSHMEM
rocSOLVER
rocSPARSE
rocThrust
@@ -1019,26 +1128,38 @@ rocWMMA
rocalution
rocblas
rocclr
rocdecode
rocfft
rocgdb
rocjpeg
rocm
rocminfo
rocpd
rocprim
rocprof
rocprofiler
rocprofv
rocr
rocrand
rocshmem
rocsolver
rocsparse
rocthrust
roctracer
rocwmma
roofline
rst
runfile
runtime
runtimes
rx
ryzen
sL
scalability
scalable
scipy
sdk
seccomp
seealso
selectattr
selectedTag
@@ -1071,11 +1192,15 @@ submatrix
submodule
submodules
subnet
subnets
supercomputing
suse
symlink
symlinks
sys
syscall
syscalls
sysdeps
sysfs
tabindex
targetContainer
td
@@ -1083,6 +1208,7 @@ tensorfloat
tf
th
threadgroups
toc
tokenization
tokenize
tokenized
@@ -1090,6 +1216,7 @@ tokenizer
tokenizes
toolchain
toolchains
toolkits
toolset
toolsets
topk
@@ -1102,9 +1229,14 @@ torchvision
tp
tqdm
tracebacks
tridiagonal
txt
typedef
uarch
ubuntu
uclk
ud
udev
uncacheable
uncached
uncorrectable
@@ -1112,12 +1244,13 @@ underoptimized
unfused
unhandled
uninstallation
unmap
unmapped
unpadded
unprefixed
unsqueeze
unstacking
unswitching
unswizzled
untar
untrusted
untuned
unwindowed
@@ -1132,6 +1265,7 @@ vectorize
vectorized
vectorizer
vectorizes
ver
verl
verl's
virtualize
@@ -1141,7 +1275,6 @@ vllm
voxel
walkthrough
walkthroughs
warmup
watchpoints
wavefront
wavefronts
@@ -1159,7 +1292,9 @@ xPacked
xargs
xcc
xdit
xplane
xe
xt
xtx
xz
yaml
ysvmadyb
+1 -1
View File
@@ -1,6 +1,6 @@
MIT License
Copyright (c) 2023 - 2025 Advanced Micro Devices, Inc. All rights reserved.
Copyright (c) 2023 - 2026 Advanced Micro Devices, Inc. All rights reserved.
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
-18
View File
@@ -1,18 +0,0 @@
Hanzo ROCm
Copyright (c) 2026 Hanzo AI, Inc.
This product includes software from AMD ROCm (https://github.com/ROCm/ROCm),
licensed under the MIT License:
Copyright (c) 2023 - 2025 Advanced Micro Devices, Inc. All rights reserved.
This repository is the AMD ROCm meta/manifest repository: documentation, build
tooling, and the repo manifest (default.xml). This meta-repository is MIT-licensed
and its upstream MIT license is retained in LICENSE.
The individual ROCm components referenced by the manifest and built by the tooling
carry their own upstream licenses, including the MIT License, the Apache License 2.0,
and the University of Illinois/NCSA Open Source License. Note that the ROCgdb
component (AMD's fork of GNU GDB, packaged by tools/rocm-build/build_rocm-gdb.sh) is
licensed under the GNU General Public License (GPL) — a copyleft license. Each
component is governed by its own license; consult that component's repository.
-2
View File
@@ -1,5 +1,3 @@
<p align="center"><img src=".github/hero.svg" alt="ROCm" width="880"></p>
<div align="center">
<img src="docs/data/amd-rocm-logo.png" width="200px" alt="ROCm logo">
-754
View File
@@ -1,754 +0,0 @@
<!-- Do not edit this file! -->
<!-- This file is autogenerated with -->
<!-- tools/autotag/tag_script.py -->
<!-- Disable lints since this is an auto-generated file. -->
<!-- markdownlint-disable blanks-around-headers -->
<!-- markdownlint-disable no-duplicate-header -->
<!-- markdownlint-disable no-blanks-blockquote -->
<!-- markdownlint-disable ul-indent -->
<!-- markdownlint-disable no-trailing-spaces -->
<!-- markdownlint-disable reference-links-images -->
<!-- markdownlint-disable no-missing-space-atx -->
<!-- spellcheck-disable -->
# ROCm 7.2.4 release notes
ROCm 7.2.4 is a quality release focused on performance and stability fixes for AI inference workloads on AMD Instinct GPUs.
## Release highlights
The following are the notable changes in ROCm 7.2.4.
### Reduced hipGraphLaunch latency for multi-list graphs
The HIP runtime's graph dispatch mechanism has been optimized, reducing launch latency for workloads using `hipGraphLaunch` with multi-list graph topologies.
### Fixed H2D memory copy latency regression in CPX mode
HIP runtime synchronization behavior has been corrected on AMD Instinct MI300 Series GPUs in CPX mode, restoring latency to previous levels for inference workloads that run multiple HIP streams with concurrent memory copies.
### Reduced ROCprofiler-SDK profiling overhead
Profiling stability has been improved for vLLM workloads traced with PyTorch `torch.profiler` using the ROCprofiler-SDK backend. The large, sporadic idle gaps that previously appeared between GPU kernels in the trace have been substantially reduced in common configurations, and the traces now more accurately reflect actual runtime behavior. Coverage may vary depending on model and parallelism settings.
### Reduced copy overhead in MIGraphX concat operations
MIGraphX now recognizes ONNX models that concatenate the same tensor multiple times and avoids redundant device-side copies, improving inference throughput at small batch sizes for the affected model class on AMD Instinct MI300X GPUs.
### User space, driver, and firmware dependent changes
The software for AMD Data Center GPU products requires maintaining a hardware
and software stack with interdependencies among the GPU and baseboard
firmware, AMD GPU drivers, and the ROCm user space software. While AMD publishes drivers and ROCm user space components, your server or infrastructure provider publishes the GPU and baseboard firmware by bundling AMDs firmware releases via the AMD Platform Level Data Model (PLDM) bundle, which includes the Integrated Firmware Image (IFWI).
GPU and baseboard firmware versioning might differ across GPU families.
<div class="pst-scrollable-table-container">
<table class="table table--middle-left">
<thead>
<tr>
<th class="head">
<p>ROCm Version</p>
</th>
<th class="head">
<p>GPU</p>
</th>
<th class="head">
<p>PLDM Bundle (Firmware)</p>
</th>
<th class="head">
<p>AMD GPU Driver (amdgpu)</p>
</th>
<th class="head">
<p>AMD GPU <br>
Virtualization Driver (GIM)</p>
</th>
</tr>
</thead>
<style>
tbody#virtualization-support-instinct tr:last-child {
border-bottom: 2px solid var(--pst-color-primary);
}
</style>
<tr>
<td rowspan="9" style="vertical-align: middle;">ROCm 7.2.4</td>
<td>MI355X</td>
<td>
01.26.00.02<br>
01.25.17.07<br>
01.25.16.03
</td>
<td>
30.30.x where x (0-4)<br>
30.20.x where x (0-1)<br>
30.10.x where x (0-2)
</td>
<td rowspan="3" style="vertical-align: middle;">8.7.1.K</td>
</tr>
<tr>
<td>MI350X</td>
<td>
01.26.00.02<br>
01.25.17.07<br>
01.25.16.03
</td>
<td>
30.30.x where x (0-4)<br>
30.20.x where x (0-1)<br>
30.10.x where x (0-2)
</td>
</tr>
<tr>
<td>MI325X<a href="#footnote1"><sup>[1]</sup></a></td>
<td>
01.25.06.08<br>
01.25.04.02
</td>
<td>30.30.x where x (0-4)<br>
30.20.x where x (0-1)<a href="#footnote1"><sup>[1]</sup></a><br>
30.10.x where x (0-2)<br>
6.4.z where z (0-3)<br>
6.3.3
</td>
</tr>
<tr>
<td>MI300X<a href="#footnote2"><sup>[2]</sup></a></td>
<td>01.25.06.04<br>
01.25.03.12<br>
01.25.02.04</td>
<td rowspan="6" style="vertical-align: middle;">
30.30.x where x (0-4)<br>
30.20.x where x (0-1)<br>
30.10.x where x (0-2)<br>
6.4.z where z (03)<br>
6.3.3
</td>
<td>8.7.1.K</td>
</tr>
<tr>
<td>MI300A</td>
<td>BKC 26.1</td>
<td rowspan="3" style="vertical-align: middle;">Not Applicable</td>
</tr>
<tr>
<td>MI250X</td>
<td>IFWI 47 (or later)</td>
</tr>
<tr>
<td>MI250</td>
<td>MU5 w/ IFWI 75 (or later)</td>
</tr>
<tr>
<td>MI210</td>
<td>MU5 w/ IFWI 75 (or later)</td>
<td>8.7.1.K</td>
</tr>
<tr>
<td>MI100</td>
<td>VBIOS D3430401-037</td>
<td>Not Applicable</td>
</tr>
</table>
</div>
<p id="footnote1">[1]: For AMD Instinct MI325X KVM SR-IOV users, don't use AMD GPU driver (amdgpu) 30.20.0.</p>
<p id="footnote2">[2]: AMD Instinct MI300X KVM SR-IOV with Multi-VF (8 VF) support requires a compatible firmware BKC bundle, which will be released in the coming months.</p>
```{note}
ROCm 7.2.4 doesn't include any other significant changes or feature additions. For comprehensive changes, new features, and enhancements in ROCm 7.2.3, refer to the [ROCm 7.2.3 release notes](#rocm-7-2-3-release-notes) below.
```
## ROCm 7.2.3 release notes
The release notes provide a summary of notable changes since the previous ROCm release.
- [Release highlights](#id1)
- [Supported hardware, operating system, and virtualization changes](#supported-hardware-operating-system-and-virtualization-changes)
- [User space, driver, and firmware dependent changes](#user-space-driver-and-firmware-dependent-changes)
- [ROCm components versioning](#rocm-components)
- [Detailed component changes](#detailed-component-changes)
- [ROCm known issues](#rocm-known-issues)
- [ROCm upcoming changes](#rocm-upcoming-changes)
```{note}
If youre using AMD Radeon™ GPUs or Ryzen™ for graphics workloads, see the [Use ROCm on Radeon and Ryzen](https://rocm.docs.amd.com/projects/radeon-ryzen/en/latest/index.html) documentation to verify compatibility and system requirements.
```
### Release highlights
The following are notable new features and improvements in ROCm 7.2.3. For changes to individual components, see
[Detailed component changes](#detailed-component-changes).
#### Supported hardware, operating system, and virtualization changes
Hardware, operating system, and virtualization support remains unchanged in this release.
For more information about:
* AMD hardware, see [Supported GPUs (Linux)](https://rocm.docs.amd.com/projects/install-on-linux/en/docs-7.2.3/reference/system-requirements.html#supported-gpus).
* Operating systems, see [Supported operating systems](https://rocm.docs.amd.com/projects/install-on-linux/en/docs-7.2.3/reference/system-requirements.html#supported-operating-systems) and [ROCm installation for Linux](https://rocm.docs.amd.com/projects/install-on-linux/en/docs-7.2.3/).
* Virtualization support, see [Virtualization support](https://rocm.docs.amd.com/projects/install-on-linux/en/docs-7.2.3/reference/system-requirements.html#virtualization-support).
#### User space, driver, and firmware dependent changes
The software for AMD Data Center GPU products requires maintaining a hardware
and software stack with interdependencies among the GPU and baseboard
firmware, AMD GPU drivers, and the ROCm user space software. While AMD publishes drivers and ROCm user space components, your server or infrastructure provider publishes the GPU and baseboard firmware by bundling AMDs firmware releases via the AMD Platform Level Data Model (PLDM) bundle, which includes the Integrated Firmware Image (IFWI).
GPU and baseboard firmware versioning might differ across GPU families.
<div class="pst-scrollable-table-container">
<table class="table table--middle-left">
<thead>
<tr>
<th class="head">
<p>ROCm Version</p>
</th>
<th class="head">
<p>GPU</p>
</th>
<th class="head">
<p>PLDM Bundle (Firmware)</p>
</th>
<th class="head">
<p>AMD GPU Driver (amdgpu)</p>
</th>
<th class="head">
<p>AMD GPU <br>
Virtualization Driver (GIM)</p>
</th>
</tr>
</thead>
<style>
tbody#virtualization-support-instinct tr:last-child {
border-bottom: 2px solid var(--pst-color-primary);
}
</style>
<tr>
<td rowspan="9" style="vertical-align: middle;">ROCm 7.2.3</td>
<td>MI355X</td>
<td>
01.26.00.02<br>
01.25.17.07<br>
01.25.16.03
</td>
<td>
30.30.x where x (0-3)<br>
30.20.x where x (0-1)<br>
30.10.x where x (0-2)
</td>
<td rowspan="3" style="vertical-align: middle;">8.7.1.K</td>
</tr>
<tr>
<td>MI350X</td>
<td>
01.26.00.02<br>
01.25.17.07<br>
01.25.16.03
</td>
<td>
30.30.x where x (0-3)<br>
30.20.x where x (0-1)<br>
30.10.x where x (0-2)
</td>
</tr>
<tr>
<td>MI325X<a href="#footnote1"><sup>[1]</sup></a></td>
<td>
01.25.06.08<br>
01.25.04.02
</td>
<td>30.30.x where x (0-3)<br>
30.20.x where x (0-1)<a href="#footnote1"><sup>[1]</sup></a><br>
30.10.x where x (0-2)<br>
6.4.z where z (0-3)<br>
6.3.3
</td>
</tr>
<tr>
<td>MI300X<a href="#footnote2"><sup>[2]</sup></a></td>
<td>01.25.06.04<br>
01.25.03.12<br>
01.25.02.04</td>
<td rowspan="6" style="vertical-align: middle;">
30.30.x where x (0-3)<br>
30.20.x where x (0-1)<br>
30.10.x where x (0-2)<br>
6.4.z where z (03)<br>
6.3.3
</td>
<td>8.7.1.K</td>
</tr>
<tr>
<td>MI300A</td>
<td>BKC 26.1</td>
<td rowspan="3" style="vertical-align: middle;">Not Applicable</td>
</tr>
<tr>
<td>MI250X</td>
<td>IFWI 47 (or later)</td>
</tr>
<tr>
<td>MI250</td>
<td>MU5 w/ IFWI 75 (or later)</td>
</tr>
<tr>
<td>MI210</td>
<td>MU5 w/ IFWI 75 (or later)</td>
<td>8.7.1.K</td>
</tr>
<tr>
<td>MI100</td>
<td>VBIOS D3430401-037</td>
<td>Not Applicable</td>
</tr>
</table>
</div>
<p id="footnote1">[1]: For AMD Instinct MI325X KVM SR-IOV users, don't use AMD GPU driver (amdgpu) 30.20.0.</p>
<p id="footnote2">[2]: AMD Instinct MI300X KVM SR-IOV with Multi-VF (8 VF) support requires a compatible firmware BKC bundle, which will be released in the coming months.</p>
#### Improved profiling accuracy for vLLM workloads
ROCm 7.2.3 improves profiling stability for vLLM workloads traced with PyTorch `torch.profiler`. The large, sporadic idle gaps that previously appeared between GPU kernels in the trace have been substantially reduced in common configurations, and the traces now more accurately reflect actual runtime behavior. Coverage may vary depending on model and parallelism settings; additional improvements are in progress.
#### MIGraphX update
[MIGraphX](https://rocm.docs.amd.com/projects/AMDMIGraphX/en/docs-7.2.3/index.html) has the following enhancements:
##### Improved performance of the Gather operator
Performance for embeddingheavy inference workloads is improved by merging multiple independent gather operations from similar embedding tables into a single batched operation. Multigather workloads now run more efficiently with fewer kernel launches and reduced memory traffic by adding horizontal fusion for cross-embedding gather operators. These gather operators have been updated to use `transpose`/`reshape`/`broadcast`/`slice`, enabling better optimization across different backends and data layouts.
##### ONNX Runtime reliability improvement
ONNX Runtime workloads accelerated with MIGraphX now provide a more reliable experience through external stream support in the MIGraphX Execution Provider, with improved memory allocation and deallocation for multi-stream inference.
#### ROCm documentation updates
ROCm documentation has been updated with ROCm XIO documentation. ROCm XIO provides an API for Accelerator-Initiated IO (XIO) for an AMD GPU `__device__` code. It enables AMD GPUs to perform direct IO operations to hardware devices without CPU intervention. ROCm XIO was initially released in April 2026 as an early-access software technology preview. Running production workloads is not recommended.
For more information, see the [ROCm XIO documentation](https://rocm.docs.amd.com/projects/rocm-xio/en/beta-0.1.0/index.html) and {fab}`github` [ROCm/rocm-xio](https://github.com/ROCm/rocm-xio) GitHub repository.
### ROCm components
The following table lists the versions of ROCm components for ROCm 7.2.3, including any version
changes from 7.2.2/7.2.1 to 7.2.3. Click the component's updated version to go to a list of its changes.
Click {fab}`github` to go to the component's source code on GitHub.
<div class="pst-scrollable-table-container">
<table id="rocm-rn-components" class="table">
<thead>
<tr>
<th>Category</th>
<th>Group</th>
<th>Name</th>
<th>Version</th>
<th></th>
</tr>
</thead>
<colgroup>
<col span="1">
<col span="1">
</colgroup>
<tbody class="rocm-components-libs rocm-components-ml">
<tr>
<th rowspan="9">Libraries</th>
<th rowspan="9">Machine learning and computer vision</th>
<td><a href="https://rocm.docs.amd.com/projects/composable_kernel/en/docs-7.2.3/index.html">Composable Kernel</a></td>
<td>1.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/composablekernel"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/AMDMIGraphX/en/docs-7.2.3/index.html">MIGraphX</a></td>
<td>2.15.0&nbsp;&Rightarrow;&nbsp;<a href="#migraphx-2-15-0">2.15.0</a></td>
<td><a href="https://github.com/ROCm/AMDMIGraphX"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/MIOpen/en/docs-7.2.3/index.html">MIOpen</a></td>
<td>3.5.1</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/miopen"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/MIVisionX/en/docs-7.2.3/index.html">MIVisionX</a></td>
<td>3.5.0</a></td>
<td><a href="https://github.com/ROCm/MIVisionX"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocAL/en/docs-7.2.3/index.html">rocAL</a></td>
<td>2.5.0</a></td>
<td><a href="https://github.com/ROCm/rocAL"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocDecode/en/docs-7.2.3/index.html">rocDecode</a></td>
<td>1.7.0</a></td>
<td><a href="https://github.com/ROCm/rocDecode"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocJPEG/en/docs-7.2.3/index.html">rocJPEG</a></td>
<td>1.4.0</a></td>
<td><a href="https://github.com/ROCm/rocJPEG"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocPyDecode/en/docs-7.2.3/index.html">rocPyDecode</a></td>
<td>0.8.0</a></td>
<td><a href="https://github.com/ROCm/rocPyDecode"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rpp/en/docs-7.2.3/index.html">RPP</a></td>
<td>2.2.1</a></td>
<td><a href="https://github.com/ROCm/rpp"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
</tbody>
<tbody class="rocm-components-libs rocm-components-communication tbody-reverse-zebra">
<tr>
<th rowspan="2"></th>
<th rowspan="2">Communication</th>
<td><a href="https://rocm.docs.amd.com/projects/rccl/en/docs-7.2.3/index.html">RCCL</a></td>
<td>2.27.7</a></td>
<td><a href="https://github.com/ROCm/rccl"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocSHMEM/en/docs-7.1.0/index.html">rocSHMEM</a></td>
<td>3.2.0</a></td>
<td><a href="https://github.com/ROCm/rocSHMEM"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
</tbody>
<tbody class="rocm-components-libs rocm-components-math tbody-reverse-zebra">
<tr>
<th rowspan="16"></th>
<th rowspan="16">Math</th>
<td><a href="https://rocm.docs.amd.com/projects/hipBLAS/en/docs-7.2.3/index.html">hipBLAS</a></td>
<td>3.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipblas"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/hipBLASLt/en/docs-7.2.3/index.html">hipBLASLt</a></td>
<td>1.2.2</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipblaslt"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/hipFFT/en/docs-7.2.3/index.html">hipFFT</a></td>
<td>1.0.22</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipfft"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/hipfort/en/docs-7.2.3/index.html">hipfort</a></td>
<td>0.7.1</a></td>
<td><a href="https://github.com/ROCm/hipfort"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/hipRAND/en/docs-7.2.3/index.html">hipRAND</a></td>
<td>3.1.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hiprand"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/hipSOLVER/en/docs-7.2.3/index.html">hipSOLVER</a></td>
<td>3.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipsolver"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/hipSPARSE/en/docs-7.2.3/index.html">hipSPARSE</a></td>
<td>4.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipsparse"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/hipSPARSELt/en/docs-7.2.3/index.html">hipSPARSELt</a></td>
<td>0.2.6</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipsparselt"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocALUTION/en/docs-7.2.3/index.html">rocALUTION</a></td>
<td>4.1.0</a></td>
<td><a href="https://github.com/ROCm/rocALUTION"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocBLAS/en/docs-7.2.3/index.html">rocBLAS</a></td>
<td>5.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocblas"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocFFT/en/docs-7.2.3/index.html">rocFFT</a></td>
<td>1.0.36</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocfft"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocRAND/en/docs-7.2.3/index.html">rocRAND</a></td>
<td>4.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocrand"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocSOLVER/en/docs-7.2.3/index.html">rocSOLVER</a></td>
<td>3.32.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocsolver"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocSPARSE/en/docs-7.2.3/index.html">rocSPARSE</a></td>
<td>4.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocsparse"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocWMMA/en/docs-7.2.3/index.html">rocWMMA</a></td>
<td>2.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocwmma"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/Tensile/en/docs-7.2.3/src/index.html">Tensile</a></td>
<td>4.45.0</td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/shared/tensile"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
</tbody>
<tbody class="rocm-components-libs rocm-components-primitives tbody-reverse-zebra">
<tr>
<th rowspan="4"></th>
<th rowspan="4">Primitives</th>
<td><a href="https://rocm.docs.amd.com/projects/hipCUB/en/docs-7.2.3/index.html">hipCUB</a></td>
<td>4.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hipcub"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/hipTensor/en/docs-7.2.3/index.html">hipTensor</a></td>
<td>2.2.0</td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/hiptensor"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocPRIM/en/docs-7.2.3/index.html">rocPRIM</a></td>
<td>4.2.0</td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocprim"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocThrust/en/docs-7.2.3/index.html">rocThrust</a></td>
<td>4.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/develop/projects/rocthrust"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
</tbody>
<tbody class="rocm-components-tools rocm-components-system tbody-reverse-zebra">
<tr>
<th rowspan="7">Tools</th>
<th rowspan="7">System management</th>
<td><a href="https://rocm.docs.amd.com/projects/amdsmi/en/docs-7.2.3/index.html">AMD SMI</a></td>
<td>26.2.2</a></td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/amdsmi"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rdc/en/docs-7.2.3/index.html">ROCm Data Center Tool</a></td>
<td>1.2.0</a></td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rdc/"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocminfo/en/docs-7.2.3/index.html">rocminfo</a></td>
<td>1.0.0</td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocminfo/"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocm_smi_lib/en/docs-7.2.3/index.html">ROCm SMI</a></td>
<td>7.8.0</td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocm-smi-lib/"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/ROCmValidationSuite/en/docs-7.2.3/index.html">ROCm Validation Suite</a></td>
<td>1.3.0</a></td>
<td><a href="https://github.com/ROCm/ROCmValidationSuite"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
</tbody>
<tbody class="rocm-components-tools rocm-components-perf">
<tr>
<th rowspan="6"></th>
<th rowspan="6">Performance</th>
<td><a href="https://rocm.docs.amd.com/projects/rocm_bandwidth_test/en/docs-7.2.3/index.html">ROCm Bandwidth
Test</a></td>
<td>2.6.0</a></td>
<td><a href="https://github.com/ROCm/rocm_bandwidth_test/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocprofiler-compute/en/docs-7.2.3/index.html">ROCm Compute Profiler</a></td>
<td>3.4.0</a></td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocprofiler-compute"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocprofiler-systems/en/docs-7.2.3/index.html">ROCm Systems Profiler</a></td>
<td>1.3.0</a></td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocprofiler-systems/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocprofiler/en/docs-7.2.3/index.html">ROCProfiler</a></td>
<td>2.0.0</a></td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocprofiler/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/docs-7.2.3/index.html">ROCprofiler-SDK</a></td>
<td>1.1.0</a></td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocprofiler-sdk/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr >
<td><a href="https://rocm.docs.amd.com/projects/roctracer/en/docs-7.2.3/index.html">ROCTracer</a></td>
<td>4.1.0</td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/roctracer/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
</tbody>
<tbody class="rocm-components-tools rocm-components-dev">
<tr>
<th rowspan="5"></th>
<th rowspan="5">Development</th>
<td><a href="https://rocm.docs.amd.com/projects/HIPIFY/en/docs-7.2.3/index.html">HIPIFY</a></td>
<td>22.0.0</td>
<td><a href="https://github.com/ROCm/HIPIFY/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/ROCdbgapi/en/docs-7.2.3/index.html">ROCdbgapi</a></td>
<td>0.77.4</a></td>
<td><a href="https://github.com/ROCm/ROCdbgapi/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/ROCmCMakeBuildTools/en/docs-7.2.3/index.html">ROCm CMake</a></td>
<td>0.14.0</td>
<td><a href="https://github.com/ROCm/rocm-cmake/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/ROCgdb/en/docs-7.2.3/index.html">ROCm Debugger (ROCgdb)</a>
</td>
<td>16.3</a></td>
<td><a href="https://github.com/ROCm/ROCgdb/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/rocr_debug_agent/en/docs-7.2.3/index.html">ROCr Debug Agent</a>
</td>
<td>2.1.0</td>
<td><a href="https://github.com/ROCm/rocr_debug_agent/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
</tbody>
<tbody class="rocm-components-compilers tbody-reverse-zebra">
<tr>
<th rowspan="2" colspan="2">Compilers</th>
<td><a href="https://rocm.docs.amd.com/projects/HIPCC/en/docs-7.2.3/index.html">HIPCC</a></td>
<td>1.1.1</td>
<td><a href="https://github.com/ROCm/llvm-project/tree/amd-staging/amd/hipcc"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/llvm-project/en/docs-7.2.3/index.html">llvm-project</a></td>
<td>22.0.0</a></td>
<td><a href="https://github.com/ROCm/llvm-project/"><i
class="fab fa-github fa-lg"></i></a></td>
</tr>
</tbody>
<tbody class="rocm-components-runtimes tbody-reverse-zebra">
<tr>
<th rowspan="2" colspan="2">Runtimes</th>
<td><a href="https://rocm.docs.amd.com/projects/HIP/en/docs-7.2.3/index.html">HIP</a></td>
<td>7.2.1</a></td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/hip"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
<tr>
<td><a href="https://rocm.docs.amd.com/projects/ROCR-Runtime/en/docs-7.2.3/index.html">ROCr Runtime</a></td>
<td>1.18.0</td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/develop/projects/rocr-runtime"><i class="fab fa-github fa-lg"></i></a></td>
</tr>
</tbody>
</table>
</div>
### Detailed component changes
The following sections describe key changes to ROCm components.
```{note}
For a historical overview of ROCm component updates, see the {doc}`ROCm consolidated changelog </release/changelog>`.
```
#### **MIGraphX** (2.15.0)
##### Added
* External stream support to the MIGraphX context, allowing external HIP streams to be used during execution.
* Ability to return a vector for output alias, supporting operators like `make_tuple`.
##### Changed
* Refactored `move_output_instructions_after` into the module class.
* Updated rocMLIR to fix `bert_squad` and `bert_tf` regressions.
##### Optimized
* Rewrote the `gather` operator to use `transpose`/`reshape`/`broadcast`/`slice` for improved performance.
* Horizontally fuse cross-embedding `gather` operators.
* Improved tuning for Split-K.
* Removed extra assignments and inserts in `find_nop_reshapes` to reduce overhead.
##### Resolved issues
The following issues have been fixed:
* `int` to `bf16`/`fp16` conversion errors.
* Comparison logic in `find_concat_op` to match the correct I/O.
* `shape_transform_descriptor::rebase` when flattening a broadcasted dimension.
* An error with `rewrite_reshapes`.
* A gather rewrite crash by validating the strided view element count.
* A bug in gather rewrite with NHWC shapes.
* A crash in rocMLIR with Inception v3 on RDNA3 architecture-based Radeon GPUs.
* Filter zero-argument operators during ONNX parsing to prevent errors.
* Conflict for missing `no_broadcast` parameter on ROCm 7.2.x.
### ROCm known issues
ROCm known issues are noted on {fab}`github` [GitHub](https://github.com/ROCm/ROCm/labels/Verified%20Issue). For known
issues related to individual components, review the [Detailed component changes](#detailed-component-changes).
#### Minor performance regression for MIGraphX with int8-quantized models
You might observe a slight performance regression when running int8-quantized models with MIGraphX. This impact is generally minimal and does not affect correctness. However, workloads sensitive to peak throughput might have reduced performance when compared to non-quantized or alternative execution paths. This issue is currently under investigation and will be fixed in a future ROCm release. See [GitHub issue #6195](https://github.com/ROCm/ROCm/issues/6195).
### ROCm upcoming changes
The following changes to the ROCm software stack are anticipated for future releases.
#### ROCTracer, ROCProfiler, rocprof, and rocprofv2 deprecation
ROCTracer, ROCProfiler, `rocprof`, and `rocprofv2` are deprecated. It's strongly recommended to upgrade to the latest version of the [ROCprofiler-SDK](https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/latest/) library and the (`rocprofv3`) tool to ensure continued support and access to new features.
To learn about key feature improvements and benefits of ROCprofiler-SDK over the deprecated ROCProfiler and ROCTracer, see [Comparing ROCprofiler-SDK to legacy ROCm profiling tools](https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/latest/conceptual/comparing-with-legacy-tools.html).
It's anticipated that ROCTracer, ROCProfiler, `rocprof`, and `rocprofv2` will reach end of support (EoS) by the end of 2026 Q2.
#### ROCm SMI deprecation
[ROCm SMI](https://github.com/ROCm/rocm_smi_lib) will be phased out in an
upcoming ROCm release and will enter maintenance mode. After this transition,
only critical bug fixes will be addressed and no further feature development
will take place.
It's strongly recommended to transition your projects to [AMD
SMI](https://github.com/ROCm/rocm-systems/tree/develop/projects/amdsmi), the successor to ROCm SMI. AMD SMI
includes all the features of the ROCm SMI and will continue to receive regular
updates, new functionality, and ongoing support. For more information on AMD
SMI, see the [AMD SMI documentation](https://rocm.docs.amd.com/projects/amdsmi/en/latest/).
#### Changes to ROCm Object Tooling
ROCm Object Tooling tools ``roc-obj-ls``, ``roc-obj-extract``, and ``roc-obj`` were
deprecated in ROCm 6.4, and will be removed in a future release. Functionality
has been added to the ``llvm-objdump --offloading`` tool option to extract all
clang-offload-bundles into individual code objects found within the objects
or executables passed as input. The ``llvm-objdump --offloading`` tool option also
supports the ``--arch-name`` option, and only extracts code objects found with
the specified target architecture. See [llvm-objdump](https://llvm.org/docs/CommandGuide/llvm-objdump.html)
for more information.
+15
View File
@@ -0,0 +1,15 @@
root = true
[*]
charset = utf-8
end_of_line = lf
indent_style = space
indent_size = 4
insert_final_newline = true
trim_trailing_whitespace = true
[*.rst]
indent_size = 3
[*.{html,md,yaml}]
indent_size = 2
@@ -0,0 +1,71 @@
<table class="rocm-docs-table table">
<thead>
<tr>
<th class="head">
<p>Framework</p>
</th>
<th class="head">
<p>Supported versions</p>
</th>
<th class="head">
<p>Supported OS</p>
</th>
<th class="head">
<p>Supported Python versions</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="2" style="vertical-align: middle;">
<p>PyTorch</p>
</td>
<td style="vertical-align: middle;">
<p>2.11.0, 2.10.0, 2.9.1</p>
</td>
<td>
<p>Linux</p>
</td>
<td rowspan="2">
<p>3.14, 3.13, 3.12, 3.11</p>
</td>
</tr>
<tr>
<td style="vertical-align: middle;">
<p>2.11.0</p>
</td>
<td style="vertical-align: middle;">
<p>Windows</p>
</td>
</tr>
<tr>
<td style="vertical-align: middle;">
<p>JAX</p>
</td>
<td style="vertical-align: middle;">
<p>0.9.1, 0.8.2</p>
</td>
<td>
<p>Linux</p>
</td>
<td>
<p>3.14, 3.13, 3.12, 3.11</p>
</td>
</tr>
<tr>
<td style="vertical-align: middle;">
<p>vLLM<br>(<a href="#release-supported-hw">gfx950, gfx942, gfx1200,<br>gfx1201, gfx1100,
gfx1101,<br>gfx1102, gfx1151 GPUs only</a>)</p>
</td>
<td>
<p>0.19.1<br>(requires PyTorch 2.10.0)</p>
</td>
<td>
<p>Linux</p>
</td>
<td>
<p>3.13</p>
</td>
</tr>
</tbody>
<table>
@@ -0,0 +1,778 @@
#### **AMD SMI (BM)** (26.4.0)
##### Added
* **Added APU metrics support (table versions 2.4 and 3.0)**.
* New `amdsmi_apu_metrics_t` struct accessible via `amdsmi_gpu_metrics_t.apu_metrics` pointer (non-null when APU-specific metrics are available).
* **v2.4 metrics**:
* `temperature_gfx`, `temperature_soc`, `temperature_core[8]`, `temperature_l3[2]`
* `average_gfx_activity`, `average_mm_activity`
* `average_socket_power`, `average_cpu_power`, `average_soc_power`, `average_gfx_power`, `average_core_power[8]`
* Average clocks: `gfxclk`, `socclk`, `uclk`, `fclk`, `vclk`, `dclk`
* Current clocks: `gfxclk`, `socclk`, `uclk`, `fclk`, `vclk`, `dclk`, `coreclk[8]`, `l3clk[2]`
* `average_temperature_gfx`, `average_temperature_soc`, `average_temperature_core[8]`, `average_temperature_l3[2]`
* `average_cpu_voltage`, `average_soc_voltage`, `average_gfx_voltage`, `average_cpu_current`, `average_soc_current`, `average_gfx_current`
* `throttle_status`, `indep_throttle_status`
* `fan_pwm`
* **v3.0 metrics**:
* `temperature_core[16]`, `temperature_skin`
* `average_vcn_activity`, `average_ipu_activity[8]`, `average_core_c0_activity[16]`
* `average_dram_reads`, `average_dram_writes`, `average_ipu_reads`, `average_ipu_writes`
* `average_apu_power`, `average_dgpu_power`, `average_all_core_power`, `average_ipu_power`, `average_sys_power`
* `stapm_power_limit`, `current_stapm_power_limit`
* `average_core_power[16]`, `current_coreclk[16]`
* `current_core_maxfreq`, `current_gfx_maxfreq`
* `average_vpeclk_frequency`, `average_ipuclk_frequency`, `average_mpipu_frequency`
* `throttle_residency_prochot`, `throttle_residency_spl`, `throttle_residency_fppt`, `throttle_residency_sppt`, `throttle_residency_thm_core`, `throttle_residency_thm_gfx`, `throttle_residency_thm_soc`
* `time_filter_alphavalue`
* Fields not applicable to the current version are set to sentinel values: `0xFFFF` for `uint16_t`, `0xFFFFFFFF` for `uint32_t`, and `UINT64_MAX` for `uint64_t` fields.
* Python bindings updated with `AmdSmiApuMetrics` ctypes structure.
* **Added `oam_id` to `amdsmi_enumeration_info_t`**.
* `amd-smi list -e` now displays `OAM_ID` (Physical XGMI ID / OAM ID).
* Added `--enumeration` as a long-form alias for `-e` in `amd-smi list`.
* **Added support for GPU metrics v1.9 new fields**.
* Added new temperature fields to `amdsmi_gpu_metrics_t`:
* `temperature_hbm_stacks` — per-stack HBM temperatures (°C)
* `temperature_mid` — per-MID temperatures (°C)
* `temperature_aid` — per-AID temperatures (°C)
* `temperature_xcd` — per-XCC compute die temperatures (°C)
* Added new per-die clock fields to `amdsmi_gpu_metrics_t`:
* `current_uclk_aid` — per-AID uclk (MHz)
* `current_socclks_mid` — per-MID SOC clock (MHz)
* Added new constants:
* `AMDSMI_MAX_NUM_HBM_STACKS` (12)
* `AMDSMI_MAX_NUM_AID` (2)
* `AMDSMI_MAX_NUM_MID` (2)
* `AMDSMI_MAX_NUM_CLKS_PER_AID` (2)
* `AMDSMI_MAX_NUM_CLKS_PER_MID` (2)
* **Added VRAM and GTT tuning interface**.
* New `amd-smi static --mem-carveout` to view VRAM carveout options.
* New `amd-smi set --mem-carveout` to change the VRAM carveout (APU).
* New `amd-smi set --gtt` and `amd-smi reset --gtt` for system-wide GTT size tuning.
* New APIs: `amdsmi_get_gpu_uma_carveout_info()`, `amdsmi_set_gpu_uma_carveout()`, `amdsmi_get_ttm_info()`, `amdsmi_set_ttm_pages_limit()`, `amdsmi_reset_ttm_pages_limit()`.
* **Added UBB power and power_limit fields to `amdsmi_power_info_t` and `amdsmi_npm_info_t`**.
* `amd-smi metric --power` now displays `ubb_power` when available.
* `amd-smi node -p` now displays UBB power threshold when available.
* **Added CPU support for family 1A Models 50h-57h**.
* New APIs: `amdsmi_get_cpu_xgmi_pstate_range()`, `amdsmi_get_cpu_core_ccd_power()`, `amdsmi_get_cpu_tdelta()`, `amdsmi_get_cpu_dimm_sb_reg()`, `amdsmi_get_cpu_svi3_vr_controller_temp()`, `amdsmi_get_cpu_pc6_enable()`, `amdsmi_get_cpu_cc6_enable()`, `amdsmi_get_cpu_sdps_limit()`, `amdsmi_get_cpu_core_floor_freq_limit()`, `amdsmi_get_cpu_core_eff_floor_freq_limit()`, and corresponding set APIs.
* **Note**: `amdsmi_get_dfc_ctrl()` renamed to `amdsmi_get_cpu_dfc_ctrl()` and `amdsmi_set_dfc_ctrl()` renamed to `amdsmi_set_cpu_dfc_ctrl()` for naming consistency.
* **Updated memory API documentation**
Added note that the sum of per-process memory usage is not expected to equal total usage.
##### Changed
* **Renamed `processor_type_t` enum typedef to `amdsmi_processor_type_t`**.
* The unprefixed typedef name did not follow the `amdsmi_*_t` convention used throughout `amdsmi.h` and was easy to collide with identifiers defined by other system-management libraries. New code should use `amdsmi_processor_type_t`. The old name is preserved as a backward-compatibility typedef alias, so existing callers continue to compile unchanged.
* **Package install no longer modifies the system-wide `logrotate` timer or cron schedule**.
* Previously, installing `amd-smi-lib` overwrote `/lib/systemd/system/logrotate.timer` (or moved `/etc/cron.daily/logrotate` to `/etc/cron.hourly/`) to force hourly rotation, which affected every other package using `logrotate`.
* The package now only ships `/etc/logrotate.d/amd_smi.conf`, which sets its own `hourly` + `size 1M` cadence. AMD-SMI logs still rotate at the same frequency; system-wide settings stay as the distribution configured them.
##### Optimized
* **Optimized `rsmi_dev_device_identifiers_get()` in the ROCm-SMI device layer**.
* Removed unnecessary iteration by directly indexing the device list.
* Added bounds checking for `device_id`, with clearer error handling/logging.
* Improves performance for device identifier queries.
##### Resolved issues
* **Fixed `amd-smi metric` crashing with `TypeError` on MI300A when no CPU flags are specified**.
* When no CPU arguments are passed, `metric_cpu()` sets all boolean CPU args to `True` to display all available data. `--cpu-svi3-vr-controller-temp` takes a TYPE argument (and optional RAIL_INDEX) rather than a boolean flag — setting it to `True` caused a `TypeError` crash when the code tried to subscript it with `[0][0]`. Added `cpu_svi3_vr_controller_temp` to the show-all exclusion list, following the existing pattern for `cpu_lclk_dpm_level`, `cpu_io_bandwidth`, `cpu_dimm_sb_reg`, and similar argument-taking flags.
* **Fixed `amdsmi_get_gpu_accelerator_partition_profile()` returning incorrect `num_partitions` when `num_partition` is unavailable from GPU metrics**.
* GPU metrics no longer always provides `num_partition`. The function now derives the partition count from the active partition type when `num_partition` is not available:
* SPX → 1, DPX → 2, TPX → 3, QPX → 4
* CPX → derived from the XCD counter via `amdsmi_get_gpu_xcd_counter()`
* **Fixed `amdsmi_topo_get_p2p_status()` returning a raw `ctypes.c_uint32` object instead of an integer for the `type` field**.
* The `'type'` key in the returned dictionary now correctly returns `type_32.value` (an `int`) rather than the unwrapped ctypes object, consistent with the pattern used in `amdsmi_topo_get_link_type()`.
* **Adjusted KFD process caching to be more responsive**.
* Updated process caching to allow cache duration adjustment via the `AMDSMI_PROCESS_INFO_CACHE_MS` environment variable for workflows with rapid metric polling.
* **Fixed CLI exit codes to use absolute values**.
* Invalid GPU parameters now return positive error codes as documented.
* **Fixed CLI breakage when `amdgpu` driver is not present**.
* Improved init to better catch driver loading issues.
* **Aligned `amdsmi_get_gpu_device_uuid()` with HIP/rocminfo UUID format**.
* Modified `amdsmi_asic_info_t.asic_serial` to report per-socket serial using KFD's `unique_id`.
* **Fixed multiple bugs in NIC/switch code and `amdsmi_init()` NIC handling**.
* Fixed `sizeof` operator precedence, `hw_mon` reset, NUMA=65535 handling, and several CLI function call errors.
* Fixed `amdsmi_init()` to succeed when no NIC hardware is present.
* **Fixed shared mutex and self-heal**.
* Improved self-heal logic to correctly identify and recover from corrupted or uninitialized mutex state.
* **Fixed `cu_occupancy` displaying `0%` instead of `N/A` when file is unavailable**.
* Process `cu_occupancy` is now initialized to `INVALID` instead of zero, so `amd-smi process` displays `N/A` rather than a misleading `0%` when the sysfs file is not accessible.
* **Fixed CLI set commands silently succeeding on invalid input values**.
* `amd-smi set --profile <INVALID>` now returns a non-zero exit code and lists available profiles in the error message; invalid profile names are rejected at parse time.
* `amd-smi set --clk-level <CLK_TYPE>` (missing performance level indices) now returns a non-zero exit code with a usage hint instead of silently succeeding.
* `amd-smi set --power-cap <OUT_OF_RANGE>` now returns a non-zero exit code.
* `amd-smi set --fan <INVALID>%` no longer prompts the out-of-spec warning before validating the percentage range; invalid values are rejected immediately.
* **Fixed `amd-smi set --profile` help text omitting `BOOTUP_DEFAULT`**.
* `BOOTUP_DEFAULT` was always accepted at runtime but was missing from the `--help` profile list. Auditing invalid-input handling exposed this gap. `amd-smi reset --profile` can also be used to return to the bootup default power profile.
* **Fixed `amd-smi monitor --brcm_nic` and `--brcm_switch` flags being registered on non-BRCM systems**.
* These flags are now only registered when BRCM hardware is present, preventing spurious failures on AMD GPU-only systems.
* **Fixed `amd-smi` default command alignment**.
* Updated default `amd-smi` output to align values to the left for improved readability.
Several items were misaligned in the default output, and this change ensures a consistent left-aligned format across all fields.
* *This change is purely cosmetic and does not affect any functionality.*
* **Renamed `lc_perf_other_end_recovery` to `lc_perf_other_end_recovery_count` in `amd-smi metric` CLI output for unification**.
* **Removed references to deprecated `amd-smi reset -r`**.
* CLI help text and memory partition change warnings no longer reference `amd-smi reset -r` for driver reloading.
* Users are now directed to use `sudo modprobe -r amdgpu && sudo modprobe amdgpu` to reload the driver after partition changes.
* **Changed CPU power APIs to return values in milliwatts (mW) for higher precision**.
* Removed lossy integer rounding (`(mW + 500) / 1000`) from 6 CPU power get APIs. Values are now
returned in milliwatts directly from the ESMI library, preserving sub-watt precision.
* **C API**: Output parameter type remains `uint32_t*`, but the unit changed from watts to milliwatts (mW).
* `amdsmi_get_cpu_socket_power`
* `amdsmi_get_cpu_socket_power_cap`
* `amdsmi_get_cpu_socket_power_cap_max`
* `amdsmi_get_cpu_pwr_efficiency_mode` (ppt_limit field)
* `amdsmi_get_cpu_core_ccd_power`
* `amdsmi_get_cpu_sdps_limit`
* **Python API (breaking)**: These functions now return `int` (milliwatts) instead of `str` (e.g., `"240 Watts"`).
Callers that parsed the string output must update to handle the numeric return value.
* **CLI output**: Power values now display with milliwatt precision (e.g., `240.500 Watts`).
* Added missing null-pointer validation for output parameters in `amdsmi_get_cpu_socket_power_cap`
and `amdsmi_get_cpu_socket_power_cap_max`.
* Updated header documentation to specify milliwatt units for all affected get and set API parameters.
* **Changed power APIs to have consistent output parameter types**.
* Modified 6 CPU power APIs to have consistent output power types. All set and get APIs have `uint32_t` output values.
* Modified get and set APIs that had double output types to have `uint32_t` output types in milliwatts (mW).
* `amdsmi_get_cpu_socket_power(amdsmi_processor_handle processor_handle, uint32_t* ppower)`
* `amdsmi_get_cpu_socket_power_cap(amdsmi_processor_handle processor_handle, uint32_t* pcap)`
* `amdsmi_get_cpu_socket_power_cap_max(amdsmi_processor_handle processor_handle, uint32_t* pmax)`
* `amdsmi_get_cpu_pwr_efficiency_mode(amdsmi_processor_handle processor_handle, uint32_t* power_efficiency_mode, uint32_t* utilization, uint32_t* ppt_limit)`
* `amdsmi_get_cpu_core_ccd_power(amdsmi_processor_handle processor_handle, uint32_t* power)`
* `amdsmi_get_cpu_sdps_limit(amdsmi_processor_handle processor_handle, uint32_t* sdps_limit)`
#### **Composable Kernel** (1.3.0)
##### Added
* Added overload of `load_tile_transpose` that takes reference to output tensor as output parameter.
* Use data type from LDS tensor view when determining tile distribution for transpose in the GEMM pipeline.
* Added `eightwarps` support for abquant mode in blockscale GEMM.
* Added `preshuffleB` support for abquant mode in blockscale GEMM.
* Added support for explicit GEMM in `CK_TILE` grouped convolution forward and backward weight.
* Added TF32 convolution support on gfx942 and gfx950 in CK. It can be enabled or disabled via `DTYPES` of `tf32`.
* Added `streamingllm` sink support for FMHA FWD, include `qr_ks_vs`, `qr_async` and `splitkv` pipelines.
* Added support for microscaling (MX) FP8/FP4 mixed data types to Flatmm pipeline.
* Added support for fp8 dynamic tensor-wise quantization of FP8 fmha fwd kernel.
* Added FP8 KV cache support for FMHA batch prefill.
* Added FMHA batch prefill kernel support for several KV cache layouts, flexible page sizes, and different lookup table configurations.
* Added gpt-oss sink support for FMHA FWD, include `qr_ks_vs`, `qr_async`, `qr_async_trload` and `splitkv` pipelines.
* Added persistent async input scheduler for CK Tile universal GEMM kernels to support asynchronous input streaming.
* Added FP8 block scale quantization for FMHA forward kernel.
* Added gfx11xx support for FMHA.
* Added microscaling (MX) FP8/FP4 support on gfx950 for FMHA forward kernel (`qr` pipeline only).
* Added FP8 per-tensor quantization support for FMHA forward V3 pipeline on gfx950.
#### **HIP** (7.13)
##### Added
* New HIP APIs
* `cooperative_groups::reduce()` allows calling reduce operators on `thread_block_tile` and `coalesced_threads`. The implementation is based on the `__reduce_*_sync` operations, so the macro `HIP_ENABLE_EXTRA_WARP_SYNC_TYPES` might be needed to unlock some optimizations.
* New device attribute `hipDeviceAttributeGPUDirectRDMAWithHipVMMSupported`, indicating support for GPU Direct RDMA when using HIP VMM. This attribute corresponds to the CUDA `CU_DEVICE_ATTRIBUTE_GPU_DIRECT_RDMA_WITH_CUDA_VMM_SUPPORTED`.
##### Resolved issues
* A segmentation fault that occurred in child graphs during the graphlaunch phase. The issue originated from the entire graph being launched solely according to the parent graphs scheduling logic. The HIP runtime now introduces a pergraph segmentscheduling control flag and propagates the parent graphs scheduling mode to its child graphs, ensuring consistent scheduling behavior (classic vs. segment) and preventing failures when the parent falls back to classic scheduling.
* A segmentation fault caused by passing a null pointer to the hipMemGetAddressRange API. The function now handles null pointers correctly, matching the behavior of the corresponding CUDA API.
##### Changed
* `__reduce_and_sync()`, `__reduce_or_sync()` and `__reduce_xor_sync()` now provide a consistent behavior for all mask values and with CUDA. Previously, some masks were translated into bitwise operations, but others were not (such as those containing "holes"). Now, all masks cause bitwise instructions to be emitted. This is a change in behavior compared to previous versions.
##### Optimized
* Improved HIP runtime error logging when an application's fat binary does not include a compatible code object for the detected GPU architecture, offering clearer guidance to rebuild with the appropriate `--offload-arch=gfxXXXX` option.
* Enables inmemory and backgroundthread asynchronous logging in the HIP runtime by default to improve overall logging capability. This behavior can be disabled by setting the environment variable `AMD_LOG_ASYNC=0`.
#### **hipBLAS** (3.4.0)
##### Added
* gfx1250 and gfx90c support to clients.
* Version and other properties to Windows `hipblas.dll`.
* Support for `OpenBLAS` ILP64-based API usage in clients.
##### Resolved issues
* Restored the fallback of using the deprecated rocBLAS API `rocblas_set_device_memory_size` if allocations are failing.
#### **hipBLASLt** (1.3.0)
##### Added
* General Batched GEMM support.
##### Changed
* Replaced `install.sh` with an invoke-based task runner (`tasks.py`) to support cross-platform builds including Windows (ROCm 7.0+).
* `gtest` and `msgpack-cxx` are now fetched automatically using CMake FetchContent if not found on the system.
#### **hipCUB** (4.4.0)
##### Optimized
* Reduced build times for unit tests.
##### Resolved issues
* Fixed more memory leak issues with some unit tests.
#### **hipFFT** (1.0.23)
##### Added
* hipFFTW plan creation functions for advanced and general plans:
* `fftw_plan_many_dft`
* `fftwf_plan_many_dft`
* `fftw_plan_many_dft_r2c`
* `fftwf_plan_many_dft_r2c`
* `fftw_plan_many_dft_c2r`
* `fftwf_plan_many_dft_c2r`
* `fftw_plan_guru_dft`
* `fftwf_plan_guru_dft`
* `fftw_plan_guru_dft_r2c`
* `fftwf_plan_guru_dft_r2c`
* `fftw_plan_guru_dft_c2r`
* `fftwf_plan_guru_dft_c2r`
* `fftw_plan_guru64_dft`
* `fftwf_plan_guru64_dft`
* `fftw_plan_guru64_dft_r2c`
* `fftwf_plan_guru64_dft_r2c`
* `fftw_plan_guru64_dft_c2r`
* `fftwf_plan_guru64_dft_c2r`
* Support for gfx1150 architecture.
##### Changed
* Moved library to C++20 standard.
* Removed Boost as a dependency for clients and samples.
* Callback functions will be deprecated in a future release.
##### Resolved issues
* Fixed potential launch failure of data generation kernels in test and benchmark programs.
#### **hipRAND** (3.3.0)
##### Added
* `hiprand.dll` now contains embedded file version metadata.
#### **hipSOLVER** (3.4.0)
##### Added
* Compatibility-only functions:
* `geev`
* `hipsolverDnXgeev_bufferSize`
* `hipsolverDnXgeev`
* `syevBatched`
* `hipsolverDnXsyevBatched_bufferSize`
* `hipsolverDnXsyevBatched`
* `syevd`
* `hipsolverDnXsyevd_bufferSize`
* `hipsolverDnXsyevd`
* `sytrs`
* `hipsolverDnXsytrs_bufferSize`
* `hipsolverDnXsytrs`
#### **hipSPARSELt** (0.2.8)
##### Added
* CTest and test categories support (`--smoke`, `--pre_checkin`, and `--nightly`).
##### Optimized
* Provided more kernels for the `FP16`, `BF16`, and `Int8` datatypes.
* Improved the performance of the `HIPSPARSELT_PRUNE_SPMMA_TILE` function.
##### Resolved issues
* Fixed incorrect behavior when retrieving the PCI chip ID.
* Fixed LDS out-of-bounds read in `prune_tile_kernel`.
* Fixed out-of-bounds access for compress function test cases.
* Fixed missing null terminator in the return value of `hipsparseLtGetArchName()`.
* Fixed incorrect CPU result when `bias_type` is `BF16` for spmm test cases.
* Fixed double-free issue in the example code `example_prune_strip`.
* Fixed symbol interposition in the hipSPARSELt library.
#### **MIOpen** (3.5.1)
##### Added
* Added `MIOPEN_LOG_BUFFER_SIZE` option: when set to non-zero, dumps recent MIOpen logs to file on error.
* [Conv] Added `ConvDepthwiseFwd3D` solver for optimizing specific 3D depthwise convolutions.
* [Conv] Added NHWC layout support for Winograd convolution solvers.
* [Conv] Added regular GEMM solver support for Conv3D forward and backward-data with 1x1x1 filters.
* [Conv] Added configurable problem size threshold (`MIOPEN_CONV_DIRECT_MAX_SIZE`) for direct solver.
* [Softmax] Added tuning support via Generic Search.
##### Changed
* [Conv] Improved default kernel selection for Composable Kernel (CK) convolution solvers with ranked shortlists.
* [Conv] Split CK grouped convolution kernels into per-architecture runtime-loaded dynamic libraries.
##### Optimized
* Optimized transpose operations with tiled and vectorized variants for NCHW/NHWC conversions.
* [BatchNorm] Optimized batchnorm reduction using warp shuffle intrinsics.
* [Conv] Added heuristic filtering of slow GEMM solver configurations during tuning.
##### Deprecated
* [Conv] Deprecated CK non-grouped convolution forward and backward solvers.
* Deprecated `miopenConvolutionBackwardBias`: the underlying OpenCL kernel (`MIOpenConvBwdBias.cl`) has been removed. The function now returns `miopenStatusNotImplemented` and will be removed in a future release.
##### Removed
* Removed GraphAPI experimental feature and related code.
##### Resolved issues
* [Conv] Fixed Winograd Fury grouped convolution correctness on gfx12xx when G > 1.
* [Conv] Fixed bf16 WrW convolution precision loss in inter-batch accumulation.
* [Conv] Fixed GPU memory fault in Winograd v3.0 WrW solver for large tensor shapes.
* Fixed BF16 `abs` function precision error caused by unnecessary cast through FP16.
* Fixed pooling kernel runtime compilation failure.
* Fixed gfx1151 inline assembly compilation errors in batchnorm kernels.
* Fixed use-after-free in HIPOCProgram binary loading.
#### **ROCm Data Center Tool (RDC)** (1.3.0)
##### Resolved issues
* **Fixed broken partition metrics**.
* Regardless of whether the GPU was partitioned, RDC only saw the GPU index and no instances due to upstream gpu_metrics changes.
#### **rocBLAS** (5.4.0)
##### Added
* gfx1250 and gfx90c enabled.
* Trace logging using `ROCBLAS_LAYER=1` for `rocblas_gemm_ex_get_solutions`, `rocblas_gemm_batched_ex_get_solutions`, `rocblas_gemm_ex_get_solutions_by_type`, and `rocblas_gemm_batched_ex_get_solutions_by_type`.
* Version and other properties to Windows `rocblas.dll`.
* Support for `OpenBLAS` ILP64 API for host reference in clients.
* Dockerfiles in the `docker` directory to assist in setting up development.
##### Optimized
* Improved the performance of Level 3 `geam` for pure transpose scale use cases.
* Improved the performance of Level 2 `tpsv`.
##### Resolved issues
* Fix for querying solutions when using the `hipBLASLt` backend with `rocblas_gemm_batched_ex_get_solutions` if using null data pointers.
#### **ROCdbgapi** (0.80.0)
##### Added
* `amd_dbgapi_process_get_info()` adds a new query to get a mask spanning
over all the bits used by all the address spaces. The query is called
`AMD_DBGAPI_PROCESS_INFO_SIGNIFICANT_ADDRESS_BITS`.
#### **rocDecode** (1.8.0)
##### Added
* Logging improvement: Added function entry and exit logs (at Info log level).
* Logging improvement: Added duration to function exit logs and optimized log message formatting to reduce runtime overhead.
* Logging improvement: Merged all logger instances into one global instance.
* Logging improvement: Unified logging format in utility classes with core library logging format.
* Logging improvement: Moved debug logging from a compile-time switch to the runtime logger level controlled by `ROCDEC_LOG_LEVEL` (debug = 4).
* Added support for user-set output surface format.
##### Changed
* Removed CPack packaging (DEB/RPM/NSIS/TGZ/ZIP generation and all related CPACK variables).
* Removed `rocDecode-setup.py` dependency installer script.
* Removed Docker files.
* Removed package install documentation; updated all documentation to reference TheRock for installation.
* Simplified libva version check (single `>= 1.22` requirement).
* Cleaned up CMake error messages.
#### **rocFFT** (1.0.37)
##### Optimized
* Allow plans to share hipModules if they use the same kernels. This reduces time spent and memory used when
creating plans that exist concurrently.
* Improved performance of unit-strided, interleaved, complex-to-complex and real-to-complex FFTs on gfx1201, gfx90a, gfx942, and gfx950.
Single-precision lengths:
* (160,72,72)
* (160,80,72)
* (160,80,80)
* (72,72,72)
* (80,80,80)
* (84,84,72)
* (96,96,96)
* (108,108,80)
Double-precision lengths:
* (72,72,52)
* (60,60,60)
* (64,64,52)
* (64,64,64)
##### Changed
* Moved library to C++20 standard.
* Removed Boost as a dependency for clients and samples.
* Split the precompiled kernel cache file (`rocfft_kernel_cache.db`) into per-architecture files (`rocfft_kernel_cache_gfx950.db`, `rocfft_kernel_cache_gfx1201.db`, etc).
* `rocfft_plan_create` returns `rocfft_status_invalid_offset` for any usage of non-zero offsets in plan descriptions. The feature is not supported yet.
* Callback functions will be deprecated in a future release.
##### Resolved issues
* Potential issue with data generation for multi-dimensional transforms in rocfft-tests and rocfft-bench.
* An issue that sometimes blocked complex-to-complex FFT plan creation when using noncontiguous strides in multiple dimensions.
* An issue that sometimes blocked complex-to-real FFT plan creation when using noncontiguous strides in multiple dimensions.
* An issue that sometimes blocked complex-to-real FFT plan creation when using noncontiguous strides with small lengths on the two fastest dimensions.
* Potential launch failure of data generation kernels in test and benchmark programs.
* Incorrect results on some strided real-complex FFTs on gfx90a.
* Incorrect results on some even-length real FFTs that have odd-length strides on higher dimensions.
* Callbacks on MPI transforms when not all ranks have the same number of data bricks.
* Functional issues for multi-device, in-place real transforms.
* Functional issues for multi-dimensional, multi-device transforms involving some unit length(s).
* Functional issues for multi-device transforms involving data divisions along the slowest-varying axis (only) for some bricks but not all.
* Functional issues for multi-device transforms setting no field on input or output.
* Automatic allocation of work memory at plan execution time, when work memory is required on multiple devices.
#### **rocJPEG** (1.5.0)
##### Changed
* rocJPEG is now delivered as part of [TheRock](https://github.com/ROCm/TheRock). All core dependencies are provided by the TheRock build.
* Removed CPack packaging (DEB/RPM/NSIS/TGZ/ZIP generation and all related CPACK variables).
* Removed `rocJPEG-setup.py` dependency installer script.
* Removed Docker files.
* Removed package install documentation; updated all documentation to reference TheRock for installation.
* Simplified libva version check (single `>= 1.22` requirement).
* Cleaned up CMake error messages.
#### **ROCm Compute Profiler** (3.6.0)
##### Added
* Added L2 memory bandwidth derived metrics under `--membw-analysis` to allow L2 memory bandwidth specific profiling and analysis metric block 30.
* Added AMD Ryzen AI Max 300 series (gfx1151) support.
* New memory hierarchy visualization for RDNA 3.5 (gfx115X) in analyze CLI mode.
* Introduced support for AMD Instinct MI350P GPU.
* ``--view table`` option in analyze mode to force all TTY output to plain tables and ignore ``cli_style`` from YAML config (for example, mem_chart, Roofline charts render as tables). The ``--view`` argument is reserved for future TTY views (for example, other chart styles).
* Added EA memory bandwidth derived metrics under `--membw-analysis` to allow EA memory bandwidth specific profiling and analysis metric block 30.
##### Changed
* Standalone roofline (`--roof-only` option) in profile mode now creates `roofline.csv` only. HTML roofline charts are generated via `rocprof-compute analyze`. The `calc_ai_profile()` function has been removed; `calc_ai_analyze()` is the single source of truth for arithmetic intensity calculation.
* Roofline visualization options (`--sort`, `--mem-level`, `--roofline-data-type`) have moved from profile mode to analyze mode.
* Standardized unit naming in analysis configs and Python utilities: `pct`/`Pct``Percent`, `instr``Instructions`.
* Profile mode output format:
* Profile mode now creates separate counter collection files for each application replay (pmc_perf_*.csv or results_*.csv).
* Analyze mode automatically merges these files into a unified pmc_perf.csv containing information from all application replays during pre-processing.
* ROCm Compute Profiler now builds and runs profile mode with vanilla Python without requiring any Python dependencies to be installed via `pip`.
* Note that analysis mode will still require Python dependencies and will report any missing packages.
##### Removed
* Removed HIP API tracing since it's out-of-scope for ROCm Compute Profiler and the trace files were not being analyzed.
##### Optimized
* Filtering for block 21 (`-b 21`) in profile mode now only performs pc sampling and skips unnecessary counter collection.
* Filtering for block 21 in analysis mode now skips metrics calculations and only shows kernel/dispatch/system statistics and pc sampling table.
##### Resolved issues
* Fixed roofline benchmark MFMA FP16/BF16/INT8 peaks for MI350.
* Fixed an issue where pc sampling profiling failed with multi-argument commands and live process attachment.
##### Upcoming changes
* `--path` and `--subpath` options are deprecated and will be removed in a future release.
* Intermediate CSV generation (`results_*.csv`) from rocpd databases during profiling is deprecated and will be removed in a future release. The analyze step will read `.db` files directly.
* `--retain-rocpd-output` is deprecated and will be removed in a future release. `.db` files will be retained by default.
##### Known issues
* For AMD Ryzen AI Max 300 series, the roofline metrics table will have N/A values for "peak" field.
* This is planned to be addressed by adding empirical benchmark support for AMD Ryzen AI Max 300 series in a future release.
#### **ROCm Systems Profiler** (1.6.0)
##### Added
* Kernel Fusion Driver (KFD) event tracing support to capture page faults, page migrations, queue evictions, GPU unmap events, and dropped events. Requires ROCprofiler-SDK 1.2.1 or later. Enable with `ROCPROFSYS_ROCM_DOMAINS=kfd_events`.
* Support for pause and resume of profiling via `roctxProfilerPause` and `roctxProfilerResume`.
* Support for selective region tracing via the `ROCPROFSYS_SELECTED_REGIONS` environment variable, limiting tracing to specified regions.
* `--selected-regions` CLI argument to `rocprof-sys-sample`, `rocprof-sys-run`, and `rocprof-sys-instrument` for specifying selective region tracing from the command line.
* Support for re-attaching to a previously profiled process. After detaching, `rocprof-sys-attach` can re-attach to the same PID for a new profiling session.
* MPI-rank-based file output filtering feature controlled with two new CLI arguments: `--rank-filter-output` and `--rank-filter-id`.
* JSON-based configurable preset system with `--preset=<name>` flag, replacing the old `--<preset-name>` flags. Presets are now loaded from JSON files in `source/bin/common/presets/`, making them extensible and exportable. Use `--list-presets` to see available presets and `--explain=<name>` for detailed preset information.
* Domain flags for composable configuration: `--gpu[=metrics]`, `--rocm[=domains]`, `--cpu[=hz]`, `--parallel[=runtimes]`. Domain flags can be combined with presets to customize profiling without editing configuration files.
* Configuration export via `--export-config[=file]` to save resolved settings as reusable JSON configuration files. Exported configs can be loaded back with `--preset=./config.json`.
* Topic-based help system: `--help` now shows a compact summary with essential options and a list of help topics. Use `--help=<topic>` (e.g., `--help=sampling`, `--help=gpu`, `--help=tracing`) to see only relevant options. Use `--help=all` for the full option listing.
* Post-run output summary during library finalization showing result file locations.
* JSON schema file (`share/rocprofiler-systems/presets/schema.json`) for preset validation.
* Documentation (`docs/how-to/instrumenting-rewriting-binary-application.rst`) describing what to do when Dyninst reports a "Failed to transform trace" error during instrumentation.
##### Changed
* `rocprof-sys-avail` no longer queries GPU devices or hardware counters unless `--hw-counters` or `--all` is requested, reducing startup time and allowing settings/component queries in environments without GPU/ROCm.
* `rocprof-sys-instrument` diagnostic file dumps (available, instrumented, excluded, coverage, overlapping) are now gated behind the `--dump-info` flag instead of being generated unconditionally.
* Preset flags changed from `--balanced` to `--preset=balanced` syntax. The old `--<preset-name>` flags are still supported and handled within `preset_registry`.
* Removed the `ROCPROFSYS_USE_ROCM` CMake option. ROCm is now required for building the ROCm Systems Profiler.
##### Resolved issues
* Fixed an issue where the `--rocm-domains` CLI option for `rocprof-sys-run` was not recognized.
#### **rocminfo** (1.0.0)
##### Resolved issues
* Fixed BDF (Bus:Device.Function) ID truncation issue that caused incorrect display of PCI device identifiers. The `bdf_id` field was incorrectly declared as `uint16_t` instead of `uint32_t`, causing silent truncation when HSA runtime returned the full 32-bit BDF ID value. This has been corrected to properly display complete BDF information for all GPU agents.
#### **rocPRIM** (4.4.0)
##### Added
* Added type trait definitions for `__hip_bfloat16`. This should resolve issues where this type did not work with radix-based algorithms.
* Unit tests for config_types.
##### Optimized
* Reduced build times for unit tests.
* Reduced memory usage in unit tests.
##### Resolved issues
* Fixed a silent overflow in `rocprim::device_segmented_reduce` where it could exceed the maximum number of HIP threads, resulting in missing output.
* Certain large unit tests now properly detect if insufficient system memory is present and skip the test case accordingly.
* Fixed out-of-bounds memory access in block run length decode.
* Fixed memory leak in unit tests.
#### **ROCprofiler-SDK** (1.3.0)
##### Added
**API:**
* Late-start profiling support: Enables profiling when `rocprofiler-sdk` is loaded after HSA/HIP runtimes have already initialized.
* `rocprofiler_force_configure()` now automatically detects and profiles runtimes initialized before the SDK loads.
* Integrates with `rocprofiler-register` to retrieve the registered API tables.
* Supports all runtime types (HSA, HIP, ROCTX, RCCL, rocDecode, rocJPEG, and more) automatically.
* No explicit late-start API calls required; works transparently.
* KFD (Kernel Fusion Driver) event tracing support:
* Buffer service configurations for each KFD buffer tracing type.
* New type `tool_buffer_tracing_kfd_record_t` using `std::variant` to wrap 8 different KFD buffer tracing types.
* Each KFD event generates `rocpd_info_pmc`, `rocpd_event`, `rocpd_region`, and `rocpd_pmc_event` rows.
* Fixed handling for special SVM location in KFD prefetch location reporting.
* Fixed parsing for queue restore events to handle both correct format (character '0') and broken driver format (NULL character '\0').
**rocprofv3 (CLI):**
* Multi-pass counter collection support: Support for multiple `--pmc` flags to define separate counter groups for different profiling passes.
* Ability to combine command-line `--pmc` flags with input file counter groups.
* Each pass generates output in a separate `pass_n` subdirectory.
* Example: `rocprofv3 --pmc SQ_WAVES --pmc GRBM_COUNT -- <app>` creates two profiling passes.
* KFD (Kernel Fusion Driver) event tracing support:
* KFD record dumping to `rocpd` with support for 8 main KFD event types.
* Support for `rocpd` to Perfetto conversion for KFD events.
* `--kfd-trace` flag to enable KFD event tracing.
* ROCTx support for ATT: Added ROCtx support to device thread trace when using `--att --selected-regions`.
* Allows `roctxProfilerPause` and `roctxProfilerResume` to explicitly control when ATT data collection starts and stops.
* Enables more precise, region-focused ATT tracing with reduced overhead and noise.
* Supports multiple resume/pause cycles, each producing separate trace output files.
* Incompatible with `--att-consecutive-kernels`.
* PC sampling support for dynamic attach: Allows users to attach to a running application and collect PC samples without restarting the workload.
* Enables profiling long-running or production-style jobs at the point of interest.
* Results integrate with the existing PC sampling analysis flow.
**Documentation:**
* Added marker-controlled thread tracing section to the thread trace how-to guide.
* Added cross-reference from ROCTx documentation to ATT with `selected-regions`.
##### Changed
**Implementation:**
* Late-start architecture redesign: Removed direct runtime symbol access in favor of proper rocprofiler-register integration.
* Replaced ~600 lines of `dlopen`/`dlsym` bypass logic with ~80 lines by using `rocprofiler_register_invoke_all_registrations()`.
* Late-start now works by requesting `rocprofiler-register` to re-propagate stored API tables.
* Extensible design. Automatically supports new runtimes without SDK code changes.
* Provides a proper separation of concerns. `rocprofiler-register` manages the table storage while SDK manages the table wrapping.
* Counter dimension encoding changed from fixed-width to variable-width allocation per dimension type.
* Dimension selection and reduction logic now uses explicit dimension masks and single-index selection.
* HSA queue interception extended to handle AMD extended kernel dispatch packets.
##### Removed
* Counter collection support for plain text (`.txt`) input files. Only structured file formats (JSON and YAML) with schema validation are now supported.
##### Resolved issues
* Fixed rocpd OTF2 output to add `ACCELERATOR_DEVICE` as system tree node domain for AMD devices.
* Fixed `rocprofv3` input file parsing where comment lines containing `pmc:` were incorrectly processed as valid counter collection directives, causing unintended profiling passes.
#### **rocRAND** (4.4.0)
##### Added
* gfx1150 and gfx1152 support.
* rocrand.dll now contains embedded file version metadata.
##### Resolved issues
* Fixed memory leak in unit tests.
#### **rocSHMEM** (3.4.0)
##### Added
* Added new APIs:
* `rocshmem_quiet_on_stream`
* `rocshmem_sync_all_on_stream`
* `rocshmem_TYPENAME_alltoall_wg`
* `rocshmem_TYPENAME_alltoallv_wg`
* `rocshmem_team_my_pe`
* `rocshmem_team_n_pes`
* `rocshmem_barrier`
* `rocshmem_barrier_wave`
* `rocshmem_barrier_wg`
* `rocshmem_buffer_register`
* `rocshmem_buffer_unregister`
* `rocshmem_info_get_version`
* `rocshmem_info_get_name`
* `rocshmem_vendor_get_version_info`
* Added library constants: `ROCSHMEM_MAJOR_VERSION`, `ROCSHMEM_MINOR_VERSION`,
`ROCSHMEM_MAX_NAME_LEN`, `ROCSHMEM_VENDOR_STRING`, `ROCSHMEM_VERSION`,
`ROCSHMEM_VENDOR_MAJOR_VERSION`, `ROCSHMEM_VENDOR_MINOR_VERSION`,
`ROCSHMEM_VENDOR_PATCH_VERSION`.
* Added vendor string and backend metadata to the `rocshmem_info` output.
* Added `ROCSHMEM_TEAM_WORLD` for device code.
* Added `ROCSHMEM_TEAM_SHARED` predefined team for PEs sharing a common memory domain (same node).
* Added new environment variables:
* `ROCSHMEM_GDA_OVERRIDE_NIC_FIRMWARE_CHECK`
* `ROCSHMEM_GDA_NUM_QPS_PER_PE_DEFAULT_CTX`
* `ROCSHMEM_GDA_NUM_QPS_PER_PE_USR_CTX`
* Added VMM POSIX memory allocator (`USE_HEAP_DEVICE_VMM_POSIX`):
* Uses HIP Virtual Memory Management (VMM) APIs for fine-grained memory control.
* Requires ROCm 7.0+ and Linux kernel 5.6+.
* Not compatible with MPI-based initialization (use `ROCSHMEM_INIT_WITH_UNIQUEID` instead).
##### Changed
* Use CQ collapsing for the Mellanox MLX5 GDA conduit.
#### **rocSOLVER** (3.34.0)
##### Added
* Computation of solution for LU factorization without pivoting:
* GETRS_NPVT (with batched and strided\_batched versions)
* GETRS_NPVT_64 (with batched and strided\_batched versions)
* Linear solver routines for symmetric matrices:
* SYTRS (with batched and strided\_batched versions)
* SYTRS_64 (with batched and strided\_batched versions)
##### Optimized
* Improved the performance of POTF2 and downstream functions such as POTRF.
##### Resolved issues
* Fixed a memory access error in SYTRF and synchronization issues in LASYF and SYTF2.
#### **rocSPARSE** (4.6.0)
##### Added
* `rocsparse_create_const_bsr_descr` routine for creating a const sparse BSR matrix descriptor.
* `rocsparse_spic0` and `rocsparse_spilu0` routines for incomplete factorizations, with strided batched computations enabled.
* `rocsparse_sptrsv_descr_create` and `rocsparse_sptrsv_descr_destroy` routines.
* `rocsparse_singularity` enumeration.
* `rocsparse_sptrsv_output_singularity` and `rocsparse_sptrsv_output_singularity_position` in `rocsparse_sptrsv_output`.
* Strided batched computations for `rocsparse_sptrsv`.
##### Optimized
* Significant performance improvement for `rocsparse_Xgtsv_no_pivot_strided_batch`.
* Significant performance improvement for `rocsparse_Xgtsv_no_pivot`.
##### Resolved issues
* Fixed incorrect usage of `__syncthreads` in `bsrmm`, `csrmm` (row_split), and `csritilu0x`.
* Fixed incorrect usage of `__syncthreads` in `csx2dense`, `dense2csx`, `prune_dense2csr`, `csrcolor`, and `csrmm` (`nnz_split`).
* Fixed `rocsparse_[s|d|c|z]csric0` where `rocsparse_status_invalid_value` was being returned when the maximum number of non-zeros in any row is between 513 and 1024.
* Fixed compilation when using `--rocsparse_ILP64`.
* Fixed off-by-one heap-buffer-overflow in temporary buffer allocation for `rocsparse_csrsort`, `rocsparse_check_matrix_csr`, and `rocsparse_check_matrix_gebsr` (and their delegating routines `rocsparse_cscsort`, `rocsparse_coosort`, `rocsparse_check_matrix_csc`, and `rocsparse_check_matrix_gebsc`) where the `shift_offsets_kernel` temp buffer was sized for `m` elements instead of `m+1`.
##### Removed
* The deprecated C++14 support, which is no longer supported by the rocPRIM dependency.
#### **rocThrust** (4.4.0)
##### Resolved issues
* Fixed memory leak in unit test.
* Fixed unit test compatibility with ASAN.
#### **rocWMMA** (2.2.1)
##### Added
* Added the following community samples for external contributions, with build support and documentation:
* `simple_gemm_silu`: demonstrates a GEMM + SiLU fused operator using the rocWMMA API.
* `simple_gemm_fusion`: demonstrates block-tile-level dual-GEMM fusion using the rocWMMA API.
* `simple_gemm_swiglu`: demonstrates a SwiGLU fused dual-GEMM kernel (LLaMA/Mistral FFN gate layer) using the rocWMMA API.
##### Changed
* Updated the `find_package` search for OpenMP to prefer the `openmp-config.cmake` provided by ROCm, with a fallback to module search mode.
* Updated `INSTALL_RPATH` and added `BUILD_RPATH` for OpenMP.
##### Resolved issues
* Improved HIP RTC regression test portability when deployed outside the default path.
@@ -0,0 +1,193 @@
<table class="rocm-docs-table table">
<thead>
<tr>
<th class="head"><p>Component group</p></th>
<th class="head"><p>Component name</p></th>
<th class="head"><p>Version</p></th>
<th class="head"><p>Supported platforms</p></th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="18" style="vertical-align: middle;">
<p>Math and compute libraries</p>
</td>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipblas">hipBLAS</a></td>
<td><a href="#hipblas-3-4-0">3.4.0</a></td>
<td rowspan="16" style="vertical-align: middle;">Linux/Windows · Instinct/Radeon/Ryzen</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipblaslt">hipBLASLt</a></td>
<td><a href="#hipblaslt-1-3-0">1.3.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipcub">hipCUB</a></td>
<td><a href="#hipcub-4-4-0">4.4.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipfft">hipFFT</a></td>
<td><a href="#hipfft-1-0-23">1.0.23</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hiprand">hipRAND</a></td>
<td><a href="#hiprand-3-3-0">3.3.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipsolver">hipSOLVER</a></td>
<td><a href="#hipsolver-3-4-0">3.4.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipsparse">hipSPARSE</a></td>
<td>4.5.0</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/miopen">MIOpen</a></td>
<td><a href="#miopen-3-5-1">3.5.1</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocblas">rocBLAS</a></td>
<td><a href="#rocblas-5-4-0">5.4.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocfft">rocFFT</a></td>
<td><a href="#rocfft-1-0-37">1.0.37</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocrand">rocRAND</a></td>
<td><a href="#rocrand-4-4-0">4.4.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocsolver">rocSOLVER</a></td>
<td><a href="#rocsolver-3-34-0">3.34.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocsparse">rocSPARSE</a></td>
<td><a href="#rocsparse-4-6-0">4.6.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocprim">rocPRIM</a></td>
<td><a href="#rocprim-4-4-0">4.4.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocthrust">rocThrust</a></td>
<td><a href="#rocthrust-4-4-0">4.4.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/rocwmma">rocWMMA</a></td>
<td><a href="#rocwmma-2-2-1">2.2.1</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/composablekernel">Composable
Kernel</a></td>
<td><a href="#composable-kernel-1-3-0">1.3.0</a></td>
<td>Linux/Windows · Instinct/Radeon</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-libraries/tree/therock-7.13/projects/hipsparselt">hipSPARSELt</a></td>
<td><a href="#hipsparselt-0-2-8">0.2.8</a></td>
<td>Linux/Windows · Instinct (gfx950/gfx942)</td>
</tr>
<tr>
<td rowspan="2" style="vertical-align: middle;">
<p>Communication libraries</p>
</td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rccl">RCCL</a></td>
<td>2.28.3</td>
<td>Linux · Instinct/Radeon/Ryzen</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocshmem">rocSHMEM</a></td>
<td><a href="#rocshmem-3-4-0">3.4.0</a></td>
<td>Linux · Instinct (gfx950/gfx942/gfx90a) · Radeon (gfx1201/gfx1200/gfx1100/gfx1101/gfx1102)</td>
</tr>
<tr>
<td rowspan="2" style="vertical-align: middle;">
<p>Media libraries</p>
</td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocdecode">rocDecode</a></td>
<td><a href="#rocdecode-1-8-0">1.8.0</a></td>
<td rowspan="2" style="vertical-align: middle;">Linux · Instinct/Radeon · Ryzen (gfx1150/gfx1151/gfx1152)</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocjpeg">rocJPEG</a></td>
<td><a href="#rocjpeg-1-5-0">1.5.0</a></td>
</tr>
<tr>
<td rowspan="5" style="vertical-align: middle;">
<p>Runtimes and compilers</p>
</td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/hip">HIP</a></td>
<td><a href="#hip-7-13">7.13</a></td>
<td rowspan="4" style="vertical-align: middle;">Linux/Windows · Instinct/Radeon/Ryzen</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/HIPIFY/tree/therock-7.13">HIPIFY</a></td>
<td>7.13</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/llvm-project/tree/therock-7.13">LLVM</a></td>
<td>23.0.0</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/SPIRV-LLVM-Translator/tree/therock-7.13">SPIRV-LLVM-Translator</a></td>
<td>23.0.0</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocr-runtime">ROCr Runtime</a></td>
<td>1.21.0</td>
<td>Linux · Instinct/Radeon/Ryzen</td>
</tr>
<tr>
<td rowspan="6" style="vertical-align: middle;">
<p>Profiling and debugging tools</p>
</td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocprofiler-compute">ROCm
Compute Profiler (rocprofiler-compute)</a></td>
<td><a href="#rocm-compute-profiler-3-6-0">3.6.0</a></td>
<td rowspan="2" style="vertical-align: middle;">Linux · Instinct · Ryzen (gfx1150/gfx1151/gfx1152)</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocprofiler-systems">ROCm
Systems Profiler (rocprofiler-systems)</a></td>
<td><a href="#rocm-systems-profiler-1-6-0">1.6.0</a></td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocprofiler-sdk">ROCprofiler-SDK</a></td>
<td><a href="#rocprofiler-sdk-1-3-0">1.3.0</a></td>
<td>Linux · Instinct/Radeon · Ryzen (gfx1150/gfx1151/gfx1152)</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocdbgapi">ROCdbgapi</a></td>
<td><a href="#rocdbgapi-0-80-0">0.80.0</a></td>
<td rowspan="3" style="vertical-align: middle;">Linux · Instinct/Radeon</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/ROCgdb/tree/therock-7.13">ROCm Debugger (ROCgdb)</a></td>
<td>16.3</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocr-debug-agent">ROCr Debug
Agent</a></td>
<td>2.1.0</td>
</tr>
<tr>
<td rowspan="3" style="vertical-align: middle;">
<p>Control and monitoring tools</p>
</td>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/amdsmi">AMD SMI (BM)</a></td>
<td><a href="#amd-smi-bm-26-4-0">26.4.0</a></td>
<td>Linux · Instinct/Radeon</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rocminfo">rocminfo</a></td>
<td><a href="#rocminfo-1-0-0">1.0.0</a></td>
<td>Linux · Instinct/Radeon/Ryzen</td>
</tr>
<tr>
<td><a href="https://github.com/ROCm/rocm-systems/tree/therock-7.13/projects/rdc">ROCm Data Center Tool
(RDC)</a></td>
<td><a href="#rocm-data-center-tool-rdc-1-3-0">1.3.0</a></td>
<td>Linux · Instinct</td>
</tr>
</tbody>
</table>
@@ -0,0 +1,268 @@
::::{tab-set}
:::{tab-item} Instinct
:sync: instinct
<table class="rocm-docs-table table">
<colgroup style="width: 25%;">
<thead>
<tr>
<th class="head">
<p>AMD device</p>
</th>
<th class="head">
<p>Firmware</p>
</th>
<th class="head">
<p>Linux driver</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<td>
<p>Instinct MI355X</p>
</td>
<td rowspan="2" style="vertical-align: middle">
<p>PLDM bundle 01.26.00.02</p>
</td>
<td rowspan="10" style="vertical-align: middle">
<p>
<strong>AMD GPU Driver (amdgpu)</strong><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/31.30.0-preview/documentation/release-notes.html"
target="_blank"
>31.30.0</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/31.20.0-preview/documentation/release-notes.html"
target="_blank"
>31.20.0</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/31.10.0-preview/documentation/release-notes.html"
target="_blank"
>31.10.0</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.3/documentation/release-notes.html"
target="_blank"
>30.30.3</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.2/documentation/release-notes.html"
target="_blank"
>30.30.2</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.1/documentation/release-notes.html"
target="_blank"
>30.30.1</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.0/documentation/release-notes.html"
target="_blank"
>30.30.0</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.20.1/documentation/release-notes.html"
target="_blank"
>30.20.1</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.20.0/documentation/release-notes.html"
target="_blank"
>30.20.0</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.10.2/documentation/release-notes.html"
target="_blank"
>30.10.2</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.10.1/documentation/release-notes.html"
target="_blank"
>30.10.1</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.10/documentation/release-notes.html"
target="_blank"
>30.10.0</a><br>
</p>
</td>
</tr>
<tr>
<td>
<p>Instinct MI350X</p>
</td>
</tr>
<tr>
<td>
<p>Instinct MI350P</p>
</td>
<td style="vertical-align: middle">
<p>IFWI 00185129</p>
</td>
</tr>
<tr>
<td>
<p>Instinct MI325X</p>
</td>
<td style="vertical-align: middle">
<p>PLDM bundle 01.25.04.02</p>
</td>
</tr>
<tr>
<td>
<p>Instinct MI300X</p>
</td>
<td>
<p>PLDM bundle 01.26.00.02</p>
</td>
</tr>
<tr>
<td>
<p>Instinct MI300A</p>
</td>
<td>
<p>BKC 26.1</p>
</td>
</tr>
<tr>
<td>
<p>Instinct MI250X</p>
</td>
<td>
<p>IFWI 75 (or later)</p>
</td>
</tr>
<tr>
<td>
<p>Instinct MI250</p>
</td>
<td rowspan="2">
<p>Maintenance update (MU) 5 with IFWI 75 (or later)</p>
</td>
</tr>
<tr>
<td>
<p>Instinct MI210</p>
</td>
</tr>
<tr>
<td>
<p>Instinct MI100</p>
</td>
<td>
<p>VBIOS D3430401-037</p>
</td>
</tr>
</tbody>
</table>
:::
:::{tab-item} Radeon
:sync: radeon
<table class="rocm-docs-table table">
<colgroup style="width: 50%;">
<thead>
<tr>
<th class="head">
<p>Linux driver</p>
</th>
<th class="head">
<p>Windows driver</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<td style="vertical-align: middle">
<p>
<strong>AMD GPU Driver (amdgpu)</strong><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/31.30.0-preview/documentation/release-notes.html"
target="_blank"
>31.30.0</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/31.20.0-preview/documentation/release-notes.html"
target="_blank"
>31.20.0</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/31.10.0-preview/documentation/release-notes.html"
target="_blank"
>31.10.0</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.3/documentation/release-notes.html"
target="_blank"
>30.30.3</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.2/documentation/release-notes.html"
target="_blank"
>30.30.2</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.1/documentation/release-notes.html"
target="_blank"
>30.30.1</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.30.0/documentation/release-notes.html"
target="_blank"
>30.30.0</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.20.1/documentation/release-notes.html"
target="_blank"
>30.20.1</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.20.0/documentation/release-notes.html"
target="_blank"
>30.20.0</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.10.2/documentation/release-notes.html"
target="_blank"
>30.10.2</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.10.1/documentation/release-notes.html"
target="_blank"
>30.10.1</a><br>
<a
href="https://instinct.docs.amd.com/projects/amdgpu-docs/en/docs-30.10/documentation/release-notes.html"
target="_blank"
>30.10.0</a><br>
</p>
</td>
<td style="vertical-align: middle">
<p>
<strong>AMD Software: Adrenalin Edition</strong>
<a
href="https://www.amd.com/en/resources/support-articles/release-notes/RN-RAD-WIN-26-5-1.html"
target="_blank"
>26.5.1</a>
</td>
</tr>
</tbody>
</table>
:::
:::{tab-item} Ryzen
:sync: ryzen
<table class="rocm-docs-table table">
<colgroup style="width: 50%;">
<thead>
<tr>
<th class="head">
<p>Linux driver</p>
</th>
<th class="head">
<p>Windows driver</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<td style="vertical-align: middle">
<p>Inbox kernel driver in Ubuntu 26.04 or 24.04.4</p>
</td>
<td rowspan="30" style="vertical-align: middle">
<p>
<strong>AMD Software: Adrenalin Edition</strong>
<a
href="https://www.amd.com/en/resources/support-articles/release-notes/RN-RAD-WIN-26-5-1.html"
target="_blank"
>26.5.1</a>
</p>
</td>
</tr>
</tbody>
</table>
:::
::::
@@ -0,0 +1,535 @@
::::{tab-set}
:::{tab-item} Instinct
:sync: instinct
<table class="rocm-docs-table table">
<thead>
<colgroup style="width: 33%;">
<colgroup style="width: 32%;">
<tr>
<th class="head">
<p>Device series</p>
</th>
<th class="head">
<p>Device</p>
</th>
<th class="head">
<p>LLVM target</p>
</th>
<th class="head">
<p>Architecture</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<td class="stub">
<a href="https://www.amd.com/en/products/accelerators/instinct/mi350.html" target="_blank">AMD Instinct MI350
Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi350/mi355x.html" target="_blank">Instinct
MI355X</a></p>
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi350/mi350x.html" target="_blank">Instinct
MI350X</a></p>
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi350/mi350p.html" target="_blank">Instinct
MI350P</a></p>
</td>
<td>
<p>gfx950</p>
</td>
<td>
<a href="https://www.amd.com/en/technologies/cdna.html#cdna4" target="_blank">CDNA 4</a>
</td>
</tr>
<tr>
<td class="stub">
<a href="https://www.amd.com/en/products/accelerators/instinct/mi300.html" target="_blank">AMD Instinct MI300
Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi300/mi325x.html" target="_blank">Instinct
MI325X</a></p>
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi300/mi300x.html" target="_blank">Instinct
MI300X</a></p>
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi300/mi300a.html" target="_blank">Instinct
MI300A</a></p>
</td>
<td>
<p>gfx942</p>
</td>
<td>
<a href="https://www.amd.com/en/technologies/cdna.html#cdna3" target="_blank">CDNA 3</a>
</td>
</tr>
<tr>
<td class="stub">
<a href="https://www.amd.com/en/products/accelerators/instinct/mi200.html" target="_blank">AMD Instinct MI200
Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi200/mi250x.html" target="_blank">Instinct
MI250X</a></p>
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi200/mi250.html" target="_blank">Instinct
MI250</a></p>
<p><a href="https://www.amd.com/en/products/accelerators/instinct/mi200/mi210.html" target="_blank">Instinct
MI210</a></p>
</td>
<td>
<p>gfx90a</p>
</td>
<td>
<a href="https://www.amd.com/en/technologies/cdna.html#cdna2" target="_blank">CDNA 2</a>
</td>
</tr>
<tr>
<td class="stub">
<a href="https://www.amd.com/en/products/accelerators/instinct/mi100.html" target="_blank">AMD Instinct MI100
Series</a>
</td>
<td>
<a href="https://www.amd.com/en/products/accelerators/instinct/mi100.html" target="_blank">Instinct MI100</a>
</td>
<td>
<p>gfx908</p>
</td>
<td>
<a href="https://www.amd.com/en/technologies/cdna.html#cdna" target="_blank">CDNA</a>
</td>
</tr>
</tbody>
</table>
:::
:::{tab-item} Radeon
:sync: radeon
<table class="rocm-docs-table table">
<thead>
<colgroup style="width: 33%;">
<colgroup style="width: 32%;">
<tr>
<th class="head">
<p>Device series</p>
</th>
<th class="head">
<p>Device</p>
</th>
<th class="head">
<p>LLVM target</p>
</th>
<th class="head">
<p>Architecture</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<td>
<a href="https://www.amd.com/en/products/graphics/workstations/radeon-ai-pro.html#tabs-95fa144b96-item-b95ec9e1ca-tab"
target="_blank">AMD Radeon AI PRO R9000 Series</a>
</td>
<td>
<p><a
href="https://www.amd.com/en/products/graphics/workstations/radeon-ai-pro/ai-9000-series/amd-radeon-ai-pro-r9700.html"
target="_blank">Radeon AI PRO R9700</a></p>
<p><a
href="https://www.amd.com/en/products/graphics/workstations/radeon-ai-pro/ai-9000-series/amd-radeon-ai-pro-r9600d.html"
target="_blank">Radeon AI PRO R9600D</a></p>
</td>
<td>
<p>gfx1201</p>
</td>
<td rowspan="3">
<a href="https://www.amd.com/en/technologies/rdna.html#tabs-1fabb91c39-item-330ee548f0-tab" target="_blank">RDNA
4</p>
</td>
</tr>
<tr>
<td rowspan="2" class="stub">
<a href="https://www.amd.com/en/products/graphics/desktops/radeon.html#tabs-ff9c5c3863-item-37fb38a236-tab"
target="_blank">AMD Radeon RX 9000 Series</p>
</td>
<td>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9070xt.html"
target="_blank">Radeon RX 9070 XT</a></p>
<p><a href="https://www.amd.com/en/support/downloads/drivers.html/graphics/radeon-rx/radeon-rx-9000-series/amd-radeon-rx-9070-gre.html"
target="_blank">Radeon RX 9070 GRE</a></p>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9070.html"
target="_blank">Radeon RX 9070</a></p>
</td>
<td>
<p>gfx1201</p>
</td>
</tr>
<tr>
<td>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9060xt-lp.html"
target="_blank">Radeon RX 9060 XT LP</a></p>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9060xt.html"
target="_blank">Radeon RX 9060 XT</a></p>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9060.html"
target="_blank">Radeon RX 9060</a></p>
</td>
<td>
<p>gfx1200</p>
</td>
</tr>
<tr>
<td rowspan="2" class="stub">
<a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro.html#tabs-990fdead92-item-20daa37284-tab"
target="_blank">AMD Radeon PRO W7000 Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w7900-dual-slot.html"
target="_blank">Radeon PRO W7900 Dual Slot</a></p>
<p><a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w7900.html" target="_blank">Radeon
PRO W7900</a></p>
<p><a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w7800-48gb.html"
target="_blank">Radeon PRO W7800 48GB</a></p>
<p><a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w7800.html" target="_blank">Radeon
PRO W7800</a></p>
</td>
<td>
<p>gfx1100</p>
</td>
<td rowspan="6">
<a href="https://www.amd.com/en/technologies/rdna.html#tabs-1fabb91c39-item-05915f6044-tab" target="_blank">RDNA
3</p>
</td>
</tr>
<tr>
<td>
<p><a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w7700.html" target="_blank">Radeon
PRO W7700</a></p>
</td>
<td>
<p>gfx1101</p>
</td>
</tr>
<tr>
<td rowspan="3" class="stub">
<a href="https://www.amd.com/en/products/graphics/desktops/radeon.html#tabs-ff9c5c3863-item-b55a56bf12-tab"
target="_blank">AMD Radeon RX 7000 Series</p>
</td>
<td>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7900xtx.html"
target="_blank">Radeon RX 7900 XTX</a></p>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7900xt.html"
target="_blank">Radeon RX 7900 XT</a></p>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7900-gre.html"
target="_blank">Radeon RX 7900 GRE</a></p>
</td>
<td>
<p>gfx1100</p>
</td>
</tr>
<tr>
<td>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7800-xt.html"
target="_blank">Radeon RX 7800 XT</a></p>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7700-xt.html"
target="_blank">Radeon RX 7700 XT</a></p>
<p>Radeon RX 7700 XE</p>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7700.html"
target="_blank">Radeon RX 7700</a></p>
</td>
<td>
<p>gfx1101</p>
</td>
</tr>
<tr>
<td>
<p><a href="https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7600.html"
target="_blank">Radeon RX 7600</a></p>
</td>
<td>
<p>gfx1102</p>
</td>
</tr>
<tr>
<td rowspan="2">
<a href="https://www.amd.com/en/products/accelerators/radeon-pro.html" target="_blank">AMD Radeon PRO V
Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/accelerators/radeon-pro/amd-radeon-pro-v710.html"
target="_blank">Radeon PRO V710</a></p>
</td>
<td>
<p>gfx1101</p>
</td>
</tr>
<tr>
<td>
<p><a href="https://www.amd.com/en/products/accelerators/radeon-pro/amd-radeon-pro-v620.html"
target="_blank">Radeon PRO V620</a></p>
</td>
<td>
<p>gfx1030</p>
</td>
<td rowspan="2">
<a href="https://www.amd.com/en/technologies/rdna.html#tabs-1fabb91c39-item-9ed969eddf-tab" target="_blank">RDNA
2</p>
</td>
</tr>
<tr>
<td class="stub">
<a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w6800.html" target="_blank">AMD Radeon
PRO W6000 Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/graphics/workstations/radeon-pro/w6800.html" target="_blank">Radeon
PRO W6800</a></p>
</td>
<td>
<p>gfx1030</p>
</td>
</tr>
</tbody>
</table>
:::
:::{tab-item} Ryzen
:sync: ryzen
<table class="rocm-docs-table table">
<thead>
<colgroup style="width: 26%;">
<colgroup style="width: 40%;">
<tr>
<th class="head">
<p>Device series</p>
</th>
<th class="head">
<p>Device</p>
</th>
<th class="head">
<p>LLVM target</p>
</th>
<th class="head">
<p>Architecture</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<td class="stub">
<a href="https://www.amd.com/en/products/processors/workstations/mobile.html#tabs-7f0c432fb2-item-5116ab7a74-tab"
target="_blank">AMD Ryzen AI Max PRO<br>300 Series</a>
</td>
<td>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-max-pro-300-series/amd-ryzen-ai-max-plus-pro-395.html"
target="_blank">Ryzen AI Max+ PRO 395</a> (Radeon 8060S)</p>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-max-pro-300-series/amd-ryzen-ai-max-pro-390.html"
target="_blank">Ryzen AI Max PRO 390</a> (Radeon 8050S)</p>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-max-pro-300-series/amd-ryzen-ai-max-pro-385.html"
target="_blank">Ryzen AI Max PRO 385</a> (Radeon 8050S)</p>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-max-pro-300-series/amd-ryzen-ai-max-pro-380.html"
target="_blank">Ryzen AI Max PRO 380</a> (Radeon 8040S)</p>
</td>
<td>
<p>gfx1151</p>
</td>
<td rowspan="10">
<p>RDNA 3.5</p>
</td>
</tr>
<tr>
<td class="stub">
<a href="https://www.amd.com/en/products/processors/laptop/ryzen.html#tabs-1181ea0b44-item-6ccfea5f65-tab"
target="_blank">AMD Ryzen AI Max<br>300 Series</a>
</td>
<td>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-plus-395.html"
target="_blank">Ryzen AI Max+ 395</a> (Radeon 8060S)</p>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-plus-392.html"
target="_blank">Ryzen AI Max+ 392</a> (Radeon 8060S)</p>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-plus-388.html"
target="_blank">Ryzen AI Max+ 388</a> (Radeon 8060S)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-390.html"
target="_blank">Ryzen AI Max 390</a> (Radeon 8050S)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-385.html"
target="_blank">Ryzen AI Max 385</a> (Radeon 8050S)</p>
</td>
<td>
<p>gfx1151</p>
</td>
</tr>
<tr>
<td rowspan="2" class="stub">
<a href="https://www.amd.com/en/products/processors/laptop/ryzen-for-business.html#tabs-0d174caf43-item-87690677fc-tab"
target="_blank">AMD Ryzen AI PRO<br>400 Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-400-series/amd-ryzen-ai-9-hx-pro-475.html"
target="_blank">Ryzen AI 9 HX PRO 475</a> (Radeon 890M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-400-series/amd-ryzen-ai-9-hx-pro-470.html"
target="_blank">Ryzen AI 9 HX PRO 470</a> (Radeon 890M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-400-series/amd-ryzen-ai-9-pro-465.html"
target="_blank">Ryzen AI 9 PRO 465</a> (Radeon 880M)</p>
</td>
<td>
<p>gfx1150</p>
</td>
</tr>
<tr>
<td>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-400-series/amd-ryzen-ai-7-pro-450.html"
target="_blank">Ryzen AI 7 PRO 450</a> (Radeon 860M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-400-series/amd-ryzen-ai-5-pro-440.html"
target="_blank">Ryzen AI 5 PRO 440</a> (Radeon 840M)</p>
</td>
<td>
<p>gfx1152</p>
</td>
</tr>
<tr>
<td rowspan="2" class="stub">
<a href="https://www.amd.com/en/products/processors/consumer/ryzen-ai.html#tabs-f556098628-item-808b56dca3-tab"
target="_blank">AMD Ryzen AI<br>400 Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-400-series/amd-ryzen-ai-9-hx-475.html"
target="_blank">Ryzen AI 9 HX 475</a> (Radeon 890M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-400-series/amd-ryzen-ai-9-hx-470.html"
target="_blank">Ryzen AI 9 HX 470</a> (Radeon 890M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-400-series/amd-ryzen-ai-9-465.html"
target="_blank">Ryzen AI 9 465</a> (Radeon 880M)</p>
</td>
<td>
<p>gfx1150</p>
</td>
</tr>
<tr>
<td>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-400-series/amd-ryzen-ai-7-450.html"
target="_blank">Ryzen AI 7 450</a> (Radeon 860M)</p>
</td>
<td>
<p>gfx1152</p>
</td>
</tr>
<tr>
<td rowspan="2" class="stub">
<a href="https://www.amd.com/en/products/processors/workstations/mobile.html#tabs-7f0c432fb2-item-387526c6cc-tab"
target="_blank">AMD Ryzen AI PRO<br>300 Series</a>
</td>
<td>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-300-series/amd-ryzen-ai-9-hx-pro-375.html"
target="_blank">Ryzen AI 9 HX PRO 375</a> (Radeon 890M)</p>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-300-series/amd-ryzen-ai-9-hx-pro-370.html"
target="_blank">Ryzen AI 9 HX PRO 370</a> (Radeon 890M)</p>
</td>
<td>
<p>gfx1150</p>
</td>
</tr>
<tr>
<td>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-300-series/amd-ryzen-ai-7-pro-350.html"
target="_blank">Ryzen AI 7 PRO 350</a> (Radeon 860M)</p>
<p><a
href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-300-series/amd-ryzen-ai-5-pro-340.html"
target="_blank">Ryzen AI 5 PRO 340</a> (Radeon 840M)</p>
</td>
<td>
<p>gfx1152</p>
</td>
</tr>
<tr>
<td rowspan="2" class="stub">
<a href="https://www.amd.com/en/products/processors/consumer/ryzen-ai.html#tabs-f556098628-item-54e149d850-tab"
target="_blank">AMD Ryzen AI<br>300 Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-9-hx-375.html"
target="_blank">Ryzen AI 9 HX 375</a> (Radeon 890M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-9-hx-370.html"
target="_blank">Ryzen AI 9 HX 370</a> (Radeon 890M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-9-365.html"
target="_blank">Ryzen AI 9 365</a> (Radeon 880M)</p>
</td>
<td>
<p>gfx1150</p>
</td>
</tr>
<tr>
<td>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-7-350.html"
target="_blank">Ryzen AI 7 350</a> (Radeon 860M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-7-345.html"
target="_blank">Ryzen AI 7 345</a> (Radeon 840M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-5-340.html"
target="_blank">Ryzen AI 5 340</a> (Radeon 840M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-5-330.html"
target="_blank">Ryzen AI 5 330</a> (Radeon 820M)</p>
</td>
<td>
<p>gfx1152</p>
</td>
</tr>
<tr>
<td class="stub">
<a href="https://www.amd.com/en/products/processors/laptop/ryzen-for-business.html#tabs-0d174caf43-item-a8ec88d07e-tab"
target="_blank">AMD Ryzen PRO<br>200 Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/200-series/amd-ryzen-7-pro-250.html"
target="_blank">Ryzen 7 PRO 250</a> (Radeon 780M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/200-series/amd-ryzen-5-pro-230.html"
target="_blank">Ryzen 5 PRO 230</a> (Radeon 760M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/200-series/amd-ryzen-5-pro-220.html"
target="_blank">Ryzen 5 PRO 220</a> (Radeon 740M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/200-series/amd-ryzen-5-pro-215.html"
target="_blank">Ryzen 5 PRO 215</a> (Radeon 740M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen-pro/200-series/amd-ryzen-3-pro-210.html"
target="_blank">Ryzen 3 PRO 210</a> (Radeon 740M)</p>
</td>
<td>
<p>gfx1103</p>
</td>
<td rowspan="2">
<a href="https://www.amd.com/en/technologies/rdna.html#tabs-1fabb91c39-item-05915f6044-tab" target="_blank">RDNA
3</a>
</td>
</tr>
<tr>
<td class="stub">
<a href="https://www.amd.com/en/products/processors/laptop/ryzen.html#tabs-1181ea0b44-item-895d56feed-tab"
target="_blank">AMD Ryzen<br>200 Series</a>
</td>
<td>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-9-270.html"
target="_blank">Ryzen 9 270</a> (Radeon 780M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-7-260.html"
target="_blank">Ryzen 7 260</a> (Radeon 780M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-7-250.html"
target="_blank">Ryzen 7 250</a> (Radeon 780M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-5-240.html"
target="_blank">Ryzen 5 240</a> (Radeon 760M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-5-230.html"
target="_blank">Ryzen 5 230</a> (Radeon 760M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-5-220.html"
target="_blank">Ryzen 5 220</a> (Radeon 740M)</p>
<p><a href="https://www.amd.com/en/products/processors/laptop/ryzen/200-series/amd-ryzen-3-210.html"
target="_blank">Ryzen 3 210</a> (Radeon 740M)</p>
</td>
<td>
<p>gfx1103</p>
</td>
</tr>
</tbody>
</table>
:::
::::
+311
View File
@@ -0,0 +1,311 @@
::::{tab-set}
:::{tab-item} Instinct
:sync: instinct
<table class="rocm-docs-table table">
<thead>
<colgroup style="width: 33%;">
<tr>
<th class="head">
<p>Linux distribution</p>
</th>
<th class="head">
<p>Supported versions</p>
</th>
<th class="head">
<p>Linux kernel version</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<th rowspan="3" class="stub" style="vertical-align: middle">
<p>Ubuntu</p>
</th>
<td>
<p>26.04</p>
</td>
<td>
<p>GA 7.0</p>
</td>
</tr>
<tr>
<td>
<p>24.04.4</p>
</td>
<td>
<p>GA 6.8</p>
</td>
</tr>
<tr>
<td>
<p>22.04.5</p>
</td>
<td>
<p>GA 5.15</p>
</td>
</tr>
<tr>
<th rowspan="2" class="stub" style="vertical-align: middle">
<p>Debian</p>
</th>
<td>
<p>13</p>
</td>
<td>
<p>6.12</p>
</td>
</tr>
<tr>
<td>
<p>12</p>
</td>
<td>
<p>6.1.0</p>
</td>
</tr>
<tr>
<th rowspan="6" class="stub" style="vertical-align: middle">
<p>Red Hat Enterprise Linux (RHEL)</p>
</th>
<td>
<p>10.1</p>
</td>
<td>
<p>6.12.0-124</p>
</td>
</tr>
<tr>
<td>
<p>10.0</p>
</td>
<td>
<p>6.12.0-55</p>
</td>
</tr>
<tr>
<td>
<p>9.7</p>
</td>
<td>
<p>5.14.0-611</p>
</td>
</tr>
<tr>
<td>
<p>9.6</p>
</td>
<td>
<p>5.14.0-570</p>
</td>
</tr>
<tr>
<td>
<p>9.4</p>
</td>
<td>
<p>5.14.0-427</p>
</td>
</tr>
<tr>
<td>
<p>8.10</p>
</td>
<td>
<p>4.18.0-553</p>
</td>
</tr>
<tr>
<th rowspan="3" class="stub" style="vertical-align: middle">
<p>Oracle Linux</p>
</th>
<td>
<p>10</p>
</td>
<td>
<p>UEK 8.1</p>
</td>
</tr>
<tr>
<td>
<p>9</p>
</td>
<td>
<p>UEK 8</p>
</td>
</tr>
<tr>
<td>
<p>8</p>
</td>
<td>
<p>UEK 7</p>
</td>
</tr>
<tr>
<th class="stub" style="vertical-align: middle">
<p>Rocky Linux</p>
</th>
<td>
<p>9</p>
</td>
<td>
<p>5.14.0-570</p>
</td>
</tr>
<tr>
<th rowspan="2" class="stub" style="vertical-align: middle">
<p>SUSE Linux Enterprise Server (SLES)</p>
</th>
<td>
<p>16.0</p>
</td>
<td>
<p>6.12</p>
</td>
</tr>
<tr>
<td>
<p>15.7</p>
</td>
<td>
<p>6.4.0-150700.51</p>
</td>
</tr>
</tbody>
</table>
:::
:::{tab-item} Radeon
:sync: radeon
<table class="rocm-docs-table table">
<thead>
<colgroup style="width: 33%;">
<tr>
<th class="head">
<p>Operating system</p>
</th>
<th class="head">
<p>Supported versions</p>
</th>
<th class="head">
<p>Linux kernel version</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<th rowspan="3" class="stub" style="vertical-align: middle">
<p>Ubuntu</p>
</th>
<td>
<p>26.04</p>
</td>
<td>
<p>GA 7.0</p>
</td>
</tr>
<tr>
<td>
<p>24.04.4</p>
</td>
<td>
<p>GA 6.8</p>
</td>
</tr>
<tr>
<td>
<p>22.04.5</p>
</td>
<td>
<p>GA 5.15</p>
</td>
</tr>
<tr>
<th rowspan="2" class="stub" style="vertical-align: middle">
<p>Red Hat Enterprise Linux (RHEL)</p>
</th>
<td>
<p>10.1</p>
</td>
<td>
<p>6.12.0-124</p>
</td>
</tr>
<tr>
<td>
<p>9.7</p>
</td>
<td>
<p>5.14.0-611</p>
</td>
</tr>
<tr>
<th class="stub" style="vertical-align: middle">
<p>Windows</p>
</th>
<td>
<p>11 25H2</p>
</td>
<td>
<p style="text-align: center;"> — </p>
</td>
</tr>
</tbody>
</table>
:::
:::{tab-item} Ryzen
:sync: ryzen
<table class="rocm-docs-table table">
<thead>
<colgroup style="width: 33%;">
<tr>
<th class="head">
<p>Operating system</p>
</th>
<th class="head">
<p>Supported versions</p>
</th>
<th class="head">
<p>Linux kernel version</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<th rowspan="2" class="stub" style="vertical-align: middle">
<p>Ubuntu</p>
</th>
<td>
<p>26.04</p>
</td>
<td>
<p>GA 7.0</p>
</td>
</tr>
<tr>
<td>
<p>24.04.4</p>
</td>
<td>
<p>HWE 6.17</p>
</td>
</tr>
<tr>
<th class="stub" style="vertical-align: middle">
<p>Windows</p>
</th>
<td>
<p>11 25H2</p>
</td>
<td>
<p style="text-align: center;"> — </p>
</td>
</tr>
</tbody>
</table>
:::
::::
@@ -0,0 +1,69 @@
<table class="rocm-docs-table table">
<thead>
<tr>
<th class="head">
<p>Device</p>
</th>
<th class="head">
<p>Compute partition mode</p>
</th>
<th class="head">
<p>NPS mode</p>
</th>
<th class="head">
<p>Deployment</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="3" style="vertical-align: middle;">
<p>Instinct MI355X, MI350X</p>
</td>
<td>
<p>CPX</p>
</td>
<td>
<p>NPS 2</p>
</td>
<td rowspan="5" style="vertical-align: middle;">
<p>Bare metal</p>
</td>
</tr>
<tr>
<td>
<p>DPX</p>
</td>
<td>
<p>NPS 2</p>
</td>
</tr>
<tr>
<td>
<p>QPX</p>
</td>
<td>
<p>NPS 2</p>
</td>
</tr>
<tr>
<td rowspan="2" style="vertical-align: middle;">
<p>Instinct MI300X</p>
</td>
<td>
<p>CPX</p>
</td>
<td>
<p>NPS 4</p>
</td>
</tr>
<tr>
<td>
<p>DPX</p>
</td>
<td>
<p>NPS 2</p>
</td>
</tr>
</tbody>
</table>
@@ -0,0 +1,227 @@
<table class="rocm-docs-table table">
<colgroup style="width: 14%;">
<colgroup style="width: 14%;">
<colgroup style="width: 17%;">
<colgroup style="width: 17%;">
<colgroup style="width: 19%;">
<colgroup style="width: 19%;">
<thead>
<tr>
<th class="head">
<p>AMD GPU</p>
</th>
<th class="head">
<p>Hypervisor</p>
</th>
<th class="head">
<p>Virtualization technology</p>
</th>
<th class="head">
<a>Virtualization driver</a>
</th>
<th class="head">
<p>Host OS</p>
</th>
<th class="head">
<p>Guest OS</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="5" style="vertical-align: middle">
<p>Instinct MI355X</p>
</td>
<td rowspan="4" style="vertical-align: middle">
<p>KVM</p>
</td>
<td style="vertical-align: middle">
<p>Passthrough</p>
</td>
<td>
<p style="text-align: center"></p>
</td>
<td rowspan="4" style="vertical-align: middle">
<p>Ubuntu 24.04</p>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 24.04</p>
</td>
</tr>
<tr>
<td rowspan="3" style="vertical-align: middle">
<p>SR-IOV</p>
</td>
<td rowspan="3" style="vertical-align: middle">
<a href="https://github.com/amd/MxGPU-Virtualization/releases/tag/9.0.0.K" target="_blank">GIM 9.0.0K
</a>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 24.04</p>
</td>
</tr>
<tr>
<td style="vertical-align: middle">
<p>RHEL 10.0</p>
</td>
</tr>
<tr>
<td style="vertical-align: middle">
<p>RHEL 9.6</p>
</td>
</tr>
<tr>
<td style="vertical-align: middle">
<p>ESXi</p>
</td>
<td>
<p style="text-align: center"></p>
</td>
<td>
<p style="text-align: center"></p>
</td>
<td style="vertical-align: middle">
<p>VMware ESXi 9.1</p>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 24.04</p>
</td>
</tr>
<tr>
<td rowspan="3" style="vertical-align: middle">
<p>Instinct MI350X</p>
</td>
<td rowspan="3" style="vertical-align: middle">
<p>KVM</p>
</td>
<td style="vertical-align: middle">
<p>Passthrough</p>
</td>
<td>
<p style="text-align: center"></p>
</td>
<td rowspan="3" style="vertical-align: middle">
<p>Ubuntu 24.04</p>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 24.04</p>
</td>
</tr>
<tr>
<td rowspan="2" style="vertical-align: middle">
<p>SR-IOV</p>
</td>
<td rowspan="2" style="vertical-align: middle">
<a href="https://github.com/amd/MxGPU-Virtualization/releases/tag/9.0.0.K" target="_blank">GIM 9.0.0K
</a>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 24.04</p>
</td>
</tr>
<tr>
<td style="vertical-align: middle">RHEL 9.6</td>
</tr>
<tr>
<td style="vertical-align: middle">
<p>Instinct MI325X</p>
</td>
<td style="vertical-align: middle">
<p>KVM</p>
</td>
<td style="vertical-align: middle">
<p>SR-IOV</p>
</td>
<td style="vertical-align: middle">
<a href="https://github.com/amd/MxGPU-Virtualization/releases/tag/9.0.0.K" target="_blank">GIM 9.0.0K
</a>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 22.04</p>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 22.04</p>
</td>
</tr>
<tr>
<td rowspan="3" style="vertical-align: middle">
<p>Instinct MI300X</p>
</td>
<td rowspan="3" style="vertical-align: middle">
<p>KVM</p>
</td>
<td style="vertical-align: middle">
<p>Passthrough</p>
</td>
<td>
<p style="text-align: center"></p>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 22.04</p>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 22.04</p>
</td>
</tr>
<tr>
<td rowspan="2" style="vertical-align: middle">
<p>SR-IOV</p>
</td>
<td rowspan="2" style="vertical-align: middle">
<a href="https://github.com/amd/MxGPU-Virtualization/releases/tag/9.0.0.K" target="_blank">GIM 9.0.0K
</a>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 24.04</p>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 24.04</p>
</td>
</tr>
<tr>
<td style="vertical-align: middle">
<p>Ubuntu 22.04</p>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 22.04</p>
</td>
</tr>
<tr>
<td rowspan="3" style="vertical-align: middle">
<p>Instinct MI210</p>
</td>
<td rowspan="3" style="vertical-align: middle">
<p>KVM</p>
</td>
<td style="vertical-align: middle">
<p>Passthrough</p>
</td>
<td>
<p style="text-align: center"></p>
</td>
<td rowspan="3" style="vertical-align: middle">
<p>RHEL 9.4</p>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 22.04</p>
</td>
</tr>
<tr>
<td rowspan="2" style="vertical-align: middle">
<p>SR-IOV</p>
</td>
<td rowspan="2" style="vertical-align: middle">
<a href="https://github.com/amd/MxGPU-Virtualization/releases/tag/9.0.0.K" target="_blank">GIM 9.0.0K
</a>
</td>
<td style="vertical-align: middle">
<p>Ubuntu 22.04</p>
</td>
</tr>
<tr>
<td style="vertical-align: middle">
<p>RHEL 9.4</p>
</td>
</tr>
</tbody>
</table>
+1 -1
View File
@@ -4,7 +4,7 @@
<meta name="keywords" content="license, licensing terms">
</head>
# ROCm license
# ROCm licenses
```{include} ../../LICENSE
```
+608
View File
@@ -0,0 +1,608 @@
# ROCm Core SDK {{ ROCM_VERSION }} release notes
ROCm Core SDK {{ ROCM_VERSION }} continues the technology preview release stream
that began with ROCm 7.9.0, advancing the transition to the new
[TheRock](https://github.com/rocm/therock) build and release system. To learn
more, see the [transition guide](/about/transition-guide-TheRock).
(preview-stream-note)=
:::{important}
ROCm {{ ROCM_VERSION }} follows the
<a href="https://rocm.docs.amd.com/en/7.9.0-preview/about/release-notes.html#preview-stream-note"
target="_blank">versioning discontinuity that began with the 7.9.0 preview release</a>
and remains separate from the 7.0 to 7.2 production releases. For the latest
production stream release, see the
<a href="https://rocm.docs.amd.com/en/latest/">ROCm documentation</a>.
Maintaining parallel release streams -- preview and production -- gives
users ample time to evaluate and adopt the new build system and dependency
changes. The technology preview stream is planned to continue through
mid-2026, after which it will replace the current production stream.
For previous preview releases, see the
<a target="_blank" href="https://rocm.docs.amd.com/en/7.12.0-preview/release/versions.html">release history</a>.
:::
## Release highlights
ROCm Core SDK {{ ROCM_VERSION }} with TheRock builds upon the [7.12.0 preview
release](https://rocm.docs.amd.com/en/7.12.0-preview/about/release-notes.html).
This release expands support for AI inference, distributed workloads, and
profiling workflows across AMD Instinct™, Radeon™, and Ryzen™ AI platforms.
ROCm 7.13.0 adds inference-ready vLLM containers, expands GPU virtualization
and partitioning support, introduces new profiling and tracing capabilities,
and improves AI kernel, sparse math, and communication libraries.
### Platform and hardware support
This release expands GPU, operating system, virtualization, and partitioning support.
#### Expanded AMD GPU support
ROCm 7.13.0 adds support for the following AMD GPUs and APUs:
* AMD Instinct MI350P (gfx950)
* AMD Radeon PRO W6800 (gfx1030)
* AMD Radeon PRO V620 (gfx1030)
* AMD Ryzen AI 7 PRO 360 (gfx1152)
* AMD Ryzen AI 7 PRO 350 (gfx1152)
* AMD Ryzen AI 5 PRO 340 (gfx1152)
* AMD Ryzen AI 7 350 (gfx1152)
* AMD Ryzen AI 7 345 (gfx1152)
* AMD Ryzen AI 5 340 (gfx1152)
* AMD Ryzen AI 5 330 (gfx1152)
For the complete list of supported AMD hardware, see [AMD hardware support](#amd-hardware-support).
#### Expanded Ubuntu support
ROCm 7.13.0 adds support for Ubuntu 26.04 on Instinct, Radeon, and Ryzen
devices.
24.04.4 is now the validated Ubuntu 24 release instead of Ubuntu 24.04.3.
For the full list of supported Linux distributions, see [Operating system support](#operating-system-support).
#### Expanded GPU virtualization support for Instinct GPUs
ROCm 7.13.0 adds support for the following virtualization configurations on AMD Instinct GPUs.
* On MI355X: VMware ESXi 9.1 with Ubuntu 24.04 guest OS.
* On MI300X: KVM SR-IOV with Ubuntu 24.04 host OS and Ubuntu 24.04 guest OS.
* On MI210:
* KVM passthrough with RHEL 9.4 host OS and Ubuntu 22.04 guest OS.
* KVM SR-IOV with RHEL 9.4 host OS and Ubuntu 22.04 guest OS.
* KVM SR-IOV with RHEL 9.4 host OS and RHEL 9.4 guest OS.
Supported SR-IOV configurations require the [GIM Driver
9.0.0K](https://github.com/amd/MxGPU-Virtualization/releases/tag/9.0.0.K). For
details, see [GPU virtualization support](#gpu-virtualization-support).
#### Expanded Instinct GPU partitioning support
ROCm 7.13.0 enables the QPX compute + NPS 2 memory partition combination in
bare metal deployments.
For details, see [GPU partitioning support](#gpu-partitioning-support).
### AI inference and frameworks
This release adds inference-ready container images and improves multi-node communication for distributed workloads.
#### vLLM 0.19.1 Docker images and pip packages
With ROCm 7.13.0, Docker images for running vLLM inference workloads are
available. Images include vLLM 0.19.1, PyTorch 2.10, and Python 3.13 on Ubuntu 24.04.
Architecture-specific images are available for:
* AMD Instinct GPUs: gfx942 (MI325X, MI300X, MI300A) and gfx950 (MI355X, MI350X, MI350P)
* AMD Radeon GPUs: gfx1100, gfx1101, gfx1102, gfx1200, gfx1201
* AMD Ryzen AI APUs: gfx1150, gfx1151, gfx1152
See [](../ai-inference/vllm) to get started.
#### RCCL multi-node optimization for AMD Ryzen AI Max 300 series
RCCL improves multi-node clustering performance on systems with AMD Ryzen AI
Max 300 series connected over Ethernet. Building on the initial
multi-node enablement in ROCm 7.12.0, this release optimizes collective
communication for distributed AI inference workloads using tensor parallelism
(TP) and expert parallelism (EP) across up to 4 Ethernet-connected nodes.
#### RCCL GDA-based alltoall via rocSHMEM integration (experimental)
RCCL adds experimental support for GPU Direct Async (GDA)-based alltoall and
alltoallv collective operations through rocSHMEM integration. When enabled,
RCCL invokes rocSHMEM operations that use GDA to reduce latency for small
message alltoall patterns.
This feature requires building RCCL with the `--rocshmem` flag and setting
`RCCL_ROCSHMEM_ENABLE=1` at runtime. GDA support currently requires Broadcom
NICs with GDA capability.
### Developer tools and profiling
This release adds new profiling capabilities, introduces the open-source ROCprof Trace Decoder, and extends HIP programming APIs.
#### ROCprof Trace Decoder open source release
ROCprof Trace Decoder, previously delivered as a closed-source
component within ROCprofiler-SDK, is now available as the open-source
rocprof-trace-decoder library. The decoder converts raw SQTT data from AMD GPUs
into structured execution traces for performance analysis and debugging. It
supports a wide range of AMD GPUs spanning Instinct, Radeon, and Ryzen
architectures, with unit and integration tests across all supported hardware.
See [AMD hardware support](#amd-hardware-support) for the complete list.
<!-- For more information, see [ROCprof Trace Decoder and thread trace APIs](https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/docs-7.13.0/api-reference/thread_trace.html) and [Using thread trace](https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/docs-7.13.0/how-to/using-thread-trace.html) in the ROCprofiler-SDK documentation. -->
#### HIP cooperative groups reduce operations
HIP adds `cooperative_groups::reduce()` for performing reduction operations
across `thread_block_tile` and `coalesced_threads` groups. The implementation
is based on `__reduce_*_sync` operations, and the
`HIP_ENABLE_EXTRA_WARP_SYNC_TYPES` macro might be required to enable some
optimizations.
Additionally, `__reduce_and_sync()`, `__reduce_or_sync()`, and
`__reduce_xor_sync()` now provide consistent behavior for all mask values. All
masks now emit bitwise instructions, aligning behavior with NVIDIA CUDA. This
is a change from previous versions, where some masks were translated to bitwise
operations, and others were not.
#### ROCm Compute Profiler feature highlights
The following are notable enhancements to the ROCm Compute Profiler
(rocprofiler-compute).
* **RDNA 3.5 support:** ROCm Compute Profiler now supports GPU performance
profiling and analysis on AMD Ryzen AI Max 300 series processors.
<!-- An [RDNA 3 -->
<!-- section](https://rocm.docs.amd.com/projects/rocprofiler-compute/en/docs-7.13.0/conceptual/rdna/rdna-performance-model.html) -->
<!-- has been added to the performance model documentation explaining the supported -->
<!-- performance metrics for AMD Ryzen AI Max 300 series processors. A new memory -->
<!-- chart visualization accommodates the architectural differences between -->
<!-- RDNA 3.5 and CDNA GPUs. Roofline is not yet supported for AMD Ryzen AI -->
<!-- Max 300 series processors. -->
* **Removed dependency requirements for profiling:** Building ROCm Compute
Profiler and using profile mode no longer requires installing Python
dependencies from the `requirements.txt` file. Analysis mode still requires
Python dependencies.
This change moves several operations from profile mode to analysis mode,
including roofline HTML generation, roofline-related options
(`--sort`, `--mem-level`, `--roofline-data-type`), and creation of the
combined `pmc_perf.csv` file. Profile mode now only runs the roofline
empirical benchmark, creates a `roofline.csv` file, and creates per-replay
CSV files without merging them.
#### ROCm Systems Profiler feature highlights
The following are notable enhancements to the ROCm Systems Profiler
(rocprofiler-systems).
* **Pause and resume profiling:** ROCm Systems Profiler now supports pausing
and resuming profiling at runtime through the `roctxProfilerPause` and
`roctxProfilerResume` APIs. This allows you to capture profiling data only
during specific execution phases, reducing overhead and minimizing output size
for long-running workloads.
<!-- For more information, see [Configuring runtime -->
<!-- options](https://rocm.docs.amd.com/projects/rocprofiler-systems/en/docs-7.13.0/how-to/configuring-runtime-options.html) -->
<!-- in the ROCm Systems Profiler documentation. -->
* **Selective region tracing:** You can now restrict tracing to defined regions
of interest using the `ROCPROFSYS_SELECTED_REGIONS` environment variable,
reducing noise and limiting data collection to relevant workload segments.
<!-- For more -->
<!-- information, see -->
<!-- [ROCPROFSYS_SELECTED_REGIONS](https://rocm.docs.amd.com/projects/rocprofiler-systems/en/docs-7.13.0/how-to/configuring-runtime-options.html#rocprofsys-selected-regions) -->
<!-- in the ROCm Systems Profiler documentation. -->
* **KFD event tracing:** Kernel Fusion Driver (KFD) event tracing is now
available for GPU memory management analysis, including page faults, page
migrations, queue evictions, GPU unmap events, and dropped events. Requires
an XNACK-capable GPU and ROCprofiler-SDK 1.2.1 or later.
<!-- For more -->
<!-- information, see [Configuring runtime -->
<!-- options](https://rocm.docs.amd.com/projects/rocprofiler-systems/en/docs-7.13.0/how-to/configuring-runtime-options.html#exploring-gpu-metrics) -->
<!-- in the ROCm Systems Profiler documentation. -->
* **MPI file-output filtering:** You can now filter profiler output files based
on MPI rank using the `--rank-filter-output` CLI option or the
`ROCPROFSYS_RANK_FILTER_OUTPUT` configuration setting, suppressing output
from all other ranks. An optional `--rank-filter-id` option
(`ROCPROFSYS_RANK_FILTER_ID`) allows specifying a custom environment variable
for rank identification.
<!-- For more information, see [Selective rank -->
<!-- profiling](https://rocm.docs.amd.com/projects/rocprofiler-systems/en/docs-7.13.0/how-to/communication-runtime-profiling.html#selective-rank-profiling) -->
<!-- in the ROCm Systems Profiler documentation. -->
* **JSON-based profiling presets and domain flags:** You can now configure
common profiling workflows using JSON-based presets and a single
`--preset=<name>` flag instead of manually setting multiple `ROCPROFSYS_*`
environment variables. Eleven built-in presets cover common profiling scenarios, including GPU
tracing, HPC workloads, and API-level analysis. Composable domain flags
(`--gpu`, `--rocm`, `--cpu`, `--parallel`) and a topic-based
`--help=<topic>` system further simplify configuration and discoverability.
<!-- For more information, see [Using preset profiles and domain -->
<!-- flags](https://rocm.docs.amd.com/projects/rocprofiler-systems/en/docs-7.13.0/how-to/using-preset-profiles.html) -->
<!-- in the ROCm Systems Profiler documentation. -->
#### AMD SMI feature highlights
* **APU metrics and memory tuning**: New APU telemetry provides per-core
temperature, power, clock, voltage, current, and throttle monitoring, with
additional support for IPU activity and DRAM bandwidth metrics. New VRAM
carveout and GTT tuning controls enable configurable memory allocation on
supported APU platforms.
* **Per-component GPU temperature and clock monitoring**: GPU metrics table
version 1.9 adds HBM stack temperatures, per-die temperature monitoring, and
per-die memory and SOC clock reporting for data center deployments.
* **CPU power APIs report in milliwatts (breaking change)**: CPU power APIs now
return values in milliwatts (mW) instead of watts. Python bindings now return
numeric integer values instead of formatted strings. Existing applications
that parse previous string-based outputs must be updated.
For more information, see the AMD SMI section in the [ROCm component changelogs](#rocm-component-changelogs).
### Libraries
This release adds new routines, data type support, and performance improvements across ROCm math and AI libraries.
#### Composable Kernel adds quantization and attention kernel capabilities
Composable Kernel adds several capabilities for AI and large language model
workloads:
* **Microscaling (MX) FP8/FP4 support:** Mixed data type support for MX FP8 and
FP4 in GEMM and Flash Multi-Head Attention (FMHA) forward kernels on AMD
Instinct MI350 Series GPUs.
* **FP8 quantization for FMHA:** FMHA forward kernels now support multiple FP8
quantization modes, including dynamic tensor-wise quantization, block scale
quantization, per-tensor quantization, and FP8 KV cache support for batch
prefill.
* **StreamingLLM and long-context inference:** Sink token support for FMHA
forward enables StreamingLLM-style long-context inference.
* **Batch prefill enhancements:** FMHA batch prefill kernels now support
multiple KV cache layouts, flexible page sizes, and configurable lookup table
configurations.
* **RDNA 3 FMHA support:** Flash Attention kernels are now available on RDNA 3
architectures.
* **SageAttention v2 forward kernel:** Multi-granularity quantization for Q, K,
and V tensors with FP8, INT8, and INT4 data types and per-tensor, per-block,
per-warp, and per-thread scale granularities on AMD Instinct MI300 Series and
MI350 Series GPUs.
#### General Batched GEMM support in hipBLASLt
hipBLASLt adds native support for General Batched GEMM, where all matrices in
a batch share the same problem dimensions but can have independent leading
dimensions and strides. This replaces the previous implementation through the
`hipblaslt_ext` Grouped GEMM APIs, which had known limitations.
The new implementation includes support for Global Split-U (GSU) to improve
performance at large problem sizes. General Batched GEMM is important for
inference workloads that dispatch batches of same-shape GEMM operations.
<!-- For more information, see the [hipBLASLt -->
<!-- documentation](https://rocm.docs.amd.com/projects/hipBLASLt/en/docs-7.13.0/index.html). -->
#### rocSOLVER adds new solver routines and matrix analysis functions
rocSOLVER adds the following new routines, all with 64-bit index support:
* **GETRS_NPVT:** Solution of linear systems using LU factorization without
pivoting. Batched and strided-batched variants are available.
* **SYTRS:** Solution of linear systems for symmetric matrices. Batched and
strided-batched variants are available.
Additionally, POTF2 and downstream POTRF Cholesky factorization performance
have been improved.
<!-- For more information, see the [rocSOLVER -->
<!-- documentation](https://rocm.docs.amd.com/projects/rocSOLVER/en/docs-7.13.0/index.html). -->
#### rocSPARSE adds sparse factorization routines
rocSPARSE adds new generic API routines for sparse incomplete factorization and
triangular solve:
* `rocsparse_spic0` and `rocsparse_spilu0`: Generic incomplete Cholesky (IC0)
and incomplete LU (ILU0) factorization routines with strided-batched
computation support.
* `rocsparse_sptrsv`: Extended with strided-batched computation support and
singularity detection through the new `rocsparse_singularity` enumeration.
Performance of tridiagonal solvers `rocsparse_Xgtsv_no_pivot` and
`rocsparse_Xgtsv_no_pivot_strided_batch` has been improved.
<!-- For more -->
<!-- information, see the [rocSPARSE -->
<!-- documentation](https://rocm.docs.amd.com/projects/rocSPARSE/en/docs-7.13.0/index.html). -->
#### Added rocDecode and rocJPEG libraries to the ROCm Core SDK
rocDecode provides hardware-accelerated video decoding for H.264, H.265/HEVC,
AV1, and VP9 codecs, while rocJPEG provides hardware-accelerated JPEG decoding
on AMD GPUs. Together, they enable
efficient GPU-based media processing pipelines for data-intensive workloads
such as AI training.
Both libraries are supported on Linux on AMD Instinct, Radeon, and Ryzen AI. See
the projects in [ROCm/rocm-systems](https://github.com/ROCm/rocm-systems) for
more information.
#### Added ROCm Data Center Tool to the ROCm Core SDK
ROCm Data Center Tool (RDC) provides telemetry collection, health monitoring,
and job-level GPU statistics for data center deployments with AMD Instinct
accelerators. RDC enables system administrators and cluster managers to monitor
GPU health, collect telemetry data, and track per-job GPU usage across
multi-node environments.
RDC is supported on Linux with AMD Instinct GPUs.
<!-- See the -->
<!-- [RDC documentation](https://rocm.docs.amd.com/projects/rdc/en/docs-7.13.0/index.html) -->
<!-- for more information. -->
(release-supported-hw)=
## AMD hardware support
The following table lists supported AMD Instinct GPUs, Radeon GPUs, and Ryzen
APUs. Each supported device is listed with its corresponding GPU
microarchitecture and LLVM target.
:::{note}
If your GPU is not listed, it might be community-enabled through TheRock
nightly builds. For more information, see [TheRock supported
GPUs](https://github.com/ROCm/TheRock/blob/main/SUPPORTED_GPUS.md). For
installation guidance, see [TheRock
releases](https://github.com/ROCm/TheRock/blob/main/RELEASES.md).
:::
```{include} ./include/hardware-support-table.md
:parser: myst
```
(release-supported-os)=
## Operating system support
ROCm supports the following Linux distribution and Microsoft Windows versions.
If you're running ROCm on Linux, ensure your system is using a supported kernel
version.
:::{important}
The following table is a general overview of supported OSes. Actual support
might vary by AMD GPU or APU. Use the {doc}`Compatibility matrix
</compatibility/compatibility-matrix>` to verify support for your specific
setup before installation.
:::
```{include} ./include/os-support-table.md
:parser: myst
```
## Installation updates
ROCm 7.13.0 introduces several improvements to the Runfile Installer:
* Performance improvements for installing and uninstalling gfx architectures.
* ROCm component tests are now included.
* Support for prerequisite OEM kernel installation as part of the dependency install on Ryzen systems. You no longer need to install it manually.
* Auto-detection of the GPU when using the GUI or when the `gfx=` argument is not provided on the command line. If the installer cannot detect the GPU, you must specify the gfx architecture using the GUI or the `gfx=` argument.
(release-supported-fw)=
## Kernel driver and firmware bundle support
ROCm requires a coordinated stack of compatible firmware, driver, and user
space components. Maintaining version alignment between these layers ensures
correct GPU operation and performance, especially for AMD data center products.
While AMD publishes the AMD GPU driver and ROCm user space components, your
server OEM (original equipment manufacturer) or infrastructure provider
distributes the firmware packages. AMD supplies those firmware images (PLDM
bundles), which the OEM integrates and distributes.
```{include} ./include/driver-firmware-support-table.md
:parser: myst
```
(release-virtualization-support)=
## GPU virtualization support
AMD Instinct data center GPUs support virtualization in the following
configurations. Supported SR-IOV configurations require the AMD GPU
Virtualization Driver (GIM) 9.0.0K -- see the [AMD Instinct Virtualization
Driver
documentation](https://instinct.docs.amd.com/projects/virt-drv/en/mainline-9.0.0.k/)
for more information.
```{include} ./include/virtualization-support-table.html
:parser: myst
```
(release-gpu-partitioning-support)=
## GPU partitioning support
The following compute partition and NUMA-per-socket (NPS) configurations are
available on AMD Instinct GPUs in bare metal deployments.
```{include} ./include/partitioning-support-table.html
:parser: myst
```
See the [AMD GPU partitioning](https://instinct.docs.amd.com/projects/amdgpu-docs/en/latest/gpu-partitioning/index.html)
topic in the AMD GPU Driver documentation to learn more.
(release-ai-ecosystem)=
## AI ecosystem support
ROCm 7.13.0 provides optimized support for popular deep learning frameworks and
AI inference engines. The following table lists supported frameworks and
libraries, their compatible operating systems, and validated versions.
```{include} ./include/ai-ecosystem-support-table.html
:parser: myst
```
(release-components)=
## ROCm Core SDK components
The following table lists core tools and libraries included in the ROCm 7.13.0
release.
:::{important}
The following table is a general overview of ROCm Core SDK components. Actual
support for these libraries and tools can vary by GPU and OS. Use the
{doc}`Compatibility matrix </compatibility/compatibility-matrix>` to verify
support for your specific setup.
:::
```{include} ./include/core-sdk-components-table.html
:parser: myst
```
### ROCm component changelogs
The following sections describe key changes to ROCm Core SDK components.
```{include} ./include/core-sdk-components-aggregated-changelog.md
:parser: myst
```
## ROCm known issues
ROCm known issues are noted on {fab}`github` [GitHub](https://github.com/ROCm/ROCm/labels/Verified%20Issue). These issues will be fixed in a future ROCm release. For known issues related to individual components, review the [ROCm component changelogs](#rocm-component-changelogs).
### ROCm Compute Profiler might fail when profiling bash script or command
Running a bash script or command as a target for ROCm Compute Profiler might fail because bash overwrites the required environment variables. As a workaround, use `--no-native-tool` option in the profile mode. Note that this will disable iteration multiplexing.
### hipFFT and rocFFT callback examples fail to build on Windows
The hipFFT and rocFFT callback examples in [rocm-examples](https://github.com/rocm/rocm-examples) fail to build on a Windows operating system due to a linker error. CMake configuration and HIP object compilation will complete successfully, but the final link step fails with `clang: error: invalid linker name in argument '-fuse-ld=lld-link'` This issue affects all Windows configurations using Relocatable Device Code (RDC) mode. Linux builds are not affected. As a workaround, skip the hipFFT and rocFFT callback examples on Windows, and refer to the Linux builds or [callback](https://github.com/ROCm/rocm-examples/tree/amd-staging/Libraries/rocFFT/callback/) functionality documentation.
### QMCPACK might become unresponsive during DMC simulation on AMD Instinct MI300A GPUs
QMCPACK might become unresponsive when running Diffusion Monte Carlo (DMC) simulations with certain inputs on AMD Instinct MI300A GPUs. The application stops making progress after initialization and must be terminated manually.
### Resource-intensive workloads might result in GPU memory faults
Applications that pass large, complex data structures between device functions using scratch memory, and particularly rely on compiler optimization to minimize the number of copy operations, might encounter GPU memory access faults and become unresponsive.
### Increased binary size for multi-target GPU builds
Applications targeting multiple AMD GPU architectures might observe significantly larger binary sizes. Multi-target builds can produce binaries up to 54 percent larger. Single-target builds add approximately 8 MB of additional size per GPU target. As a workaround, reduce the number of GPU targets in multi-target builds, or strip the resource-usage symbols from release binaries.
### HIP cooperative groups might fail when compiled using the SPIR-V path
HIP applications that use cooperative groups might fail at kernel launch when compiled with `--offload-arch=amdgcnspirv`. The application fails at runtime with `LLVM ERROR: Cannot select: intrinsic %llvm.amdgcn.s.wait.asynccnt` error message. This
affects all GPU architectures when using the SPIR-V compilation path. As a workaround, compile using a direct GPU architecture target (for example, `--offload-arch=gfx942`) instead of `--offload-arch=amdgcnspirv`.
### Illegal memory address error when using placement new with device function returns
HIP kernels that use the placement new operators to construct objects in the `hipMalloc` device memory might crash with `hipErrorIllegalAddress` error message when you pass a `__device__` function return value as the constructor argument. This only affects non-trivially-copyable types (for example, types with user-defined or deleted copy/move constructors). Trivially-copyable types are not affected. As a workaround, assign the device function return value to a local variable before passing it to placement new.
### LLVM-based compilers might fail when compiling half-precision vector operations
LLVM-based compilers might fail, returning `Failed to find subregs!` error message in `SIInstrInfo::copyPhysReg`, when compiling half-precision vector operations with optimization enabled. The issue was observed at optimization levels `-O1` to `-O3`.
### hipBLAS test suites failure on Windows
When using hipBLAS on Windows, the test suites might return non-zero exit codes, even when all mathematical correctness tests pass. This issue can affect CI/CD pipeline validation and block automated testing workflows on Windows systems, because the test framework might fail to detect successful test completion.
### ROCm Systems Profiler overwrites ROCPD output after process re-attachment
When you use `rocprof-sys-attach` to re-attach to a previously profiled process, the `ROCPD` output database files (.db) are written to the initial session's output directory instead of a new timestamped directory. This makes it difficult to distinguish profiling data between sessions. Perfetto trace files are not affected. As a workaround, back up your output directory before re-attaching to a previously profiled process.
### Missing dependencies when installing ROCm Core SDK
Installing the ROCm Core SDK using `amdrocm-core-sdk` or `amdrocm-core-dev/devel` might succeed, but some dependencies from the dev/devel meta packages might not be installed. As a workaround, install the dev packages manually:
```bash
sudo apt install amdrocm-*
```
### Issues related to AddressSanitizer
Multiple issues associated with AddressSanitizer (ASAN) `-fsanitize=address` being enabled have been observed including:
#### ASAN reports false errors for GPU kernels using shared memory
When you compile GPU kernels with ASAN enabled, kernels that use `__shared__` memory might produce false heap-buffer-overflow errors or GPU memory faults. As a workaround, disable ASAN by removing `-fsanitize=address` setting for affected kernels.
#### GPU kernels fail to launch in ASAN builds with large thread counts
When you build GPU libraries with ASAN enabled, kernels configured with large thread counts might fail to launch with `HSA_STATUS_ERROR_INVALID_ISA` error. As a workaround, reduce the thread block sizes to 256 threads or fewer for ASAN builds. The issue is currently under investigation.
#### ASAN breaks multi-architecture HIP binary builds
HIP applications built with ASAN enabled, targeting multiple GPU architectures, might fail to launch with `RuntimeError: .hipFatBinSegment size N is not a multiple of wrapper size (24)` and `RuntimeError: Unexpected magic 0x00000000 at wrapper i` error messages. Single-architecture builds are not affected. As a workaround, build single-architecture binaries using `--offload-arch` targeting only one GPU architecture, or disable ASAN by removing `-fsanitize=address` for HIP compilation.
#### ASAN produces incorrect results with ternary operators on struct kernel arguments
When you compile GPU kernels with ASAN enabled, ternary operators with struct kernel arguments might produce incorrect results. This can mask real bugs and produce false-positive results during memory-safety validation. The issue doesn't occur when the kernel arguments are first copied to local variables, or when compiled without ASAN. As a workaround, copy kernel arguments to local variables before using them in ternary expressions:
```cpp
auto local_arg = kernel_arg;
result = condition ? local_arg : other_arg;
```
Alternatively, disable ASAN by removing `-fsanitize=address` when compiling GPU kernels.
## ROCm resolved issues
The following notable issues have been fixed in ROCm 7.13.0.
### Multi-ROCm installation failed on RPM-based distributions
Previously, installing multiple ROCm versions side by side on RPM-based distributions (RHEL and SLES) failed due to `.build-id` file conflicts between versioned packages.
### vLLM server failed to launch in ROCm Docker images
Previously, the vLLM server failed to start in ROCm 7.12.0 Docker images with an `ImportError` for `librocm_smi64.so.1` due to missing library path configuration.
### vLLM server failed to launch with tensor parallelism
Previously, the vLLM server failed to start with an invalid device pointer error when launching models with tensor parallelism set to 8 on AMD Instinct MI300 and MI355X GPUs.
### PyTorch DDP Gloo backend test failed on AMD GPUs
Previously, the PyTorch Distributed Data Parallel (DDP) test `test_ddp_apply_optim_in_backward_grad_as_bucket_view_false` failed when using the Gloo backend.
### rocWMMA header produced unknown type errors in HIP RTC
Previously, HIP RTC programs that included the `rocwmma/rocwmma.hpp` header failed to compile with unknown type name errors.
## ROCm upcoming changes
Future releases will add support for:
* Additional ROCm Core SDK components
* Domain-specific expansion toolkits (data science, life science, finance,
simulation, and other HPC domains)
* More AMD hardware support
+338
View File
@@ -0,0 +1,338 @@
# Transition guide from legacy ROCm release stream
[ROCm Core SDK 7.13.0](https://rocm-stg.amd.com/en/docs-7.13.0/index.html#rocm-core-sdk) marks a step change from the ROCm legacy release stream. It is a preview release built on our new build system, TheRock.
## Major changes
<table class="rocm-docs-table table">
<thead>
<tr>
<th class="head">Feature</th>
<th class="head">ROCm Core SDK</th>
<th class="head">ROCm Legacy</th>
<th class="head">Description</th>
</tr>
</thead>
<tbody>
<tr>
<td>Installation directory</td>
<td><code class="docutils literal notranslate"><span class="pre">/opt/rocm/core</span></code></td>
<td><code class="docutils literal notranslate"><span class="pre">/opt/rocm/</span></code></td>
<td>To support additional release streams downstream of the ROCm Core SDK</td>
</tr>
<tr>
<td>Package names</td>
<td><code class="docutils literal notranslate"><span class="pre">amdrocm-[$component]</span></code></td>
<td><code class="docutils literal notranslate"><span class="pre">rocm-[$component]</span></code> or <code class="docutils literal notranslate"><span class="pre">roc[$component]</span></code> or <code class="docutils literal notranslate"><span class="pre">hip[$component]</span></code></td>
<td>Unique package prefix to avoid conflicts with upstream packages</td>
</tr>
<tr>
<td>Extras directory</td>
<td><code class="docutils literal notranslate"><span class="pre">/opt/rocm/extras-7/</span></code></td>
<td>N/A</td>
<td>Shared install prefix scoped to each ROCm major version for projects built on the ROCm Core SDK</td>
</tr>
</tbody>
</table>
## Paths and linking
ROCm Core SDK 7.13.0 maintains ABI and API compatibility with the ROCm 7.2
legacy releases, so recompilation is not required. For installations using your
Linux distribution's package manager, the `amdrocm` meta package configures
`update-alternatives` and provides backward-compatible symlinks for
`/opt/rocm/bin`, `/opt/rocm/lib`, and other `/opt/rocm/` directories. For
tarball installs, update `PATH`, `LD_LIBRARY_PATH`, `ROCM_PATH`, or other
environment variables to reflect the new installation path (`/opt/rocm/core`).
## Software packages
ROCm Core SDK packages are more consolidated than the legacy ROCm release
stream. For example, hipBLAS and rocBLAS are now combined into one package,
`amdrocm-blas`. The table below lists new packages, their contents, and the
corresponding legacy packages.
> **Note:** ASAN packages are not available in 7.13.0 and are planned for a future release.
(linux-packages-available-in-rocm-7-13-0)=
### Linux packages available in ROCm 7.13.0
<table class="rocm-docs-table table">
<thead>
<tr>
<th class="head">ROCm Core SDK Package</th>
<th class="head">Package Contents</th>
<th class="head">ROCm Legacy Package</th>
</tr>
</thead>
<tbody>
<tr>
<td>amdrocm-amdsmi</td>
<td>amd-smi</td>
<td>amd-smi-lib, rocm-smi-lib</td>
</tr>
<tr>
<td>amdrocm-llvm</td>
<td>amdclang++, hipcc, flang</td>
<td>rocm-llvm, rocm-llvm-dev, Fortran compiler (included in rocm-llvm OpenMP runtime)</td>
</tr>
<tr>
<td>amdrocm-runtime</td>
<td>HIP, ROCR, runtime compilation</td>
<td>hip-runtime-amd, rocm-hip-runtime, rocm-language-runtime, hsa-rocr, comgr</td>
</tr>
<tr>
<td>amdrocm-fft</td>
<td>rocFFT, hipFFT, hipFFTW</td>
<td>rocfft, hipfft</td>
</tr>
<tr>
<td>amdrocm-blas</td>
<td>rocBLAS, hipBLAS, hipBLASLt, hipSPARSELt</td>
<td>rocblas, hipblas, hipblaslt, hipsparselt</td>
</tr>
<tr>
<td>amdrocm-sparse</td>
<td>rocSPARSE, hipSPARSE</td>
<td>rocsparse, hipsparse</td>
</tr>
<tr>
<td>amdrocm-solver</td>
<td>rocSOLVER, hipSOLVER</td>
<td>rocsolver, hipsolver, rocalution</td>
</tr>
<tr>
<td>amdrocm-dnn</td>
<td>hipDNN, MIOpen</td>
<td>miopen-hip</td>
</tr>
<tr>
<td>amdrocm-rand</td>
<td>rocRAND, hipRAND</td>
<td>rocrand, hiprand</td>
</tr>
<tr>
<td>amdrocm-ccl</td>
<td>rocPRIM, rocThrust, hipCUB</td>
<td>rocprim, rocthrust, hipcub, rocwmma</td>
</tr>
<tr>
<td>amdrocm-profiler</td>
<td>rocprofiler-systems, rocprofiler-compute, rocprofiler-sdk, roctracer</td>
<td>rocprofiler, rocprofiler-compute, rocprofiler-systems, rocprofiler-sdk, roctracer</td>
</tr>
<tr>
<td>amdrocm-profiler-base</td>
<td>rocprofiler-sdk, roctracer</td>
<td>rocprofiler-register, roctracer, hsa-amd-aqlprofile</td>
</tr>
<tr>
<td>amdrocm-base</td>
<td>rocminfo, rocm-core</td>
<td>rocm-core, rocminfo, rocm-cmake, half</td>
</tr>
<tr>
<td>amdrocm-ck</td>
<td>Composable Kernel</td>
<td>composablekernel</td>
</tr>
<tr>
<td>amdrocm-debugger</td>
<td>rocgdb, ROCdbgapi, ROCr Debug Agent</td>
<td>rocm-gdb, rocm-dbgapi, rocm-debug-agent</td>
</tr>
<tr>
<td>amdrocm-hipify</td>
<td>HIPIFY</td>
<td>hipify-clang</td>
</tr>
<tr>
<td>amdrocm-opencl</td>
<td>OpenCL runtime and ICD loader</td>
<td>rocm-opencl-runtime, rocm-opencl, hip-opencl</td>
</tr>
<tr>
<td>amdrocm-decode</td>
<td>rocDecode (newly included in the ROCm Core SDK)</td>
<td>rocdecode</td>
</tr>
<tr>
<td>amdrocm-jpeg</td>
<td>rocJPEG (newly included in the ROCm Core SDK)</td>
<td>rocjpeg</td>
</tr>
<tr>
<td>amdrocm-rccl</td>
<td>rccl</td>
<td>rccl</td>
</tr>
<tr>
<td>amdrocm-rocshmem</td>
<td>rocSHMEM</td>
<td>rocshmem</td>
</tr>
<tr>
<td>amdrocm-rdc</td>
<td>ROCm Data Center Tool (newly included in the ROCm Core SDK)</td>
<td>rdc</td>
</tr>
<tr>
<td>amdrocm-sysdeps</td>
<td>Bundled third-party dependencies (libdrm, libelf, numa, libVA)</td>
<td>System dependencies</td>
</tr>
</tbody>
</table>
Packages are offered in the following variants:
- **For all supported GPUs** -- works across all GPUs supported by ROCm (for example, `apt install amdrocm-core-sdk7.13`).
- **For a specific GPU architecture** -- smaller install size, but requires you to know the GPU installed in your system (for example, `apt install amdrocm-core-sdk7.13-gfx110x`).
Installing all GPU architectures is not required. You can install packages for a specific architecture, multiple architectures side by side, or all supported GPU architectures.
When redistributing software built on the ROCm Core SDK (for example, via containers), we recommend the all GPU package variant for broad hardware support. If disk footprint is a concern, you can use a single GPU architecture package variant instead.
### Architecture-specific packages available in ROCm 7.13.0
<table class="rocm-docs-table table">
<thead>
<tr>
<th class="head">Architecture Family</th>
<th class="head">Package Suffix</th>
<th class="head">Product Name (Not Exhaustive)</th>
</tr>
</thead>
<tbody>
<tr>
<td>CDNA4</td>
<td>-gfx950</td>
<td>AMD Instinct MI355X / MI350X</td>
</tr>
<tr>
<td>CDNA3</td>
<td>-gfx94x</td>
<td>AMD Instinct MI325X / MI300X / MI300A</td>
</tr>
<tr>
<td>CDNA2</td>
<td>-gfx90a</td>
<td>AMD Instinct MI250X / MI250 / MI210</td>
</tr>
<tr>
<td>CDNA</td>
<td>-gfx908</td>
<td>AMD Instinct MI100</td>
</tr>
<tr>
<td>RDNA4</td>
<td>-gfx120x</td>
<td>AMD Radeon RX 9070 / AMD Radeon RX 9060 / AMD Radeon RX 9070 XT / AMD Radeon RX 9060 XT / AMD Radeon RX 9070 GRE / AMD Radeon AI PRO R9700 / AMD Radeon AI PRO R9600D / AMD Radeon RX 9060 XT LP</td>
</tr>
<tr>
<td>RDNA3.5</td>
<td>-gfx1150<br>-gfx1151<br>-gfx1152</td>
<td>AMD Ryzen AI 9 465 / AMD Ryzen AI 9 365 / AMD Ryzen AI 9 HX 475 / AMD Ryzen AI 9 HX 470 / AMD Ryzen AI 9 HX 375 / AMD Ryzen AI 9 HX 370 / AMD Ryzen AI 9 PRO 465 / AMD Ryzen AI 9 PRO HX 475 / AMD Ryzen AI 9 PRO HX 470 / AMD Ryzen AI 9 HX PRO 375 / AMD Ryzen AI 9 HX PRO 370 / AMD Ryzen AI Max 390 / AMD Ryzen AI Max 385 / AMD Ryzen AI Max+ 395 / AMD Ryzen AI Max+ 392 / AMD Ryzen AI Max+ 388 / AMD Ryzen AI Max PRO 390 / AMD Ryzen AI Max PRO 385 / AMD Ryzen AI Max PRO 380 / AMD Ryzen AI Max+ PRO 395 / AMD Ryzen AI 7 450 / AMD Ryzen AI 7 350 / AMD Ryzen AI 7 345 / AMD Ryzen AI 5 340 / AMD Ryzen AI 5 330 / AMD Ryzen AI 7 PRO 450 / AMD Ryzen AI 5 PRO 440 / AMD Ryzen AI 7 PRO 350 / AMD Ryzen AI 5 PRO 340</td>
</tr>
<tr>
<td>RDNA3</td>
<td>-gfx110x</td>
<td>AMD Radeon RX 7700 / AMD Radeon RX 7600 / AMD Radeon PRO V710 / AMD Radeon PRO W7900 / AMD Radeon PRO W7800 / AMD Radeon PRO W7700 / AMD Radeon RX 7900 XT / AMD Radeon RX 7800 XT / AMD Radeon RX 7700 XT / AMD Radeon RX 7700 XE / AMD Radeon RX 7900 XTX / AMD Radeon RX 7900 GRE / AMD Radeon PRO W7800 48GB / AMD Radeon PRO W7900 Dual Slot</td>
</tr>
<tr>
<td>RDNA2</td>
<td>-gfx1030</td>
<td>AMD Radeon PRO V620 / AMD Radeon PRO W6800</td>
</tr>
</tbody>
</table>
## ROCm Core SDK component changes (moved or removed)
### Planned for future releases
- ROCm Core SDK: RPP
- ROCm-Extras: hipfort, rocALUTION, rocPyDecode, rocAL, MIVisionX
### Moved to ROCm-Extras
- ROCm Validation Suite
- ROCm Bandwidth Test
- TransferBench
- MIGraphX
### Moved to Standalone/ONNX
- ONNX runtime
### Removed
- [ROCm SMI](https://rocm.docs.amd.com/en/latest/about/release-notes.html#rocm-smi-deprecation) (replaced by AMD SMI)
## Notable package relocations
- rocMLIR (now included in MIGraphX)
- HIPCC (now included in `amdrocm-llvm`)
- FLANG (now included in `amdrocm-llvm`)
- ROCm CMake (now in `amdrocm-base`)
- ROCTracer (now in `amdrocm-profiler-base`)
- ROCProfiler (functionality in `amdrocm-profiler`)
## Components available in the ROCm Core SDK, ROCm-Extras, and Standalone/ONNX
<table class="rocm-docs-table table">
<thead>
<tr>
<th class="head"></th>
<th class="head">Category</th>
<th class="head">Present</th>
<th class="head">Absent/Moved</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="6" class="stub" style="vertical-align: middle"><strong>ROCm Core SDK</strong></td>
<td>Math and compute libraries</td>
<td>CK, hipBLAS, hipBLASLt, hipCUB, hipFFT, hipRAND, hipSOLVER, hipSPARSE/SPARSELt, MIOpen, rocBLAS, rocFFT, rocRAND, rocSOLVER, rocSPARSE, rocPRIM, rocThrust, rocWMMA</td>
<td>hipfort, rocALUTION</td>
</tr>
<tr>
<td>Communication libraries</td>
<td>RCCL, rocSHMEM</td>
<td>—</td>
</tr>
<tr>
<td>Media libraries</td>
<td>rocDecode, rocJPEG, ROCm Performance Primitives (RPP planned for a future release)</td>
<td>rocPyDecode, rocAL, MIVisionX, MIGraphX, CK (moved to math and compute)</td>
</tr>
<tr>
<td>Runtime, compilers, build tools</td>
<td>HIP, HIPIFY, LLVM</td>
<td>HIPCC, FLANG, ROCm CMake</td>
</tr>
<tr>
<td>Profiling and debugging tools</td>
<td>ROCm Compute Profiler, ROCm Systems Profiler, ROCprofiler-SDK, ROCdbgapi, ROCm Debugger, ROCr Debug Agent</td>
<td>ROCTracer, ROCProfiler</td>
</tr>
<tr>
<td>Control and monitoring tools</td>
<td>AMD SMI, ROCm Data Center Tool, rocminfo, hipinfo</td>
<td>ROCm SMI (removed), ROCm Validation Suite, ROCm Bandwidth Test</td>
</tr>
<tr>
<td style="vertical-align: middle"><strong>ROCm-Extras</strong></td>
<td>—</td>
<td>ROCm Validation Suite, ROCm Bandwidth Test, TransferBench, MIGraphX</td>
<td>—</td>
</tr>
<tr>
<td style="vertical-align: middle"><strong>Standalone/ONNX</strong></td>
<td>—</td>
<td>rocMLIR, ONNX runtime</td>
<td>—</td>
</tr>
</tbody>
</table>
@@ -8,7 +8,7 @@ docker:
- "Parallel VAE decode support for Wan models"
- "Batch inference and data parallel support"
components:
TheRock:
TheRock:
version: 9b611c6
url: https://github.com/ROCm/TheRock
rocm-libraries:
@@ -75,7 +75,7 @@ docker:
- '--guidance_scale 6.0 \'
- '--use_torch_compile \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- model: Hunyuan Video 1.5
model_repo: hunyuanvideo-community/HunyuanVideo-1.5-Diffusers-720p_t2v
url: https://huggingface.co/hunyuanvideo-community/HunyuanVideo-1.5-Diffusers-720p_t2v
@@ -97,7 +97,7 @@ docker:
- '--enable_tiling --enable_slicing \'
- '--use_torch_compile \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- group: Wan-AI
js_tag: wan
models:
@@ -123,7 +123,7 @@ docker:
- '--num_inference_steps 40 \'
- '--use_torch_compile \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- model: Wan2.2
model_repo: Wan-AI/Wan2.2-I2V-A14B-Diffusers
url: https://huggingface.co/Wan-AI/Wan2.2-I2V-A14B-Diffusers
@@ -146,7 +146,7 @@ docker:
- '--num_inference_steps 40 \'
- '--use_torch_compile \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- group: FLUX
js_tag: flux
models:
@@ -172,7 +172,7 @@ docker:
- '--guidance_scale 0.0 \'
- '--num_iterations 50 \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- model: FLUX.1 Kontext
model_repo: black-forest-labs/FLUX.1-Kontext-dev
url: https://huggingface.co/black-forest-labs/FLUX.1-Kontext-dev
@@ -196,7 +196,7 @@ docker:
- '--guidance_scale 2.5 \'
- '--num_iterations 25 \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- model: FLUX.2
model_repo: black-forest-labs/FLUX.2-dev
url: https://huggingface.co/black-forest-labs/FLUX.2-dev
@@ -220,7 +220,7 @@ docker:
- '--guidance_scale 4.0 \'
- '--num_iterations 25 \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- model: FLUX.2 Klein
model_repo: black-forest-labs/FLUX.2-klein-9B
url: https://huggingface.co/black-forest-labs/FLUX.2-klein-9B
@@ -242,7 +242,7 @@ docker:
- '--guidance_scale 1.0 \'
- '--num_iterations 25 \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- group: StableDiffusion
js_tag: stablediffusion
models:
@@ -263,7 +263,7 @@ docker:
- '--use_cfg_parallel \'
- '--use_torch_compile \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- group: Z-Image
js_tag: z_image
models:
@@ -289,7 +289,7 @@ docker:
- '--guidance_scale 4.0 \'
- '--num_iterations 25 \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- group: LTX
js_tag: ltx
models:
@@ -313,7 +313,7 @@ docker:
- '--guidance_scale 4.0 \'
- '--num_iterations 1 \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- group: Qwen-Image
js_tag: qwen_image
models:
@@ -336,7 +336,7 @@ docker:
- '--use_torch_compile \'
- '--num_iterations 1 \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
- model: Qwen-Image-Edit
model_repo: Qwen/Qwen-Image-Edit
url: https://huggingface.co/Qwen/Qwen-Image-Edit
@@ -358,4 +358,4 @@ docker:
- '--use_torch_compile \'
- '--num_iterations 1 \'
- '--attention_backend aiter \'
- '--output_directory results'
- "--output_directory results"
@@ -17,7 +17,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker-812:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.0_20250812-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.0-20250812.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -64,7 +64,7 @@ The following is summary of notable changes since the :doc:`previous ROCm/vLLM D
Supported models
================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.0_20250812-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.0-20250812.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -123,8 +123,6 @@ Supported models
vLLM is a toolkit and library for LLM inference and serving. AMD implements
high-performance custom kernels and modules in vLLM to enhance performance.
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
more information.
.. _vllm-benchmark-performance-measurements-812:
@@ -157,7 +155,7 @@ To test for optimal performance, consult the recommended :ref:`System health ben
<rocm-for-ai-system-health-bench>`. This suite of tests will help you verify and fine-tune your
system's configuration.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.0_20250812-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.0-20250812.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -296,21 +294,21 @@ system's configuration.
* -
- ``configs/extended.csv``
-
-
* -
- ``configs/performance.csv``
-
-
* - ``--benchmark``
- ``throughput``
- Measure offline end-to-end throughput.
* -
* -
- ``serving``
- Measure online serving performance.
* -
* -
- ``all``
- Measure both throughput and serving.
@@ -433,9 +431,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -16,7 +16,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker-909:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20250909-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.1-20250909.yaml
{% set docker = data.dockers[0] %}
@@ -57,7 +57,7 @@ The following is summary of notable changes since the :doc:`previous ROCm/vLLM D
Supported models
================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20250909-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.1-20250909.yaml
{% set docker = data.dockers[0] %}
{% set model_groups = data.model_groups %}
@@ -146,7 +146,7 @@ To test for optimal performance, consult the recommended :ref:`System health ben
<rocm-for-ai-system-health-bench>`. This suite of tests will help you verify and fine-tune your
system's configuration.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20250909-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.1-20250909.yaml
{% set docker = data.dockers[0] %}
{% set model_groups = data.model_groups %}
@@ -433,9 +433,6 @@ Further reading
- To learn more about system settings and management practices to configure your system for
AMD Instinct MI300X Series accelerators, see `AMD Instinct MI300X system optimization <https://instinct.docs.amd.com/projects/amdgpu-docs/en/latest/system-optimization/mi300x.html>`_.
- See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
a brief introduction to vLLM and optimization strategies.
- For application performance optimization strategies for HPC and AI workloads,
including inference with vLLM, see :doc:`/how-to/rocm-for-ai/inference-optimization/workload`.
@@ -16,7 +16,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker-930:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20251006-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.1-20251006.yaml
{% set docker = data.dockers[0] %}
@@ -75,7 +75,7 @@ The following is summary of notable changes since the :doc:`previous ROCm/vLLM D
Supported models
================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20251006-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.1-20251006.yaml
{% set docker = data.dockers[0] %}
{% set model_groups = data.model_groups %}
@@ -178,7 +178,7 @@ system's configuration.
Pull the Docker image
=====================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20251006-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.1-20251006.yaml
{% set docker = data.dockers[0] %}
@@ -192,7 +192,7 @@ Pull the Docker image
Benchmarking
============
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20251006-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.1-20251006.yaml
{% set docker = data.dockers[0] %}
{% set model_groups = data.model_groups %}
@@ -441,7 +441,7 @@ To reproduce this ROCm-enabled vLLM Docker image release, follow these steps:
2. Use the following command to build the image directly from the specified commit.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.10.1_20251006-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.10.1-20251006.yaml
{% set docker = data.dockers[0] %}
.. code-block:: shell
@@ -467,9 +467,6 @@ Further reading
- To learn more about system settings and management practices to configure your system for
AMD Instinct MI300X Series GPUs, see `AMD Instinct MI300X system optimization <https://instinct.docs.amd.com/projects/amdgpu-docs/en/latest/system-optimization/mi300x.html>`_.
- See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
a brief introduction to vLLM and optimization strategies.
- For application performance optimization strategies for HPC and AI workloads,
including inference with vLLM, see :doc:`/how-to/rocm-for-ai/inference-optimization/workload`.
@@ -16,7 +16,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker-1103:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.11.1_20251103-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.11.1-20251103.yaml
{% set docker = data.dockers[0] %}
@@ -61,7 +61,7 @@ The following is summary of notable changes since the :doc:`previous ROCm/vLLM D
Supported models
================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.11.1_20251103-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.11.1-20251103.yaml
{% set docker = data.dockers[0] %}
{% set model_groups = data.model_groups %}
@@ -164,7 +164,7 @@ system's configuration.
Pull the Docker image
=====================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.11.1_20251103-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.11.1-20251103.yaml
{% set docker = data.dockers[0] %}
@@ -178,7 +178,7 @@ Pull the Docker image
Benchmarking
============
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.11.1_20251103-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.11.1-20251103.yaml
{% set docker = data.dockers[0] %}
{% set model_groups = data.model_groups %}
@@ -431,7 +431,7 @@ To reproduce this ROCm-enabled vLLM Docker image release, follow these steps:
2. Use the following command to build the image directly from the specified commit.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.11.1_20251103-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.11.1-20251103.yaml
{% set docker = data.dockers[0] %}
.. code-block:: shell
@@ -457,9 +457,6 @@ Further reading
- To learn more about system settings and management practices to configure your system for
AMD Instinct MI300X Series GPUs, see `AMD Instinct MI300X system optimization <https://instinct.docs.amd.com/projects/amdgpu-docs/en/latest/system-optimization/mi300x.html>`_.
- See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
a brief introduction to vLLM and optimization strategies.
- For application performance optimization strategies for HPC and AI workloads,
including inference with vLLM, see :doc:`/how-to/rocm-for-ai/inference-optimization/workload`.
@@ -44,9 +44,7 @@ optimizing performance with popular AI models.
consumption and increases throughput by leveraging dynamic key and value
allocation in GPU memory. vLLM also incorporates many LLM acceleration
and quantization algorithms. In addition, AMD implements high-performance
custom kernels and modules in vLLM to enhance performance further. See
:ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for more
information.
custom kernels and modules in vLLM to enhance performance further.
Getting started
===============
@@ -277,7 +275,7 @@ options and their descriptions.
Latency benchmark example
^^^^^^^^^^^^^^^^^^^^^^^^^
Use this command to benchmark the latency of the Llama 3.1 8B model on one GPU with the ``float16`` data type.
.. code-block::
@@ -334,9 +332,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -68,8 +68,6 @@ optimizing performance with popular AI models.
vLLM is a toolkit and library for LLM inference and serving. AMD implements
high-performance custom kernels and modules in vLLM to enhance performance.
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
more information.
Getting started
===============
@@ -342,7 +340,7 @@ options and their descriptions.
Example 1: latency benchmark
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
Use this command to benchmark the latency of the Llama 3.1 8B model on one GPU with the ``float16`` and ``float8`` data types.
.. code-block::
@@ -404,9 +402,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -41,8 +41,6 @@ and :ref:`standalone benchmarking <vllm-benchmark-standalone-v066-options>`.
vLLM is a toolkit and library for LLM inference and serving. AMD implements
high-performance custom kernels and modules in vLLM to enhance performance.
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
more information.
Getting started
===============
@@ -387,7 +385,7 @@ options and their descriptions.
Example 1: latency benchmark
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
Use this command to benchmark the latency of the Llama 3.1 70B model on eight GPUs with the ``float16`` and ``float8`` data types.
.. code-block::
@@ -449,9 +447,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -17,7 +17,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.7.3_20250325-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.7.3-20250325.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -93,8 +93,6 @@ vLLM inference performance testing
vLLM is a toolkit and library for LLM inference and serving. AMD implements
high-performance custom kernels and modules in vLLM to enhance performance.
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
more information.
.. _vllm-benchmark-performance-measurements-v073:
@@ -317,9 +315,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -12,7 +12,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.8.3_20250415-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.8.3-20250415.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -88,8 +88,6 @@ vLLM inference performance testing
vLLM is a toolkit and library for LLM inference and serving. AMD implements
high-performance custom kernels and modules in vLLM to enhance performance.
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
more information.
.. _vllm-benchmark-performance-measurements-v083:
@@ -333,9 +331,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -17,7 +17,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.8.5_20250513-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.8.5-20250513.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -97,8 +97,6 @@ vLLM inference performance testing
vLLM is a toolkit and library for LLM inference and serving. AMD implements
high-performance custom kernels and modules in vLLM to enhance performance.
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
more information.
.. _vllm-benchmark-performance-measurements-v085-20250513:
@@ -342,9 +340,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -17,7 +17,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.8.5_20250521-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.8.5-20250521.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -97,8 +97,6 @@ vLLM inference performance testing
vLLM is a toolkit and library for LLM inference and serving. AMD implements
high-performance custom kernels and modules in vLLM to enhance performance.
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
more information.
.. _vllm-benchmark-performance-measurements-v085-20250521:
@@ -342,9 +340,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -17,7 +17,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.9.0.1_20250605-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.9.0.1-20250605.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -97,8 +97,6 @@ vLLM inference performance testing
vLLM is a toolkit and library for LLM inference and serving. AMD implements
high-performance custom kernels and modules in vLLM to enhance performance.
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
more information.
.. _vllm-benchmark-performance-measurements-v0901-20250605:
@@ -341,9 +339,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -17,7 +17,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker-702:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.9.1_20250702-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.9.1-20250702.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -97,8 +97,6 @@ vLLM inference performance testing
vLLM is a toolkit and library for LLM inference and serving. AMD implements
high-performance custom kernels and modules in vLLM to enhance performance.
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
more information.
.. _vllm-benchmark-performance-measurements-20250702:
@@ -341,9 +339,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -17,7 +17,7 @@ vLLM inference performance testing
.. _vllm-benchmark-unified-docker-715:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.9.1_20250715-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.9.1-20250715.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -70,7 +70,7 @@ The following is summary of notable changes since the :doc:`previous ROCm/vLLM D
Supported models
================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.9.1_20250715-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.9.1-20250715.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -129,8 +129,6 @@ Supported models
vLLM is a toolkit and library for LLM inference and serving. AMD implements
high-performance custom kernels and modules in vLLM to enhance performance.
See :ref:`fine-tuning-llms-vllm` and :ref:`mi300x-vllm-optimization` for
more information.
.. _vllm-benchmark-performance-measurements-715:
@@ -163,7 +161,7 @@ To test for optimal performance, consult the recommended :ref:`System health ben
<rocm-for-ai-system-health-bench>`. This suite of tests will help you verify and fine-tune your
system's configuration.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/vllm_0.9.1_20250715-benchmark-models.yaml
.. datatemplate:yaml:: ./data/vllm-0.9.1-20250715.yaml
{% set unified_docker = data.vllm_benchmark.unified_docker.latest %}
{% set model_groups = data.vllm_benchmark.model_groups %}
@@ -438,9 +436,6 @@ Further reading
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
@@ -6,13 +6,19 @@
prebuilt and optimized docker images.
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
************************
xDiT diffusion inference
************************
******************************
xDiT diffusion inference 25.10
******************************
.. caution::
This documentation does not reflect the latest version of the xDiT diffusion
inference performance documentation. See
:doc:`/ai-inference/archive/xdit-history` for the latest version.
.. _xdit-video-diffusion-2510:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.10-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.10.yaml
{% set docker = data.xdit_diffusion_inference.docker %}
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
@@ -59,7 +65,7 @@ The following models are supported for inference performance benchmarking.
Some instructions, commands, and recommendations in this documentation might
vary by model -- select one to get started.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.10-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.10.yaml
{% set docker = data.xdit_diffusion_inference.docker %}
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
@@ -122,7 +128,7 @@ guide to properly configure your system settings before starting.
Pull the Docker image
=====================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.10-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.10.yaml
{% set docker = data.xdit_diffusion_inference.docker %}
@@ -139,7 +145,7 @@ Validate and benchmark
Once the image has been downloaded you can follow these steps to
run benchmarks and generate outputs.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.10-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.10.yaml
{% set model_groups = data.xdit_diffusion_inference.model_groups %}
{% for model_group in model_groups %}
@@ -166,7 +172,7 @@ Prepare the model
You can either use an existing Hugging Face cache or download the model fresh inside the container.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.10-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.10.yaml
{% set docker = data.xdit_diffusion_inference.docker %}
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
@@ -264,7 +270,7 @@ Run inference
You can benchmark models through `MAD <https://github.com/ROCm/MAD>`__-integrated automation or standalone
torchrun commands.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.10-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.10.yaml
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
{% for model_group in model_groups %}
@@ -295,7 +301,7 @@ torchrun commands.
--tags {{model.mad_tag}} \
--keep-model-dir \
--live-output
MAD launches a Docker container with the name
``container_ci-{{model.mad_tag}}``. The throughput and serving reports of the
model are collected in the following paths: ``{{ model.mad_tag }}_throughput.csv``
@@ -395,5 +401,5 @@ Further reading
Previous versions
=================
See :doc:`xdit-history` to find documentation for previous releases
of xDiT diffusion inference performance testing.
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
releases of xDiT diffusion inference performance testing.
@@ -6,20 +6,19 @@
prebuilt and optimized docker images.
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
************************
xDiT diffusion inference
************************
******************************
xDiT diffusion inference 25.11
******************************
.. caution::
This documentation does not reflect the latest version of ROCm vLLM
This documentation does not reflect the latest version of the xDiT diffusion
inference performance documentation. See
:doc:`/how-to/rocm-for-ai/inference/xdit-diffusion-inference` for the latest
version.
:doc:`/ai-inference/archive/xdit-history` for the latest version.
.. _xdit-video-diffusion-2511:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
{% set docker = data.xdit_diffusion_inference.docker | selectattr("version", "equalto", "v25-11") | first %}
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
@@ -48,7 +47,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
What's new
==========
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
{% set docker = data.xdit_diffusion_inference.docker | selectattr("version", "equalto", "v25-11") | first %}
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
@@ -66,7 +65,7 @@ The following models are supported for inference performance benchmarking.
Some instructions, commands, and recommendations in this documentation might
vary by model -- select one to get started.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
{% set docker = data.xdit_diffusion_inference.docker | selectattr("version", "equalto", "v25-11") | first %}
{% set model_groups = data.xdit_diffusion_inference.model_groups %}
@@ -145,7 +144,7 @@ system's configuration.
Pull the Docker image
=====================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
{% set docker = data.xdit_diffusion_inference.docker | selectattr("version", "equalto", "v25-11") | first %}
@@ -162,7 +161,7 @@ Validate and benchmark
Once the image has been downloaded you can follow these steps to
run benchmarks and generate outputs.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
{% for model_group in model_groups %}
{% for model in model_group.models %}
@@ -180,7 +179,7 @@ Choose your setup method
You can either use an existing Hugging Face cache or download the model fresh inside the container.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
{% set docker = data.xdit_diffusion_inference.docker | selectattr("version", "equalto", "v25-11") | first %}
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
@@ -270,7 +269,7 @@ You can either use an existing Hugging Face cache or download the model fresh in
Run inference
=============
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.11-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.11.yaml
{% set model_groups = data.xdit_diffusion_inference.model_groups%}
{% for model_group in model_groups %}
@@ -384,7 +383,5 @@ Run inference
Previous versions
=================
See
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
to find documentation for previous releases of xDiT diffusion inference
performance testing.
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
releases of xDiT diffusion inference performance testing.
@@ -6,20 +6,19 @@
prebuilt and optimized docker images.
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
************************
xDiT diffusion inference
************************
******************************
xDiT diffusion inference 25.12
******************************
.. caution::
This documentation does not reflect the latest version of xDiT diffusion
This documentation does not reflect the latest version of the xDiT diffusion
inference performance documentation. See
:doc:`/how-to/rocm-for-ai/inference/xdit-diffusion-inference` for the latest
version.
:doc:`/ai-inference/archive/xdit-history` for the latest version.
.. _xdit-video-diffusion-2512:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
{% set docker = data.docker %}
@@ -51,7 +50,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
What's new
==========
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
{% set docker = data.docker %}
@@ -68,7 +67,7 @@ The following models are supported for inference performance benchmarking.
Some instructions, commands, and recommendations in this documentation might
vary by model -- select one to get started.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
{% set docker = data.docker %}
@@ -133,7 +132,7 @@ system's configuration.
Pull the Docker image
=====================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
{% set docker = data.docker %}
@@ -147,7 +146,7 @@ Pull the Docker image
Validate and benchmark
======================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
{% set docker = data.docker %}
@@ -170,7 +169,7 @@ Choose your setup method
You can either use an existing Hugging Face cache or download the model fresh inside the container.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
{% set docker = data.docker %}
@@ -262,7 +261,7 @@ You can either use an existing Hugging Face cache or download the model fresh in
Run inference
=============
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.12-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.12.yaml
{% set docker = data.docker %}
@@ -406,7 +405,5 @@ Run inference
Previous versions
=================
See
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
to find documentation for previous releases of xDiT diffusion inference
performance testing.
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
releases of xDiT diffusion inference performance testing.
@@ -6,20 +6,19 @@
prebuilt and optimized docker images.
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
************************
xDiT diffusion inference
************************
******************************
xDiT diffusion inference 25.13
******************************
.. caution::
This documentation does not reflect the latest version of the xDiT diffusion
inference performance documentation. See
:doc:`/how-to/rocm-for-ai/inference/xdit-diffusion-inference` for the latest
version.
:doc:`/ai-inference/archive/xdit-history` for the latest version.
.. _xdit-video-diffusion-2513:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
{% set docker = data.docker %}
@@ -52,7 +51,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
What's new
==========
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
{% set docker = data.docker %}
@@ -69,7 +68,7 @@ The following models are supported for inference performance benchmarking.
Some instructions, commands, and recommendations in this documentation might
vary by model -- select one to get started.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
{% set docker = data.docker %}
@@ -150,7 +149,7 @@ system's configuration.
Pull the Docker image
=====================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
{% set docker = data.docker %}
@@ -164,7 +163,7 @@ Pull the Docker image
Validate and benchmark
======================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
{% set docker = data.docker %}
@@ -187,7 +186,7 @@ Choose your setup method
You can either use an existing Hugging Face cache or download the model fresh inside the container.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
{% set docker = data.docker %}
@@ -279,7 +278,7 @@ You can either use an existing Hugging Face cache or download the model fresh in
Run inference
=============
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_25.13-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-25.13.yaml
{% set docker = data.docker %}
@@ -469,7 +468,5 @@ Run inference
Previous versions
=================
See
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
to find documentation for previous releases of xDiT diffusion inference
performance testing.
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
releases of xDiT diffusion inference performance testing.
@@ -6,20 +6,19 @@
prebuilt and optimized docker images.
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
************************
xDiT diffusion inference
************************
*****************************
xDiT diffusion inference 26.1
*****************************
.. caution::
This documentation does not reflect the latest version of the xDiT diffusion
inference performance documentation. See
:doc:`/how-to/rocm-for-ai/inference/xdit-diffusion-inference` for the latest
version.
:doc:`/ai-inference/archive/xdit-history` for the latest version.
.. _xdit-video-diffusion-v261-v261:
.. _xdit-video-diffusion-v261:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
{% set docker = data.docker %}
@@ -47,7 +46,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
What's new
==========
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
{% set docker = data.docker %}
@@ -64,7 +63,7 @@ The following models are supported for inference performance benchmarking.
Some instructions, commands, and recommendations in this documentation might
vary by model -- select one to get started.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
{% set docker = data.docker %}
@@ -129,7 +128,7 @@ system's configuration.
Pull the Docker image
=====================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
{% set docker = data.docker %}
@@ -143,7 +142,7 @@ Pull the Docker image
Validate and benchmark
======================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
{% set docker = data.docker %}
@@ -166,7 +165,7 @@ Choose your setup method
You can either use an existing Hugging Face cache or download the model fresh inside the container.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
{% set docker = data.docker %}
@@ -258,7 +257,7 @@ You can either use an existing Hugging Face cache or download the model fresh in
Run inference
=============
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.1-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.1.yaml
{% set docker = data.docker %}
@@ -317,7 +316,5 @@ Run inference
Previous versions
=================
See
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
to find documentation for previous releases of xDiT diffusion inference
performance testing.
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
releases of xDiT diffusion inference performance testing.
@@ -14,12 +14,11 @@ xDiT diffusion inference
This documentation does not reflect the latest version of the xDiT diffusion
inference performance documentation. See
:doc:`/how-to/rocm-for-ai/inference/xdit-diffusion-inference` for the latest
version.
:doc:`/ai-inference/archive/xdit-history` for the latest version.
.. _xdit-video-diffusion-262:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
{% set docker = data.docker %}
@@ -47,7 +46,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
What's new
==========
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
{% set docker = data.docker %}
@@ -64,7 +63,7 @@ The following models are supported for inference performance benchmarking.
Some instructions, commands, and recommendations in this documentation might
vary by model -- select one to get started.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
{% set docker = data.docker %}
@@ -129,7 +128,7 @@ system's configuration.
Pull the Docker image
=====================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
{% set docker = data.docker %}
@@ -143,7 +142,7 @@ Pull the Docker image
Validate and benchmark
======================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
{% set docker = data.docker %}
@@ -166,7 +165,7 @@ Choose your setup method
You can either use an existing Hugging Face cache or download the model fresh inside the container.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
{% set docker = data.docker %}
@@ -258,7 +257,7 @@ You can either use an existing Hugging Face cache or download the model fresh in
Run inference
=============
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.2-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.2.yaml
{% set docker = data.docker %}
@@ -315,7 +314,5 @@ Run inference
Previous versions
=================
See
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
to find documentation for previous releases of xDiT diffusion inference
performance testing.
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
releases of xDiT diffusion inference performance testing.
@@ -6,20 +6,19 @@
prebuilt and optimized docker images.
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
************************
xDiT diffusion inference
************************
*****************************
xDiT diffusion inference 26.3
*****************************
.. caution::
This documentation does not reflect the latest version of the xDiT diffusion
inference performance documentation. See
:doc:`/how-to/rocm-for-ai/inference/xdit-diffusion-inference` for the latest
version.
:doc:`/ai-inference/archive/xdit-history` for the latest version.
.. _xdit-video-diffusion-263:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
{% set docker = data.docker %}
@@ -47,7 +46,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
What's new
==========
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
{% set docker = data.docker %}
@@ -64,7 +63,7 @@ The following models are supported for inference performance benchmarking.
Some instructions, commands, and recommendations in this documentation might
vary by model -- select one to get started.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
{% set docker = data.docker %}
@@ -129,7 +128,7 @@ system's configuration.
Pull the Docker image
=====================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
{% set docker = data.docker %}
@@ -143,7 +142,7 @@ Pull the Docker image
Validate and benchmark
======================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
{% set docker = data.docker %}
@@ -166,7 +165,7 @@ Choose your setup method
You can either use an existing Hugging Face cache or download the model fresh inside the container.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
{% set docker = data.docker %}
@@ -258,7 +257,7 @@ You can either use an existing Hugging Face cache or download the model fresh in
Run inference
=============
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.3-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.3.yaml
{% set docker = data.docker %}
@@ -315,7 +314,5 @@ Run inference
Previous versions
=================
See
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
to find documentation for previous releases of xDiT diffusion inference
performance testing.
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
releases of xDiT diffusion inference performance testing.
@@ -1,25 +1,26 @@
:orphan:
:no-search:
:selector-toc2: Model
:selector-toc2-icon: fa-solid fa-robot
.. meta::
:description: Learn to validate diffusion model video generation on MI300X, MI350X and MI355X accelerators using
prebuilt and optimized docker images.
:keywords: xDiT, diffusion, video, video generation, image, image generation, validate, benchmark
************************
xDiT diffusion inference
************************
*****************************
xDiT diffusion inference 26.4
*****************************
.. caution::
This documentation does not reflect the latest version of the xDiT diffusion
inference performance documentation. See
:doc:`/how-to/rocm-for-ai/inference/xdit-diffusion-inference` for the latest
version.
:doc:`/ai-inference/archive/xdit-history` for the latest version.
.. _xdit-video-diffusion-264:
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.4-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
{% set docker = data.docker %}
@@ -47,7 +48,7 @@ For preview and development releases, see `amdsiloai/pytorch-xdit <https://hub.d
What's new
==========
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.4-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
{% set docker = data.docker %}
@@ -64,43 +65,37 @@ The following models are supported for inference performance benchmarking.
Some instructions, commands, and recommendations in this documentation might
vary by model -- select one to get started.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.4-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
{% set docker = data.docker %}
.. raw:: html
.. selector:: Model
:key: model-group
<div id="vllm-benchmark-ud-params-picker" class="container-fluid">
<div class="row gx-0">
<div class="col-2 me-1 px-2 model-param-head">Model</div>
<div class="row col-10 pe-0">
{% for model_group in docker.supported_models %}
<div class="col-6 px-2 model-param" data-param-k="model-group" data-param-v="{{ model_group.js_tag }}" tabindex="0">{{ model_group.group }}</div>
{% endfor %}
</div>
</div>
{% for model_group in docker.supported_models %}
.. selector-option:: {{ model_group.group }}
:value: {{ model_group.js_tag }}
:width: 25%
<div class="row gx-0 pt-1">
<div class="col-2 me-1 px-2 model-param-head">Variant</div>
<div class="row col-10 pe-0">
{% for model_group in docker.supported_models %}
{% set models = model_group.models %}
{% for model in models %}
{% if models|length % 3 == 0 %}
<div class="col-4 px-2 model-param" data-param-k="model" data-param-v="{{ model.js_tag }}" data-param-group="{{ model_group.js_tag }}" tabindex="0">{{ model.model }}</div>
{% else %}
<div class="col-6 px-2 model-param" data-param-k="model" data-param-v="{{ model.js_tag }}" data-param-group="{{ model_group.js_tag }}" tabindex="0">{{ model.model }}</div>
{% endif %}
{% endfor %}
{% endfor %}
</div>
</div>
</div>
{% endfor %}
{% for model_group in docker.supported_models %}
.. selector:: Variant
:key: model
:show-cond: model-group={{ model_group.js_tag }}
{% set models = model_group.models %}
{% for model in models %}
.. selector-option:: {{ model.model }}
:value: {{ model.js_tag }}
{% endfor %}
{% endfor %}
{% for model_group in docker.supported_models %}
{% for model in model_group.models %}
.. container:: model-doc {{ model.js_tag }}
.. selected:: model={{ model.js_tag }}
.. note::
@@ -129,7 +124,7 @@ system's configuration.
Pull the Docker image
=====================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.4-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
{% set docker = data.docker %}
@@ -143,7 +138,7 @@ Pull the Docker image
Validate and benchmark
======================
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.4-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
{% set docker = data.docker %}
@@ -153,7 +148,7 @@ Validate and benchmark
{% for model_group in docker.supported_models %}
{% for model in model_group.models %}
.. container:: model-doc {{model.js_tag}}
.. selected:: model={{ model.js_tag }}
The following commands are written for {{ model.model }}.
See :ref:`xdit-video-diffusion-supported-models-264` to switch to another available model.
@@ -166,13 +161,13 @@ Choose your setup method
You can either use an existing Hugging Face cache or download the model fresh inside the container.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.4-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
{% set docker = data.docker %}
{% for model_group in docker.supported_models %}
{% for model in model_group.models %}
.. container:: model-doc {{model.js_tag}}
.. selected:: model={{model.js_tag}}
.. tab-set::
@@ -258,14 +253,14 @@ You can either use an existing Hugging Face cache or download the model fresh in
Run inference
=============
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/previous-versions/xdit_26.4-inference-models.yaml
.. datatemplate:yaml:: ./data/xdit-26.4.yaml
{% set docker = data.docker %}
{% for model_group in docker.supported_models %}
{% for model in model_group.models %}
.. container:: model-doc {{ model.js_tag }}
.. selected:: model={{ model.js_tag }}
.. tab-set::
@@ -315,7 +310,5 @@ Run inference
Previous versions
=================
See
:doc:`/how-to/rocm-for-ai/inference/benchmark-docker/previous-versions/xdit-history`
to find documentation for previous releases of xDiT diffusion inference
performance testing.
See :doc:`/ai-inference/archive/xdit-history` to find documentation for previous
releases of xDiT diffusion inference performance testing.
@@ -1,8 +1,8 @@
:orphan:
************************************************************
xDiT diffusion inference performance testing version history
************************************************************
****************************************
xDiT diffusion inference version history
****************************************
This table lists previous versions of the ROCm xDiT diffusion inference performance
testing environment. For detailed information about available models for
@@ -20,7 +20,7 @@ benchmarking, see the version-specific documentation.
* ROCm 7.13.0
* TheRock cbff3d1
-
* :doc:`Documentation </how-to/rocm-for-ai/inference/xdit-diffusion-inference>`
* :doc:`Documentation <../xdit>`
* `Docker Hub <https://hub.docker.com/layers/rocm/pytorch-xdit/v26.5/images/sha256-b8ad9fd4b41bc116ac2aff07c1066bf369cf7fc110b1a323f6302191985a51fd>`__
* - ``rocm/pytorch-xdit:v26.4``
+124
View File
@@ -0,0 +1,124 @@
********************************
ComfyUI image generation on ROCm
********************************
`ComfyUI <https://github.com/comfyanonymous/ComfyUI>`__ is an open-source,
node-based interface for building and running image generation workflows with
diffusion models such as Stable Diffusion. Its modular graph-based design lets
you construct, customize, and share complex pipelines without writing code. This
page walks through installing and running ComfyUI on AMD GPUs.
Prerequisites
=============
Ensure your working environment is running ROCm-enabled PyTorch on
a :ref:`supported system <compat-matrix>`. See :ref:`pytorch-install` for
instructions.
.. important::
On Windows, ComfyUI might not start if Smart App Control is enabled in your
Windows security settings.
Installation
============
After installing ROCm and PyTorch in your Python environment, follow these
steps to install ComfyUI.
1. Clone the ComfyUI repository.
.. code-block:: shell
git clone https://github.com/comfyanonymous/ComfyUI.git
2. Activate your Python virtual environment and install dependencies.
.. tab-set::
.. tab-item:: Linux
:sync: linux
.. code-block:: bash
pip install -r ComfyUI/requirements.txt
.. tab-item:: Windows
:sync: windows
.. code-block:: bat
pip install -r ComfyUI\requirements.txt
Run ComfyUI
===========
Use the following steps for a simple example of running ComfyUI.
1. Start the ComfyUI server from the command line.
.. tab-set::
.. tab-item:: Linux
:sync: linux
.. code-block:: bash
python ComfyUI/main.py
.. tab-item:: Windows
:sync: windows
.. code-block:: bat
python ComfyUI\main.py
This starts the server, displaying a prompt like:
.. code-block:: text
To see the GUI go to: http://127.0.0.1:8188
2. Go to ``http://127.0.0.1:8188`` in your web browser. You might need to
replace ``8188`` with the appropriate port.
.. image:: ./images/comfyui/comfyui-main.png
:align: center
3. Search for one of the following templates and download any missing
models.
.. tab-set::
.. tab-item:: SD3.5 Simple
Select **Template****Model Filter****SD3.5****SD3.5 Simple**
.. image:: ./images/comfyui/sd3_5-simple-card.png
:align: center
Download required models, if missing.
.. image:: ./images/comfyui/sd3_5-missing-models.png
:align: center
.. tab-item:: Chroma1 Radiance text to image
Select **Template****Model Filter****Chroma****Chroma1 Radiance text to image**
.. image:: ./images/comfyui/chroma1-radiance-tti-card.png
:align: center
Download required models, if missing.
.. image:: ./images/comfyui/chroma1-radiance-tti-missing-models.png
:align: center
4. Click the **Run** button.
The application will use your AMD GPU to convert the prompted text to an image.
.. seealso::
To learn more about the ComfyUI interface and workflows, see the `ComfyUI
documentation <https://docs.comfy.org/development/core-concepts/workflow>`__.

Before

Width:  |  Height:  |  Size: 44 KiB

After

Width:  |  Height:  |  Size: 44 KiB

Before

Width:  |  Height:  |  Size: 28 KiB

After

Width:  |  Height:  |  Size: 28 KiB

Before

Width:  |  Height:  |  Size: 112 KiB

After

Width:  |  Height:  |  Size: 112 KiB

Before

Width:  |  Height:  |  Size: 188 KiB

After

Width:  |  Height:  |  Size: 188 KiB

Before

Width:  |  Height:  |  Size: 129 KiB

After

Width:  |  Height:  |  Size: 129 KiB

Before

Width:  |  Height:  |  Size: 80 KiB

After

Width:  |  Height:  |  Size: 80 KiB

Before

Width:  |  Height:  |  Size: 153 KiB

After

Width:  |  Height:  |  Size: 153 KiB

Before

Width:  |  Height:  |  Size: 219 KiB

After

Width:  |  Height:  |  Size: 219 KiB

Before

Width:  |  Height:  |  Size: 310 KiB

After

Width:  |  Height:  |  Size: 310 KiB

Before

Width:  |  Height:  |  Size: 342 KiB

After

Width:  |  Height:  |  Size: 342 KiB

@@ -2,35 +2,35 @@
:description: How to Use ROCm for AI inference optimization
:keywords: ROCm, LLM, AI inference, Optimization, GPUs, usage, tutorial
*******************************************
Use ROCm for AI inference optimization
*******************************************
**********************
Inference optimization
**********************
AI inference optimization is the process of improving the performance of machine learning models and speeding up the inference process. It includes:
- **Quantization**: This involves reducing the precision of model weights and activations while maintaining acceptable accuracy levels. Reduced precision improves inference efficiency because lower precision data requires less storage and better utilizes the hardware's computation power.
- **Quantization**: This involves reducing the precision of model weights and activations while maintaining acceptable accuracy levels. Reduced precision improves inference efficiency because lower precision data requires less storage and better utilizes the hardware's computation power.
- **Kernel optimization**: This technique involves optimizing computation kernels to exploit the underlying hardware capabilities. For example, the kernels can be optimized to use multiple GPU cores or utilize specialized hardware like tensor cores to accelerate the computations.
- **Kernel optimization**: This technique involves optimizing computation kernels to exploit the underlying hardware capabilities. For example, the kernels can be optimized to use multiple GPU cores or utilize specialized hardware like tensor cores to accelerate the computations.
- **Libraries**: Libraries such as Flash Attention, xFormers, and PyTorch TunableOp are used to accelerate deep learning models and improve the performance of inference workloads.
- **Libraries**: Libraries such as Flash Attention, xFormers, and PyTorch TunableOp are used to accelerate deep learning models and improve the performance of inference workloads.
- **Hardware acceleration**: Hardware acceleration techniques, like GPUs for AI inference, can significantly improve performance due to their parallel processing capabilities.
- **Hardware acceleration**: Hardware acceleration techniques, like GPUs for AI inference, can significantly improve performance due to their parallel processing capabilities.
- **Pruning**: This involves removing unnecessary connections, layers, or weights from a pre-trained model while maintaining acceptable accuracy levels, resulting in a smaller model that requires fewer computational resources to run inference.
- **Pruning**: This involves removing unnecessary connections, layers, or weights from a pre-trained model while maintaining acceptable accuracy levels, resulting in a smaller model that requires fewer computational resources to run inference.
Utilizing these optimization techniques with the ROCm™ software platform can significantly reduce inference time, improve performance, and reduce the cost of your AI applications.
Utilizing these optimization techniques with the ROCm™ software platform can significantly reduce inference time, improve performance, and reduce the cost of your AI applications.
Throughout the following topics, this guide discusses optimization techniques for inference workloads.
- :doc:`Model quantization <model-quantization>`
- :doc:`Model acceleration libraries <model-acceleration-libraries>`
- :doc:`Model acceleration libraries <model-acceleration-libs>`
- :doc:`Optimizing with Composable Kernel <optimizing-with-composable-kernel>`
- :doc:`Optimizing with Composable Kernel <optimize-with-composable-kernel>`
- :doc:`Optimizing Triton kernels <optimizing-triton-kernel>`
- :doc:`Optimizing Triton kernels <optimize-triton-kernels>`
- :doc:`Profiling and debugging <profiling-and-debugging>`
- :doc:`Workload tuning <workload-optimization>`
- :doc:`Workload tuning <workload>`
- :ref:`Profiling and debugging <mi300x-profiling-tools>`
@@ -7,9 +7,7 @@ LLM inference frameworks
************************
This section discusses how to implement `vLLM <https://docs.vllm.ai/en/latest>`_ and `Hugging Face TGI
<https://huggingface.co/docs/text-generation-inference/en/index>`_ using
:doc:`single-accelerator <../fine-tuning/single-gpu-fine-tuning-and-inference>` and
:doc:`multi-accelerator <../fine-tuning/multi-gpu-fine-tuning-and-inference>` systems.
<https://huggingface.co/docs/text-generation-inference/en/index>`_.
.. _fine-tuning-llms-vllm:
@@ -68,7 +66,7 @@ Installing vLLM
The following log message is displayed in your command line indicates that the server is listening for requests.
.. image:: ../../../data/how-to/llm-fine-tuning-optimization/vllm-single-gpu-log.png
.. image:: ./images/llm-inference-frameworks/vllm-single-gpu-log.png
:alt: vLLM API server log message
:align: center
@@ -141,7 +139,7 @@ Installing vLLM
ROCm provides a prebuilt optimized Docker image for validating the performance of LLM inference with vLLM
on the MI300X GPU. The Docker image includes ROCm, vLLM, and PyTorch.
For more information, see :doc:`/how-to/rocm-for-ai/inference/benchmark-docker/vllm`.
For more information, see :doc:`/ai-inference/vllm`.
.. _fine-tuning-llms-tgi:
@@ -20,7 +20,7 @@ Attention (GQA), and Multi-Query Attention (MQA). This reduction in memory movem
time-to-first-token (TTFT) latency for large batch sizes and long prompt sequences, thereby enhancing overall
performance.
.. image:: ../../../data/how-to/llm-fine-tuning-optimization/attention-module.png
.. image:: ./images/model-acceleration-libs/attention-module.png
:alt: Attention module of a large language module utilizing tiling
:align: center
@@ -36,7 +36,7 @@ These can be installed by following the official
`PyTorch installation guide <https://pytorch.org/get-started/locally/>`_. Alternatively, for a simpler setup, you can use a preconfigured
:ref:`ROCm PyTorch Docker image <using-docker-with-pytorch-pre-installed>`, which already includes the required libraries.
Installing Flash Attention 2
Installing Flash Attention 2
----------------------------
`Flash Attention <https://github.com/Dao-AILab/flash-attention>`_ supports two backend implementations on AMD GPUs.
@@ -61,29 +61,29 @@ To install Flash Attention 2, use the following commands:
pip install ninja
# To install the CK backend flash attention
python setup.py install
python setup.py install
# To install the Triton backend flash attention
FLASH_ATTENTION_TRITON_AMD_ENABLE="TRUE" python setup.py install
FLASH_ATTENTION_TRITON_AMD_ENABLE="TRUE" python setup.py install
# To install both CK and Triton backend flash attention
FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE && FLASH_ATTENTION_SKIP_CK_BUILD=FALSE python setup.py install
For detailed installation instructions, see `Flash Attention <https://github.com/Dao-AILab/flash-attention>`_.
Benchmarking Flash Attention 2
Benchmarking Flash Attention 2
------------------------------
Benchmark scripts to evaluate the performance of Flash Attention 2 are stored in the ``flash-attention/benchmarks/`` directory.
To benchmark the CK backend
To benchmark the CK backend
.. code-block:: shell
cd flash-attention/benchmarks
pip install transformers einops ninja
python3 benchmark_flash_attention.py
python3 benchmark_flash_attention.py
To benchmark the Triton backend
@@ -91,7 +91,7 @@ To benchmark the Triton backend
FLASH_ATTENTION_TRITON_AMD_ENABLE="TRUE" python3 benchmark_flash_attention.py
Using Flash Attention 2
Using Flash Attention 2
-----------------------
.. code-block:: python
@@ -128,13 +128,13 @@ xFormers also improves the performance of attention modules. Although xFormers a
similarly to Flash Attention 2 due to its tiling behavior of query, key, and value, its widely used for LLMs and
Stable Diffusion models with the Hugging Face Diffusers library.
Installing CK xFormers
Installing CK xFormers
----------------------
Use the following commands to install CK xFormers.
.. code-block:: shell
# Install from source
git clone https://github.com/ROCm/xformers.git
cd xformers/
@@ -175,20 +175,20 @@ of the PyTorch compilation.
os.environ["TOKENIZERS_PARALLELISM"] = "false"
model_name = "NousResearch/Meta-Llama-3-8B"
prompts = []
for b in range(1):
prompts.append("New york city is where "
)
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.float16).to(device).eval()
inputs = tokenizer(prompts, return_tensors="pt").to(model.device)
def decode_one_tokens(model, cur_token, input_pos, cache_position):
logits = model(cur_token, position_ids=input_pos, cache_position=cache_position, return_dict=False, use_cache=True)[0]
new_token = torch.argmax(logits[:, -1], dim=-1)[:, None]
return new_token
batch_size, seq_length = inputs["input_ids"].shape
# Static key-value cache
@@ -198,16 +198,16 @@ of the PyTorch compilation.
cache_position = torch.arange(seq_length, device=device)
generated_ids = torch.zeros(batch_size, seq_length + max_new_tokens + 1, dtype=torch.int, device=device)
generated_ids[:, cache_position] = inputs["input_ids"].to(device).to(torch.int)
logits = model(**inputs, cache_position=cache_position, return_dict=False, use_cache=True)[0]
next_token = torch.argmax(logits[:, -1], dim=-1)[:, None]
# torch compilation
decode_one_tokens = torch.compile(decode_one_tokens, mode="max-autotune-no-cudagraphs",fullgraph=True)
generated_ids[:, seq_length] = next_token[:, 0]
cache_position = torch.tensor([seq_length + 1], device=device)
with torch.no_grad():
for _ in range(1, max_new_tokens):
with torch.backends.cuda.sdp_kernel(enable_flash=False, enable_mem_efficient=False, enable_math=True):
@@ -235,7 +235,7 @@ page describes the options.
# To turn on TunableOp, simply set this environment variable
export PYTORCH_TUNABLEOP_ENABLED=1
# Python
import torch
import torch.nn as nn
@@ -244,7 +244,7 @@ page describes the options.
W = torch.rand(200, 20, device="cuda")
Out = F.linear(A, W)
print(Out.size())
# tunableop_results0.csv
Validator,PT_VERSION,2.4.0
Validator,ROCM_VERSION,6.1.0.0-82-5fabb4c
@@ -253,7 +253,7 @@ page describes the options.
Validator,ROCBLAS_VERSION,4.1.0-cefa4a9b-dirty
GemmTunableOp_float_TN,tn_200_100_20,Gemm_Rocblas_32323,0.00669595
.. image:: ../../../data/how-to/llm-fine-tuning-optimization/tunableop.png
.. image:: ./images/model-acceleration-libs/tunableop.png
:alt: GEMM and TunableOp
:align: center
@@ -270,7 +270,7 @@ and as a back end for PyTorch quantized operators. FBGEMM offers optimized on-CP
strong performance on native tensor formats, and the ability to generate
high-performance shape- and size-specific kernels at runtime.
FBGEMM_GPU collects several high-performance PyTorch GPU operator libraries
FBGEMM_GPU collects several high-performance PyTorch GPU operator libraries
for use in training and inference. It provides efficient table-batched embedding functionality,
data layout transformation, and quantization support.
@@ -288,7 +288,7 @@ Installing FBGEMM_GPU consists of the following steps:
* Install ROCm using Docker or the :doc:`package manager <rocm-install-on-linux:install/install-methods/package-manager-index>`
* Install the nightly `PyTorch <https://pytorch.org/>`_ build
* Complete the pre-build and build tasks
.. note::
FBGEMM_GPU doesn't require the installation of FBGEMM. To optionally install
@@ -375,7 +375,7 @@ and run the ROCm Docker image, use this command:
You can also install ROCm using the package manager. FBGEMM_GPU requires the installation of the full ROCm package.
For more information, see :doc:`the ROCm installation guide <rocm-install-on-linux:install/detailed-install>`.
The ROCm package also requires the :doc:`MIOpen <miopen:index>` component as a dependency.
The ROCm package also requires the :doc:`MIOpen <miopen:index>` component as a dependency.
To install MIOpen, use the ``apt install`` command.
.. code-block:: shell
@@ -407,7 +407,7 @@ Install `PyTorch <https://pytorch.org/>`_ using ``pip`` for the most reliable an
Perform the prebuild and build
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
#. Clone the FBGEMM repository and the relevant submodules. Use ``pip`` to install the
#. Clone the FBGEMM repository and the relevant submodules. Use ``pip`` to install the
components in ``requirements.txt``. Run the following commands inside the Miniconda environment.
.. code-block:: shell
@@ -447,7 +447,7 @@ Perform the prebuild and build
# Set the Python platform name for the Linux case
export python_plat_name="manylinux2014_${ARCH}"
#. Build FBGEMM_GPU for the ROCm platform. Set ``ROCM_PATH`` to the path to your ROCm installation.
#. Build FBGEMM_GPU for the ROCm platform. Set ``ROCM_PATH`` to the path to your ROCm installation.
Run these commands from the ``fbgemm_gpu/`` directory inside the Miniconda environment.
.. code-block:: shell
@@ -474,7 +474,7 @@ Perform the prebuild and build
--package_variant=rocm \
-DHIP_ROOT_DIR="${ROCM_PATH}" \
-DCMAKE_C_FLAGS="-DTORCH_USE_HIP_DSA" \
-DCMAKE_CXX_FLAGS="-DTORCH_USE_HIP_DSA"
-DCMAKE_CXX_FLAGS="-DTORCH_USE_HIP_DSA"
Post-build validation
----------------------
@@ -533,8 +533,8 @@ follow these instructions:
# Run the test
python -m pytest -v -rsx -s -W ignore::pytest.PytestCollectionWarning split_table_batched_embeddings_test.py
To run the FBGEMM_GPU ``uvm`` test, use these commands. These tests only support the AMD MI210 and
more recent GPUs.
To run the FBGEMM_GPU ``uvm`` test, use these commands. These tests only support the AMD MI210 and
more recent GPUs.
.. code-block:: shell
@@ -6,14 +6,14 @@
Optimizing Triton kernels
*************************
This section introduces the general steps for
This section introduces the general steps for
`Triton <https://openai.com/index/triton/>`_ kernel optimization. Broadly,
Triton kernel optimization is similar to :doc:`HIP <hip:how-to/performance_guidelines>`
and CUDA kernel optimization.
Refer to the
:ref:`Triton kernel performance optimization <mi300x-triton-kernel-performance-optimization>`
section of the :doc:`workload` guide
section of the :doc:`workload-optimization` guide
for detailed information.
Triton kernel performance optimization includes the following topics.
@@ -29,11 +29,11 @@ The template parameters of the instance are grouped into four parameter types:
- [Parameters for determining extra operations on matrix elements](matrix-element-operation)
- [Performance-oriented tunable parameters](tunable-parameters)
<!--
<!--
================
### Figure 2
================ -->
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-template_parameters.jpg
```{figure} ./images/optimize-with-composable-kernel/ck-template_parameters.jpg
The template parameters of the selected GEMM kernel are classified into four groups. These template parameter groups should be defined properly before running the instance.
```
@@ -100,7 +100,7 @@ struct AddRelu
(tunable-parameters)=
#### Tunable parameters
#### Tunable parameters
The CK instance includes a series of tunable template parameters to control the parallel granularity of the workload to achieve load balancing on different hardware platforms.
@@ -123,11 +123,11 @@ After determining the template parameters, we instantiate the kernel with actual
The row and column, and stride information of input matrices are also passed to the instance. For batched GEMM, you must pass in additional batch count and batch stride values. The extra operations for pre and post-processing are also passed with an actual argument; for example, α and β for GEMM scaling operations. Afterward, the instantiated kernel is launched by the invoker, as illustrated in Figure 3.
<!--
<!--
================
### Figure 3
================ -->
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-kernel_launch.jpg
```{figure} ./images/optimize-with-composable-kernel/ck-kernel_launch.jpg
Templated kernel launching consists of kernel instantiation, making arguments by passing in actual application parameters, creating an invoker, and running the instance through the invoker.
```
@@ -152,11 +152,11 @@ The following section discusses the analysis of the operation flow of `Linear_Re
The first operation in the process is to perform the multiplication of input matrices A and B. The resulting matrix C is then scaled with α to obtain T1. At the same time, the process performs a scaling operation on D elements to obtain T2. Afterward, the process performs matrix addition between T1 and T2, element activation calculation using ReLU, and element rounding sequentially. The operations to generate E1, E2, and E are encapsulated and completed by a user-defined template function in CK (given in the next sub-section). This template function is integrated into the fundamental instance directly during the compilation phase so that all these steps can be fused in a single GPU kernel.
<!--
<!--
================
### Figure 4
================ -->
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-operation_flow.jpg
```{figure} ./images/optimize-with-composable-kernel/ck-operation_flow.jpg
Operation flow.
```
@@ -168,11 +168,11 @@ Third, consider the platform for implementing CK instances. The instances suffix
Here, we use [DeviceBatchedGemmMultiD_Xdl](https://github.com/ROCm/composable_kernel/tree/develop/example/24_batched_gemm) as the fundamental instance to implement the functionalities in the previous table.
<!--
<!--
================
### Figure 5
================ -->
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-root_instance.jpg
```{figure} ./images/optimize-with-composable-kernel/ck-root_instance.jpg
Use the DeviceBatchedGemmMultiD_Xdl instance as a root.
```
@@ -189,7 +189,7 @@ The inference of SQ quantized models relies on using PyTorch and Transformer lib
In GEMM, the A and B inputs are two-dimensional matrices, and the required input matrices of the selected fundamental CK instance are three-dimensional matrices. Therefore, we must convert the input 2-D tensors to 3-D tensors, by using `tensor`'s `unsqueeze()` method before passing these matrices to the instance. For batched GEMM in the preceding table, ignore this step.
```c++
// Function input and output
// Function input and output
torch::Tensor linear_relu_abde_i8(
torch::Tensor A_,
torch::Tensor B_,
@@ -197,10 +197,10 @@ torch::Tensor linear_relu_abde_i8(
float alpha,
float beta)
{
// Convert torch::Tensor A_ (M, K) to torch::Tensor A (1, M, K)
// Convert torch::Tensor A_ (M, K) to torch::Tensor A (1, M, K)
auto A = A_.unsqueeze(0);
// Convert torch::Tensor B_ (K, N) to torch::Tensor A (1, K, N)
// Convert torch::Tensor B_ (K, N) to torch::Tensor A (1, K, N)
auto B = B_.unsqueeze(0);
...
```
@@ -232,7 +232,7 @@ As shown in the following code block, we obtain M, N, and K values using input t
auto D = D_.view({1,-1}).repeat({M, 1});
// Allocate memory for E
auto E = torch::empty({batch_count, M, N},
auto E = torch::empty({batch_count, M, N},
torch::dtype(torch::kInt8).device(A.device()));
```
@@ -241,7 +241,7 @@ In the following code block, `ADataType`, `BDataType` and `D0DataType` are used
`AccDataType` determines the data precision used to represent the multiply-add results of A and B elements. Generally, a larger range data type is applied to store the multiply-add results of A and B to avoid result overflow; `I32` is applied in this case. The `CShuffleDataType I32` data type indicates that the multiply-add results continue to be stored in LDS as an `I32` data format. All of this is implemented through the following code block.
```c++
// Data precision
// Data precision
using ADataType = I8;
using BDataType = I8;
using AccDataType = I32;
@@ -265,7 +265,7 @@ Following the convention of various linear algebra libraries, row-major and colu
In CK, `PassThrough` is a struct denoting if an operation is applied to the tensor it binds to. To fuse the operations between E1, E2, and E introduced in section [Operation flow analysis](#operation-flow-analysis), we define a custom C++ struct, `ScaleScaleAddRelu`, and bind it to `CDEELementOp`. It determines the operations that will be applied to `CShuffle` (A×B results), tensor D, α, and β.
```c++
// No operations bound to the elements of A and B
// No operations bound to the elements of A and B
using AElementOp = PassThrough;
using BElementOp = PassThrough;
@@ -290,17 +290,17 @@ struct ScaleScaleAddRelu {
// Perform addition operation
F32 temp = c_scale + d_scale;
// Perform RELU operation
temp = temp > 0 ? temp : 0;
// Perform rounding operation
// Perform rounding operation
temp = temp > 127 ? 127 : temp;
// Return to E
e = ck::type_convert<I8>(temp);
}
F32 alpha;
F32 beta;
};
@@ -315,16 +315,16 @@ static constexpr auto GemmDefault = ck::tensor_operation::device::GemmSpecializa
The template parameters of the target fundamental instance are initialized with the above parameters and includes default tunable parameters. For specific tuning methods, see [Tunable parameters](#tunable-parameters).
```c++
using DeviceOpInstance = ck::tensor_operation::device::DeviceBatchedGemmMultiD_Xdl<
using DeviceOpInstance = ck::tensor_operation::device::DeviceBatchedGemmMultiD_Xdl<
// Tensor layout
ALayout, BLayout, DsLayout, ELayout,
ALayout, BLayout, DsLayout, ELayout,
// Tensor data type
ADataType, BDataType, AccDataType, CShuffleDataType, DsDataType, EDataType,
ADataType, BDataType, AccDataType, CShuffleDataType, DsDataType, EDataType,
// Tensor operation
AElementOp, BElementOp, CDEElementOp,
// Padding strategy
AElementOp, BElementOp, CDEElementOp,
// Padding strategy
GemmDefault,
// Tunable parameters
// Tunable parameters
tunable parameters>;
```
@@ -356,13 +356,13 @@ invoker.Run(argument, StreamConfig{nullptr, 0});
The output of the fundamental instance is a calculated batched matrix E (batch, M, N). Before the return, it needs to be converted to a 2-D matrix if a normal GEMM result is required.
```c++
// Convert (1, M, N) to (M, N)
// Convert (1, M, N) to (M, N)
return E.squeeze(0);
```
### Binding to Python
Since these functions are written in C++ and `torch::Tensor`, you can use `pybind11` to bind the functions and import them as Python modules. For the example, the necessary binding code for exposing the functions in the table spans but a few lines.
Since these functions are written in C++ and `torch::Tensor`, you can use `pybind11` to bind the functions and import them as Python modules. For the example, the necessary binding code for exposing the functions in the table spans but a few lines.
```c++
#include <torch/extension.h>
@@ -390,7 +390,7 @@ os.environ["CXX"] = "hipcc"
sources = [
'torch_int/kernels/linear.cpp',
'torch_int/kernels/bmm.cpp',
'torch_int/kernels/pybind.cpp',
'torch_int/kernels/pybind.cpp',
]
include_dirs = ['torch_int/kernels/include']
@@ -418,11 +418,11 @@ setup(
Run `python setup.py install` to build and install the extension. It should look something like Figure 6:
<!--
<!--
================
### Figure 6
================ -->
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-compilation.jpg
```{figure} ./images/optimize-with-composable-kernel/ck-compilation.jpg
Compilation and installation of the INT8 kernels.
```
@@ -430,11 +430,11 @@ Compilation and installation of the INT8 kernels.
The implementation architecture of running SmoothQuant models on MI300X GPUs is illustrated in Figure 7, where (a) shows the decoder layer composition components of the target model, (b) shows the major implementation class for the decoder layer components, and \(c\) denotes the underlying GPU kernels implemented by CK instance.
<!--
<!--
================
### Figure 7
================ -->
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-inference_flow.jpg
```{figure} ./images/optimize-with-composable-kernel/ck-inference_flow.jpg
The implementation architecture of running SmoothQuant models on AMD MI300X GPUs.
```
@@ -456,11 +456,11 @@ Note that since the default values were used for the tunable parameters of the f
Figure 8 shows the performance comparisons between the original FP16 and the SmoothQuant-quantized INT8 models on a single MI300X GPU. The GPU memory footprints of SmoothQuant-quantized models are significantly reduced. It also indicates the per-sample inference latency is significantly reduced for all SmoothQuant-quantized OPT models (illustrated in (b)). Notably, the performance of the CK instance-based INT8 kernel steadily improves with an increase in model size.
<!--
<!--
================
### Figure 8
================ -->
```{figure} ../../../data/how-to/llm-fine-tuning-optimization/ck-comparisons.jpg
```{figure} ./images/optimize-with-composable-kernel/ck-comparisons.jpg
Performance comparisons between the original FP16 and the SmoothQuant-quantized INT8 models on a single MI300X GPU.
```
@@ -1,6 +1,6 @@
.. meta::
:description: Learn about vLLM V1 inference tuning on AMD Instinct GPUs for optimal performance.
:keywords: AMD, Instinct, MI300X, MI325X, MI350X, MI355X, HPC, tuning, BIOS settings, NBIO, ROCm,
:keywords: AMD, Instinct, MI300X, HPC, tuning, BIOS settings, NBIO, ROCm,
environment variable, performance, HIP, Triton, PyTorch TunableOp, vLLM, RCCL,
MIOpen, GPU, resource utilization
@@ -25,18 +25,14 @@ Instinct MI300X, MI325X, MI350X, and MI355X GPUs. Learn how to:
Performance environment variables
=================================
The following variables are generally useful for Instinct MI300X/MI325X/MI350X/MI355X GPUs and vLLM:
The following variables are generally useful for Instinct MI300X/MI355X GPUs and vLLM:
* **HIP and math libraries**
* ``export HIP_FORCE_DEV_KERNARG=1`` — improves kernel launch performance by
forcing device kernel arguments. This is already set by default in
:doc:`vLLM ROCm Docker images
</how-to/rocm-for-ai/inference/benchmark-docker/vllm>`. Bare-metal users
:doc:`vLLM ROCm Docker images </ai-inference/vllm>`. Bare-metal users
should set this manually.
* ``export SAFETENSORS_FAST_GPU=1`` — enables GPU-accelerated safetensors
loading, significantly reducing model load time for large models. Already
set in vLLM ROCm Docker images. Bare-metal users should set this manually.
* ``export TORCH_BLAS_PREFER_HIPBLASLT=1`` — explicitly prefers hipBLASLt
over hipBLAS for GEMM operations. By default, PyTorch uses heuristics to
choose the best BLAS library. Setting this can improve linear layer
@@ -45,7 +41,7 @@ The following variables are generally useful for Instinct MI300X/MI325X/MI350X/M
* **RCCL (collectives for multi-GPU)**
* ``export NCCL_MIN_NCHANNELS=112`` — increases RCCL channels from default
(typically 32-64) to 112 on the Instinct MI300X/MI325X. **Only beneficial for
(typically 32-64) to 112 on the Instinct MI300X. **Only beneficial for
multi-GPU distributed workloads** (tensor parallelism, pipeline
parallelism). Single-GPU inference does not need this.
@@ -54,31 +50,31 @@ The following variables are generally useful for Instinct MI300X/MI325X/MI350X/M
AITER (AI Tensor Engine for ROCm) switches
==========================================
AITER (AI Tensor Engine for ROCm) provides ROCm-specific fused kernels optimized for Instinct MI350 Series and MI300X/MI325X GPUs in vLLM V1.
AITER (AI Tensor Engine for ROCm) provides ROCm-specific fused kernels optimized for Instinct MI350 Series and MI300X GPUs in vLLM V1.
Enable all AITER optimizations with a single master switch:
How AITER flags work:
* ``VLLM_ROCM_USE_AITER`` is the master switch (defaults to ``False``/``0``).
* Individual feature flags (``VLLM_ROCM_USE_AITER_LINEAR``, ``VLLM_ROCM_USE_AITER_MOE``, and so on) default to ``True`` but only activate when the master switch is enabled.
* To enable a specific AITER feature, you must set both ``VLLM_ROCM_USE_AITER=1`` and the specific feature flag to ``1``.
Quick start examples:
.. code-block:: bash
# Enable all AITER optimizations (recommended for most workloads)
export VLLM_ROCM_USE_AITER=1
vllm serve MODEL_NAME
Most individual AITER sub-flags default to ``1`` when the master switch is on,
while specialized features retain the defaults listed below. You rarely need to
change them. To select a specific attention backend, use ``--attention-backend``
(see :ref:`backend selection <vllm-optimization-aiter-backend-selection>`).
# Enable AITER Fused MoE and enable Triton Prefill-Decode (split) attention
export VLLM_ROCM_USE_AITER=1
export VLLM_V1_USE_PREFILL_DECODE_ATTENTION=1
export VLLM_ROCM_USE_AITER_MHA=0
vllm serve MODEL_NAME
**Flags you might adjust:**
* ``VLLM_ROCM_SHUFFLE_KV_CACHE_LAYOUT=1`` — Set for high-concurrency MHA workloads (≥32 concurrent requests) with ``ROCM_AITER_FA``. Defaults to ``0``.
* ``VLLM_ROCM_USE_AITER_MOE=0`` — Disable only if you hit ``RuntimeError: wrong! device_gemm ...``. Try ``AITER_ONLINE_TUNE=1`` first. See :ref:`AITER MoE requirements <vllm-optimization-aiter-moe-requirements>`.
* ``VLLM_ROCM_USE_AITER=0`` — Disable AITER entirely to fall back to Triton kernels (for debugging).
Advanced: individual AITER flags
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
The following table lists AITER-related sub-flags for fine-grained control. Most
users do not need to modify these; the default behavior for each flag is listed below.
# Disable AITER entirely (i.e, use vLLM Triton Unified Attention Kernel)
export VLLM_ROCM_USE_AITER=0
vllm serve MODEL_NAME
.. list-table::
:header-rows: 1
@@ -88,135 +84,232 @@ users do not need to modify these; the default behavior for each flag is listed
- Description (default behavior)
* - ``VLLM_ROCM_USE_AITER``
- Master switch to enable AITER kernels (``0`` by default). All other ``VLLM_ROCM_USE_AITER_*`` flags require this to be set to ``1``.
- Master switch to enable AITER kernels (``0``/``False`` by default). All other ``VLLM_ROCM_USE_AITER_*`` flags require this to be set to ``1``.
* - ``VLLM_ROCM_USE_AITER_LINEAR``
- Use AITER quantization operators + GEMM for linear layers (defaults to ``1`` when AITER is on). Accelerates matrix multiplications in all transformer layers. **Recommended to keep enabled**.
- Use AITER quantization operators + GEMM for linear layers (defaults to ``True`` when AITER is on). Accelerates matrix multiplications in all transformer layers. **Recommended to keep enabled**.
* - ``VLLM_ROCM_USE_AITER_MOE``
- Use AITER fused-MoE kernels (defaults to ``1`` when AITER is on). Accelerates Mixture-of-Experts routing and computation. See the note on :ref:`AITER MoE requirements <vllm-optimization-aiter-moe-requirements>`.
- Use AITER fused-MoE kernels (defaults to ``True`` when AITER is on). Accelerates Mixture-of-Experts routing and computation. See the note on :ref:`AITER MoE requirements <vllm-optimization-aiter-moe-requirements>`.
* - ``VLLM_ROCM_USE_AITER_RMSNORM``
- Use AITER RMSNorm kernels (defaults to ``1`` when AITER is on). Accelerates normalization layers. **Recommended: keep enabled.**
- Use AITER RMSNorm kernels (defaults to ``True`` when AITER is on). Accelerates normalization layers. **Recommended: keep enabled.**
* - ``VLLM_ROCM_USE_AITER_MLA``
- Use AITER Multi-head Latent Attention for supported models, for example, DeepSeek-V3/R1 (defaults to ``1`` when AITER is on). See the section on :ref:`AITER MLA requirements <vllm-optimization-aiter-mla-requirements>`.
- Use AITER Multi-head Latent Attention for supported models, for example, DeepSeek-V3/R1 (defaults to ``True`` when AITER is on). See the section on :ref:`AITER MLA requirements <vllm-optimization-aiter-mla-requirements>`.
* - ``VLLM_ROCM_USE_AITER_MHA``
- Use AITER Multi-Head Attention kernels (defaults to ``1`` when AITER is on; set to ``0`` to use Triton attention backends or ``ROCM_ATTN`` backend instead). See :ref:`attention backend selection <vllm-optimization-aiter-backend-selection>`.
- Use AITER Multi-Head Attention kernels (defaults to ``True`` when AITER is on; set to ``0`` to use Triton attention backends and Prefill-Decode attention backend instead). See :ref:`attention backend selection <vllm-optimization-aiter-backend-selection>`.
* - ``VLLM_ROCM_USE_AITER_UNIFIED_ATTENTION``
- Enable AITER's optimized unified attention kernel (defaults to ``0``). Only takes effect when AITER is enabled and AITER MHA is disabled (``VLLM_ROCM_USE_AITER_MHA=0``). When set to ``0``, falls back to vLLM's Triton unified attention. Can also be enabled via ``--attention-backend ROCM_AITER_UNIFIED_ATTN``.
- Enable AITER's optimized unified attention kernel (defaults to ``False``). Only takes effect when: AITER is enabled; unified attention mode is active (``VLLM_V1_USE_PREFILL_DECODE_ATTENTION=0``); and AITER MHA is disabled (``VLLM_ROCM_USE_AITER_MHA=0``). When disabled, falls back to vLLM's Triton unified attention.
* - ``VLLM_ROCM_USE_AITER_FP8BMM``
- Use AITER ``FP8`` batched matmul (defaults to ``1`` when AITER is on). Fuses ``FP8`` per-token quantization with batched GEMM (used in MLA models like DeepSeek-V3).
* - ``VLLM_ROCM_USE_AITER_FP4BMM``
- Use AITER ``FP4`` batched matmul (defaults to ``1`` when AITER is on). Fuses ``FP4`` per-token quantization with batched GEMM (used in MLA models like DeepSeek-V3). Requires an Instinct MI350X/MI355X GPU.
* - ``VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS``
- Fuse shared expert computation into the AITER fused-MoE kernel (defaults to ``0``). Applies to MoE models with shared experts (for example, DeepSeek-V3/R1 with 1 shared expert). Requires SiLU/GELU activation (``is_act_and_mul``). Incompatible with `MoRI <https://github.com/ROCm/mori#mori>`__ scheduling — disable with ``VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS=0`` if using MoRI.
* - ``VLLM_ROCM_USE_AITER_FP4_ASM_GEMM``
- Enable AITER assembly (HIP) FP4 GEMM kernels for MXFP4-quantized models (defaults to ``0``). When set to ``1``, uses hand-tuned ASM kernels instead of Triton for ``FP4×FP4`` weight GEMMs — faster at small batch sizes (``M`` ≤ 64). Requires Instinct MI350X/MI355X (``supports_mx()``). Combine with ``--quantization quark`` for MXFP4 models.
- Use AITER ``FP8`` batched matmul (defaults to ``True`` when AITER is on). Fuses ``FP8`` per-token quantization with batched GEMM (used in MLA models like DeepSeek-V3). Requires an Instinct MI300X/MI355X GPU.
* - ``VLLM_ROCM_USE_SKINNY_GEMM``
- Prefer skinny-GEMM kernel variants for small batch sizes (defaults to ``1``). Improves performance when ``M`` dimension is small. **Recommended to keep enabled**.
- Prefer skinny-GEMM kernel variants for small batch sizes (defaults to ``True``). Improves performance when ``M`` dimension is small. **Recommended to keep enabled**.
* - ``VLLM_ROCM_FP8_PADDING``
- Pad ``FP8`` linear weight tensors to improve memory locality (defaults to ``1``). Minor memory overhead for better performance.
- Pad ``FP8`` linear weight tensors to improve memory locality (defaults to ``True``). Minor memory overhead for better performance.
* - ``VLLM_ROCM_MOE_PADDING``
- Pad MoE weight tensors for better memory access patterns (defaults to ``1``). Same memory/performance tradeoff as ``FP8`` padding.
* - ``VLLM_ROCM_SHUFFLE_KV_CACHE_LAYOUT``
- This only affects the ``ROCM_AITER_FA`` backend. When set to ``0``, it uses the HIP Paged Attention implementation ``torch.ops.aiter.paged_attention_v1`` from AITER. **Good for low concurrency (for example, concurrency ≤32).** When set to ``1``, it uses the ASM Paged Attention Kernel ``pa_fwd_asm`` from AITER. **Good for high concurrency (for example, concurrency ≥32).** (Defaults to ``0`` when AITER is on.)
- Pad MoE weight tensors for better memory access patterns (defaults to ``True``). Same memory/performance tradeoff as ``FP8`` padding.
* - ``VLLM_ROCM_CUSTOM_PAGED_ATTN``
- Use custom paged-attention decode kernel when ``ROCM_ATTN`` backend is selected (defaults to ``1``). See :ref:`Attention backend selection with AITER <vllm-optimization-aiter-backend-selection>`.
- Use custom paged-attention decode kernel when Prefill-Decode attention backend is selected (defaults to ``True``). See :ref:`Attention backend selection with AITER <vllm-optimization-aiter-backend-selection>`.
.. note::
When ``VLLM_ROCM_USE_AITER=1``, most AITER component flags (``LINEAR``,
``MOE``, ``RMSNORM``, ``MLA``, ``MHA``, ``FP8BMM``) automatically default to
``True``. You typically only need to set the master switch
``VLLM_ROCM_USE_AITER=1`` to enable all optimizations. ROCm provides a
prebuilt optimized Docker image for validating the performance of LLM
inference with vLLM on MI300X Series GPUs. The Docker image includes ROCm,
vLLM, and PyTorch. For more information, see :doc:`/ai-inference/vllm`.
.. _vllm-optimization-aiter-moe-requirements:
AITER MoE requirements (Mixtral, DeepSeek-V2/V3, Qwen-MoE models)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
``VLLM_ROCM_USE_AITER_MOE`` enables AITER's optimized Mixture-of-Experts kernels, such as expert routing (topk selection) and expert computation for better performance.
Applicable models:
* Mixtral series: for example, Mixtral-8x7B / Mixtral-8x22B
* Llama-4 family: for example, Llama-4-Scout-17B-16E / Llama-4-Maverick-17B-128E
* DeepSeek family: DeepSeek-V2 / DeepSeek-V3 / DeepSeek-R1
* Qwen family: Qwen1.5-MoE / Qwen2-MoE / Qwen2.5-MoE series
* Other MoE architectures
When to enable:
* **Enable (default):** For all MoE models on the Instinct MI300X/MI355X for best throughput
* **Disable:** Only for debugging or if you encounter numerical issues
Example usage:
.. code-block:: bash
# Standard MoE model (Mixtral)
VLLM_ROCM_USE_AITER=1 vllm serve mistralai/Mixtral-8x7B-Instruct-v0.1
# Hybrid MoE+MLA model (DeepSeek-V3) - requires both MOE and MLA flags
VLLM_ROCM_USE_AITER=1 vllm serve deepseek-ai/DeepSeek-V3 \
--block-size 1 \
--tensor-parallel-size 8
.. _vllm-optimization-aiter-mla-requirements:
.. _vllm-optimization-aiter-mla-sparse-requirements:
AITER MLA requirements (DeepSeek-V3/R1 models)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
``VLLM_ROCM_USE_AITER_MLA`` enables AITER MLA (Multi-head Latent Attention) optimization for supported models. Defaults to **True** when AITER is on.
Critical requirement:
* **Must** explicitly set ``--block-size 1``
.. important::
If you omit ``--block-size 1``, vLLM will raise an error rather than defaulting to 1.
Applicable models:
* DeepSeek-V3 / DeepSeek-R1
* DeepSeek-V2
* Other models using multi-head latent attention (MLA) architecture
Example usage:
.. code-block:: bash
# DeepSeek-R1 with AITER MLA (requires 8 GPUs)
VLLM_ROCM_USE_AITER=1 vllm serve deepseek-ai/DeepSeek-R1 \
--block-size 1 \
--tensor-parallel-size 8
.. _vllm-optimization-aiter-backend-selection:
Attention backend selection with AITER
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
Most models work out of the box with ``VLLM_ROCM_USE_AITER=1`` — vLLM auto-selects
the optimal backend. Use ``--attention-backend`` to override the auto-selected backend.
Understanding which attention backend to use helps optimize your deployment.
.. code-block:: bash
Quick reference: Which attention backend will I get?
export VLLM_ROCM_USE_AITER=1
vllm serve <your-model> --tensor-parallel-size <tp>
Default behavior (no configuration)
.. note::
Always set ``VLLM_ROCM_USE_AITER=1`` even when using ``--attention-backend`` explicitly.
``--attention-backend`` only overrides the attention kernel; ``VLLM_ROCM_USE_AITER=1``
is still required to enable AITER for GEMM, RMSNorm, and MoE kernels.
The Radeon/fallback backends (``ROCM_ATTN``, ``TRITON_MLA``) are the exception —
they do not use AITER and do not require the env var.
Without setting any environment variables, vLLM uses:
The table below shows which backend is selected per model type and how to tune it.
* **vLLM Triton Unified Attention** — A single Triton kernel handling both prefill and decode phases
* Works on all ROCm platforms
* Good baseline performance
**Recommended**: Enable AITER (set ``VLLM_ROCM_USE_AITER=1``)
When you enable AITER, the backend is automatically selected based on your model:
.. code-block:: text
Is your model using MLA architecture? (DeepSeek-V3/R1/V2)
├─ YES → AITER MLA Backend
│ • Requires --block-size 1
│ • Best performance for MLA models
│ • Automatically selected
└─ NO → AITER MHA Backend
• For standard transformer models (Llama, Mistral, etc.)
• Optimized for Instinct MI300X/MI355X
• Automatically selected
**Advanced**: Manual backend selection
Most users won't need this, but you can override the defaults:
.. list-table::
:widths: 40 60
:header-rows: 1
:widths: 15 18 35 32
* - Model type
- Backend
- How to enable
- Tuning tips
* - To use this backend
- Set these flags
* - **MHA models** (Llama, Mistral, Qwen, Mixtral, MiniMax-M2.5)
- **ROCM_AITER_FA** (recommended, auto-selected)
- ``VLLM_ROCM_USE_AITER=1`` (auto-selected). To override: ``VLLM_ROCM_USE_AITER=1 --attention-backend ROCM_AITER_FA``. Add ``VLLM_ROCM_SHUFFLE_KV_CACHE_LAYOUT=1`` for shuffled KV cache.
- **2.74.4x TPS** over legacy ``ROCM_ATTN``. Set ``VLLM_ROCM_SHUFFLE_KV_CACHE_LAYOUT=1`` for **1520% decode improvement** at high concurrency. Low TTFT: ``--max-num-batched-tokens`` ≤ 8k16k. High throughput: ≥ 32k with ``cudagraph_mode=FULL``.
* - AITER MLA (MLA models only)
- ``VLLM_ROCM_USE_AITER=1`` (auto-selects for DeepSeek-V3/R1)
* - **MLA models** (DeepSeek-V3/R1/V2, Kimi-K2.5, Mistral-Large-3-675B)
- **ROCM_AITER_MLA** (recommended, auto-selected)
- ``VLLM_ROCM_USE_AITER=1`` (auto-selected). To override: ``VLLM_ROCM_USE_AITER=1 --attention-backend ROCM_AITER_MLA``. ``--block-size 1`` is no longer mandatory (vLLM ≥0.14); still recommended for prefix-caching workloads.
- **1.21.5x higher TPS** over ``TRITON_MLA``. Supports uniform-batch CUDA graphs and MTP. On MI300X/MI325X (gfx942), ``ROCM_AITER_TRITON_MLA`` may show 23% higher TPS. On MI355X (gfx950), ``ROCM_AITER_MLA`` is preferred (uses AITER assembly MHA for prefill).
* - AITER MHA (standard models)
- ``VLLM_ROCM_USE_AITER=1`` (auto-selects for non-MLA models)
* - **DSA models** (DeepSeek-V3.2, GLM-5)
- **ROCM_AITER_MLA_SPARSE** (auto-selected)
- ``VLLM_ROCM_USE_AITER=1`` — auto-detected from ``index_topk`` in model config. Requires ``--block-size 1``.
- Instinct MI300X/MI325X/MI350X/MI355X only.
* - vLLM Triton Unified (default)
- ``VLLM_ROCM_USE_AITER=0`` (or unset)
* - **gpt-oss models** (gpt-oss-120b/20b)
- **ROCM_AITER_UNIFIED_ATTN**
- ``VLLM_ROCM_USE_AITER=1 --attention-backend ROCM_AITER_UNIFIED_ATTN``
-
* - Triton Prefill-Decode (split) without AITER
- | ``VLLM_V1_USE_PREFILL_DECODE_ATTENTION=1``
* - **Radeon / fallback**
- **ROCM_ATTN** (MHA) or **TRITON_MLA** (MLA)
- ``--attention-backend ROCM_ATTN`` or ``TRITON_MLA``
- ``ROCM_ATTN`` is preferred over ``TRITON_ATTN`` — it uses a custom HIP paged-attention kernel for decode when the model's KV head size is supported, and falls back to Triton only when not. If ``ROCM_ATTN`` is slow for your model (unsupported head size triggers Triton decode), try ``TRITON_ATTN``. Both work on Radeon GPUs.
* - Triton Prefill-Decode (split) along with AITER Fused-MoE
- | ``VLLM_ROCM_USE_AITER=1``
| ``VLLM_ROCM_USE_AITER_MHA=0``
| ``VLLM_V1_USE_PREFILL_DECODE_ATTENTION=1``
.. note::
**MoE models** (Mixtral, Llama-4-Scout/Maverick, DeepSeek-V2/V3/R1, Kimi-K2.5, MiniMax-M2.5, GLM-5, Qwen-MoE): AITER MoE kernels activate automatically with ``VLLM_ROCM_USE_AITER=1`` — no extra attention backend flags needed. If you hit ``RuntimeError: wrong! device_gemm ...``, set ``AITER_ONLINE_TUNE=1`` and retry. Only disable MoE kernels (``VLLM_ROCM_USE_AITER_MOE=0``) if that also fails.
Once AITER is configured, see `Parallelism strategies (run vLLM on multiple GPUs)`_ for TP/DP/EP choices — especially for MLA and MoE models where the wrong strategy wastes memory or throughput.
* - AITER Unified Attention
- | ``VLLM_ROCM_USE_AITER=1``
| ``VLLM_ROCM_USE_AITER_MHA=0``
| ``VLLM_ROCM_USE_AITER_UNIFIED_ATTENTION=1``
**Quick start examples**:
.. code-block:: bash
# DSA model (DeepSeek-V3.2) — backend auto-selected from model config
VLLM_ROCM_USE_AITER=1 vllm serve deepseek-ai/DeepSeek-V3.2 \
# Recommended: Standard model with AITER (Llama, Mistral, Qwen, etc.)
VLLM_ROCM_USE_AITER=1 vllm serve meta-llama/Llama-3.3-70B-Instruct
# MLA model with AITER (DeepSeek-V3/R1)
VLLM_ROCM_USE_AITER=1 vllm serve deepseek-ai/DeepSeek-R1 \
--block-size 1 \
--tensor-parallel-size 8
# Explicitly select a backend for MLA models
VLLM_ROCM_USE_AITER=1 vllm serve deepseek-ai/DeepSeek-R1-0528 \
--tensor-parallel-size 8 \
--attention-backend ROCM_AITER_MLA
# Advanced: Use Prefill-Decode split (for short input cases) with AITER Fused-MoE
VLLM_ROCM_USE_AITER=1 \
VLLM_ROCM_USE_AITER_MHA=0 \
VLLM_V1_USE_PREFILL_DECODE_ATTENTION=1 \
vllm serve meta-llama/Llama-4-Scout-17B-16E
# MHA model with shuffled KV cache layout for high concurrency
VLLM_ROCM_USE_AITER=1 VLLM_ROCM_SHUFFLE_KV_CACHE_LAYOUT=1 \
vllm serve meta-llama/Llama-3.3-70B-Instruct \
--attention-backend ROCM_AITER_FA
**Which backend should I choose?**
.. list-table::
:widths: 30 70
:header-rows: 1
* - Your use case
- Recommended backend
* - **Standard transformer models** (Llama, Mistral, Qwen, Mixtral)
- **AITER MHA** (``VLLM_ROCM_USE_AITER=1``) — **Recommended for most workloads** on Instinct MI300X/MI355X. Provides optimized attention kernels for both prefill and decode phases.
* - **MLA models** (DeepSeek-V3/R1/V2)
- **AITER MLA** (auto-selected with ``VLLM_ROCM_USE_AITER=1``) — Required for optimal performance, must use ``--block-size 1``
* - **gpt-oss models** (gpt-oss-120b/20b)
- **AITER Unified Attention** (``VLLM_ROCM_USE_AITER=1``, ``VLLM_ROCM_USE_AITER_MHA=0``, ``VLLM_ROCM_USE_AITER_UNIFIED_ATTENTION=1``) — Required for optimal performance
* - **Debugging or compatibility**
- **vLLM Triton Unified** (default with ``VLLM_ROCM_USE_AITER=0``) — Generic fallback, works everywhere
**Important notes:**
* **AITER MHA and AITER MLA are mutually exclusive** — vLLM automatically detects MLA models and selects the appropriate backend
* **For 95% of users:** Simply set ``VLLM_ROCM_USE_AITER=1`` and let vLLM choose the right backend
* When in doubt, start with AITER enabled (the recommended configuration) and profile your specific workload
Backend choice quick recipes
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
* **Standard transformers (any prompt length):** Start with ``VLLM_ROCM_USE_AITER=1`` → AITER MHA. For CUDA graph modes, see architecture-specific guidance below (Dense vs MoE models have different optimal modes).
* **Latency-sensitive chat (low TTFT):** keep ``--max-num-batched-tokens`` ≤ **8k16k** with AITER.
* **Streaming decode (low ITL):** raise ``--max-num-batched-tokens`` to **32k64k**.
* **Offline max throughput:** ``--max-num-batched-tokens`` ≥ **32k** with ``cudagraph_mode=FULL``.
**How to verify which backend is active**
@@ -225,13 +318,68 @@ Check vLLM's startup logs to confirm which attention backend is being used:
.. code-block:: bash
# Start vLLM and check logs
VLLM_ROCM_USE_AITER=1 vllm serve meta-llama/Llama-3.3-70B-Instruct 2>&1 | grep -i "using.*backend"
VLLM_ROCM_USE_AITER=1 vllm serve meta-llama/Llama-3.3-70B-Instruct 2>&1 | grep -i attention
Look for ``Using <backend_name> backend.`` in the startup output — for example,
``Using ROCM_AITER_FA backend.``
**Expected log messages:**
For in-depth architecture and benchmarks of all 7 ROCm attention backends, see the
`ROCm Attention Backend blog post <https://vllm.ai/blog/rocm-attention-backend>`_.
* AITER MHA: ``Using Aiter Flash Attention backend on V1 engine.``
* AITER MLA: ``Using AITER MLA backend on V1 engine.``
* vLLM Triton MLA: ``Using Triton MLA backend on V1 engine.``
* vLLM Triton Unified: ``Using Triton Attention backend on V1 engine.``
* AITER Triton Unified: ``Using Aiter Unified Attention backend on V1 engine.``
* AITER Triton Prefill-Decode: ``Using Rocm Attention backend on V1 engine.``
Attention backend technical details
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
This section provides technical details about vLLM's attention backends on ROCm.
vLLM V1 on ROCm provides these attention implementations:
1. **vLLM Triton Unified Attention** (default when AITER is **off**)
* Single unified Triton kernel handling both chunked prefill and decode phases
* Generic implementation that works across all ROCm platforms
* Good baseline performance
* Automatically selected when ``VLLM_ROCM_USE_AITER=0`` (or unset)
* Supports GPT-OSS
2. **AITER Triton Unified Attention** (advanced, requires manual configuration)
* The AMD optimized unified Triton kernel
* Enable with ``VLLM_ROCM_USE_AITER=1``, ``VLLM_ROCM_USE_AITER_MHA=0``, and ``VLLM_ROCM_USE_AITER_UNIFIED_ATTENTION=1``.
* Only useful for specific workloads. Most users should use AITER MHA instead.
* Recommended this backend when running GPT-OSS.
3. **AITER Triton PrefillDecode Attention** (hybrid, Instinct MI300X-optimized)
* Enable with ``VLLM_V1_USE_PREFILL_DECODE_ATTENTION=1``
* Uses separate kernels for prefill and decode phases:
* **Prefill**: ``context_attention_fwd`` Triton kernel
* **Primary decode**: ``torch.ops._rocm_C.paged_attention`` (custom ROCm kernel optimized for head sizes 64/128, block sizes 16/32, GQA 116, context ≤131k; sliding window not supported)
* **Fallback decode**: ``kernel_paged_attention_2d`` Triton kernel when shapes don't meet primary decode requirements
* Usually better compared to unified Triton kernels
* Performance vs AITER MHA varies: AITER MHA is typically faster overall, but Prefill-Decode split may win in short input scenarios
* The custom paged attention decode kernel is controlled by ``VLLM_ROCM_CUSTOM_PAGED_ATTN`` (default **True**)
4. **AITER Multi-Head Attention (MHA)** (default when AITER is **on**)
* Controlled by ``VLLM_ROCM_USE_AITER_MHA`` (**1** = enabled)
* Best all-around performance for standard transformer models
* Automatically selected when ``VLLM_ROCM_USE_AITER=1`` and model is not MLA
5. **vLLM Triton Multi-head Latent Attention (MLA)** (for DeepSeek-V3/R1/V2)
* Automatically selected when ``VLLM_ROCM_USE_AITER=0`` (or unset)
6. **AITER Multi-head Latent Attention (MLA)** (for DeepSeek-V3/R1/V2)
* Controlled by ``VLLM_ROCM_USE_AITER_MLA`` (``1`` = enabled)
* Required for optimal performance on MLA architecture models
* Automatically selected when ``VLLM_ROCM_USE_AITER=1`` and model uses MLA
* Requires ``--block-size 1``
Quick Reduce (large all-reduces on ROCm)
========================================
@@ -246,7 +394,7 @@ It supports FP16/BF16 as well as symmetric INT8/INT6/INT4 quantized all-reduce (
Control via:
* ``VLLM_ROCM_QUICK_REDUCE_QUANTIZATION````["NONE","FP","INT8","INT6","INT4"]`` (default ``NONE``).
* ``VLLM_ROCM_QUICK_REDUCE_CAST_BF16_TO_FP16``: cast BF16 input to FP16 (``1`` by default for performance).
* ``VLLM_ROCM_QUICK_REDUCE_CAST_BF16_TO_FP16``: cast BF16 input to FP16 (``1/True`` by default for performance).
* ``VLLM_ROCM_QUICK_REDUCE_MAX_SIZE_BYTES_MB``: cap the preset buffer (default ``NONE````2048`` MB).
Quick Reduce tends to help **throughput** at higher TP counts (for example, 48) with many concurrent requests.
@@ -263,34 +411,12 @@ vLLM supports the following parallelism strategies:
For more details, see `Parallelism and scaling <https://docs.vllm.ai/en/stable/serving/parallelism_scaling.html>`_.
**Quick-reference decision table:**
**Choosing the right strategy:**
.. list-table::
:header-rows: 1
:widths: 30 35 35
* - Model type
- Low concurrency (≤128 requests)
- High concurrency (≥512 requests)
* - **Dense** (for example, Llama, Qwen-dense, Mistral-dense)
- TP only
- TP + independent DP replicas (your own load balancer)
* - **MoE, standard density ≥3%** (for example, Qwen3-235B-A22B, DeepSeek-V3/R1)
- TP + EP
- DP + EP
* - **MoE, ultra-sparse <1%** (for example, Llama-4-Maverick at 0.78%)
- TP only — **no EP** (AllToAll overhead exceeds benefit)
- DP only — **no EP**
* - **MLA models** (for example, DeepSeek-V2/V3/R1, Kimi-K2.5, Mistral-Large-3-675B)
- TP + EP
- **DP + EP** — TP alone duplicates the full KV cache on every GPU; use DP Attention to partition it
EP = ``--enable-expert-parallel``. DP = ``--data-parallel-size N``.
See `Data Parallel Attention (advanced)`_ for the MLA memory explanation and `Expert parallelism`_ for EP details.
* **Tensor Parallelism (TP)**: Use when model doesn't fit on one GPU. Prefer staying within a single XGMI island (≤8 GPUs on the Instinct MI300X).
* **Pipeline Parallelism (PP)**: Use for very large models across nodes. Set TP to GPUs per node, scale with PP across nodes.
* **Data Parallelism (DP)**: Use when model fits on single GPU or TP group, and you need higher throughput. Combine with TP/PP for large models.
* **Expert Parallelism (EP)**: Use for MoE models with ``--enable-expert-parallel``. More efficient than TP for MoE layers.
Tensor parallelism
^^^^^^^^^^^^^^^^^^
@@ -319,9 +445,6 @@ Tensor parallelism splits each layer of the model weights across multiple GPUs w
.. tip::
For structured data parallelism deployments with load balancing, see :ref:`data-parallelism-section`.
.. note::
**MLA models (DeepSeek, Kimi-K2.5, Mistral-Large-3-675B):** TP alone replicates the full KV cache on every GPU, which wastes memory at high concurrency. See `Data Parallel Attention (advanced)`_ for the DP+EP configuration that partitions the KV cache instead.
Pipeline parallelism
^^^^^^^^^^^^^^^^^^^^
@@ -344,7 +467,7 @@ Pipeline parallelism splits the model's layers across multiple GPUs or nodes, wi
--pipeline-parallel-size 2
.. note::
**ROCm best practice**: On Instinct MI300X/MI325X/MI350X/MI355X, prefer staying within a single XGMI island (≤8 GPUs) using TP only. Use PP when scaling beyond eight GPUs or across nodes.
**ROCm best practice**: On the Instinct MI300X, prefer staying within a single XGMI island (≤8 GPUs) using TP only. Use PP when scaling beyond eight GPUs or across nodes.
.. _data-parallelism-section:
@@ -413,21 +536,29 @@ For more technical details, see `vLLM Data Parallel Deployment <https://docs.vll
Data Parallel Attention (advanced)
""""""""""""""""""""""""""""""""""
For MLA models (DeepSeek V2/V3/R1, Kimi-K2.5), **DP+EP is the recommended configuration at high concurrency** (≥512 concurrent requests). Unlike traditional DP which replicates model weights, Data Parallel Attention uses inter-GPU AllToAll communication to partition KV cache across GPUs, avoiding the KV cache duplication that occurs with tensor parallelism.
For models with Multi-head Latent Attention (MLA) architecture like DeepSeek V2, V3, and R1, vLLM supports **Data Parallel Attention**,
which provides request-level parallelism instead of model replication. This avoids KV cache duplication across tensor parallel ranks,
significantly reducing memory usage and enabling larger batch sizes.
* At **≤128 concurrent requests**, TP=8 provides 4086% higher throughput
* At **≥512 concurrent requests**, DP=8+EP provides 1647% higher throughput
* Crossover typically occurs around **256512 concurrent requests**
**Key benefits for MLA models:**
* Eliminates KV cache duplication when using tensor parallelism
* Enables higher throughput for high-QPS serving scenarios
* Better memory efficiency for large context windows
**Usage with Expert Parallelism:**
Data parallel attention works seamlessly with Expert Parallelism for MoE models:
.. code-block:: bash
# DeepSeek-R1 with DP attention and expert parallelism (high concurrency)
# DeepSeek-R1 with DP attention and expert parallelism
VLLM_ALL2ALL_BACKEND="allgather_reducescatter" vllm serve deepseek-ai/DeepSeek-R1 \
--data-parallel-size 8 \
--enable-expert-parallel \
--disable-nccl-for-dp-synchronization
For more technical details, see `vLLM RFC #16037 <https://github.com/vllm-project/vllm/issues/16037>`_ and the `vLLM MoE Playbook <https://rocm.blogs.amd.com/software-tools-optimization/vllm-moe-guide/README.html>`_.
For more technical details, see `vLLM RFC #16037 <https://github.com/vllm-project/vllm/issues/16037>`_.
Expert parallelism
^^^^^^^^^^^^^^^^^^
@@ -435,47 +566,23 @@ Expert parallelism
Expert parallelism (EP) distributes expert layers of Mixture-of-Experts (MoE) models across multiple GPUs,
where tokens are routed to the GPUs holding the experts they need.
**Performance considerations:**
Expert parallelism is designed primarily for cross-node MoE deployments where high-bandwidth interconnects (like InfiniBand) between nodes make EP communication efficient. For single-node Instinct MI300X/MI355X deployments with XGMI connectivity, tensor parallelism typically provides better performance due to optimized all-to-all collectives on XGMI.
**When to use EP:**
.. list-table::
:header-rows: 1
:widths: 30 35 35
* Multi-node MoE deployments with fast inter-node networking
* Models with very large numbers of experts that benefit from expert distribution
* Workloads where EP's reduced data movement outweighs communication overhead
* - Scenario
- Recommended config
- Rationale
* - **Low concurrency** (≤128 requests)
- TP=8 (EP optional)
- 4086% higher throughput than DP at low concurrency.
* - **High concurrency** (≥512 requests)
- DP=8 + EP
- 1647% higher throughput at scale (for example, 7,114 TPS for DeepSeek-R1 at 1024 concurrent requests).
* - **MLA/MQA models** (DeepSeek-V2/V3/R1, Kimi-K2.5)
- DP + EP
- Avoids KV cache duplication across TP ranks. Mandatory for optimal memory at high concurrency.
* - **Ultra-sparse MoE** (<1% activation density, for example, Llama-4-Maverick)
- DP or TP **without** EP
- EP adds AllToAll overhead that exceeds the benefit — EP is 712% *slower* for these models.
* - **Standard MoE** (≥3% activation density, for example, DeepSeek-R1, Qwen3-235B)
- EP flag
- Improves expert routing efficiency.
**Single-node recommendation:** For Instinct MI300X/MI355X within a single node (≤8 GPUs), prefer tensor parallelism over expert parallelism for MoE models to leverage XGMI's high bandwidth and low latency.
**Basic usage:**
.. code-block:: bash
# DP + EP for MLA+MoE models (DeepSeek-R1, high concurrency)
VLLM_ALL2ALL_BACKEND="allgather_reducescatter" vllm serve deepseek-ai/DeepSeek-R1 \
--data-parallel-size 8 \
--enable-expert-parallel \
--disable-nccl-for-dp-synchronization
# TP + EP (low concurrency, non-MLA models)
# Enable expert parallelism for MoE models (DeepSeek example with 8 GPUs)
vllm serve deepseek-ai/DeepSeek-R1 \
--tensor-parallel-size 8 \
--enable-expert-parallel
@@ -487,37 +594,17 @@ When EP is enabled alongside tensor parallelism:
* Fused MoE layers use expert parallelism
* Non-fused MoE layers use tensor parallelism
Multimodal model optimization (vision-language)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
**Combining with Data Parallelism:**
For multimodal models (Qwen3-VL, InternVL, step3), use batch-level data parallelism
for the vision encoder instead of the default tensor parallelism:
EP works seamlessly with Data Parallel Attention for optimal memory efficiency in MLA+MoE models (for example, DeepSeek V3):
.. code-block:: bash
vllm serve Qwen/Qwen3-VL-235B-A22B-Instruct \
--tensor-parallel-size 8 \
--mm-encoder-tp-mode data \
# DP attention + EP for DeepSeek-R1
VLLM_ALL2ALL_BACKEND="allgather_reducescatter" vllm serve deepseek-ai/DeepSeek-R1 \
--data-parallel-size 8 \
--enable-expert-parallel \
--max-model-len 32768
``--mm-encoder-tp-mode data`` replaces per-layer all-reduce synchronization (58126 ops
in TP mode) with a single all-gather after encoding, yielding **1045% throughput
improvement** with negligible memory overhead (0.22.3% model size increase).
**When it helps most:**
* High-resolution images (1024×1024 px): **+16% average** throughput
* 13 images per request: **+1316%** throughput
* Deep vision encoders (for example, InternVL 45 blocks, step3 63 blocks)
**When to skip it:**
* Very small vision encoders (<1% of total model parameters)
* 10+ small images per request (diminishing returns)
* Memory-constrained deployments (encoder weights are replicated per GPU)
For more details, see the `vLLM Multimodal DP blog post <https://rocm.blogs.amd.com/software-tools-optimization/vllm-dp-vision/README.html>`_.
--disable-nccl-for-dp-synchronization
Throughput benchmarking
=======================
@@ -556,22 +643,22 @@ Maximizing instances per node
To maximize **per-node throughput**, run as many vLLM instances as model memory allows,
balancing KV-cache capacity.
* **HBM capacities**: MI300X = 192GB HBM3; MI325X = 256 GB HBM3E; MI350X/MI355X = 288GB HBM3E.
* **HBM capacities**: MI300X = 192GB HBM3; MI355X = 288GB HBM3E.
* Up to **eight** single-GPU vLLM instances can run in parallel on an 8×GPU node (one per GPU):
.. code-block:: bash
for i in $(seq 0 7); do
CUDA_VISIBLE_DEVICES="$i" vllm bench throughput
-tp 1 --model /path/to/model
CUDA_VISIBLE_DEVICES="$i" vllm bench throughput
-tp 1 --model /path/to/model
--dataset /path/to/ShareGPT_V3_unfiltered_cleaned_split.json &
done
Total throughput from **N** single-GPU instances usually exceeds one instance stretched across **N** GPUs (``-tp N``).
**Model coverage**: Llama 2 (7B/13B/70B), Llama 3 (8B/70B), Qwen2 (7B/72B), Mixtral-8x7B/8x22B, and others Llama270B
and Llama370B can fit a single MI300X/MI325X/MI350X/MI355X; Llama3.1405B fits on a single 8×MI300X/MI325X/MI350X/MI355X node.
and Llama370B can fit a single MI300X/MI355X; Llama3.1405B fits on a single 8×MI300X/MI355X node.
Configure the gpu-memory-utilization parameter
==================================================
@@ -681,15 +768,11 @@ CUDA graphs reduce kernel launch overhead by capturing and replaying GPU operati
* - Attention backend
- CUDA graph support
* - ``TRITON_ATTN``
* - vLLM/AITER Triton Unified Attention, vLLM Prefill-Decode Attention
- Full support (prefill + decode)
* - ``ROCM_ATTN``, ``ROCM_AITER_UNIFIED_ATTN``
- Full support (prefill + decode)
* - ``ROCM_AITER_FA``, ``ROCM_AITER_MLA``, ``ROCM_AITER_TRITON_MLA``
* - AITER MHA, AITER MLA
- Uniform batches only
* - ``ROCM_AITER_MLA_SPARSE``
- Uniform single-token decode only
* - ``TRITON_MLA``
* - vLLM Triton MLA
- Must exclude attention from graph — ``PIECEWISE`` required
**Usage examples:**
@@ -716,20 +799,20 @@ CUDA graphs reduce kernel launch overhead by capturing and replaying GPU operati
Quantization support
====================
vLLM supports FP4/FP8 (4-bit/8-bit floating point) weight and activation quantization using hardware acceleration on the Instinct MI300X, MI325X, MI350X, and MI355X.
Quantization of models with FP4/FP8 allows for a **2x-4x** reduction in model memory requirements and up to a **1.6x**
improvement in throughput with minimal impact on accuracy.
vLLM supports FP4/FP8 (4-bit/8-bit floating point) weight and activation quantization using hardware acceleration on the Instinct MI300X and MI355X.
Quantization of models with FP4/FP8 allows for a **2x-4x** reduction in model memory requirements and up to a **1.6x**
improvement in throughput with minimal impact on accuracy.
vLLM ROCm supports a variety of quantization demands:
vLLM ROCm supports a variety of quantization demands:
* On-the-fly quantization
* On-the-fly quantization
* Pre-quantized model through Quark and llm-compressor
* Pre-quantized model through Quark and llm-compressor
Supported quantization methods
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
vLLM on ROCm supports the following quantization methods for the AMD Instinct MI300 series and Instinct MI350 series GPUs:
vLLM on ROCm supports the following quantization methods for the AMD Instinct MI300 series and Instinct MI355X GPUs:
.. list-table::
:header-rows: 1
@@ -851,23 +934,40 @@ For models without pre-quantization, vLLM can quantize ``FP16``/``BF16`` models
GPTQ
^^^^
GPTQ (4-bit/8-bit weight quantization) is fully supported on ROCm via HIP-compiled kernels.
Pre-quantized GPTQ models from Hugging Face work out of the box. For better throughput on AMD Instinct GPUs,
consider **AWQ with Triton kernels** or **FP8 quantization** instead.
GPTQ is a 4-bit/8-bit weight quantization method that compresses models with minimal accuracy loss. GPTQ
is fully supported on ROCm via HIP-compiled kernels in vLLM.
**ROCm support status**:
- **Fully supported** - GPTQ kernels compile and run on ROCm via HIP
- **Pre-quantized models work** with standard GPTQ kernels
**Recommendation**: For the AMD Instinct MI300X, **AWQ with Triton kernels** or **FP8 quantization** might provide better
performance due to ROCm-specific optimizations, but GPTQ is a viable alternative.
**Using pre-quantized GPTQ models**:
.. code-block:: bash
# Using pre-quantized GPTQ model on ROCm
vllm serve RedHatAI/Meta-Llama-3.1-70B-Instruct-quantized.w4a16 \
--quantization gptq \
--dtype auto \
--tensor-parallel-size 1
**Important notes**:
- **Kernel support:** GPTQ uses standard HIP-compiled kernels on ROCm
- **Performance:** AWQ with Triton kernels might offer better throughput on AMD GPUs due to ROCm optimizations
- **Compatibility:** GPTQ models from Hugging Face work on ROCm with standard performance
- **Use case:** GPTQ is suitable when pre-quantized GPTQ models are readily available
AWQ (Activation-aware Weight Quantization)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
AWQ (Activation-aware Weight Quantization) is a 4-bit weight quantization technique that provides excellent
model compression with minimal accuracy loss (<1%). ROCm supports AWQ quantization on the AMD Instinct MI300 series and
MI350 series GPUs with vLLM.
MI355X GPUs with vLLM.
**Using pre-quantized AWQ models:**
@@ -880,7 +980,7 @@ Many AWQ-quantized models are available on Hugging Face. Use them directly with
vllm serve hugging-quants/Meta-Llama-3.1-70B-Instruct-AWQ-INT4 \
--quantization awq \
--tensor-parallel-size 1 \
--dtype auto
--dtype auto
**Important Notes:**
@@ -998,7 +1098,7 @@ Speculative decoding (experimental)
===================================
Recent vLLM versions add support for speculative decoding backends (for example, Eaglev3). Evaluate for your model and latency/throughput goals.
Speculative decoding is a technique to reduce latency when max number of concurrency is low.
Speculative decoding is a technique to reduce latency when max number of concurrency is low.
Depending on the methods, the effective concurrency varies, for example, from 16 to 64.
Example command:
@@ -1026,7 +1126,7 @@ Example command:
It has been observed that more ``num_speculative_tokens`` causes less
acceptance rate of draft model tokens and a decline in throughput. As a
workaround, set ``num_speculative_tokens`` to <= 2.
workaround, set ``num_speculative_tokens`` to <= 2.
Multi-node checklist and troubleshooting
@@ -1037,16 +1137,8 @@ Multi-node checklist and troubleshooting
3. For GPUDirect RDMA, set ``RCCL_NET_GDR_LEVEL=2`` and verify links (``ibstat``). Requires supported NICs (for example, ConnectX6+).
4. Collect RCCL logs: ``RCCL_DEBUG=INFO`` and optionally ``RCCL_DEBUG_SUBSYS=INIT,GRAPH`` for init/graph stalls.
Deprecated terms
================
* **Prefill-Decode attention** has been renamed to **ROCM_ATTN** (ROCm attention). Use ``--attention-backend ROCM_ATTN`` to select this backend.
Further reading
===============
* :doc:`workload`
* :doc:`/how-to/rocm-for-ai/inference/benchmark-docker/vllm`
* `ROCm Attention Backend deep-dive <https://vllm.ai/blog/rocm-attention-backend>`_ — architecture and benchmarks for all 7 backends
* `vLLM MoE Playbook - A Practical Guide to TP, DP, PP and Expert Parallelism <https://rocm.blogs.amd.com/software-tools-optimization/vllm-moe-guide/README.html>`_ — DP+EP tuning for MoE models
* `Multimodal DP optimization <https://rocm.blogs.amd.com/software-tools-optimization/vllm-dp-vision/README.html>`_ — batch-level DP for vision encoders
* :doc:`workload-optimization`
* :doc:`/ai-inference/vllm`
@@ -7,7 +7,7 @@ docker:
- "Support fp8 MLA for MI355"
- "Block wise sparsity support for AMD triton FAv3 Sage attention"
components:
TheRock:
TheRock:
version: cbff3d1
url: https://github.com/ROCm/TheRock
rocm-libraries:
Binary file not shown.

After

Width:  |  Height:  |  Size: 167 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 47 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 778 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 39 KiB

Some files were not shown because too many files have changed in this diff Show More