docs: add mi350 workload tuning specific and consolidate with mi300x (#6318)

* add mi350 workload tuning specific and consolidate with mi300x

* update content and keep old content as much as possible

* Remove TunableOp + max-autotune note.

* Update Inductor compiler url.

* Mention hipBLASLt.

* Update Inductor tuning knobs.

* Lint

* More changes to Inductor section.

* add gluon perf section

* update and cleanup

* address feedback

* add words to `.wordlist.txt` to satisfy linter

* Update docs/how-to/rocm-for-ai/inference-optimization/workload.rst

Co-authored-by: peterjunpark <peter.park@amd.com>

* Update docs/how-to/rocm-for-ai/inference-optimization/workload.rst

Co-authored-by: peterjunpark <peter.park@amd.com>

* Update docs/how-to/rocm-for-ai/inference-optimization/workload.rst

Co-authored-by: peterjunpark <peter.park@amd.com>

* address feedback and add gluon tutorial public link

---------

Co-authored-by: Hongxia Yang <hongxiay.yang@amd.com>
Co-authored-by: Nichols A. Romero <nick.romero@amd.com>
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com>
Co-authored-by: Hongxia Yang <62075498+hongxiayang@users.noreply.github.com>
This commit is contained in:
peterjunpark
2026-06-01 15:36:10 -04:00
committed by GitHub
co-authored by Hongxia Yang Nichols A. Romero Hongxia Yang Hongxia Yang
parent 60761656b4
commit 54788f54d2
2 changed files with 673 additions and 120 deletions
+17
View File
@@ -7,6 +7,10 @@ AITER
ALU
AllReduce
AllToAll
AGPR
AGPRs
AITER
ALU
AMD
AMDGPU
AMDGPUs
@@ -146,6 +150,7 @@ FHS
FIFOs
FIXME
FMA
FNUZ
FP
FX
FiLM
@@ -182,6 +187,8 @@ GIM
GL
Glibc
GLM
GIM
GL
GLXT
GMI
GNN
@@ -206,6 +213,7 @@ GitHub
Gitpod
Glibc
Gloo
Gluon
GraphBolt
GraphSage
HBM
@@ -426,6 +434,7 @@ Pensando
PerfDb
Perfetto
PipelineParallel
Pipelining
PnP
Pollara
PowerEdge
@@ -636,6 +645,7 @@ allocator
allocators
amdgpu
api
async
aten
atmi
atomicRMW
@@ -643,6 +653,7 @@ atomics
autogenerated
autograd
autotune
autotuning
avx
awk
az
@@ -660,6 +671,7 @@ blit
bootloader
boson
bosons
bottlenecked
br
btn
buildable
@@ -771,6 +783,7 @@ ffmpeg
filesystem
flashinfer
forEach
foreach
fortran
fp
framebuffer
@@ -924,6 +937,8 @@ perfcounter
performant
perl
piecewise
pipelined
pipelining
pragma
pre
prebuild
@@ -1096,9 +1111,11 @@ unfused
unhandled
uninstallation
unmapped
unpadded
unsqueeze
unstacking
unswitching
unswizzled
untrusted
untuned
unwindowed
File diff suppressed because it is too large Load Diff