Compare commits

...
7 Commits
Author SHA1 Message Date
Hanzo Dev 001fa7c0ce chore: OSS attribution — AMD ROCm (MIT) LICENSE + NOTICE 2026-07-02 12:44:34 -07:00
z 582cacea47 docs(brand): add hero banner 2026-06-28 20:15:26 -07:00
z 5209dccdb2 chore(brand): dynamic hero banner 2026-06-28 20:15:25 -07:00
peterjunparkandGitHub 876343e53b Update vLLM and SGLang + MoRI distributed inference docs (#6329) 2026-06-03 16:23:58 -04:00
peterjunparkandGitHub 9f0ec581e2 Remove extra CODEOWNERS (#6310) 2026-06-02 20:21:17 -04:00
peterjunparkandGitHub 6a6dc9136f docs: rm sglang DI w/ mooncake doc (#6320) 2026-06-01 22:09:40 -04:00
54788f54d2 docs: add mi350 workload tuning specific and consolidate with mi300x (#6318)
* add mi350 workload tuning specific and consolidate with mi300x

* update content and keep old content as much as possible

* Remove TunableOp + max-autotune note.

* Update Inductor compiler url.

* Mention hipBLASLt.

* Update Inductor tuning knobs.

* Lint

* More changes to Inductor section.

* add gluon perf section

* update and cleanup

* address feedback

* add words to `.wordlist.txt` to satisfy linter

* Update docs/how-to/rocm-for-ai/inference-optimization/workload.rst

Co-authored-by: peterjunpark <peter.park@amd.com>

* Update docs/how-to/rocm-for-ai/inference-optimization/workload.rst

Co-authored-by: peterjunpark <peter.park@amd.com>

* Update docs/how-to/rocm-for-ai/inference-optimization/workload.rst

Co-authored-by: peterjunpark <peter.park@amd.com>

* address feedback and add gluon tutorial public link

---------

Co-authored-by: Hongxia Yang <hongxiay.yang@amd.com>
Co-authored-by: Nichols A. Romero <nick.romero@amd.com>
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com>
Co-authored-by: Hongxia Yang <62075498+hongxiayang@users.noreply.github.com>
2026-06-01 15:36:10 -04:00
11 changed files with 1486 additions and 809 deletions
+4 -4
View File
@@ -1,8 +1,8 @@
* @amd-aakash @jlgreathouse @samjwu @yhuiYH @ROCm/rocm-documentation
* @ROCm/rocm-documentation
# Documentation files
docs/ @amd-aakash @jlgreathouse @samjwu @yhuiYH @ROCm/rocm-documentation
*.md @amd-aakash @jlgreathouse @samjwu @yhuiYH @ROCm/rocm-documentation
*.rst @amd-aakash @jlgreathouse @samjwu @yhuiYH @ROCm/rocm-documentation
docs/ @ROCm/rocm-documentation
*.md @ROCm/rocm-documentation
*.rst @ROCm/rocm-documentation
# External CI
/.azuredevops/ @ROCm/external-ci
tools/rocm-build/ @ROCm/rocm-devops
+9
View File
@@ -0,0 +1,9 @@
<svg xmlns="http://www.w3.org/2000/svg" width="1280" height="640" viewBox="0 0 1280 640" role="img" aria-label="ROCm">
<rect width="1280" height="640" fill="#0A0A0A"/>
<svg x="96" y="215" width="210" height="210" viewBox="0 0 67 67"><path d="M22.21 67V44.6369H0V67H22.21Z" fill="#fff"/><path d="M66.7038 22.3184H22.2534L0.0878906 44.6367H44.4634L66.7038 22.3184Z" fill="#fff"/><path d="M22.21 0H0V22.3184H22.21V0Z" fill="#fff"/><path d="M66.7198 0H44.5098V22.3184H66.7198V0Z" fill="#fff"/><path d="M66.7198 67V44.6369H44.5098V67H66.7198Z" fill="#fff"/></svg>
<text x="378" y="276" font-family="Inter,system-ui,-apple-system,sans-serif" font-size="78" font-weight="800" letter-spacing="-2" fill="#ffffff">ROCm</text>
<text x="378" y="322" font-family="Inter,system-ui,sans-serif" font-size="30" fill="#ffffff" opacity=".66">AMD ROCm™ Software - GitHub Home</text>
<rect x="378" y="338" width="806" height="3" rx="1.5" fill="#ffffff" opacity=".9"/>
<text x="378" y="390" font-family="Inter,system-ui,sans-serif" font-size="24" font-weight="600" fill="#ffffff" opacity=".5">github.com/hanzoai</text>
<text x="1184" y="390" text-anchor="end" font-family="Inter,system-ui,sans-serif" font-size="24" font-weight="600" fill="#ffffff" opacity=".5">hanzo.ai</text>
</svg>

After

Width:  |  Height:  |  Size: 1.2 KiB

+19
View File
@@ -7,6 +7,10 @@ AITER
ALU
AllReduce
AllToAll
AGPR
AGPRs
AITER
ALU
AMD
AMDGPU
AMDGPUs
@@ -40,6 +44,7 @@ BARs
BKC
BLAS
BMC
BNXT
BabelStream
Backported
BatchNorm
@@ -146,6 +151,7 @@ FHS
FIFOs
FIXME
FMA
FNUZ
FP
FX
FiLM
@@ -182,6 +188,8 @@ GIM
GL
Glibc
GLM
GIM
GL
GLXT
GMI
GNN
@@ -206,6 +214,7 @@ GitHub
Gitpod
Glibc
Gloo
Gluon
GraphBolt
GraphSage
HBM
@@ -426,6 +435,7 @@ Pensando
PerfDb
Perfetto
PipelineParallel
Pipelining
PnP
Pollara
PowerEdge
@@ -636,6 +646,7 @@ allocator
allocators
amdgpu
api
async
aten
atmi
atomicRMW
@@ -643,6 +654,7 @@ atomics
autogenerated
autograd
autotune
autotuning
avx
awk
az
@@ -660,6 +672,7 @@ blit
bootloader
boson
bosons
bottlenecked
br
btn
buildable
@@ -771,6 +784,7 @@ ffmpeg
filesystem
flashinfer
forEach
foreach
fortran
fp
framebuffer
@@ -924,6 +938,8 @@ perfcounter
performant
perl
piecewise
pipelined
pipelining
pragma
pre
prebuild
@@ -1055,6 +1071,7 @@ submatrix
submodule
submodules
subnet
subnets
supercomputing
symlink
symlinks
@@ -1096,9 +1113,11 @@ unfused
unhandled
uninstallation
unmapped
unpadded
unsqueeze
unstacking
unswitching
unswizzled
untrusted
untuned
unwindowed
+18
View File
@@ -0,0 +1,18 @@
Hanzo ROCm
Copyright (c) 2026 Hanzo AI, Inc.
This product includes software from AMD ROCm (https://github.com/ROCm/ROCm),
licensed under the MIT License:
Copyright (c) 2023 - 2025 Advanced Micro Devices, Inc. All rights reserved.
This repository is the AMD ROCm meta/manifest repository: documentation, build
tooling, and the repo manifest (default.xml). This meta-repository is MIT-licensed
and its upstream MIT license is retained in LICENSE.
The individual ROCm components referenced by the manifest and built by the tooling
carry their own upstream licenses, including the MIT License, the Apache License 2.0,
and the University of Illinois/NCSA Open Source License. Note that the ROCgdb
component (AMD's fork of GNU GDB, packaged by tools/rocm-build/build_rocm-gdb.sh) is
licensed under the GNU General Public License (GPL) — a copyleft license. Each
component is governed by its own license; consult that component's repository.
+2
View File
@@ -1,3 +1,5 @@
<p align="center"><img src=".github/hero.svg" alt="ROCm" width="880"></p>
<div align="center">
<img src="docs/data/amd-rocm-logo.png" width="200px" alt="ROCm logo">
File diff suppressed because it is too large Load Diff
@@ -1,257 +0,0 @@
.. meta::
:description: SGLang multi-node disaggregated distributed inference using Mooncake
:keywords: model, sglang, mooncake, disagg, disaggregated, distributed, multi-node, docker
******************************************
SGLang distributed inference with Mooncake
******************************************
As LLM inference increasingly demands handling massive models and dynamic workloads, efficient
distributed inference becomes essential. Traditional co-located architectures face bottlenecks due
to tightly coupled memory and compute resources, which limits scalability and flexibility.
Disaggregated inference refers to the process of splitting the inference of LLMs into distinct
phases. This architecture, facilitated by libraries like Mooncake, uses high-bandwidth
RDMA to transfer the Key-Value (KV) cache between prefill and decode nodes.
This allows for independent resource scaling and optimization, resulting in
improved efficiency and throughput.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/sglang-distributed-benchmark-models.yaml
{% set docker = data.dockers[0] %}
`SGLang <https://docs.sglang.ai>`__ is a high-performance inference and
serving engine for large language models (LLMs) and vision models. The
ROCm-enabled `SGLang base Docker image <{{ docker.docker_hub_url }}>`__
bundles SGLang with PyTorch, which is optimized for AMD Instinct MI300X Series
GPUs. It includes the following software components:
.. list-table::
:header-rows: 1
* - Software component
- Version
{% for component_name, component_version in docker.components.items() %}
* - {{ component_name }}
- {{ component_version }}
{% endfor %}
The following guides on setting up and running SGLang and Mooncake for disaggregated
distributed inference on a Slurm cluster using AMD Instinct MI300X Series GPUs backed by
Mellanox CX-7 NICs.
Prerequisites
=============
Before starting, ensure you have:
* A Slurm cluster with at least three nodes: one for the proxy, one for prefill (``xP``), and one for decode (``yD``).
``Nodes -> xP + yD + 1``
* A Dockerized environment with SGLang, Mooncake, etcd, and NIC drivers built in. See :ref:`sglang-disagg-inf-build-docker-image` for instructions.
* A shared filesystem for storing models, scripts, and logs (cluster-specific).
Supported models
================
The following models are supported for SGLang disaggregated prefill/decode
inference. Some instructions, commands, and recommendations in this
documentation might vary by selected model.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/sglang-distributed-benchmark-models.yaml
{% set model_groups = data.model_groups %}
.. raw:: html
<div id="vllm-benchmark-ud-params-picker" class="container-fluid">
<div class="row gx-0">
<div class="col-2 me-1 px-2 model-param-head">Model type</div>
<div class="row col-10 pe-0">
{% for model_group in model_groups %}
<div class="col-6 px-2 model-param" data-param-k="model-group" data-param-v="{{ model_group.tag }}" tabindex="0">{{ model_group.group }}</div>
{% endfor %}
</div>
</div>
<div class="row gx-0 pt-1">
<div class="col-2 me-1 px-2 model-param-head">Model</div>
<div class="row col-10 pe-0">
{% for model_group in model_groups %}
{% set models = model_group.models %}
{% for model in models %}
{% if models|length % 3 == 0 %}
<div class="col-4 px-2 model-param" data-param-k="model" data-param-v="{{ model.model_repo | lower }}" data-param-group="{{ model_group.tag }}" tabindex="0">{{ model.model }}</div>
{% else %}
<div class="col-6 px-2 model-param" data-param-k="model" data-param-v="{{ model.model_repo | lower }}" data-param-group="{{ model_group.tag }}" tabindex="0">{{ model.model }}</div>
{% endif %}
{% endfor %}
{% endfor %}
</div>
</div>
</div>
{% for model_group in model_groups %}
{% for model in model_group.models %}
.. container:: model-doc {{ model.model_repo }}
.. note::
See the `{{ model.model }} model card on Hugging Face <{{ model.url }}>`__ to learn more about this model.
Some models require access authorization prior to use through an external license agreement with a third party.
{% endfor %}
{% endfor %}
.. _sglang-disagg-inf-build-docker-image:
Build the Docker image
----------------------
Get the Dockerfile located in
`<https://github.com/ROCm/MAD/blob/develop/docker/sglang_disagg_inference.ubuntu.amd.Dockerfile>`__.
It uses `lmsysorg/sglang:v0.5.2rc1-rocm700-mi30x
<https://hub.docker.com/layers/lmsysorg/sglang/v0.4.9.post1-rocm630/images/sha256-2f6b1748e4bcc70717875a7da76c87795fd8aa46a9646e08d38aa7232fc78538>`__
as the base Docker image and installs the necessary components for Mooncake, etcd, and Mellanox network
drivers.
.. code-block:: shell
git clone https://github.com/ROCm/MAD.git
cd MAD/docker
docker build \
-t sglang_disagg_pd_image \
-f sglang_disagg_inference.ubuntu.amd.Dockerfile .
Benchmarking
============
The `<https://github.com/ROCm/MAD/tree/develop/scripts/sglang_disagg>`__
repository contains scripts to launch SGLang inference with prefill/decode
disaggregation via Mooncake for supported models.
* `scripts/sglang_dissag/run_xPyD_models.slurm <https://github.com/ROCm/MAD/blob/develop/scripts/sglang_disagg/run_xPyD_models.slurm>`__
-- the main Slurm batch script to launch Docker containers on all nodes using ``sbatch`` or ``salloc``.
* `scripts/sglang_dissag/sglang_disagg_server.sh <https://github.com/ROCm/MAD/blob/develop/scripts/sglang_disagg/sglang_disagg_server.sh>`__
-- the entrypoint script that runs inside each container to start the correct service -- proxy, prefill, or decode.
* `scripts/sglang_dissag/benchmark_xPyD.sh <https://github.com/ROCm/MAD/blob/develop/scripts/sglang_disagg/benchmark_xPyD.sh>`__
-- the benchmark script to run the GSM8K accuracy benchmark and the SGLang benchmarking tool for performance measurement.
* `scripts/sglang_dissag/benchmark_parser.py <https://github.com/ROCm/MAD/blob/develop/scripts/sglang_disagg/benchmark_parser.py>`__
-- the log parser script to be run on the concurrency benchmark log file to generate tabulated data.
Launch the service
------------------
The service is deployed using a Slurm batch script that orchestrates the containers across the
allocated nodes.
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/sglang-distributed-benchmark-models.yaml
{% set model_groups = data.model_groups %}
{% for model_group in model_groups %}
{% for model in model_group.models %}
.. container:: model-doc {{ model.model_repo }}
.. code-block:: shell
# Clone the MAD repo if you haven't already and
# navigate to the scripts directory
git clone https://github.com/ROCm/MAD.git
cd MAD/scripts/sglang_disagg/
# Slurm sbatch run command
export DOCKER_IMAGE_NAME=sglang_disagg_pd_image
export xP=<num_prefill_nodes>
export yD=<num_decode_nodes>
export MODEL_NAME={{ model.model_repo }}
# num_nodes = xP + yD + 1
sbatch -N <num_nodes> -n <num_nodes> --nodelist=<Nodes> run_xPyD_models.slurm
{% endfor %}
{% endfor %}
Post-run logs and testing
-------------------------
Logs are stored in your shared filesystem in the directory specified by the ``LOG_PATH`` variable in the Slurm script.
A new directory named after the Slurm job ID is created for each run.
Inside that directory, you can access various logs:
* ``pd_sglang_bench_serving.sh_NODE<...>.log`` -- the main log for each server node.
* ``etcd_NODE<...>.log`` -- logs for etcd services.
* ``prefill_NODE<...>.log`` -- logs for the prefill services.
* ``decode_NODE<...>.log`` -- logs for the decode services.
Use the benchmark parser script for concurrency logs to tabulate different data.
.. code-block:: shell
python3 benchmark_parser.py <log_path/benchmark_XXX_CONCURRENCY.log>
To verify the service is responsive, you can try sending a ``curl`` request to test the launched
server from the Docker container on the proxy node. For example:
.. code-block:: shell
curl -X POST http://127.0.0.1:30000/generate \
-H "Content-Type: application/json" \
-d '{ "text": "Let me tell you a story ", "sampling_params": { "temperature": 0.3 } }'
Known issues
============
When running larger models, such as DeepSeek-V3 and Llama-3.1-405B-Instruct-FP8-KV, at
higher concurrency levels (512+), the following error might occur:
.. code-block:: shell-session
<TransferEncodingError: 400, message:
Not enough data to satisfy transfer length header.
The above exception was the direct cause of the following exception:
Traceback (most recent call last):
...
This leads to dropping requests and lower throughput.
Further reading
===============
- To learn about Mooncake, see `Welcome to Mooncake <https://kvcache-ai.github.io/Mooncake/>`__.
- To learn more about the options for latency and throughput benchmark scripts,
see `<https://github.com/sgl-project/sglang/tree/main/benchmark/blog_v0_2>`__.
- See the base upstream Docker image on `Docker Hub <https://hub.docker.com/layers/lmsysorg/sglang/v0.5.2rc1-rocm700-mi30x/images/sha256-10c4ee502ddba44dd8c13325e6e03868bfe7f43d23d0a44780a8ee8b393f4729>`__.
- To learn more about system settings and management practices to configure your system for
MI300X Series GPUs, see `AMD Instinct MI300X system optimization <https://instinct.docs.amd.com/projects/amdgpu-docs/en/latest/system-optimization/mi300x.html>`__.
- For application performance optimization strategies for HPC and AI workloads,
including inference with vLLM, see :doc:`/how-to/rocm-for-ai/inference-optimization/workload`.
- To learn how to run community models from Hugging Face on AMD GPUs, see
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
- To learn how to fine-tune LLMs and optimize inference, see
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
- For a list of other ready-made Docker images for AI with ROCm, see
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
Previous versions
=================
See :doc:`previous-versions/sglang-history` to find documentation for previous releases
of SGLang inference performance testing.
@@ -12,14 +12,15 @@ scripts.
The following configuration is required to implement this setup:
* **Nodes:** A minimum of three GPU nodes (Virtual machines or Physical
machines) for wide expert parallelism (EP) evaluation.
* **GPUs** 8x AMD Instinct MI355X GPU cards per node.
* **Networking:** 8x AMD Pensando Pollara 400 AI NICs per node, providing
a dedicated 1:1 mapping between GPUs and network interfaces for optimal
inter-node communication.
* **Orchestration:** A Slurm cluster with at least three nodes -- one for
prefill service and two for decode services (EP16)
* **Nodes**: A minimum of three GPU nodes (virtual machines or physical
machines) for wide expert parallelism (EP) evaluation.
* **GPUs**: 8x AMD Instinct MI355X GPU cards per node.
* **Networking**: 8x RDMA-capable NICs per node (AMD Pensando Pollara 400,
NVIDIA Mellanox ConnectX-7, or Broadcom Thor 2), providing a dedicated 1:1
mapping between GPUs and network interfaces for optimal inter-node
communication.
* **Orchestration**: A Slurm cluster with at least three nodes — one for
prefill service and two for decode services (EP16).
## System configuration
@@ -29,8 +30,6 @@ baselines and firmware versions, configuring the AMD Pensando Pollara 400 AI
NICs for high-bandwidth networking, and applying thermal and Quality of Service
(QoS) tunings to ensure a stable, lossless RDMA fabric.
(sglang-mori-verify-baseline)=
### Verify baseline software
The following table outlines the validated software stack. Use the provided
@@ -81,18 +80,19 @@ Redfish API:
Before proceeding with software deployment, verify that all cluster nodes
comply with the [MI355X Basic Health
Checks](https://instinct.docs.amd.com/projects/system-acceptance/en/latest/gpus/mi355x.html#basic-health-checks)
Checks](https://instinct.docs.amd.com/projects/system-acceptance/en/latest/gpus/mi355x.html#basic-health-checks).
Key requirements include specific kernel boot arguments, minimum system memory
thresholds, PCIe Gen5 link stability, and so on.
### Install AMD Pensando Pollara 400 AI NIC drivers
### NIC installation
#### AMD Pensando Pollara 400 AI NIC installation
For detailed instructions on upgrading the firmware and installing drivers for
the AMD Pensando Pollara 400 AI NIC, refer to the [AMD Instinct System
Acceptance
Guide](https://instinct.docs.amd.com/projects/system-acceptance/en/latest/network/nic-installation.html#amd-pensando-pollara-400-ai-nic).
Acceptance Guide](https://instinct.docs.amd.com/projects/system-acceptance/en/latest/network/nic-installation.html#amd-pensando-pollara-400-ai-nic).
After installation, verify the active firmware version on all NICs to ensure it
matches the software baseline. See [Verify baseline software](#verify-best-known-configuration-bkc).
matches the software baseline. See [Verify baseline software](#verify-baseline-software).
To display the current firmware version for all AI NICs, use the following command.
@@ -100,6 +100,26 @@ To display the current firmware version for all AI NICs, use the following comma
sudo nicctl show version firmware
```
#### CX7 driver and firmware installation
1. Download and install the `DOCA 2.9.3` driver following the instructions in
[NVIDIA DOCA 2.9.3 Downloads](https://developer.nvidia.com/doca-downloads).
2. Download the appropriate firmware for your hardware PSID from the
[ConnectX-7 Firmware Download
Center](https://network.nvidia.com/support/firmware/connectx7/) and flash
the device.
3. To verify driver and firmware versions, use the following command. Replace
`IB Device` with your specific backend interface.
```bash
ethtool -i <IB Device>
```
#### Broadcom BNXT driver and firmware installation
Refer to your Broadcom representative for driver and firmware installation
instructions specific to your NIC model.
### Configure thermal management (fan speed)
For systems equipped with 400G optics, standard fan profiles are often
@@ -140,8 +160,10 @@ the addresses `192.168.1.36`, `192.168.2.36`, and so on. Another node would
have `192.168.1.37`, `192.168.2.37`, and so on. Ensure MTU is set to `9000`.
```{note}
Ensure you identify the correct interface names for your system using ip link
before applying this configuration.
Ensure you identify the correct interface names for your system using `ip link`
before applying this configuration. The `macaddress:` values in the example
below are illustrative only and must be replaced with the actual MAC addresses
of your NICs, which you can find using `ip link show <interface>`.
```
For example, your `/etc/netplan/70-backend.yaml` should look like the
@@ -154,7 +176,7 @@ network:
addresses:
- 192.168.8.38/31
match:
macaddress: 04:90:81:2a:34:08
macaddress: 04:90:81:00:00:08
mtu: 9000
routes:
- table: 108
@@ -168,7 +190,7 @@ network:
addresses:
- 192.168.7.38/31
match:
macaddress: 04:90:81:2b:82:40
macaddress: 04:90:81:00:00:07
mtu: 9000
routes:
- table: 107
@@ -182,7 +204,7 @@ network:
addresses:
- 192.168.6.38/31
match:
macaddress: 04:90:81:30:c9:30
macaddress: 04:90:81:00:00:06
mtu: 9000
routes:
- table: 106
@@ -196,7 +218,7 @@ network:
addresses:
- 192.168.5.38/31
match:
macaddress: 04:90:81:2a:23:40
macaddress: 04:90:81:00:00:05
mtu: 9000
routes:
- table: 105
@@ -210,7 +232,7 @@ network:
addresses:
- 192.168.4.38/31
match:
macaddress: 04:90:81:2d:69:60
macaddress: 04:90:81:00:00:04
mtu: 9000
routes:
- table: 104
@@ -224,7 +246,7 @@ network:
addresses:
- 192.168.3.38/31
match:
macaddress: 04:90:81:2a:2c:40
macaddress: 04:90:81:00:00:03
mtu: 9000
routes:
- table: 103
@@ -238,7 +260,7 @@ network:
addresses:
- 192.168.2.38/31
match:
macaddress: 04:90:81:30:d5:30
macaddress: 04:90:81:00:00:02
mtu: 9000
routes:
- table: 102
@@ -252,7 +274,7 @@ network:
addresses:
- 192.168.1.38/31
match:
macaddress: 04:90:81:30:e4:00
macaddress: 04:90:81:00:00:01
mtu: 9000
routes:
- table: 101
@@ -276,17 +298,18 @@ To verify your configuration, use the following command.
sudo apt install -y net-tools && ip -br a
```
### Configure Quality of Service (QoS) and Congestion Control (DCQCN)
### Configure quality of service (QoS) and congestion control (DCQCN)
To ensure lossless communication and optimal performance for RDMA traffic, the
network must be configured with specific QoS and Data Center Quantized
Congestion Notification (DCQCN) settings.
The following configuration achieves:
• It enables RX and TX Pause frames on the ports
• Maps DSCP 24 (Data) to Q3 and DSCP 46 (CNP) to Q6, all other DSCP to Q0
• Enables PFC for Q3
• Scheduling : 99% to Q3, 1% to Q0 and strict priority for Q6
The following configuration:
* Enables RX and TX pause frames on the ports.
* Maps DSCP 24 (Data) to Q3 and DSCP 46 (CNP) to Q6, with all other DSCP to Q0.
* Enables PFC for Q3.
* Scheduling: 99% to Q3, 1% to Q0, and strict priority for Q6.
#### Configure DCQCN
@@ -294,7 +317,7 @@ Create and run a `/nfsdata/enable_dcqcn.sh` script to initialize congestion
control parameters.
``` bash
# !/bin/bash
#!/bin/bash
TOKEN_BUCKET_SIZE=800000
AI_RATE=160
@@ -361,7 +384,7 @@ sudo nicctl update qos pfc --priority $data_prio --no-drop enable
sudo nicctl update qos scheduling --priority $data_prio,$default_prio,$cts_prio --dwrr 99,1,0 --rate-limit 0,0,10
```
#### Verification your configuration
#### Verify your configuration
Verify the configuration using `nicctl`.
@@ -374,9 +397,9 @@ Verify the configuration using `nicctl`.
Expected QoS output:
``` bash
NIC : 42424650-4c32-3531-3230-303443000000 (0000:f6:00.0)
NIC : 00000000-0000-0000-0000-000000000001 (0000:f6:00.0)
Port : 04908130-a7a0-4242-4242-000011010000
Port : 00000000-0001-4242-4242-000000000000
Classification type : DSCP
@@ -398,10 +421,10 @@ Verify the configuration using `nicctl`.
Expected DCQCN and scheduling output:
``` bash
NIC : 42424650-4c32-3531-3230-303443000000 (0000:f6:00.0)
NIC : 00000000-0000-0000-0000-000000000001 (0000:f6:00.0)
------------------------------------------------------------------------------------------
Lif id : 43000070-0100-0000-4242-04908130a7a0
Lif id : 00000000-0100-0000-4242-000000000000
ROCE device : ionic_7
DCQCN profile id : 1
Status : Enabled
@@ -497,9 +520,10 @@ the cluster interconnects.
### Verify network connectivity
Verify that all network interfaces are reachable across the cluster nodes.
Assuming `eth0` is the management interface, and `benic1p1` through `benic8p1` are the
dedicated RoCE backend interfaces, use the following loop to test reachability
to a remote node (for instance, a target node with host IP suffix `.38`).
Assuming `benic1p1` through `benic8p1` are the dedicated RoCE backend
interfaces, use the following ping loop to verify reachability across the
backend subnets (for instance, a target node at host IP suffix
`.38`).
```bash
# Test connectivity for RoCE subnets 192.168.x.38 (node B) through 192.168.x.37 (node A)
@@ -522,9 +546,9 @@ The output should look something like this:
```bash
-------------------------------------------------------------------------------------
NIC : 42424650-4c32-3531-3530-314343000000 (0000:f6:00.0)
NIC : 00000000-0000-0000-0000-000000000002 (0000:f6:00.0)
Port : 04908132-5d88-4242-4242-000011010000 (eth1/1)
Port : 00000000-0002-4242-4242-000000000000 (eth1/1)
Spec:
Ifindex : 0x11010000
Type : ETH
@@ -548,7 +572,7 @@ Port : 04908132-5d88-4242-4242-000011010000 (eth1/1)
Auto negotiation : disabled
MAC ID : 0
MAC channel : 0
MAC address : 04:90:81:32:5d:88
MAC address : 04:90:81:00:00:00
Transceiver type : QSFP_CMIS
Transceiver state : SPROM-READ
Transceiver PID : QSFP-400G-DR4
@@ -569,7 +593,7 @@ ibv_devinfo -v | grep GID
The output should look something like this:
```bash
GID[ 0]: fe80::690:81ff:fe30:a7a0, RoCE v2
GID[ 0]: fe80::6a00:00ff:fe00:0001, RoCE v2
GID[ 1]: ::ffff:192.168.7.36, RoCE v2
```
@@ -601,17 +625,17 @@ appropriate IP.
```bash
# On Server Node
./ib_write_bw --use_rocm=0 -d mlx5_0 --report_gbits -a
./ib_write_bw --use_rocm=0 -d ionic_0 --report_gbits -a
# On Client Node
./ib_write_bw --use_rocm=0 -d mlx5_0 --report_gbits -a <SERVER_IP>
./ib_write_bw --use_rocm=0 -d ionic_0 --report_gbits -a <SERVER_IP>
```
## SGLang serving and MoRI unit tests
### Install Docker Engine
Install the Docker engine to manage the containerized vLLM and MoRI serving
Install the Docker engine to manage the containerized SGLang and MoRI serving
environments.
```bash
@@ -629,7 +653,7 @@ IMAGE_NAME=rocm/sgl-dev:sglang-0.5.6.post1-rocm700-mi35x-mori-0113
docker run -it \
--rm \
--device /dev/dri --device /dev/kfd --device=/dev/infiniBand \
--device /dev/dri --device /dev/kfd -v /dev/infiniband:/dev/infiniband \
--network host --ipc host \
--group-add video \
--cap-add SYS_PTRACE \
@@ -642,7 +666,7 @@ docker run -it \
### Run MoRI inter-node unit tests
Before starting the vLLM service, run the MoRI unit test to verify that the
Before starting the SGLang service, run the MoRI unit test to verify that the
inter-node communication backend is correctly configured.
MoRI unit test uses 2 nodes as a minimal validation before running the full
@@ -650,18 +674,34 @@ MoRI unit test uses 2 nodes as a minimal validation before running the full
The key configuration variables are:
* `GLOO_SOCKET_IFNAME`: The network interface used for backend initialization such as `eth2`.
* `GLOO_SOCKET_IFNAME`: The network interface used for backend initialization (for example, `benic1p1`).
* `MORI_SOCKET_IFNAME`: The network interface used by MoRI's own bootstrap. Set it to the same backend interface as `GLOO_SOCKET_IFNAME`.
* `MORI_GPU_ARCHS`: The target GPU architecture. Set to `gfx950` for MI355X; otherwise the test may auto-select the wrong arch (for example, `gfx942`).
* `<MASTER_IP>`: The IP address of the primary node's backend interface.
```{note}
You can find reference performance data in the [ROCm/MoRI
Performance reference data can be found in the [ROCm/MoRI
repository](https://github.com/ROCm/mori?tab=readme-ov-file#mori-ep).
```{note}
The `rocm/sgl-dev:sglang-0.5.6.post1-rocm700-mi35x-mori-0113` image ships MoRI
under `/sgl-workspace/mori`. If `/sgl-workspace/mori` or the example test
scripts (for example, `test_dispatch_combine_internode.py`) are missing, clone
the repository:
`git clone https://github.com/ROCm/mori.git /sgl-workspace/mori`
```
```bash
# Set up environment inside the container
export PYTHONPATH=/app/mori:$PYTHONPATH
export GLOO_SOCKET_IFNAME=<BACKEND_INTERFACE>
cd /sgl-workspace/mori
# prettytable is required to render the benchmark result table:
pip install prettytable
export PYTHONPATH=/sgl-workspace/mori:$PYTHONPATH
export MORI_GPU_ARCHS=gfx950 # MI355X arch; avoids auto-selecting gfx942
export GLOO_SOCKET_IFNAME=<BACKEND_INTERFACE> # e.g. benic1p1
export MORI_SOCKET_IFNAME=<BACKEND_INTERFACE> # MoRI bootstrap interface; same as GLOO_SOCKET_IFNAME
# Node 0 (Primary)
torchrun --nnodes=2 --node_rank=0 --nproc_per_node=1 \
@@ -679,9 +719,7 @@ torchrun --nnodes=2 --node_rank=1 --nproc_per_node=1 \
## End-to-end 1P2D performance testing
This section guides you through running distributed inference benchmarks using
the SGLang disagg recipe. For detailed implementation details, refer to the
[SGLang Disaggregation
Recipe](https://github.com/billishyahao/sglang_disagg/blob/9n_cluster/README.md).
the SGLang disagg recipe.
### Download the model and setup your run environment
@@ -711,13 +749,12 @@ hf download --token <your_hf_token> \
### Clone the SGLang disaggregation recipe
Clone the SGLang disaggregation repository to the shared file system and switch
to the appropriate branch:
Clone the [ROCm/distributed_inference](https://github.com/ROCm/distributed_inference)
repository to the shared file system:
```bash
git clone https://github.com/billishyahao/sglang_disagg.git
git checkout 9n_cluster
cd sglang_disagg
git clone https://github.com/ROCm/distributed_inference.git
cd distributed_inference
```
```{note}
@@ -749,28 +786,54 @@ Identify and configure the available InfiniBand devices.
ionic_7
```
2. Update environment variables. Edit `set_env_vars.sh` and add the
comma-separated list of your system's IB devices. For example:
2. Update environment variables. Edit `set_env_vars.sh` and set the
following variables:
```bash
export IBDEVICES=ionic_0,ionic_1,ionic_2,ionic_3,ionic_4,ionic_5,ionic_6,ionic_7
# Must be >= chunked_prefill_size / dp_size.
# Default recipe: 262144 / 8 = 32768. set_env_vars.sh's value (16384) is
# too small and causes an AssertionError on prefill startup.
export SGLANG_MORI_NUM_MAX_DISPATCH_TOKENS_PER_RANK=32768
# Must be large enough for the dispatch buffer. The default 4 GB heap
# causes an out-of-memory error at first inference with a 32768-token budget.
export MORI_SHMEM_HEAP_SIZE=16G
```
### Configure the script and submit the job
```{important}
Two `set_env_vars.sh` variables must be set before submitting the job, or
the prefill service crashes at startup:
* `SGLANG_MORI_NUM_MAX_DISPATCH_TOKENS_PER_RANK` — SGLang asserts this value is
≥ `chunked_prefill_size / dp_size`. For the default DeepSeek-R1 recipe
(262144 / 8 = 32768), `set_env_vars.sh`'s shipped value of 16384 triggers an
`AssertionError` on the prefill node only (decode is exempt). Raising this to
32768 fixes the crash.
* `MORI_SHMEM_HEAP_SIZE` — Raising the dispatch budget to 32768 tokens exceeds
MoRI's default 4 GB static heap and causes an out-of-memory error at first
inference. Set this to `16G` (MoRI's own inter-node test default).
These values are set in the preceding [Configure InfiniBand
devices](#configure-infiniBand-devices) step.
```
1. To set the required configuration parameters, update the following
environment variables in `run_submit_disagg.sh` to match your cluster setup:
```bash
# SLURM Job Configuration
export SLURM_ACCOUNT="amd" # The account name for SLURM job accounting and resource allocation
export SLURM_ACCOUNT="<your_slurm_account>" # The account name for SLURM job accounting and resource allocation
export SLURM_PARTITION="compute" # The specific cluster partition (queue) to submit the job to
export TIME_LIMIT="24:00:00" # Maximum wall time for the job (Hours:Minutes:Seconds)
# Model Configuration
export MODEL_PATH="/nfsdata" # Base directory where the model weights are stored
export MODEL_NAME="DeepSeek-R1" # Specific model directory name (joined with MODEL_PATH)
export CONTAINER_IMAGE="rocm/sgl-dev:sglang-0.5.6.post1-rocm700-mi35x-mori-1224" # Docker image to use for the environment
export CONTAINER_IMAGE="lmsysorg/sglang-rocm:v0.5.12.post1-rocm720-mi35x-20260529" # Docker image to use for the environment
# Cluster Topology (Disaggregation Setup)
export PREFILL_NODES=1 # Number of prefill nodes
@@ -827,7 +890,7 @@ Identify and configure the available InfiniBand devices.
```{note}
The following benchmark utility output is provided for reference only and
should not be used to compare performance. See the
[InferenceMAX](https://inferencemax.semianalysis.com/) website for validated
[InferenceX](https://inferencex.semianalysis.com/) website for validated
performance results.
```
@@ -863,7 +926,7 @@ Identify and configure the available InfiniBand devices.
The following section outlines common issues and their solutions.
### Bandwidth test fails with error
### Bandwidth test failures
1. Use ROCm-optimized `rdma-perftest`, not the generic `perftest`
File diff suppressed because it is too large Load Diff
@@ -30,8 +30,6 @@ training, fine-tuning, and inference. It leverages popular machine learning fram
- :doc:`SGLang distributed inference with MoRI <benchmark-docker/sglang-mori-distributed>`
- :doc:`SGLang distributed inference with Mooncake <benchmark-docker/sglang-distributed>`
- :doc:`xDiT diffusion inference <xdit-diffusion-inference>`
- :doc:`Deploying your model <deploy-your-model>`
-2
View File
@@ -105,8 +105,6 @@ subtrees:
title: vLLM distributed inference with MoRI
- file: how-to/rocm-for-ai/inference/benchmark-docker/sglang-mori-distributed.md
title: SGLang distributed inference with MoRI
- file: how-to/rocm-for-ai/inference/benchmark-docker/sglang-distributed.rst
title: SGLang distributed inference with Mooncake
- file: how-to/rocm-for-ai/inference/xdit-diffusion-inference.rst
title: xDiT diffusion inference
- file: how-to/rocm-for-ai/inference/deploy-your-model.rst