Compare commits
7
Commits
rocm-7.2.4
...
develop
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
001fa7c0ce | ||
|
|
582cacea47 | ||
|
|
5209dccdb2 | ||
|
|
876343e53b | ||
|
|
9f0ec581e2 | ||
|
|
6a6dc9136f | ||
|
|
54788f54d2 |
+4
-4
@@ -1,8 +1,8 @@
|
||||
* @amd-aakash @jlgreathouse @samjwu @yhuiYH @ROCm/rocm-documentation
|
||||
* @ROCm/rocm-documentation
|
||||
# Documentation files
|
||||
docs/ @amd-aakash @jlgreathouse @samjwu @yhuiYH @ROCm/rocm-documentation
|
||||
*.md @amd-aakash @jlgreathouse @samjwu @yhuiYH @ROCm/rocm-documentation
|
||||
*.rst @amd-aakash @jlgreathouse @samjwu @yhuiYH @ROCm/rocm-documentation
|
||||
docs/ @ROCm/rocm-documentation
|
||||
*.md @ROCm/rocm-documentation
|
||||
*.rst @ROCm/rocm-documentation
|
||||
# External CI
|
||||
/.azuredevops/ @ROCm/external-ci
|
||||
tools/rocm-build/ @ROCm/rocm-devops
|
||||
|
||||
@@ -0,0 +1,9 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" width="1280" height="640" viewBox="0 0 1280 640" role="img" aria-label="ROCm">
|
||||
<rect width="1280" height="640" fill="#0A0A0A"/>
|
||||
<svg x="96" y="215" width="210" height="210" viewBox="0 0 67 67"><path d="M22.21 67V44.6369H0V67H22.21Z" fill="#fff"/><path d="M66.7038 22.3184H22.2534L0.0878906 44.6367H44.4634L66.7038 22.3184Z" fill="#fff"/><path d="M22.21 0H0V22.3184H22.21V0Z" fill="#fff"/><path d="M66.7198 0H44.5098V22.3184H66.7198V0Z" fill="#fff"/><path d="M66.7198 67V44.6369H44.5098V67H66.7198Z" fill="#fff"/></svg>
|
||||
<text x="378" y="276" font-family="Inter,system-ui,-apple-system,sans-serif" font-size="78" font-weight="800" letter-spacing="-2" fill="#ffffff">ROCm</text>
|
||||
<text x="378" y="322" font-family="Inter,system-ui,sans-serif" font-size="30" fill="#ffffff" opacity=".66">AMD ROCm™ Software - GitHub Home</text>
|
||||
<rect x="378" y="338" width="806" height="3" rx="1.5" fill="#ffffff" opacity=".9"/>
|
||||
<text x="378" y="390" font-family="Inter,system-ui,sans-serif" font-size="24" font-weight="600" fill="#ffffff" opacity=".5">github.com/hanzoai</text>
|
||||
<text x="1184" y="390" text-anchor="end" font-family="Inter,system-ui,sans-serif" font-size="24" font-weight="600" fill="#ffffff" opacity=".5">hanzo.ai</text>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 1.2 KiB |
@@ -7,6 +7,10 @@ AITER
|
||||
ALU
|
||||
AllReduce
|
||||
AllToAll
|
||||
AGPR
|
||||
AGPRs
|
||||
AITER
|
||||
ALU
|
||||
AMD
|
||||
AMDGPU
|
||||
AMDGPUs
|
||||
@@ -40,6 +44,7 @@ BARs
|
||||
BKC
|
||||
BLAS
|
||||
BMC
|
||||
BNXT
|
||||
BabelStream
|
||||
Backported
|
||||
BatchNorm
|
||||
@@ -146,6 +151,7 @@ FHS
|
||||
FIFOs
|
||||
FIXME
|
||||
FMA
|
||||
FNUZ
|
||||
FP
|
||||
FX
|
||||
FiLM
|
||||
@@ -182,6 +188,8 @@ GIM
|
||||
GL
|
||||
Glibc
|
||||
GLM
|
||||
GIM
|
||||
GL
|
||||
GLXT
|
||||
GMI
|
||||
GNN
|
||||
@@ -206,6 +214,7 @@ GitHub
|
||||
Gitpod
|
||||
Glibc
|
||||
Gloo
|
||||
Gluon
|
||||
GraphBolt
|
||||
GraphSage
|
||||
HBM
|
||||
@@ -426,6 +435,7 @@ Pensando
|
||||
PerfDb
|
||||
Perfetto
|
||||
PipelineParallel
|
||||
Pipelining
|
||||
PnP
|
||||
Pollara
|
||||
PowerEdge
|
||||
@@ -636,6 +646,7 @@ allocator
|
||||
allocators
|
||||
amdgpu
|
||||
api
|
||||
async
|
||||
aten
|
||||
atmi
|
||||
atomicRMW
|
||||
@@ -643,6 +654,7 @@ atomics
|
||||
autogenerated
|
||||
autograd
|
||||
autotune
|
||||
autotuning
|
||||
avx
|
||||
awk
|
||||
az
|
||||
@@ -660,6 +672,7 @@ blit
|
||||
bootloader
|
||||
boson
|
||||
bosons
|
||||
bottlenecked
|
||||
br
|
||||
btn
|
||||
buildable
|
||||
@@ -771,6 +784,7 @@ ffmpeg
|
||||
filesystem
|
||||
flashinfer
|
||||
forEach
|
||||
foreach
|
||||
fortran
|
||||
fp
|
||||
framebuffer
|
||||
@@ -924,6 +938,8 @@ perfcounter
|
||||
performant
|
||||
perl
|
||||
piecewise
|
||||
pipelined
|
||||
pipelining
|
||||
pragma
|
||||
pre
|
||||
prebuild
|
||||
@@ -1055,6 +1071,7 @@ submatrix
|
||||
submodule
|
||||
submodules
|
||||
subnet
|
||||
subnets
|
||||
supercomputing
|
||||
symlink
|
||||
symlinks
|
||||
@@ -1096,9 +1113,11 @@ unfused
|
||||
unhandled
|
||||
uninstallation
|
||||
unmapped
|
||||
unpadded
|
||||
unsqueeze
|
||||
unstacking
|
||||
unswitching
|
||||
unswizzled
|
||||
untrusted
|
||||
untuned
|
||||
unwindowed
|
||||
|
||||
@@ -0,0 +1,18 @@
|
||||
Hanzo ROCm
|
||||
Copyright (c) 2026 Hanzo AI, Inc.
|
||||
|
||||
This product includes software from AMD ROCm (https://github.com/ROCm/ROCm),
|
||||
licensed under the MIT License:
|
||||
|
||||
Copyright (c) 2023 - 2025 Advanced Micro Devices, Inc. All rights reserved.
|
||||
|
||||
This repository is the AMD ROCm meta/manifest repository: documentation, build
|
||||
tooling, and the repo manifest (default.xml). This meta-repository is MIT-licensed
|
||||
and its upstream MIT license is retained in LICENSE.
|
||||
|
||||
The individual ROCm components referenced by the manifest and built by the tooling
|
||||
carry their own upstream licenses, including the MIT License, the Apache License 2.0,
|
||||
and the University of Illinois/NCSA Open Source License. Note that the ROCgdb
|
||||
component (AMD's fork of GNU GDB, packaged by tools/rocm-build/build_rocm-gdb.sh) is
|
||||
licensed under the GNU General Public License (GPL) — a copyleft license. Each
|
||||
component is governed by its own license; consult that component's repository.
|
||||
@@ -1,3 +1,5 @@
|
||||
<p align="center"><img src=".github/hero.svg" alt="ROCm" width="880"></p>
|
||||
|
||||
<div align="center">
|
||||
<img src="docs/data/amd-rocm-logo.png" width="200px" alt="ROCm logo">
|
||||
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -1,257 +0,0 @@
|
||||
.. meta::
|
||||
:description: SGLang multi-node disaggregated distributed inference using Mooncake
|
||||
:keywords: model, sglang, mooncake, disagg, disaggregated, distributed, multi-node, docker
|
||||
|
||||
******************************************
|
||||
SGLang distributed inference with Mooncake
|
||||
******************************************
|
||||
|
||||
As LLM inference increasingly demands handling massive models and dynamic workloads, efficient
|
||||
distributed inference becomes essential. Traditional co-located architectures face bottlenecks due
|
||||
to tightly coupled memory and compute resources, which limits scalability and flexibility.
|
||||
Disaggregated inference refers to the process of splitting the inference of LLMs into distinct
|
||||
phases. This architecture, facilitated by libraries like Mooncake, uses high-bandwidth
|
||||
RDMA to transfer the Key-Value (KV) cache between prefill and decode nodes.
|
||||
This allows for independent resource scaling and optimization, resulting in
|
||||
improved efficiency and throughput.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/sglang-distributed-benchmark-models.yaml
|
||||
|
||||
{% set docker = data.dockers[0] %}
|
||||
|
||||
`SGLang <https://docs.sglang.ai>`__ is a high-performance inference and
|
||||
serving engine for large language models (LLMs) and vision models. The
|
||||
ROCm-enabled `SGLang base Docker image <{{ docker.docker_hub_url }}>`__
|
||||
bundles SGLang with PyTorch, which is optimized for AMD Instinct MI300X Series
|
||||
GPUs. It includes the following software components:
|
||||
|
||||
.. list-table::
|
||||
:header-rows: 1
|
||||
|
||||
* - Software component
|
||||
- Version
|
||||
|
||||
{% for component_name, component_version in docker.components.items() %}
|
||||
* - {{ component_name }}
|
||||
- {{ component_version }}
|
||||
{% endfor %}
|
||||
|
||||
The following guides on setting up and running SGLang and Mooncake for disaggregated
|
||||
distributed inference on a Slurm cluster using AMD Instinct MI300X Series GPUs backed by
|
||||
Mellanox CX-7 NICs.
|
||||
|
||||
Prerequisites
|
||||
=============
|
||||
|
||||
Before starting, ensure you have:
|
||||
|
||||
* A Slurm cluster with at least three nodes: one for the proxy, one for prefill (``xP``), and one for decode (``yD``).
|
||||
|
||||
``Nodes -> xP + yD + 1``
|
||||
|
||||
* A Dockerized environment with SGLang, Mooncake, etcd, and NIC drivers built in. See :ref:`sglang-disagg-inf-build-docker-image` for instructions.
|
||||
|
||||
* A shared filesystem for storing models, scripts, and logs (cluster-specific).
|
||||
|
||||
Supported models
|
||||
================
|
||||
|
||||
The following models are supported for SGLang disaggregated prefill/decode
|
||||
inference. Some instructions, commands, and recommendations in this
|
||||
documentation might vary by selected model.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/sglang-distributed-benchmark-models.yaml
|
||||
|
||||
{% set model_groups = data.model_groups %}
|
||||
.. raw:: html
|
||||
|
||||
<div id="vllm-benchmark-ud-params-picker" class="container-fluid">
|
||||
<div class="row gx-0">
|
||||
<div class="col-2 me-1 px-2 model-param-head">Model type</div>
|
||||
<div class="row col-10 pe-0">
|
||||
{% for model_group in model_groups %}
|
||||
<div class="col-6 px-2 model-param" data-param-k="model-group" data-param-v="{{ model_group.tag }}" tabindex="0">{{ model_group.group }}</div>
|
||||
{% endfor %}
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div class="row gx-0 pt-1">
|
||||
<div class="col-2 me-1 px-2 model-param-head">Model</div>
|
||||
<div class="row col-10 pe-0">
|
||||
{% for model_group in model_groups %}
|
||||
{% set models = model_group.models %}
|
||||
{% for model in models %}
|
||||
{% if models|length % 3 == 0 %}
|
||||
<div class="col-4 px-2 model-param" data-param-k="model" data-param-v="{{ model.model_repo | lower }}" data-param-group="{{ model_group.tag }}" tabindex="0">{{ model.model }}</div>
|
||||
{% else %}
|
||||
<div class="col-6 px-2 model-param" data-param-k="model" data-param-v="{{ model.model_repo | lower }}" data-param-group="{{ model_group.tag }}" tabindex="0">{{ model.model }}</div>
|
||||
{% endif %}
|
||||
{% endfor %}
|
||||
{% endfor %}
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
{% for model_group in model_groups %}
|
||||
{% for model in model_group.models %}
|
||||
|
||||
.. container:: model-doc {{ model.model_repo }}
|
||||
|
||||
.. note::
|
||||
|
||||
See the `{{ model.model }} model card on Hugging Face <{{ model.url }}>`__ to learn more about this model.
|
||||
Some models require access authorization prior to use through an external license agreement with a third party.
|
||||
|
||||
{% endfor %}
|
||||
{% endfor %}
|
||||
|
||||
.. _sglang-disagg-inf-build-docker-image:
|
||||
|
||||
Build the Docker image
|
||||
----------------------
|
||||
|
||||
Get the Dockerfile located in
|
||||
`<https://github.com/ROCm/MAD/blob/develop/docker/sglang_disagg_inference.ubuntu.amd.Dockerfile>`__.
|
||||
It uses `lmsysorg/sglang:v0.5.2rc1-rocm700-mi30x
|
||||
<https://hub.docker.com/layers/lmsysorg/sglang/v0.4.9.post1-rocm630/images/sha256-2f6b1748e4bcc70717875a7da76c87795fd8aa46a9646e08d38aa7232fc78538>`__
|
||||
as the base Docker image and installs the necessary components for Mooncake, etcd, and Mellanox network
|
||||
drivers.
|
||||
|
||||
.. code-block:: shell
|
||||
|
||||
git clone https://github.com/ROCm/MAD.git
|
||||
cd MAD/docker
|
||||
docker build \
|
||||
-t sglang_disagg_pd_image \
|
||||
-f sglang_disagg_inference.ubuntu.amd.Dockerfile .
|
||||
|
||||
Benchmarking
|
||||
============
|
||||
|
||||
The `<https://github.com/ROCm/MAD/tree/develop/scripts/sglang_disagg>`__
|
||||
repository contains scripts to launch SGLang inference with prefill/decode
|
||||
disaggregation via Mooncake for supported models.
|
||||
|
||||
* `scripts/sglang_dissag/run_xPyD_models.slurm <https://github.com/ROCm/MAD/blob/develop/scripts/sglang_disagg/run_xPyD_models.slurm>`__
|
||||
-- the main Slurm batch script to launch Docker containers on all nodes using ``sbatch`` or ``salloc``.
|
||||
|
||||
* `scripts/sglang_dissag/sglang_disagg_server.sh <https://github.com/ROCm/MAD/blob/develop/scripts/sglang_disagg/sglang_disagg_server.sh>`__
|
||||
-- the entrypoint script that runs inside each container to start the correct service -- proxy, prefill, or decode.
|
||||
|
||||
* `scripts/sglang_dissag/benchmark_xPyD.sh <https://github.com/ROCm/MAD/blob/develop/scripts/sglang_disagg/benchmark_xPyD.sh>`__
|
||||
-- the benchmark script to run the GSM8K accuracy benchmark and the SGLang benchmarking tool for performance measurement.
|
||||
|
||||
* `scripts/sglang_dissag/benchmark_parser.py <https://github.com/ROCm/MAD/blob/develop/scripts/sglang_disagg/benchmark_parser.py>`__
|
||||
-- the log parser script to be run on the concurrency benchmark log file to generate tabulated data.
|
||||
|
||||
Launch the service
|
||||
------------------
|
||||
|
||||
The service is deployed using a Slurm batch script that orchestrates the containers across the
|
||||
allocated nodes.
|
||||
|
||||
.. datatemplate:yaml:: /data/how-to/rocm-for-ai/inference/sglang-distributed-benchmark-models.yaml
|
||||
|
||||
{% set model_groups = data.model_groups %}
|
||||
{% for model_group in model_groups %}
|
||||
{% for model in model_group.models %}
|
||||
|
||||
.. container:: model-doc {{ model.model_repo }}
|
||||
|
||||
.. code-block:: shell
|
||||
|
||||
# Clone the MAD repo if you haven't already and
|
||||
# navigate to the scripts directory
|
||||
git clone https://github.com/ROCm/MAD.git
|
||||
cd MAD/scripts/sglang_disagg/
|
||||
|
||||
# Slurm sbatch run command
|
||||
export DOCKER_IMAGE_NAME=sglang_disagg_pd_image
|
||||
export xP=<num_prefill_nodes>
|
||||
export yD=<num_decode_nodes>
|
||||
export MODEL_NAME={{ model.model_repo }}
|
||||
# num_nodes = xP + yD + 1
|
||||
sbatch -N <num_nodes> -n <num_nodes> --nodelist=<Nodes> run_xPyD_models.slurm
|
||||
|
||||
{% endfor %}
|
||||
{% endfor %}
|
||||
|
||||
Post-run logs and testing
|
||||
-------------------------
|
||||
|
||||
Logs are stored in your shared filesystem in the directory specified by the ``LOG_PATH`` variable in the Slurm script.
|
||||
A new directory named after the Slurm job ID is created for each run.
|
||||
|
||||
Inside that directory, you can access various logs:
|
||||
|
||||
* ``pd_sglang_bench_serving.sh_NODE<...>.log`` -- the main log for each server node.
|
||||
|
||||
* ``etcd_NODE<...>.log`` -- logs for etcd services.
|
||||
|
||||
* ``prefill_NODE<...>.log`` -- logs for the prefill services.
|
||||
|
||||
* ``decode_NODE<...>.log`` -- logs for the decode services.
|
||||
|
||||
Use the benchmark parser script for concurrency logs to tabulate different data.
|
||||
|
||||
.. code-block:: shell
|
||||
|
||||
python3 benchmark_parser.py <log_path/benchmark_XXX_CONCURRENCY.log>
|
||||
|
||||
To verify the service is responsive, you can try sending a ``curl`` request to test the launched
|
||||
server from the Docker container on the proxy node. For example:
|
||||
|
||||
.. code-block:: shell
|
||||
|
||||
curl -X POST http://127.0.0.1:30000/generate \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{ "text": "Let me tell you a story ", "sampling_params": { "temperature": 0.3 } }'
|
||||
|
||||
Known issues
|
||||
============
|
||||
|
||||
When running larger models, such as DeepSeek-V3 and Llama-3.1-405B-Instruct-FP8-KV, at
|
||||
higher concurrency levels (512+), the following error might occur:
|
||||
|
||||
.. code-block:: shell-session
|
||||
|
||||
<TransferEncodingError: 400, message:
|
||||
Not enough data to satisfy transfer length header.
|
||||
|
||||
The above exception was the direct cause of the following exception:
|
||||
|
||||
Traceback (most recent call last):
|
||||
...
|
||||
|
||||
This leads to dropping requests and lower throughput.
|
||||
|
||||
Further reading
|
||||
===============
|
||||
|
||||
- To learn about Mooncake, see `Welcome to Mooncake <https://kvcache-ai.github.io/Mooncake/>`__.
|
||||
|
||||
- To learn more about the options for latency and throughput benchmark scripts,
|
||||
see `<https://github.com/sgl-project/sglang/tree/main/benchmark/blog_v0_2>`__.
|
||||
|
||||
- See the base upstream Docker image on `Docker Hub <https://hub.docker.com/layers/lmsysorg/sglang/v0.5.2rc1-rocm700-mi30x/images/sha256-10c4ee502ddba44dd8c13325e6e03868bfe7f43d23d0a44780a8ee8b393f4729>`__.
|
||||
|
||||
- To learn more about system settings and management practices to configure your system for
|
||||
MI300X Series GPUs, see `AMD Instinct MI300X system optimization <https://instinct.docs.amd.com/projects/amdgpu-docs/en/latest/system-optimization/mi300x.html>`__.
|
||||
|
||||
- For application performance optimization strategies for HPC and AI workloads,
|
||||
including inference with vLLM, see :doc:`/how-to/rocm-for-ai/inference-optimization/workload`.
|
||||
|
||||
- To learn how to run community models from Hugging Face on AMD GPUs, see
|
||||
:doc:`Running models from Hugging Face </how-to/rocm-for-ai/inference/hugging-face-models>`.
|
||||
|
||||
- To learn how to fine-tune LLMs and optimize inference, see
|
||||
:doc:`Fine-tuning LLMs and inference optimization </how-to/rocm-for-ai/fine-tuning/fine-tuning-and-inference>`.
|
||||
|
||||
- For a list of other ready-made Docker images for AI with ROCm, see
|
||||
`AMD Infinity Hub <https://www.amd.com/en/developer/resources/infinity-hub.html#f-amd_hub_category=AI%20%26%20ML%20Models>`_.
|
||||
|
||||
Previous versions
|
||||
=================
|
||||
|
||||
See :doc:`previous-versions/sglang-history` to find documentation for previous releases
|
||||
of SGLang inference performance testing.
|
||||
@@ -12,14 +12,15 @@ scripts.
|
||||
|
||||
The following configuration is required to implement this setup:
|
||||
|
||||
* **Nodes:** A minimum of three GPU nodes (Virtual machines or Physical
|
||||
machines) for wide expert parallelism (EP) evaluation.
|
||||
* **GPUs** 8x AMD Instinct MI355X GPU cards per node.
|
||||
* **Networking:** 8x AMD Pensando™ Pollara 400 AI NICs per node, providing
|
||||
a dedicated 1:1 mapping between GPUs and network interfaces for optimal
|
||||
inter-node communication.
|
||||
* **Orchestration:** A Slurm cluster with at least three nodes -- one for
|
||||
prefill service and two for decode services (EP16)
|
||||
* **Nodes**: A minimum of three GPU nodes (virtual machines or physical
|
||||
machines) for wide expert parallelism (EP) evaluation.
|
||||
* **GPUs**: 8x AMD Instinct MI355X GPU cards per node.
|
||||
* **Networking**: 8x RDMA-capable NICs per node (AMD Pensando Pollara 400,
|
||||
NVIDIA Mellanox ConnectX-7, or Broadcom Thor 2), providing a dedicated 1:1
|
||||
mapping between GPUs and network interfaces for optimal inter-node
|
||||
communication.
|
||||
* **Orchestration**: A Slurm cluster with at least three nodes — one for
|
||||
prefill service and two for decode services (EP16).
|
||||
|
||||
## System configuration
|
||||
|
||||
@@ -29,8 +30,6 @@ baselines and firmware versions, configuring the AMD Pensando Pollara 400 AI
|
||||
NICs for high-bandwidth networking, and applying thermal and Quality of Service
|
||||
(QoS) tunings to ensure a stable, lossless RDMA fabric.
|
||||
|
||||
(sglang-mori-verify-baseline)=
|
||||
|
||||
### Verify baseline software
|
||||
|
||||
The following table outlines the validated software stack. Use the provided
|
||||
@@ -81,18 +80,19 @@ Redfish API:
|
||||
|
||||
Before proceeding with software deployment, verify that all cluster nodes
|
||||
comply with the [MI355X Basic Health
|
||||
Checks](https://instinct.docs.amd.com/projects/system-acceptance/en/latest/gpus/mi355x.html#basic-health-checks)
|
||||
Checks](https://instinct.docs.amd.com/projects/system-acceptance/en/latest/gpus/mi355x.html#basic-health-checks).
|
||||
Key requirements include specific kernel boot arguments, minimum system memory
|
||||
thresholds, PCIe Gen5 link stability, and so on.
|
||||
|
||||
### Install AMD Pensando Pollara 400 AI NIC drivers
|
||||
### NIC installation
|
||||
|
||||
#### AMD Pensando Pollara 400 AI NIC installation
|
||||
|
||||
For detailed instructions on upgrading the firmware and installing drivers for
|
||||
the AMD Pensando Pollara 400 AI NIC, refer to the [AMD Instinct System
|
||||
Acceptance
|
||||
Guide](https://instinct.docs.amd.com/projects/system-acceptance/en/latest/network/nic-installation.html#amd-pensando-pollara-400-ai-nic).
|
||||
Acceptance Guide](https://instinct.docs.amd.com/projects/system-acceptance/en/latest/network/nic-installation.html#amd-pensando-pollara-400-ai-nic).
|
||||
After installation, verify the active firmware version on all NICs to ensure it
|
||||
matches the software baseline. See [Verify baseline software](#verify-best-known-configuration-bkc).
|
||||
matches the software baseline. See [Verify baseline software](#verify-baseline-software).
|
||||
|
||||
To display the current firmware version for all AI NICs, use the following command.
|
||||
|
||||
@@ -100,6 +100,26 @@ To display the current firmware version for all AI NICs, use the following comma
|
||||
sudo nicctl show version firmware
|
||||
```
|
||||
|
||||
#### CX7 driver and firmware installation
|
||||
|
||||
1. Download and install the `DOCA 2.9.3` driver following the instructions in
|
||||
[NVIDIA DOCA 2.9.3 Downloads](https://developer.nvidia.com/doca-downloads).
|
||||
2. Download the appropriate firmware for your hardware PSID from the
|
||||
[ConnectX-7 Firmware Download
|
||||
Center](https://network.nvidia.com/support/firmware/connectx7/) and flash
|
||||
the device.
|
||||
3. To verify driver and firmware versions, use the following command. Replace
|
||||
`IB Device` with your specific backend interface.
|
||||
|
||||
```bash
|
||||
ethtool -i <IB Device>
|
||||
```
|
||||
|
||||
#### Broadcom BNXT driver and firmware installation
|
||||
|
||||
Refer to your Broadcom representative for driver and firmware installation
|
||||
instructions specific to your NIC model.
|
||||
|
||||
### Configure thermal management (fan speed)
|
||||
|
||||
For systems equipped with 400G optics, standard fan profiles are often
|
||||
@@ -140,8 +160,10 @@ the addresses `192.168.1.36`, `192.168.2.36`, and so on. Another node would
|
||||
have `192.168.1.37`, `192.168.2.37`, and so on. Ensure MTU is set to `9000`.
|
||||
|
||||
```{note}
|
||||
Ensure you identify the correct interface names for your system using ip link
|
||||
before applying this configuration.
|
||||
Ensure you identify the correct interface names for your system using `ip link`
|
||||
before applying this configuration. The `macaddress:` values in the example
|
||||
below are illustrative only and must be replaced with the actual MAC addresses
|
||||
of your NICs, which you can find using `ip link show <interface>`.
|
||||
```
|
||||
|
||||
For example, your `/etc/netplan/70-backend.yaml` should look like the
|
||||
@@ -154,7 +176,7 @@ network:
|
||||
addresses:
|
||||
- 192.168.8.38/31
|
||||
match:
|
||||
macaddress: 04:90:81:2a:34:08
|
||||
macaddress: 04:90:81:00:00:08
|
||||
mtu: 9000
|
||||
routes:
|
||||
- table: 108
|
||||
@@ -168,7 +190,7 @@ network:
|
||||
addresses:
|
||||
- 192.168.7.38/31
|
||||
match:
|
||||
macaddress: 04:90:81:2b:82:40
|
||||
macaddress: 04:90:81:00:00:07
|
||||
mtu: 9000
|
||||
routes:
|
||||
- table: 107
|
||||
@@ -182,7 +204,7 @@ network:
|
||||
addresses:
|
||||
- 192.168.6.38/31
|
||||
match:
|
||||
macaddress: 04:90:81:30:c9:30
|
||||
macaddress: 04:90:81:00:00:06
|
||||
mtu: 9000
|
||||
routes:
|
||||
- table: 106
|
||||
@@ -196,7 +218,7 @@ network:
|
||||
addresses:
|
||||
- 192.168.5.38/31
|
||||
match:
|
||||
macaddress: 04:90:81:2a:23:40
|
||||
macaddress: 04:90:81:00:00:05
|
||||
mtu: 9000
|
||||
routes:
|
||||
- table: 105
|
||||
@@ -210,7 +232,7 @@ network:
|
||||
addresses:
|
||||
- 192.168.4.38/31
|
||||
match:
|
||||
macaddress: 04:90:81:2d:69:60
|
||||
macaddress: 04:90:81:00:00:04
|
||||
mtu: 9000
|
||||
routes:
|
||||
- table: 104
|
||||
@@ -224,7 +246,7 @@ network:
|
||||
addresses:
|
||||
- 192.168.3.38/31
|
||||
match:
|
||||
macaddress: 04:90:81:2a:2c:40
|
||||
macaddress: 04:90:81:00:00:03
|
||||
mtu: 9000
|
||||
routes:
|
||||
- table: 103
|
||||
@@ -238,7 +260,7 @@ network:
|
||||
addresses:
|
||||
- 192.168.2.38/31
|
||||
match:
|
||||
macaddress: 04:90:81:30:d5:30
|
||||
macaddress: 04:90:81:00:00:02
|
||||
mtu: 9000
|
||||
routes:
|
||||
- table: 102
|
||||
@@ -252,7 +274,7 @@ network:
|
||||
addresses:
|
||||
- 192.168.1.38/31
|
||||
match:
|
||||
macaddress: 04:90:81:30:e4:00
|
||||
macaddress: 04:90:81:00:00:01
|
||||
mtu: 9000
|
||||
routes:
|
||||
- table: 101
|
||||
@@ -276,17 +298,18 @@ To verify your configuration, use the following command.
|
||||
sudo apt install -y net-tools && ip -br a
|
||||
```
|
||||
|
||||
### Configure Quality of Service (QoS) and Congestion Control (DCQCN)
|
||||
### Configure quality of service (QoS) and congestion control (DCQCN)
|
||||
|
||||
To ensure lossless communication and optimal performance for RDMA traffic, the
|
||||
network must be configured with specific QoS and Data Center Quantized
|
||||
Congestion Notification (DCQCN) settings.
|
||||
|
||||
The following configuration achieves:
|
||||
• It enables RX and TX Pause frames on the ports
|
||||
• Maps DSCP 24 (Data) to Q3 and DSCP 46 (CNP) to Q6, all other DSCP to Q0
|
||||
• Enables PFC for Q3
|
||||
• Scheduling : 99% to Q3, 1% to Q0 and strict priority for Q6
|
||||
The following configuration:
|
||||
|
||||
* Enables RX and TX pause frames on the ports.
|
||||
* Maps DSCP 24 (Data) to Q3 and DSCP 46 (CNP) to Q6, with all other DSCP to Q0.
|
||||
* Enables PFC for Q3.
|
||||
* Scheduling: 99% to Q3, 1% to Q0, and strict priority for Q6.
|
||||
|
||||
#### Configure DCQCN
|
||||
|
||||
@@ -294,7 +317,7 @@ Create and run a `/nfsdata/enable_dcqcn.sh` script to initialize congestion
|
||||
control parameters.
|
||||
|
||||
``` bash
|
||||
# !/bin/bash
|
||||
#!/bin/bash
|
||||
|
||||
TOKEN_BUCKET_SIZE=800000
|
||||
AI_RATE=160
|
||||
@@ -361,7 +384,7 @@ sudo nicctl update qos pfc --priority $data_prio --no-drop enable
|
||||
sudo nicctl update qos scheduling --priority $data_prio,$default_prio,$cts_prio --dwrr 99,1,0 --rate-limit 0,0,10
|
||||
```
|
||||
|
||||
#### Verification your configuration
|
||||
#### Verify your configuration
|
||||
|
||||
Verify the configuration using `nicctl`.
|
||||
|
||||
@@ -374,9 +397,9 @@ Verify the configuration using `nicctl`.
|
||||
Expected QoS output:
|
||||
|
||||
``` bash
|
||||
NIC : 42424650-4c32-3531-3230-303443000000 (0000:f6:00.0)
|
||||
NIC : 00000000-0000-0000-0000-000000000001 (0000:f6:00.0)
|
||||
|
||||
Port : 04908130-a7a0-4242-4242-000011010000
|
||||
Port : 00000000-0001-4242-4242-000000000000
|
||||
|
||||
Classification type : DSCP
|
||||
|
||||
@@ -398,10 +421,10 @@ Verify the configuration using `nicctl`.
|
||||
Expected DCQCN and scheduling output:
|
||||
|
||||
``` bash
|
||||
NIC : 42424650-4c32-3531-3230-303443000000 (0000:f6:00.0)
|
||||
NIC : 00000000-0000-0000-0000-000000000001 (0000:f6:00.0)
|
||||
------------------------------------------------------------------------------------------
|
||||
|
||||
Lif id : 43000070-0100-0000-4242-04908130a7a0
|
||||
Lif id : 00000000-0100-0000-4242-000000000000
|
||||
ROCE device : ionic_7
|
||||
DCQCN profile id : 1
|
||||
Status : Enabled
|
||||
@@ -497,9 +520,10 @@ the cluster interconnects.
|
||||
### Verify network connectivity
|
||||
|
||||
Verify that all network interfaces are reachable across the cluster nodes.
|
||||
Assuming `eth0` is the management interface, and `benic1p1` through `benic8p1` are the
|
||||
dedicated RoCE backend interfaces, use the following loop to test reachability
|
||||
to a remote node (for instance, a target node with host IP suffix `.38`).
|
||||
Assuming `benic1p1` through `benic8p1` are the dedicated RoCE backend
|
||||
interfaces, use the following ping loop to verify reachability across the
|
||||
backend subnets (for instance, a target node at host IP suffix
|
||||
`.38`).
|
||||
|
||||
```bash
|
||||
# Test connectivity for RoCE subnets 192.168.x.38 (node B) through 192.168.x.37 (node A)
|
||||
@@ -522,9 +546,9 @@ The output should look something like this:
|
||||
```bash
|
||||
-------------------------------------------------------------------------------------
|
||||
|
||||
NIC : 42424650-4c32-3531-3530-314343000000 (0000:f6:00.0)
|
||||
NIC : 00000000-0000-0000-0000-000000000002 (0000:f6:00.0)
|
||||
|
||||
Port : 04908132-5d88-4242-4242-000011010000 (eth1/1)
|
||||
Port : 00000000-0002-4242-4242-000000000000 (eth1/1)
|
||||
Spec:
|
||||
Ifindex : 0x11010000
|
||||
Type : ETH
|
||||
@@ -548,7 +572,7 @@ Port : 04908132-5d88-4242-4242-000011010000 (eth1/1)
|
||||
Auto negotiation : disabled
|
||||
MAC ID : 0
|
||||
MAC channel : 0
|
||||
MAC address : 04:90:81:32:5d:88
|
||||
MAC address : 04:90:81:00:00:00
|
||||
Transceiver type : QSFP_CMIS
|
||||
Transceiver state : SPROM-READ
|
||||
Transceiver PID : QSFP-400G-DR4
|
||||
@@ -569,7 +593,7 @@ ibv_devinfo -v | grep GID
|
||||
The output should look something like this:
|
||||
|
||||
```bash
|
||||
GID[ 0]: fe80::690:81ff:fe30:a7a0, RoCE v2
|
||||
GID[ 0]: fe80::6a00:00ff:fe00:0001, RoCE v2
|
||||
GID[ 1]: ::ffff:192.168.7.36, RoCE v2
|
||||
```
|
||||
|
||||
@@ -601,17 +625,17 @@ appropriate IP.
|
||||
|
||||
```bash
|
||||
# On Server Node
|
||||
./ib_write_bw --use_rocm=0 -d mlx5_0 --report_gbits -a
|
||||
./ib_write_bw --use_rocm=0 -d ionic_0 --report_gbits -a
|
||||
|
||||
# On Client Node
|
||||
./ib_write_bw --use_rocm=0 -d mlx5_0 --report_gbits -a <SERVER_IP>
|
||||
./ib_write_bw --use_rocm=0 -d ionic_0 --report_gbits -a <SERVER_IP>
|
||||
```
|
||||
|
||||
## SGLang serving and MoRI unit tests
|
||||
|
||||
### Install Docker Engine
|
||||
|
||||
Install the Docker engine to manage the containerized vLLM and MoRI serving
|
||||
Install the Docker engine to manage the containerized SGLang and MoRI serving
|
||||
environments.
|
||||
|
||||
```bash
|
||||
@@ -629,7 +653,7 @@ IMAGE_NAME=rocm/sgl-dev:sglang-0.5.6.post1-rocm700-mi35x-mori-0113
|
||||
|
||||
docker run -it \
|
||||
--rm \
|
||||
--device /dev/dri --device /dev/kfd --device=/dev/infiniBand \
|
||||
--device /dev/dri --device /dev/kfd -v /dev/infiniband:/dev/infiniband \
|
||||
--network host --ipc host \
|
||||
--group-add video \
|
||||
--cap-add SYS_PTRACE \
|
||||
@@ -642,7 +666,7 @@ docker run -it \
|
||||
|
||||
### Run MoRI inter-node unit tests
|
||||
|
||||
Before starting the vLLM service, run the MoRI unit test to verify that the
|
||||
Before starting the SGLang service, run the MoRI unit test to verify that the
|
||||
inter-node communication backend is correctly configured.
|
||||
|
||||
MoRI unit test uses 2 nodes as a minimal validation before running the full
|
||||
@@ -650,18 +674,34 @@ MoRI unit test uses 2 nodes as a minimal validation before running the full
|
||||
|
||||
The key configuration variables are:
|
||||
|
||||
* `GLOO_SOCKET_IFNAME`: The network interface used for backend initialization such as `eth2`.
|
||||
* `GLOO_SOCKET_IFNAME`: The network interface used for backend initialization (for example, `benic1p1`).
|
||||
* `MORI_SOCKET_IFNAME`: The network interface used by MoRI's own bootstrap. Set it to the same backend interface as `GLOO_SOCKET_IFNAME`.
|
||||
* `MORI_GPU_ARCHS`: The target GPU architecture. Set to `gfx950` for MI355X; otherwise the test may auto-select the wrong arch (for example, `gfx942`).
|
||||
* `<MASTER_IP>`: The IP address of the primary node's backend interface.
|
||||
|
||||
```{note}
|
||||
You can find reference performance data in the [ROCm/MoRI
|
||||
Performance reference data can be found in the [ROCm/MoRI
|
||||
repository](https://github.com/ROCm/mori?tab=readme-ov-file#mori-ep).
|
||||
|
||||
```{note}
|
||||
The `rocm/sgl-dev:sglang-0.5.6.post1-rocm700-mi35x-mori-0113` image ships MoRI
|
||||
under `/sgl-workspace/mori`. If `/sgl-workspace/mori` or the example test
|
||||
scripts (for example, `test_dispatch_combine_internode.py`) are missing, clone
|
||||
the repository:
|
||||
|
||||
`git clone https://github.com/ROCm/mori.git /sgl-workspace/mori`
|
||||
```
|
||||
|
||||
```bash
|
||||
# Set up environment inside the container
|
||||
export PYTHONPATH=/app/mori:$PYTHONPATH
|
||||
export GLOO_SOCKET_IFNAME=<BACKEND_INTERFACE>
|
||||
cd /sgl-workspace/mori
|
||||
|
||||
# prettytable is required to render the benchmark result table:
|
||||
pip install prettytable
|
||||
|
||||
export PYTHONPATH=/sgl-workspace/mori:$PYTHONPATH
|
||||
export MORI_GPU_ARCHS=gfx950 # MI355X arch; avoids auto-selecting gfx942
|
||||
export GLOO_SOCKET_IFNAME=<BACKEND_INTERFACE> # e.g. benic1p1
|
||||
export MORI_SOCKET_IFNAME=<BACKEND_INTERFACE> # MoRI bootstrap interface; same as GLOO_SOCKET_IFNAME
|
||||
|
||||
# Node 0 (Primary)
|
||||
torchrun --nnodes=2 --node_rank=0 --nproc_per_node=1 \
|
||||
@@ -679,9 +719,7 @@ torchrun --nnodes=2 --node_rank=1 --nproc_per_node=1 \
|
||||
## End-to-end 1P2D performance testing
|
||||
|
||||
This section guides you through running distributed inference benchmarks using
|
||||
the SGLang disagg recipe. For detailed implementation details, refer to the
|
||||
[SGLang Disaggregation
|
||||
Recipe](https://github.com/billishyahao/sglang_disagg/blob/9n_cluster/README.md).
|
||||
the SGLang disagg recipe.
|
||||
|
||||
### Download the model and setup your run environment
|
||||
|
||||
@@ -711,13 +749,12 @@ hf download --token <your_hf_token> \
|
||||
|
||||
### Clone the SGLang disaggregation recipe
|
||||
|
||||
Clone the SGLang disaggregation repository to the shared file system and switch
|
||||
to the appropriate branch:
|
||||
Clone the [ROCm/distributed_inference](https://github.com/ROCm/distributed_inference)
|
||||
repository to the shared file system:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/billishyahao/sglang_disagg.git
|
||||
git checkout 9n_cluster
|
||||
cd sglang_disagg
|
||||
git clone https://github.com/ROCm/distributed_inference.git
|
||||
cd distributed_inference
|
||||
```
|
||||
|
||||
```{note}
|
||||
@@ -749,28 +786,54 @@ Identify and configure the available InfiniBand devices.
|
||||
ionic_7
|
||||
```
|
||||
|
||||
2. Update environment variables. Edit `set_env_vars.sh` and add the
|
||||
comma-separated list of your system's IB devices. For example:
|
||||
2. Update environment variables. Edit `set_env_vars.sh` and set the
|
||||
following variables:
|
||||
|
||||
```bash
|
||||
export IBDEVICES=ionic_0,ionic_1,ionic_2,ionic_3,ionic_4,ionic_5,ionic_6,ionic_7
|
||||
|
||||
# Must be >= chunked_prefill_size / dp_size.
|
||||
# Default recipe: 262144 / 8 = 32768. set_env_vars.sh's value (16384) is
|
||||
# too small and causes an AssertionError on prefill startup.
|
||||
export SGLANG_MORI_NUM_MAX_DISPATCH_TOKENS_PER_RANK=32768
|
||||
|
||||
# Must be large enough for the dispatch buffer. The default 4 GB heap
|
||||
# causes an out-of-memory error at first inference with a 32768-token budget.
|
||||
export MORI_SHMEM_HEAP_SIZE=16G
|
||||
```
|
||||
|
||||
### Configure the script and submit the job
|
||||
|
||||
```{important}
|
||||
Two `set_env_vars.sh` variables must be set before submitting the job, or
|
||||
the prefill service crashes at startup:
|
||||
|
||||
* `SGLANG_MORI_NUM_MAX_DISPATCH_TOKENS_PER_RANK` — SGLang asserts this value is
|
||||
≥ `chunked_prefill_size / dp_size`. For the default DeepSeek-R1 recipe
|
||||
(262144 / 8 = 32768), `set_env_vars.sh`'s shipped value of 16384 triggers an
|
||||
`AssertionError` on the prefill node only (decode is exempt). Raising this to
|
||||
32768 fixes the crash.
|
||||
* `MORI_SHMEM_HEAP_SIZE` — Raising the dispatch budget to 32768 tokens exceeds
|
||||
MoRI's default 4 GB static heap and causes an out-of-memory error at first
|
||||
inference. Set this to `16G` (MoRI's own inter-node test default).
|
||||
|
||||
These values are set in the preceding [Configure InfiniBand
|
||||
devices](#configure-infiniBand-devices) step.
|
||||
```
|
||||
|
||||
1. To set the required configuration parameters, update the following
|
||||
environment variables in `run_submit_disagg.sh` to match your cluster setup:
|
||||
|
||||
```bash
|
||||
# SLURM Job Configuration
|
||||
export SLURM_ACCOUNT="amd" # The account name for SLURM job accounting and resource allocation
|
||||
export SLURM_ACCOUNT="<your_slurm_account>" # The account name for SLURM job accounting and resource allocation
|
||||
export SLURM_PARTITION="compute" # The specific cluster partition (queue) to submit the job to
|
||||
export TIME_LIMIT="24:00:00" # Maximum wall time for the job (Hours:Minutes:Seconds)
|
||||
|
||||
# Model Configuration
|
||||
export MODEL_PATH="/nfsdata" # Base directory where the model weights are stored
|
||||
export MODEL_NAME="DeepSeek-R1" # Specific model directory name (joined with MODEL_PATH)
|
||||
export CONTAINER_IMAGE="rocm/sgl-dev:sglang-0.5.6.post1-rocm700-mi35x-mori-1224" # Docker image to use for the environment
|
||||
export CONTAINER_IMAGE="lmsysorg/sglang-rocm:v0.5.12.post1-rocm720-mi35x-20260529" # Docker image to use for the environment
|
||||
|
||||
# Cluster Topology (Disaggregation Setup)
|
||||
export PREFILL_NODES=1 # Number of prefill nodes
|
||||
@@ -827,7 +890,7 @@ Identify and configure the available InfiniBand devices.
|
||||
```{note}
|
||||
The following benchmark utility output is provided for reference only and
|
||||
should not be used to compare performance. See the
|
||||
[InferenceMAX](https://inferencemax.semianalysis.com/) website for validated
|
||||
[InferenceX](https://inferencex.semianalysis.com/) website for validated
|
||||
performance results.
|
||||
```
|
||||
|
||||
@@ -863,7 +926,7 @@ Identify and configure the available InfiniBand devices.
|
||||
|
||||
The following section outlines common issues and their solutions.
|
||||
|
||||
### Bandwidth test fails with error
|
||||
### Bandwidth test failures
|
||||
|
||||
1. Use ROCm-optimized `rdma-perftest`, not the generic `perftest`
|
||||
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -30,8 +30,6 @@ training, fine-tuning, and inference. It leverages popular machine learning fram
|
||||
|
||||
- :doc:`SGLang distributed inference with MoRI <benchmark-docker/sglang-mori-distributed>`
|
||||
|
||||
- :doc:`SGLang distributed inference with Mooncake <benchmark-docker/sglang-distributed>`
|
||||
|
||||
- :doc:`xDiT diffusion inference <xdit-diffusion-inference>`
|
||||
|
||||
- :doc:`Deploying your model <deploy-your-model>`
|
||||
|
||||
@@ -105,8 +105,6 @@ subtrees:
|
||||
title: vLLM distributed inference with MoRI
|
||||
- file: how-to/rocm-for-ai/inference/benchmark-docker/sglang-mori-distributed.md
|
||||
title: SGLang distributed inference with MoRI
|
||||
- file: how-to/rocm-for-ai/inference/benchmark-docker/sglang-distributed.rst
|
||||
title: SGLang distributed inference with Mooncake
|
||||
- file: how-to/rocm-for-ai/inference/xdit-diffusion-inference.rst
|
||||
title: xDiT diffusion inference
|
||||
- file: how-to/rocm-for-ai/inference/deploy-your-model.rst
|
||||
|
||||
Reference in New Issue
Block a user