Introduction
GLM-5.1 is a Mixture-of-Experts (MoE) large language model developed by Z.ai, featuring 744B total parameters with 40B active parameters. It uses 256 routed experts (top-8) plus one shared expert, with Multi-head Latent Attention (MLA) and DeepSeek Sparse Attention (DSA), and a built-in multi-token prediction (MTP) head for speculative decoding. The model features built-in bilingual (Chinese-English) capabilities with a unified pre-training framework, excelling at reasoning, math, code, and tool calling tasks. GLM-5.1 supports both Thinking mode (step-by-step reasoning) and Instruct mode (direct response), with a native context window of approximately 200k tokens. This document demonstrates the deployment of GLM-5.1 on Ascend NPUs using SGLang, including multi-node PD mixed mode, multi-node PD disaggregation mode, feature configuration, and performance optimization. This document is validated and written based on SGLang v0.5.13. The current model (GLM-5.1) is fully supported in this version. To use the latest features (e.g., speculative decoding, multi-node deployment), it is recommended to use v0.5.13 or a later version.Supported features
The values in the Example usage column are for illustration only. Adjust them according to your hardware, deployment
mode, and workload. For parameter details, see
Feature descriptions; for
recommended configurations for each deployment scenario, see Best practices.
Prerequisites
Environment
Before following this tutorial, complete the environment setup in the documents below:- Ascend NPU Quickstart — the fastest way to get started. It walks you through launching the official container image, starting the SGLang server, and sending a test request. Recommended if you are new to SGLang on Ascend.
- SGLang Installation with NPU Support — the full installation guide. It covers the component version mapping (CANN, TorchNPU, Triton, kernels, etc.), building from source or from a Dockerfile, and recommended system settings (CPU power scheme, NUMA, swap). Use it when you need to install or customize the environment instead of using the official image.
Model weights
Before downloading model weights, check the model size to reserve enough disk space. For multi-node deployment, download the weights to a shared directory accessible to all nodes.- GLM-5.1 (BF16, 1.51TB)
- GLM-5.1-w4a8 (Quantized version, 420.17GB)
- You can use msmodelslim to quantize the model naively.
We recommend deploying the W4A8 variant for reduced resource usage and higher throughput.
It (420.17GB) can be deployed on 8 × 64GB of device memory (
--tp-size 8), which corresponds to one full A2 node or 8 dies on A3 (4 cards).Installation
The dependencies required for the NPU runtime environment have been integrated into a Docker image and uploaded to the online platform. You can directly pull it. Both stable releases and daily builds are available. The following command is based on the stable release tag. For details, see Docker image versions.- Atlas 800I A3
- Atlas 800I A2
Command
Online service deployment
Multi-node PD mixed deployment
Multi-node deployment distributes the model across multiple Atlas 800I A3 nodes using tensor parallelism while keeping prefill and decode on the same nodes (PD mixed mode), suitable for scenarios that need more device memory than a single node can provide. This scenario is already covered in the best practice. For the complete, optimized deployment commands and benchmark data, see GLM-5.1 Best Practice — Multi-node PD Mixed On A3.Multi-node PD disaggregation deployment
PD disaggregation splits the prefill and decode stages onto separate nodes, reducing interference and improving throughput for high-concurrency scenarios. This scenario is already covered in the best practice. For the complete, optimized deployment commands and benchmark data, see GLM-5.1 Best Practice — PD Disaggregation On A3.Functional verification
After the service is started, you can invoke the model by sending a prompt:The server is fired up and ready to roll! in the logs, it is ready to accept requests. For more
testing examples (Health Check, Generate, Chat Completions, and port usage guidance),
see Testing the Service.
