> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# GLM-5

> Fine-tune GLM-5 through GLM-5.3 mixture-of-experts checkpoints with NeMo AutoModel using checked-in MLA and DSA recipes.

[GLM-5](https://huggingface.co/zai-org/GLM-5),
[GLM-5.1](https://huggingface.co/zai-org/GLM-5.1),
[GLM-5.2](https://huggingface.co/zai-org/GLM-5.2), and
[GLM-5.3](https://huggingface.co/zai-org/GLM-5.3) are Z.ai's open-weight
Mixture-of-Experts language models with Multi-head Latent Attention (MLA) and
Dynamic Sparse Attention (DSA). GLM-5.3 uses the same base model and
`GlmMoeDsaForCausalLM` architecture as GLM-5.2, with its gains coming from
post-training.

## Model Reference

### Model Architecture

| Property                   | Value                                                                       |
| -------------------------- | --------------------------------------------------------------------------- |
| **Task**                   | Text Generation                                                             |
| **Architecture**           | `GlmMoeDsaForCausalLM`                                                      |
| **GLM-5.3 decoder**        | 78 layers, 6,144 hidden size, 256 routed experts                            |
| **GLM-5.3 context length** | 1,048,576 tokens in the model configuration; validated here at 4,096 tokens |
| **HF Org**                 | [zai-org](https://huggingface.co/zai-org)                                   |

## Key Features

GLM-5 family models in NeMo AutoModel support:

* **Mixture-of-Experts (MoE)** with 256 routed experts, top-8 routing, and one
  shared expert in the GLM-5.2/5.3 configuration. The first three layers use
  dense feed-forward networks.
* **IndexShare DSA for GLM-5.2 and GLM-5.3**. Shared DSA layers reuse the
  previous full layer's top-k sparse-attention selection.
* **Optional cuDNN DSA and FlashMLA sparse attention** on SM90 or later through
  `backend.attn: cudnn`.
* **Optional TileLang kernels** for the DSA indexer and sparse MLA path through
  `backend.attn: tilelang`.
* **Packed-sequence training and distributed execution** with FSDP2, expert
  parallelism, HybridEP dispatch, and optional context parallelism.

## Available Models

* **GLM-5** (`GlmMoeDsaForCausalLM`)
* **GLM-5.1** (`GlmMoeDsaForCausalLM`): Updated weights
* **GLM-5.2** (`GlmMoeDsaForCausalLM`): IndexShare DSA with cuDNN and
  TileLang backends
* **GLM-5.3** (`GlmMoeDsaForCausalLM`): Same base architecture as GLM-5.2;
  updated post-training

## Example HF Models

| Model   | HF ID                                                       |
| ------- | ----------------------------------------------------------- |
| GLM-5   | [`zai-org/GLM-5`](https://huggingface.co/zai-org/GLM-5)     |
| GLM-5.1 | [`zai-org/GLM-5.1`](https://huggingface.co/zai-org/GLM-5.1) |
| GLM-5.2 | [`zai-org/GLM-5.2`](https://huggingface.co/zai-org/GLM-5.2) |
| GLM-5.3 | [`zai-org/GLM-5.3`](https://huggingface.co/zai-org/GLM-5.3) |

## Example Recipes

| Recipe                                                                                                                                                       | Description                                                                                 |
| ------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------- |
| [glm\_5.3\_tulu3\_4k\_cudnn\_100step.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/llm_finetune/glm/glm_5.3_tulu3_4k_cudnn_100step.yaml) | Full-parameter GLM-5.3 SFT on Tulu3 with packed 4K sequences, cuDNN DSA, HybridEP, and EP64 |
| [glm\_5.2\_tulu3\_32k\_tilelang\_cp8.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/llm_finetune/glm/glm_5.2_tulu3_32k_tilelang_cp8.yaml) | GLM-5.2 SFT on Tulu3 with packed 32K sequences, CP8, TileLang DSA, and EP64                 |
| [glm\_5.2\_tulu3\_4k\_tilelang\_100k.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/llm_finetune/glm/glm_5.2_tulu3_4k_tilelang_100k.yaml) | GLM-5.2 SFT on Tulu3 with packed 4K sequences and TileLang DSA                              |
| [glm\_5.2\_tulu3\_4k\_cudnn\_100k.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/llm_finetune/glm/glm_5.2_tulu3_4k_cudnn_100k.yaml)       | GLM-5.2 SFT on Tulu3 with packed 4K sequences and cuDNN DSA + FlashMLA                      |

## Validated GLM-5.3 Configuration

| Area              | Validated configuration                                                    |
| ----------------- | -------------------------------------------------------------------------- |
| Training          | Full-parameter SFT with FSDP2 and activation checkpointing                 |
| Distributed setup | EP64 / CP1 with HybridEP dispatch and expert `reshard_after_forward: true` |
| Attention         | cuDNN DSA with FlashMLA sparse forward                                     |
| Workload          | `allenai/tulu-3-sft-mixture`, packed 4,096-token sequences                 |
| Batch             | Local batch size 4; global batch size 256                                  |
| Scale             | 32 nodes / 256 GPUs                                                        |
| Duration          | 100 optimizer steps                                                        |

The relevant model and expert-distribution settings are:

```yaml
distributed:
  strategy: fsdp2
  ep_size: 64
  activation_checkpointing: true
  moe:
    reshard_after_forward: true
    wrap_outer_model: false
    ignore_router_for_ac: true

model:
  pretrained_model_name_or_path: zai-org/GLM-5.3
  backend:
    attn: cudnn
    dispatcher: hybridep
    gate_precision: float32
```

## Install and Run

Clone and install NeMo AutoModel from source:

```bash
git clone https://github.com/NVIDIA-NeMo/Automodel.git
cd Automodel
uv sync --locked --all-groups --all-extras
```

The cuDNN backend also requires FlashMLA's `flash_mla_sparse_fwd`, which is not
published as a complete source distribution on PyPI. Install the tested
revision with its submodules:

```bash
git clone --recursive https://github.com/deepseek-ai/FlashMLA.git /tmp/FlashMLA
git -C /tmp/FlashMLA checkout b7643bd54521f563b839b98289b5cd048c062ba2
git -C /tmp/FlashMLA submodule update --init --recursive
uv pip install --no-build-isolation /tmp/FlashMLA
```

The published GLM-5.3 recipe requires 32 nodes with 8 GPUs per node. Launch it
through the cluster launcher from inside the repository:

```bash
uv run automodel examples/llm_finetune/glm/glm_5.3_tulu3_4k_cudnn_100step.yaml \\
  --nproc-per-node=8 \\
  --wandb.enable=true
```

> **Note**
>
> Use the [Slurm Launcher Guide](/job-launchers/slurm-cluster) to configure the
> multi-node launch. The published recipe is a distributed full-model workflow.

## Numerical Validation

### Hugging Face Logit Parity

The first four GLM-5.3 layers were compared with the Hugging Face reference.
Representative loaded weights matched their source tensors exactly.

| Metric                   | Result |
| ------------------------ | -----: |
| Logits relative L2 error | 0.639% |
| Top-1 token agreement    | 96.61% |

### 100-Step Training

The full 32-node / 256-GPU configuration completed all 100 optimizer steps.
Throughput statistics exclude the first 10 warmup steps.

| Metric                       |          Result |
| ---------------------------- | --------------: |
| Step 0 loss                  |         1.33347 |
| Step 99 loss                 |         0.53756 |
| Mean loss, steps 10-99       |         0.55548 |
| Mean throughput, steps 10-99 | 41,664 tokens/s |

The run completed without non-finite metrics or critical rank errors.

## Current Scope

* The published GLM-5.3 path covers full-parameter SFT with cuDNN DSA,
  FlashMLA, HybridEP, and EP64 at 4K sequence length.
* GLM-5.3 context-parallel and TileLang configurations were not validated by
  this recipe. Use the GLM-5.2 recipes above for the established CP8 and
  TileLang paths.
* The cuDNN sparse-attention path requires SM90 or later, cuDNN Frontend, and
  FlashMLA.

See the [LLM Fine-Tuning Guide](/recipes-e2e-examples/sft-peft) and the
[Large MoE Fine-Tuning Guide](/recipes-e2e-examples/large-moe-fine-tuning) for
dataset and training configuration details.

## References

* [zai-org/GLM-5](https://huggingface.co/zai-org/GLM-5)
* [zai-org/GLM-5.1](https://huggingface.co/zai-org/GLM-5.1)
* [zai-org/GLM-5.2](https://huggingface.co/zai-org/GLM-5.2)
* [zai-org/GLM-5.3](https://huggingface.co/zai-org/GLM-5.3)
* [GLM-5 technical report](https://arxiv.org/abs/2602.15763)
* [FlashMLA](https://github.com/deepseek-ai/FlashMLA)