GLM-5, GLM-5.1, and GLM-5.2#
GLM-5, GLM-5.1, and GLM-5.2 are large sparse MoE language models with Multi-Latent Attention and Dynamic Sparse Attention. Megatron Bridge supports these checkpoints through the shared GLM5Bridge.
Supported Variants#
Variant |
Hugging Face ID |
Notes |
|---|---|---|
GLM-5 |
|
MoE + MLA + DSA architecture |
GLM-5.1 |
|
Same architecture and mapping shape as GLM-5 |
GLM-5.2 |
|
Adds IndexShare-style DSA index reuse settings |
Architecture Notes#
GlmMoeDsaForCausalLMarchitecture with 78 transformer layers.First 3 layers are dense; remaining layers use MoE.
256 routed experts with top-8 routing and one shared expert per MoE layer.
Uses MLA plus DSA indexer parameters (
index_head_dim,index_n_heads,index_topk).Requires
transformers >= 5.2.0.DSA requires the
fast-hadamard-transformCUDA extension and MCore support for the DSA experimental attention variant.
Examples#
For the pinned checkpoint revision, tested topology, commands, and expected results for verified GLM-5.2 workflows, see the GLM-5.2 model verification card. For the GLM-5 conversion wrapper, dependency notes, and architecture constraints, see the GLM-5 examples README.