MiMo-V2.5-Pro
MiMo-V2.5-Pro
MiMo-V2.5-Pro is Xiaomi’s hybrid attention Mixture-of-Experts (MoE) language model. It alternates full and sliding-window attention layers, uses a sigmoid_with_bias router with group-limited expert routing, and ships as an FP8 HF checkpoint.
Model Reference
Model Architecture
Available Models
- MiMo-V2.5-Pro: Hybrid full and sliding-window attention with FP8 weights.
Architecture
- The checkpoint declares
MiMoV2ForCausalLM; the recipe selects the dedicatedmimo_v25implementation witharchitectures: [MiMoV25ForCausalLM]. - Sliding-window attention using the
MiMoV2Attention(is_swa=True)path. - MoE blocks use
nemo_automodel.components.moe.layers.MoEwithscore_func="sigmoid_with_bias"andgate_precision=fp32. MiMoV2StateDictAdapterdequantizes the pretrained FP8 checkpoint and restores canonical QKV row order. Unquantized export preserves weight precision; FP8 export is unsupported.
Example HF Models
Example Recipes
Try with NeMo AutoModel
1. Install (full instructions):
2. Clone the repo to get the example recipes:
3. Launch on the cluster with an 18-node allocation and eight GPU workers per node, using the SLURM cluster guide. Set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True in each worker’s environment and select examples/llm_finetune/mimo_v25/mimo_v25_pro_hellaswag.yaml.
The recipe runs a 20-step pretrained training sanity check with validation every five steps and checkpoint writing disabled. It uses BF16 weights and AdamW state; its measured memory requirement applies to this precision choice. The validated run used PP18 / EP8 with expert dispatch confined to each node. A single-node launch cannot run this recipe.
See the Installation Guide and LLM Fine-Tuning Guide.
Fine-Tuning
See the LLM Fine-Tuning Guide.