MiMo-V2.6-Pro-RL
MiMo-V2.6-Pro-RL
MiMo-V2.6-Pro-RL is Xiaomi’s hybrid-attention Mixture-of-Experts (MoE) model. It uses the registered MiMoV2ForCausalLM model class. When vision_config is not None, the implementation sets self.visual to a MiMoVisionTransformer.
Set up NeMo AutoModel with the latest container or follow the installation instructions.
Choose a Workflow
The recipes target GB200 nodes with 4 GPUs per node. Keep each expert-parallel group within one NVLink domain.
Follow the launcher guide to run a multi-node recipe.