MiMo-V2.6-Pro-RL

View as Markdown

MiMo-V2.6-Pro-RL is Xiaomi’s hybrid-attention Mixture-of-Experts (MoE) model. It uses the registered MiMoV2ForCausalLM model class. When vision_config is not None, the implementation sets self.visual to a MiMoVisionTransformer.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Choose a Workflow

The recipes target GB200 nodes with 4 GPUs per node. Keep each expert-parallel group within one NVLink domain.

GoalStart Here
Fine-tune on packed 64K-token Tulu3 text (32 nodes, EP64 / PP2 / CP8)Use mimo_v2_6_pro_rl_tulu3_packed64k_ep64pp2cp8_100steps.yaml.
Fine-tune on MedPix-VQA (32 nodes, EP64 / PP2)Use mimo_v2_6_pro_rl_medpix_nonpacked4k_ep64pp2_100steps.yaml.
Apply LoRA on MedPix-VQA (8 nodes, EP16 / PP2)Use mimo_v2_6_pro_rl_medpix_nonpacked4k_lora_ep16pp2_100steps.yaml.

Follow the launcher guide to run a multi-node recipe.

Model Reference

Model Architecture

PropertyValue
Model ImplementationMiMoV2ForCausalLM backs MiMo-V2.6-Pro-RL.
Vision EncoderMiMoVisionTransformer is attached as self.visual when vision_config is not None.
Recipe HardwareGB200 nodes with 4 GPUs per node; 32 nodes for full fine-tuning and 8 nodes for LoRA.

Available Models

ModelHugging Face ID
MiMo-V2.6-Pro-RLXiaomiMiMo/MiMo-V2.6-Pro-RL