MiMo-V2-Flash

View as Markdown

MiMo-V2-Flash is a Mixture-of-Experts (MoE) language model with a hybrid attention architecture. It interleaves sliding-window and global attention.

MiMo-V2.6-Flash-RL uses the registered MiMoV2ForCausalLM model class. When vision_config is not None, the implementation sets self.visual to a MiMoVisionTransformer.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Fine-Tune MiMo-V2-Flash

Use the checked-in recipes to fine-tune either supported checkpoint.

Choose a Workflow

Choose a recipe for your checkpoint and workload. The MiMo-V2-Flash recipe is configured for 16 nodes with 8 H100 GPUs per node. The MiMo-V2.6-Flash-RL recipes target 8 nodes with 8 H100 GPUs per node.

GoalStart Here
Fine-tune MiMo-V2-Flash on HellaSwagUse mimo_v2_flash_hellaswag.yaml.
Fine-tune MiMo-V2.6-Flash-RL on packed Tulu3 textUse mimo_v2_6_flash_rl_tulu3_packed4k_ep64_cp2_100steps.yaml.
Fine-tune MiMo-V2.6-Flash-RL on MedPix-VQAUse mimo_v2_6_flash_rl_medpix_nonpacked4k_ep64_100steps.yaml.

Follow the launcher guide to run a multi-node recipe.

Model Reference

Model Architecture

PropertyValue
Model ImplementationMiMoV2FlashForCausalLM backs MiMo-V2-Flash.
Model ImplementationMiMoV2ForCausalLM backs MiMo-V2.6-Flash-RL.
Sliding-Window AttentionSliding-window attention uses the MiMoV2FlashAttention(is_swa=True) path.
Expert RoutingExpert routing maps scoring_func="sigmoid" to score_func="sigmoid_with_bias"; the recipes set gate_precision=float32.

Available Models

ModelHugging Face ID
MiMo-V2-FlashXiaomiMiMo/MiMo-V2-Flash
MiMo-V2.6-Flash-RLXiaomiMiMo/MiMo-V2.6-Flash-RL