MiMo-V2-Flash
MiMo-V2-Flash
MiMo-V2-Flash is a Mixture-of-Experts (MoE) language model with a hybrid attention architecture. It interleaves sliding-window and global attention.
MiMo-V2.6-Flash-RL uses the registered MiMoV2ForCausalLM model class. When vision_config is not None, the implementation sets self.visual to a MiMoVisionTransformer.
Set up NeMo AutoModel with the latest container or follow the installation instructions.
Fine-Tune MiMo-V2-Flash
Use the checked-in recipes to fine-tune either supported checkpoint.
Choose a Workflow
Choose a recipe for your checkpoint and workload. The MiMo-V2-Flash recipe is configured for 16 nodes with 8 H100 GPUs per node. The MiMo-V2.6-Flash-RL recipes target 8 nodes with 8 H100 GPUs per node.
Follow the launcher guide to run a multi-node recipe.