> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# MiMo-V2.6-Pro-RL

> Fine-tune MiMo-V2.6-Pro-RL with NeMo AutoModel using the checked-in GB200 recipes for packed Tulu3 text and MedPix-VQA vision-language data.

[MiMo-V2.6-Pro-RL](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL) is Xiaomi's hybrid-attention Mixture-of-Experts (MoE) model. It uses the registered `MiMoV2ForCausalLM` model class. When `vision_config` is not `None`, the implementation sets `self.visual` to a `MiMoVisionTransformer`.

Set up NeMo AutoModel with the [latest container](https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo-automodel) or follow the [installation instructions](/get-started/installation).

## Choose a Workflow

The recipes target GB200 nodes with 4 GPUs per node. Keep each expert-parallel group within one NVLink domain.

| Goal                                                                  | Start Here                                                                                                                                                                                                                            |
| --------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Fine-tune on packed 64K-token Tulu3 text (32 nodes, EP64 / PP2 / CP8) | Use [mimo\_v2\_6\_pro\_rl\_tulu3\_packed64k\_ep64pp2cp8\_100steps.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/llm_finetune/mimo_v2_flash/mimo_v2_6_pro_rl_tulu3_packed64k_ep64pp2cp8_100steps.yaml).            |
| Fine-tune on MedPix-VQA (32 nodes, EP64 / PP2)                        | Use [mimo\_v2\_6\_pro\_rl\_medpix\_nonpacked4k\_ep64pp2\_100steps.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/vlm_finetune/mimo_v2_flash/mimo_v2_6_pro_rl_medpix_nonpacked4k_ep64pp2_100steps.yaml).            |
| Apply LoRA on MedPix-VQA (8 nodes, EP16 / PP2)                        | Use [mimo\_v2\_6\_pro\_rl\_medpix\_nonpacked4k\_lora\_ep16pp2\_100steps.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/vlm_finetune/mimo_v2_flash/mimo_v2_6_pro_rl_medpix_nonpacked4k_lora_ep16pp2_100steps.yaml). |

Follow the [launcher guide](/job-launchers/overview) to run a multi-node recipe.

## Model Reference

### Model Architecture

| Property             | Value                                                                                    |
| -------------------- | ---------------------------------------------------------------------------------------- |
| Model Implementation | `MiMoV2ForCausalLM` backs MiMo-V2.6-Pro-RL.                                              |
| Vision Encoder       | `MiMoVisionTransformer` is attached as `self.visual` when `vision_config` is not `None`. |
| Recipe Hardware      | GB200 nodes with 4 GPUs per node; 32 nodes for full fine-tuning and 8 nodes for LoRA.    |

### Available Models

| Model            | Hugging Face ID                                                                     |
| ---------------- | ----------------------------------------------------------------------------------- |
| MiMo-V2.6-Pro-RL | [`XiaomiMiMo/MiMo-V2.6-Pro-RL`](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL) |

## Related Resources

* [MiMo-V2-Flash](/model-coverage/large-language-models/xiaomimimo/MiMo-V2-Flash)
* [LLM Fine-Tuning Guide](/recipes-e2e-examples/sft-peft)
* [Launcher Guide](/job-launchers/overview)