> This page is for version Nightly (default).
> For other versions, use one of these documentation indexes:
> - Nightly (default): https://docs.nvidia.com/nemo/automodel/nightly/llms.txt
> - Latest: https://docs.nvidia.com/nemo/automodel/latest/llms.txt
> - 0.5.0 · 26.06: https://docs.nvidia.com/nemo/automodel/v0.5/llms.txt
> - 0.4.0 · 26.04: https://docs.nvidia.com/nemo/automodel/v0.4/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# nemo_automodel.components.models.mimo_v25

## Submodules

* **[`nemo_automodel.components.models.mimo_v25.config`](/nemo-automodel/nemo_automodel/components/models/mimo_v25/config)**
* **[`nemo_automodel.components.models.mimo_v25.model`](/nemo-automodel/nemo_automodel/components/models/mimo_v25/model)**
* **[`nemo_automodel.components.models.mimo_v25.state_dict_adapter`](/nemo-automodel/nemo_automodel/components/models/mimo_v25/state_dict_adapter)**

## Package Contents

### Classes

| Name                                                                                      | Description                                         |
| ----------------------------------------------------------------------------------------- | --------------------------------------------------- |
| [`MiMoV2Config`](#nemo_automodel-components-models-mimo_v25-config-MiMoV2Config)          | Configuration for XiaomiMiMo/MiMo-V2.5-Pro.         |
| [`MiMoV2ForCausalLM`](#nemo_automodel-components-models-mimo_v25-model-MiMoV2ForCausalLM) | NeMo AutoModel causal LM wrapper for MiMo-V2.5-Pro. |
| [`MiMoV2Model`](#nemo_automodel-components-models-mimo_v25-model-MiMoV2Model)             | -                                                   |

### API

```python
class nemo_automodel.components.models.mimo_v25.config.MiMoV2Config(
    vocab_size: int = 151936,
    hidden_size: int = 4096,
    intermediate_size: int = 22016,
    num_hidden_layers: int = 32,
    num_attention_heads: int = 32,
    num_key_value_heads: int = 32,
    hidden_act: str = 'silu',
    max_position_embeddings: int = 32768,
    initializer_range: float = 0.02,
    layernorm_epsilon: float = 1e-06,
    rms_norm_eps: float | None = None,
    use_cache: bool = True,
    tie_word_embeddings: bool = False,
    rope_theta: float = 10000.0,
    rope_scaling: dict | None = None,
    attention_dropout: float = 0.0,
    attention_bias: bool = False,
    attention_value_scale: float | None = None,
    head_dim: int | None = None,
    v_head_dim: int | None = None,
    swa_num_attention_heads: int | None = None,
    swa_num_key_value_heads: int | None = None,
    swa_head_dim: int | None = None,
    swa_v_head_dim: int | None = None,
    swa_rope_theta: float | None = None,
    sliding_window: int | None = None,
    sliding_window_size: int | None = None,
    attention_chunk_size: int | None = None,
    add_full_attention_sink_bias: bool = False,
    add_swa_attention_sink_bias: bool = False,
    hybrid_block_size: int | None = None,
    hybrid_layer_pattern: list[int] | None = None,
    partial_rotary_factor: float = 1.0,
    n_routed_experts: int | None = None,
    n_shared_experts: int | None = None,
    moe_intermediate_size: int | None = None,
    num_experts_per_tok: int | None = None,
    routed_scaling_factor: float | None = None,
    scoring_func: str = 'sigmoid',
    topk_method: str = 'noaux_tc',
    n_group: int | None = None,
    topk_group: int | None = None,
    norm_topk_prob: bool = True,
    moe_layer_freq: list[int] | None = None,
    attention_projection_layout: str = 'split',
    torch_dtype: str = 'bfloat16',
    kwargs = {}
)
```

**Bases:** `PretrainedConfig`

Configuration for XiaomiMiMo/MiMo-V2.5-Pro.

**`attribute_map`** `= {'num_local_experts': 'n_routed_experts'}`

---

**`head_dim`**

---

**`keys_to_ignore_at_inference`** `= ['past_key_values']`

---

**`model_type`** `= 'mimo_v2'`

---

**`moe_intermediate_size`**

---

**`rms_norm_eps`**

---

**`sliding_window_size`**

---

**`swa_head_dim`**

---

**`swa_num_attention_heads`**

---

**`swa_num_key_value_heads`**

---

**`swa_rope_theta`**

---

**`swa_v_head_dim`**

---

**`v_head_dim`**

---

```python
class nemo_automodel.components.models.mimo_v25.model.MiMoV2ForCausalLM(
    config: nemo_automodel.components.models.mimo_v25.config.MiMoV2Config,
    moe_config: nemo_automodel.components.moe.config.MoEConfig | None = None,
    backend: nemo_automodel.components.models.common.BackendConfig | None = None,
    kwargs = {}
)
```

**Bases:** [HFCheckpointingMixin](/nemo-automodel/nemo_automodel/components/models/common/hf_checkpointing_mixin#nemo_automodel-components-models-common-hf_checkpointing_mixin-HFCheckpointingMixin), `Module`, [MoEFSDPSyncMixin](/nemo-automodel/nemo_automodel/components/moe/fsdp_mixin#nemo_automodel-components-moe-fsdp_mixin-MoEFSDPSyncMixin)

NeMo AutoModel causal LM wrapper for MiMo-V2.5-Pro.

**`_keep_in_fp32_modules_strict`**

---

**`backend`** `= backend or BackendConfig()`

---

**`lm_head`**

---

**`model`**

---

**`state_dict_adapter`**

---

**`tie_word_embeddings_support`** `TieSupport = TieSupport.UNTIED_ONLY`

---

```python
nemo_automodel.components.models.mimo_v25.model.MiMoV2ForCausalLM.customize_pipeline_stage_modules(
    module_names_per_stage: list[list[str]],
    layers_prefix: str,
    text_model: torch.nn.Module | None = None
) -> list[list[str]]
```

Keep the SWA rotary embedding on every PP stage.

```python
nemo_automodel.components.models.mimo_v25.model.MiMoV2ForCausalLM.forward(
    input_ids: torch.Tensor | None = None,
    inputs_embeds: torch.FloatTensor | None = None,
    position_ids: torch.LongTensor | None = None,
    attention_mask: torch.Tensor | None = None,
    padding_mask: torch.Tensor | None = None,
    logits_to_keep: typing.Union[int, torch.Tensor] = 0,
    output_hidden_states: bool | None = None,
    kwargs: typing.Any = {}
) -> transformers.modeling_outputs.CausalLMOutputWithPast
```

Compute logits or pass hidden states to the next pipeline stage.

**Parameters:**

**`input_ids`** `torch.Tensor | None` — default: None

Token IDs of shape \[batch, sequence] on the first stage,
or hidden states of shape \[batch, sequence, hidden] thereafter.

---

**`inputs_embeds`** `torch.FloatTensor | None` — default: None

Optional embeddings of shape \[batch, sequence, hidden].

---

**`position_ids`** `torch.LongTensor | None` — default: None

Optional positions of shape \[batch, sequence] or
\[1, sequence] for broadcasting across the batch.

---

**`attention_mask`** `torch.Tensor | None` — default: None

Optional padding mask of shape \[batch, sequence] or
additive attention mask of shape \[batch, 1, sequence, sequence].

---

**`padding_mask`** `torch.Tensor | None` — default: None

Optional padding indicators of shape \[batch, sequence].

---

**`logits_to_keep`** `Union[int, torch.Tensor]` — default: 0

Number of trailing token positions to retain, or
indices of shape \[retained\_sequence]. Zero retains all positions.

---

**`output_hidden_states`** `bool | None` — default: None

Whether to include the decoder hidden states.

---

**`**kwargs`** `Any` — default: \{}

Additional decoder arguments, including optional
cache\_position of shape \[sequence].

---

**Returns:** `CausalLMOutputWithPast`

A model output containing logits of shape \[batch, retained\_sequence,

```python
nemo_automodel.components.models.mimo_v25.model.MiMoV2ForCausalLM.from_config(
    config: nemo_automodel.components.models.mimo_v25.config.MiMoV2Config,
    moe_config: nemo_automodel.components.moe.config.MoEConfig | None = None,
    backend: nemo_automodel.components.models.common.BackendConfig | None = None,
    kwargs = {}
) -> nemo_automodel.components.models.mimo_v25.model.MiMoV2ForCausalLM
```

classmethod

```python
nemo_automodel.components.models.mimo_v25.model.MiMoV2ForCausalLM.from_pretrained(
    pretrained_model_name_or_path: str,
    model_args = (),
    kwargs = {}
) -> nemo_automodel.components.models.mimo_v25.model.MiMoV2ForCausalLM
```

classmethod

```python
nemo_automodel.components.models.mimo_v25.model.MiMoV2ForCausalLM.get_input_embeddings() -> torch.nn.Embedding
```

```python
nemo_automodel.components.models.mimo_v25.model.MiMoV2ForCausalLM.get_output_embeddings() -> torch.nn.Linear
```

```python
nemo_automodel.components.models.mimo_v25.model.MiMoV2ForCausalLM.initialize_weights(
    buffer_device: torch.device | None = None,
    dtype: torch.dtype = torch.bfloat16
) -> None
```

```python
nemo_automodel.components.models.mimo_v25.model.MiMoV2ForCausalLM.set_input_embeddings(
    value: torch.nn.Embedding
) -> None
```

```python
nemo_automodel.components.models.mimo_v25.model.MiMoV2ForCausalLM.set_output_embeddings(
    new_embeddings: torch.nn.Linear
) -> None
```

```python
class nemo_automodel.components.models.mimo_v25.model.MiMoV2Model(
    config: nemo_automodel.components.models.mimo_v25.config.MiMoV2Config,
    moe_config: nemo_automodel.components.moe.config.MoEConfig,
    backend: nemo_automodel.components.models.common.BackendConfig
)
```

**Bases:** `Module`

**`embed_tokens`**

---

**`has_sliding_layers`**

---

**`layers`**

---

**`norm`**

---

**`rotary_emb`** `= MiMoV2RotaryEmbedding(config=config, is_swa=False)`

---

**`swa_rotary_emb`** `= MiMoV2RotaryEmbedding(config=config, is_swa=True)`

---

```python
nemo_automodel.components.models.mimo_v25.model.MiMoV2Model._build_causal_mask_mapping(
    inputs_embeds: torch.Tensor,
    attention_mask: torch.Tensor | dict[str, torch.Tensor] | None,
    position_ids: torch.Tensor
) -> dict[str, torch.Tensor]
```

Build full and sliding masks for the uncached training sequence.

**Parameters:**

**`inputs_embeds`** `torch.Tensor`

Embeddings of shape \[batch, sequence, hidden].

---

**`attention_mask`** `torch.Tensor | dict[str, torch.Tensor] | None`

Padding mask of shape \[batch, sequence], an additive
mask of shape \[batch, 1, sequence, sequence], or a mapping of
attention types to masks with that four-dimensional layout.

---

**`position_ids`** `torch.Tensor`

Positions of shape \[batch, sequence] or \[1, sequence].

---

**Returns:** `dict[str, torch.Tensor]`

Full and sliding additive masks of shape \[batch, 1, sequence, sequence].

```python
nemo_automodel.components.models.mimo_v25.model.MiMoV2Model.forward(
    input_ids: torch.Tensor | None = None,
    inputs_embeds: torch.FloatTensor | None = None,
    position_ids: torch.LongTensor | None = None,
    attention_mask: torch.Tensor | None = None,
    padding_mask: torch.Tensor | None = None,
    cache_position: torch.LongTensor | None = None,
    kwargs: typing.Any = {}
) -> torch.Tensor
```

Run the decoder layers owned by this pipeline stage.

**Parameters:**

**`input_ids`** `torch.Tensor | None` — default: None

Token IDs of shape \[batch, sequence] on the embedding
stage, or hidden states of shape \[batch, sequence, hidden] on
later stages.

---

**`inputs_embeds`** `torch.FloatTensor | None` — default: None

Optional embeddings of shape \[batch, sequence, hidden].

---

**`position_ids`** `torch.LongTensor | None` — default: None

Optional positions of shape \[batch, sequence] or
\[1, sequence] for broadcasting across the batch.

---

**`attention_mask`** `torch.Tensor | None` — default: None

Optional padding mask of shape \[batch, sequence] or
additive attention mask of shape \[batch, 1, sequence, sequence].

---

**`padding_mask`** `torch.Tensor | None` — default: None

Optional padding indicators of shape \[batch, sequence].

---

**`cache_position`** `torch.LongTensor | None` — default: None

Optional token positions of shape \[sequence].

---

**`**kwargs`** `Any` — default: \{}

Unused compatibility arguments.

---

**Returns:** `torch.Tensor`

Hidden states of shape \[batch, sequence, hidden], normalized only

```python
nemo_automodel.components.models.mimo_v25.model.MiMoV2Model.init_weights(
    buffer_device: torch.device | None = None
) -> None
```