> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# nemo_automodel.components.models.mimo_v25.config

## Module Contents

### Classes

| Name                                                                             | Description                                 |
| -------------------------------------------------------------------------------- | ------------------------------------------- |
| [`MiMoV2Config`](#nemo_automodel-components-models-mimo_v25-config-MiMoV2Config) | Configuration for XiaomiMiMo/MiMo-V2.5-Pro. |

### Data

[`_MIMOV2_ATTENTION_PROJECTION_LAYOUTS`](#nemo_automodel-components-models-mimo_v25-config-_MIMOV2_ATTENTION_PROJECTION_LAYOUTS)

### API

```python
class nemo_automodel.components.models.mimo_v25.config.MiMoV2Config(
    vocab_size: int = 151936,
    hidden_size: int = 4096,
    intermediate_size: int = 22016,
    num_hidden_layers: int = 32,
    num_attention_heads: int = 32,
    num_key_value_heads: int = 32,
    hidden_act: str = 'silu',
    max_position_embeddings: int = 32768,
    initializer_range: float = 0.02,
    layernorm_epsilon: float = 1e-06,
    rms_norm_eps: float | None = None,
    use_cache: bool = True,
    tie_word_embeddings: bool = False,
    rope_theta: float = 10000.0,
    rope_scaling: dict | None = None,
    attention_dropout: float = 0.0,
    attention_bias: bool = False,
    attention_value_scale: float | None = None,
    head_dim: int | None = None,
    v_head_dim: int | None = None,
    swa_num_attention_heads: int | None = None,
    swa_num_key_value_heads: int | None = None,
    swa_head_dim: int | None = None,
    swa_v_head_dim: int | None = None,
    swa_rope_theta: float | None = None,
    sliding_window: int | None = None,
    sliding_window_size: int | None = None,
    attention_chunk_size: int | None = None,
    add_full_attention_sink_bias: bool = False,
    add_swa_attention_sink_bias: bool = False,
    hybrid_block_size: int | None = None,
    hybrid_layer_pattern: list[int] | None = None,
    partial_rotary_factor: float = 1.0,
    n_routed_experts: int | None = None,
    n_shared_experts: int | None = None,
    moe_intermediate_size: int | None = None,
    num_experts_per_tok: int | None = None,
    routed_scaling_factor: float | None = None,
    scoring_func: str = 'sigmoid',
    topk_method: str = 'noaux_tc',
    n_group: int | None = None,
    topk_group: int | None = None,
    norm_topk_prob: bool = True,
    moe_layer_freq: list[int] | None = None,
    attention_projection_layout: str = 'split',
    torch_dtype: str = 'bfloat16',
    kwargs = {}
)
```

**Bases:** `PretrainedConfig`

Configuration for XiaomiMiMo/MiMo-V2.5-Pro.

**`attribute_map`** `= {'num_local_experts': 'n_routed_experts'}`

---

**`head_dim`**

---

**`keys_to_ignore_at_inference`** `= ['past_key_values']`

---

**`model_type`** `= 'mimo_v2'`

---

**`moe_intermediate_size`**

---

**`rms_norm_eps`**

---

**`sliding_window_size`**

---

**`swa_head_dim`**

---

**`swa_num_attention_heads`**

---

**`swa_num_key_value_heads`**

---

**`swa_rope_theta`**

---

**`swa_v_head_dim`**

---

**`v_head_dim`**

---

```python
nemo_automodel.components.models.mimo_v25.config._MIMOV2_ATTENTION_PROJECTION_LAYOUTS = {'split', 'fused_qkv'}
```