nemo_automodel.components.models.mimo_v25

View as Markdown

Submodules

Package Contents

Classes

NameDescription
MiMoV2ConfigConfiguration for XiaomiMiMo/MiMo-V2.5-Pro.
MiMoV2ForCausalLMNeMo AutoModel causal LM wrapper for MiMo-V2.5-Pro.
MiMoV2Model-

API

class nemo_automodel.components.models.mimo_v25.config.MiMoV2Config(
vocab_size: int = 151936,
hidden_size: int = 4096,
intermediate_size: int = 22016,
num_hidden_layers: int = 32,
num_attention_heads: int = 32,
num_key_value_heads: int = 32,
hidden_act: str = 'silu',
max_position_embeddings: int = 32768,
initializer_range: float = 0.02,
layernorm_epsilon: float = 1e-06,
rms_norm_eps: float | None = None,
use_cache: bool = True,
tie_word_embeddings: bool = False,
rope_theta: float = 10000.0,
rope_scaling: dict | None = None,
attention_dropout: float = 0.0,
attention_bias: bool = False,
attention_value_scale: float | None = None,
head_dim: int | None = None,
v_head_dim: int | None = None,
swa_num_attention_heads: int | None = None,
swa_num_key_value_heads: int | None = None,
swa_head_dim: int | None = None,
swa_v_head_dim: int | None = None,
swa_rope_theta: float | None = None,
sliding_window: int | None = None,
sliding_window_size: int | None = None,
attention_chunk_size: int | None = None,
add_full_attention_sink_bias: bool = False,
add_swa_attention_sink_bias: bool = False,
hybrid_block_size: int | None = None,
hybrid_layer_pattern: list[int] | None = None,
partial_rotary_factor: float = 1.0,
n_routed_experts: int | None = None,
n_shared_experts: int | None = None,
moe_intermediate_size: int | None = None,
num_experts_per_tok: int | None = None,
routed_scaling_factor: float | None = None,
scoring_func: str = 'sigmoid',
topk_method: str = 'noaux_tc',
n_group: int | None = None,
topk_group: int | None = None,
norm_topk_prob: bool = True,
moe_layer_freq: list[int] | None = None,
attention_projection_layout: str = 'split',
torch_dtype: str = 'bfloat16',
kwargs = {}
)

Bases: PretrainedConfig

Configuration for XiaomiMiMo/MiMo-V2.5-Pro.

attribute_map
= {'num_local_experts': 'n_routed_experts'}
head_dim
keys_to_ignore_at_inference
= ['past_key_values']
model_type
= 'mimo_v2'
moe_intermediate_size
rms_norm_eps
sliding_window_size
swa_head_dim
swa_num_attention_heads
swa_num_key_value_heads
swa_rope_theta
swa_v_head_dim
v_head_dim
class nemo_automodel.components.models.mimo_v25.model.MiMoV2ForCausalLM(
kwargs = {}
)

Bases: HFCheckpointingMixin, Module, MoEFSDPSyncMixin

NeMo AutoModel causal LM wrapper for MiMo-V2.5-Pro.

_keep_in_fp32_modules_strict
backend
= backend or BackendConfig()
lm_head
model
state_dict_adapter
tie_word_embeddings_support
TieSupport = TieSupport.UNTIED_ONLY
nemo_automodel.components.models.mimo_v25.model.MiMoV2ForCausalLM.customize_pipeline_stage_modules(
module_names_per_stage: list[list[str]],
layers_prefix: str,
text_model: torch.nn.Module | None = None
) -> list[list[str]]

Keep the SWA rotary embedding on every PP stage.

nemo_automodel.components.models.mimo_v25.model.MiMoV2ForCausalLM.forward(
input_ids: torch.Tensor | None = None,
inputs_embeds: torch.FloatTensor | None = None,
position_ids: torch.LongTensor | None = None,
attention_mask: torch.Tensor | None = None,
padding_mask: torch.Tensor | None = None,
logits_to_keep: typing.Union[int, torch.Tensor] = 0,
output_hidden_states: bool | None = None,
kwargs: typing.Any = {}
) -> transformers.modeling_outputs.CausalLMOutputWithPast

Compute logits or pass hidden states to the next pipeline stage.

Parameters:

input_ids
torch.Tensor | NoneDefaults to None

Token IDs of shape [batch, sequence] on the first stage, or hidden states of shape [batch, sequence, hidden] thereafter.

inputs_embeds
torch.FloatTensor | NoneDefaults to None

Optional embeddings of shape [batch, sequence, hidden].

position_ids
torch.LongTensor | NoneDefaults to None

Optional positions of shape [batch, sequence] or [1, sequence] for broadcasting across the batch.

attention_mask
torch.Tensor | NoneDefaults to None

Optional padding mask of shape [batch, sequence] or additive attention mask of shape [batch, 1, sequence, sequence].

padding_mask
torch.Tensor | NoneDefaults to None

Optional padding indicators of shape [batch, sequence].

logits_to_keep
Union[int, torch.Tensor]Defaults to 0

Number of trailing token positions to retain, or indices of shape [retained_sequence]. Zero retains all positions.

output_hidden_states
bool | NoneDefaults to None

Whether to include the decoder hidden states.

**kwargs
AnyDefaults to {}

Additional decoder arguments, including optional cache_position of shape [sequence].

Returns: CausalLMOutputWithPast

A model output containing logits of shape [batch, retained_sequence,

classmethod
pretrained_model_name_or_path: str,
model_args = (),
kwargs = {}
classmethod
nemo_automodel.components.models.mimo_v25.model.MiMoV2ForCausalLM.get_input_embeddings() -> torch.nn.Embedding
nemo_automodel.components.models.mimo_v25.model.MiMoV2ForCausalLM.get_output_embeddings() -> torch.nn.Linear
nemo_automodel.components.models.mimo_v25.model.MiMoV2ForCausalLM.initialize_weights(
buffer_device: torch.device | None = None,
dtype: torch.dtype = torch.bfloat16
) -> None
nemo_automodel.components.models.mimo_v25.model.MiMoV2ForCausalLM.set_input_embeddings(
value: torch.nn.Embedding
) -> None
nemo_automodel.components.models.mimo_v25.model.MiMoV2ForCausalLM.set_output_embeddings(
new_embeddings: torch.nn.Linear
) -> None

Bases: Module

embed_tokens
has_sliding_layers
layers
norm
rotary_emb
= MiMoV2RotaryEmbedding(config=config, is_swa=False)
swa_rotary_emb
= MiMoV2RotaryEmbedding(config=config, is_swa=True)
nemo_automodel.components.models.mimo_v25.model.MiMoV2Model._build_causal_mask_mapping(
inputs_embeds: torch.Tensor,
attention_mask: torch.Tensor | dict[str, torch.Tensor] | None,
position_ids: torch.Tensor
) -> dict[str, torch.Tensor]

Build full and sliding masks for the uncached training sequence.

Parameters:

inputs_embeds
torch.Tensor

Embeddings of shape [batch, sequence, hidden].

attention_mask
torch.Tensor | dict[str, torch.Tensor] | None

Padding mask of shape [batch, sequence], an additive mask of shape [batch, 1, sequence, sequence], or a mapping of attention types to masks with that four-dimensional layout.

position_ids
torch.Tensor

Positions of shape [batch, sequence] or [1, sequence].

Returns: dict[str, torch.Tensor]

Full and sliding additive masks of shape [batch, 1, sequence, sequence].

nemo_automodel.components.models.mimo_v25.model.MiMoV2Model.forward(
input_ids: torch.Tensor | None = None,
inputs_embeds: torch.FloatTensor | None = None,
position_ids: torch.LongTensor | None = None,
attention_mask: torch.Tensor | None = None,
padding_mask: torch.Tensor | None = None,
cache_position: torch.LongTensor | None = None,
kwargs: typing.Any = {}
) -> torch.Tensor

Run the decoder layers owned by this pipeline stage.

Parameters:

input_ids
torch.Tensor | NoneDefaults to None

Token IDs of shape [batch, sequence] on the embedding stage, or hidden states of shape [batch, sequence, hidden] on later stages.

inputs_embeds
torch.FloatTensor | NoneDefaults to None

Optional embeddings of shape [batch, sequence, hidden].

position_ids
torch.LongTensor | NoneDefaults to None

Optional positions of shape [batch, sequence] or [1, sequence] for broadcasting across the batch.

attention_mask
torch.Tensor | NoneDefaults to None

Optional padding mask of shape [batch, sequence] or additive attention mask of shape [batch, 1, sequence, sequence].

padding_mask
torch.Tensor | NoneDefaults to None

Optional padding indicators of shape [batch, sequence].

cache_position
torch.LongTensor | NoneDefaults to None

Optional token positions of shape [sequence].

**kwargs
AnyDefaults to {}

Unused compatibility arguments.

Returns: torch.Tensor

Hidden states of shape [batch, sequence, hidden], normalized only

nemo_automodel.components.models.mimo_v25.model.MiMoV2Model.init_weights(
buffer_device: torch.device | None = None
) -> None