nemo_automodel.components.models.muse_glimmer.config
nemo_automodel.components.models.muse_glimmer.config
Configuration for the native AutoModel MuseGlimmer implementation.
Module Contents
Classes
| Name | Description |
|---|---|
MuseGlimmerConfig | Configuration accepting both legacy flat and canonical nested MuseGlimmer checkpoints. |
MuseGlimmerTextConfig | Canonical nested configuration for the MuseGlimmer language backbone. |
MuseGlimmerVisionConfig | Canonical nested configuration for the MuseGlimmer vision tower. |
API
class nemo_automodel.components.models.muse_glimmer.config.MuseGlimmerConfig( hidden_size: int = 6656, num_hidden_layers: int = 52, num_attention_heads: int = 32, num_key_value_heads: int = 2, head_dim: int = 128, intermediate_size: int = 19968, vocab_size: int = 202048, rms_norm_eps: float = 1e-05, post_norm_eps: float = 1e-08, rope_theta: float = 500000.0, max_position_embeddings: int = 16384, use_qk_norm: bool = True, qk_scale_factor: float = 43.7840518911, use_attn_output_gate: bool = True, output_multiplier: float = 0.19611613513818404, output_soft_cap_temp: float | None = 20.0, normalize_tok_embeddings: bool = True, sliding_window: int = 2048, sliding_window_pattern: list[int] | None = None, every_n_layers_nope: int = 4, no_rope_layers: list[int] | None = None, layer_types: list[str] | None = None, vision_latent_dim: int = 1536, vision_output_dim: int = 6144, vision_layers: int = 50, vision_heads: int = 16, vision_mlp_ratio: float = 8960 / 1536, vision_patch_size: int = 14, vision_patch_temporal: int = 2, vision_downsample_factor: int = 2, vision_sparse_attention_factor: int = 4, vision_pos_emb_grid_h: int = 32, vision_pos_emb_grid_w: int = 32, vision_adapter_dim: int = 4096, patch_token_id: int = 200092, image_token_id: int | None = None, video_token_id: int = 200091, vid_start_id: int = 200082, vid_end_id: int = 200083, vid_frame_sep_id: int = 200087, video_num_frames: int = 96, video_sampling_fps: float = 2.0, has_vision: bool = True, hidden_act: str = 'silu', attention_dropout: float = 0.0, attention_bias: bool = False, mlp_bias: bool = False, tie_word_embeddings: bool = False, bos_token_id: int = 200000, eos_token_id: int = 200001, text_config: dict[str, typing.Any] | nemo_automodel.components.models.muse_glimmer.config.MuseGlimmerTextConfig | None = None, vision_config: dict[str, typing.Any] | nemo_automodel.components.models.muse_glimmer.config.MuseGlimmerVisionConfig | None = None, out_hidden_size: int | None = None, projector_hidden_size: int | None = None, projector_hidden_act: str = 'gelu', kwargs: typing.Any = {} )
Bases: PretrainedConfig
Configuration accepting both legacy flat and canonical nested MuseGlimmer checkpoints.
model_type
= 'muse_glimmer'
scale_query_by
sub_configs
class nemo_automodel.components.models.muse_glimmer.config.MuseGlimmerTextConfig( vocab_size: int = 202048, hidden_size: int = 6656, intermediate_size: int = 19968, num_hidden_layers: int = 52, num_attention_heads: int = 32, num_key_value_heads: int = 2, head_dim: int = 128, hidden_activation: str = 'silu', max_position_embeddings: int = 131072, initializer_range: float = 0.02, rms_norm_eps: float = 1e-05, post_norm_eps: float = 1e-08, use_cache: bool = True, attention_bias: bool = False, attention_dropout: float = 0.0, sliding_window: int = 2048, layer_types: list[str] | None = None, layer_rope_theta: list[float] | None = None, rope_parameters: dict[str, typing.Any] | None = None, qk_scale_factor: float = 3.87, output_multiplier: float = 0.19611613513818404, final_logit_softcapping: float | None = 20.0, tie_word_embeddings: bool = False, bos_token_id: int = 200000, eos_token_id: int = 200001, pad_token_id: int | None = None, kwargs: typing.Any = {} )
Bases: PretrainedConfig
Canonical nested configuration for the MuseGlimmer language backbone.
keys_to_ignore_at_inference
= ['past_key_values']
model_type
= 'muse_glimmer_text'
class nemo_automodel.components.models.muse_glimmer.config.MuseGlimmerVisionConfig( hidden_size: int = 1536, intermediate_size: int = 8960, num_hidden_layers: int = 50, num_attention_heads: int = 16, hidden_act: str = 'gelu', patch_size: int = 14, patch_temporal: int = 2, merge_size: int = 2, pos_emb_height: int = 32, pos_emb_width: int = 32, max_position_embeddings: int = 1024, layer_norm_eps: float = 1e-05, layer_types: list[str] | None = None, rope_parameters: dict[str, typing.Any] | None = None, kwargs: typing.Any = {} )
Bases: PretrainedConfig
Canonical nested configuration for the MuseGlimmer vision tower.
model_type
= 'muse_glimmer_vision'