Muse Glimmer 30B#

Muse Glimmer 30B is a dense vision-language model with a roughly 2B-parameter vision encoder and a 28B-parameter decoder. Megatron Bridge models the complete checkpoint: the vision tower, adapter, projection, and text decoder all participate in checkpoint conversion.

Verified configurations#

Choose an exact recorded configuration to see its command and expected result. These selectors are generated from the authoritative verification cards and never synthesize combinations.

Run a configuration#

Choose a workflow, precision, and exact recorded combination. The command and expected result update below.

Import · CPU

✓ Verified
Hardware
not specified
Precision
BF16
Last verified
2026-08-14
Exact command
Command
./scripts/conversion/convert.sh import --executor slurm --device cpu --nodes 1 --hf-model meta-models/Muse-Glimmer-30B --hf-revision f84ecc3a0ea984a4c04542a84269e3d065350a6e --megatron-path work/model-verification/muse-glimmer-30b/cpu-megatron --torch-dtype bfloat16 --overwrite
Expected result

On a CPU-only node with no CUDA runtime, the command builds a native 52-layer HybridModel, completes all 1,228 mappings for 29,776,626,688 parameters, and persists iter_0000000. The checkpoint reloads for CPU export, whose exact audit covers all 1,436 source tensors and 29,776,626,688 elements with zero missing, unexpected, shape, dtype, or value mismatches.

Import · GPU

✓ Verified
Hardware
not specified
Precision
BF16
Last verified
2026-08-14
Exact command
Command
./scripts/conversion/convert.sh import --executor slurm --device gpu --nodes 1 --gpus-per-node 8 --hf-model meta-models/Muse-Glimmer-30B --hf-revision f84ecc3a0ea984a4c04542a84269e3d065350a6e --megatron-path work/model-verification/muse-glimmer-30b/gpu-megatron --torch-dtype bfloat16 --tp 8 --overwrite
Expected result

The eight-GPU command builds the complete native HybridModel, completes all 1,228 mappings, validates the TP8 distributed state, and persists iter_0000000. The checkpoint reloads at TP8 for strict GPU export; its complete round-trip audit is bitwise exact across all 1,436 tensors and 29,776,626,688 elements.

Export · CPU

✓ Verified
Hardware
not specified
Precision
BF16
Last verified
2026-08-14
Exact command
Command
./scripts/conversion/convert.sh export --executor slurm --device cpu --nodes 1 --hf-model meta-models/Muse-Glimmer-30B --hf-revision f84ecc3a0ea984a4c04542a84269e3d065350a6e --megatron-path work/model-verification/muse-glimmer-30b/cpu-megatron/iter_0000000 --hf-path work/model-verification/muse-glimmer-30b/cpu-hf-export --torch-dtype bfloat16 --overwrite
Expected result

Strict export exits successfully. All 1,436 BF16 tensors and 29,776,626,688 elements match the pinned source exactly in keys, shapes, dtypes, and values (maximum absolute difference 0). Transformers 5.15.0 reloads the result as MuseGlimmerForConditionalGeneration with no missing, unexpected, mismatched, or error keys and the same parameter count.

Export · GPU

✓ Verified
Hardware
not specified
Precision
BF16
Last verified
2026-08-14
Exact command
Command
./scripts/conversion/convert.sh export --executor slurm --device gpu --nodes 1 --gpus-per-node 8 --hf-model meta-models/Muse-Glimmer-30B --hf-revision f84ecc3a0ea984a4c04542a84269e3d065350a6e --megatron-path work/model-verification/muse-glimmer-30b/gpu-megatron/iter_0000000 --hf-path work/model-verification/muse-glimmer-30b/gpu-hf-export --torch-dtype bfloat16 --export-weight-dtype bfloat16 --distributed-save --tp 8 --overwrite
Expected result

Strict distributed export exits successfully. Every language, vision, adapter, and projection tensor matches the pinned source bitwise: 1,436 tensors and 29,776,626,688 elements with zero missing, unexpected, shape, dtype, or value mismatches. Transformers 5.15.0 strictly reloads the result with the same class and parameter count.

Pretrain · H100

✓ Verified
Hardware
H100
Precision
BF16
Last verified
2026-08-12
Recorded metrics
Initial loss
12.23977
Final loss
1.4103
Step time · last 10 avg
1,035.340 ms
Model throughput · last 10 avg
42.490 TFLOP/s/GPU
Token throughput · last 10 avg
247.262 tokens/s/GPU
Exact command
Command
./scripts/training/train.sh --nodes 4 --gpus-per-node 8 --recipe muse_glimmer_30b_pretrain_32gpu_h100_bf16_multimodal_config --mode pretrain --max_steps 100 --save_dir work/model-verification/muse-glimmer-30b/pretrain-reference --save_interval 50 checkpoint.load=null checkpoint.finetune=false checkpoint.save_optim=true checkpoint.save_rng=true checkpoint.async_save=false ddp.check_for_large_grads=true rerun_state_machine.check_for_nan_in_loss=true logger.log_interval=1 logger.log_throughput=true logger.tensorboard_dir=null logger.save_config_filepath=work/model-verification/muse-glimmer-30b/pretrain-reference-resolved.yaml
Expected result

On 32 H100 GPUs at TP8/PP2/DP2, the command trains the complete randomly initialized HybridModel for exactly 100 optimizer steps on the revision-pinned public CORD-v2 multimodal dataset. All losses and gradient norms remain finite, with zero skipped and zero NaN iterations. Loss decreases from 12.23977 to 1.410300; steps 91-100 average 1,035.34 ms, 42.49 model TFLOP/s/GPU, and 247.262 token slots/s/GPU. The persisted post-setup config records the exact runtime. Complete step-50 and step-100 checkpoints each contain 32 nonempty distributed shards and 1,550 state entries, including model, optimizer, scheduler, and RNG state, plus metadata, train state, and the native HybridModel run config.

SFT · H100

✓ Verified
Hardware
H100
Precision
BF16
Last verified
2026-08-12
Recorded metrics
Initial loss
0.9772241
Final loss
0.02695329
Step time · last 10 avg
1,826.640 ms
Model throughput · last 10 avg
94.030 TFLOP/s/GPU
Token throughput · last 10 avg
560.592 tokens/s/GPU
Exact command
Command
./scripts/training/train.sh --nodes 4 --gpus-per-node 8 --recipe muse_glimmer_30b_sft_32gpu_h100_bf16_config --mode sft --pretrained_checkpoint work/model-verification/muse-glimmer-30b/gpu-megatron/iter_0000000 --max_steps 100 --save_dir work/model-verification/muse-glimmer-30b/sft-checkpoints --save_interval 100 checkpoint.load=null checkpoint.save_optim=false checkpoint.save_rng=false checkpoint.async_save=false ddp.check_for_large_grads=true rerun_state_machine.check_for_nan_in_loss=true logger.log_interval=1 logger.log_throughput=true logger.tensorboard_dir=null logger.save_config_filepath=work/model-verification/muse-glimmer-30b/sft-resolved.yaml
Expected result

On 32 H100 GPUs at TP8/PP2/DP2, the command loads the complete imported HybridModel and performs exactly 100 full-model image-conditioned SFT optimizer steps on the revision-pinned public CORD-v2 dataset. All losses and gradient norms remain finite, with zero skipped and zero NaN iterations. Loss decreases from 0.9772241 to 0.02695329; steps 91-100 average 1,826.64 ms, 94.03 model TFLOP/s/GPU, and 560.592 token slots/s/GPU. The complete 32-shard step-100 full-model checkpoint, metadata, train state, and resolved config are saved and reloadable.

Long Context · H100

✓ Verified
Hardware
H100
Precision
BF16
Last verified
2026-08-12
Recorded metrics
Initial loss
1.034497
Final loss
0.03382061
Step time · last 10 avg
1,419.910 ms
Model throughput · last 10 avg
16.070 TFLOP/s/GPU
Token throughput · last 10 avg
1,442.345 tokens/s/GPU
Exact command
Command
./scripts/training/train.sh --nodes 4 --gpus-per-node 8 --recipe muse_glimmer_30b_sft_32gpu_h100_bf16_long_context_config --mode sft --pretrained_checkpoint work/model-verification/muse-glimmer-30b/gpu-megatron/iter_0000000 --max_steps 100 --save_dir work/model-verification/muse-glimmer-30b/long-sft-checkpoints --save_interval 100 checkpoint.load=null checkpoint.save_optim=false checkpoint.save_rng=false checkpoint.async_save=false ddp.check_for_large_grads=true rerun_state_machine.check_for_nan_in_loss=true logger.log_interval=1 logger.log_throughput=true logger.tensorboard_dir=null logger.save_config_filepath=work/model-verification/muse-glimmer-30b/long-sft-resolved.yaml
Expected result

On 32 H100 GPUs, the command loads the complete imported HybridModel and completes exactly 100 full-model SFT optimizer steps at sequence length 8192 with in-batch packing, TP1/PP4/CP2, A2A context-parallel communication, and selective core-attention recompute. Every loss and gradient norm is finite, with zero skipped and zero NaN iterations. Loss decreases from 1.034497 to 0.03382061; steps 91-100 average 1,419.91 ms, 16.07 model TFLOP/s/GPU, and 1,442.345 token slots/s/GPU. The complete step-100 full-model checkpoint, metadata, train state, and resolved config are saved and reloadable.

LoRA · H100

✓ Verified
Hardware
H100
Precision
BF16
Last verified
2026-08-12
Recorded metrics
Initial loss
0.7263421
Final loss
0.03298394
Step time · last 10 avg
4,823.070 ms
Model throughput · last 10 avg
283.360 TFLOP/s/GPU
Token throughput · last 10 avg
1,698.503 tokens/s/GPU
Exact command
Command
./scripts/training/train.sh --nodes 1 --gpus-per-node 8 --recipe muse_glimmer_30b_peft_8gpu_h100_bf16_config --mode lora --pretrained_checkpoint work/model-verification/muse-glimmer-30b/gpu-megatron/iter_0000000 --max_steps 100 --save_dir work/model-verification/muse-glimmer-30b/peft-checkpoints --save_interval 100 checkpoint.load=null checkpoint.async_save=false ddp.check_for_large_grads=true rerun_state_machine.check_for_nan_in_loss=true logger.log_interval=1 logger.log_throughput=true logger.tensorboard_dir=null logger.save_config_filepath=work/model-verification/muse-glimmer-30b/peft-resolved.yaml
Expected result

The command loads the complete imported native HybridModel checkpoint, freezes the base model, and inserts LoRA dim 8 / alpha 16 adapters only into the native linear_qkv and linear_proj attention projections. All 100 optimizer-step rows are present exactly once with finite loss, zero skipped iterations, and zero NaN iterations. Loss decreases from 0.7263421 to 0.03298394; steps 91-100 average 4,823.07 ms, 283.36 model TFLOP/s/GPU, and 1,698.503 token slots/s/GPU. The persisted post-setup config records the exact runtime. The complete eight-shard iter_0000100 checkpoint contains all 208 expected adapter tensors, its run config, metadata, and train state. A direct reload restores model and optimizer at step 100, completes finite step 101 with zero skipped or NaN iterations, and saves a distinct complete eight-shard checkpoint.

Benchmark · H100

✓ Verified
Hardware
H100
Precision
BF16
Last verified
2026-08-12
Recorded metrics
Initial loss
12.26882
Final loss
2.442728
Step time · last 10 avg
9,373.060 ms
Model throughput · last 10 avg
428.240 TFLOP/s/GPU
Token throughput · last 10 avg
2,621.983 tokens/s/GPU
Exact command
Command
./scripts/training/train.sh --nodes 4 --gpus-per-node 8 --recipe muse_glimmer_30b_pretrain_32gpu_h100_fp8cs_config --mode pretrain --max_steps 50 checkpoint.async_save=false ddp.check_for_large_grads=true rerun_state_machine.check_for_nan_in_loss=true logger.log_interval=1 logger.log_throughput=true logger.tensorboard_dir=null logger.save_config_filepath=work/model-verification/muse-glimmer-30b/performance-resolved.yaml
Expected result

The exact 32-H100 command completes all 50 fixed-shape mock-token dense decoder steps at TP4/PP4/CP2, a 9|15|15|13 split of all 52 language layers, sequence length 4096, MBS6/GBS192, BF16 model and gradient state, and FP8 current-scaling Transformer GEMMs. The vision and projection towers are frozen and excluded from the fixed-shape decoder workload. Loss remains finite from 12.26882 to 2.442728 with zero skipped or NaN iterations. Steps 41-50 average 9,373.06 ms, 428.24 model TFLOP/s/GPU, and 2,621.983 token slots/s/GPU. The persisted post-setup config records the exact runtime, and the process exits successfully without checkpoint output.