Muse Glimmer 30B#
Muse Glimmer 30B is a dense vision-language model with a roughly 2B-parameter vision encoder and a 28B-parameter decoder. Megatron Bridge models the complete checkpoint: the vision tower, adapter, projection, and text decoder all participate in checkpoint conversion.
Verified configurations#
Choose an exact recorded configuration to see its command and expected result. These selectors are generated from the authoritative verification cards and never synthesize combinations.
Run a configuration#
Choose a workflow, precision, and exact recorded combination. The command and expected result update below.
Import · CPU
✓ VerifiedExact command
./scripts/conversion/convert.sh import --executor slurm --device cpu --nodes 1 --hf-model meta-models/Muse-Glimmer-30B --hf-revision f84ecc3a0ea984a4c04542a84269e3d065350a6e --megatron-path work/model-verification/muse-glimmer-30b/cpu-megatron --torch-dtype bfloat16 --overwrite
Expected result
On a CPU-only node with no CUDA runtime, the command builds a native 52-layer HybridModel, completes all 1,228 mappings for 29,776,626,688 parameters, and persists iter_0000000. The checkpoint reloads for CPU export, whose exact audit covers all 1,436 source tensors and 29,776,626,688 elements with zero missing, unexpected, shape, dtype, or value mismatches.
Import · GPU
✓ VerifiedExact command
./scripts/conversion/convert.sh import --executor slurm --device gpu --nodes 1 --gpus-per-node 8 --hf-model meta-models/Muse-Glimmer-30B --hf-revision f84ecc3a0ea984a4c04542a84269e3d065350a6e --megatron-path work/model-verification/muse-glimmer-30b/gpu-megatron --torch-dtype bfloat16 --tp 8 --overwrite
Expected result
The eight-GPU command builds the complete native HybridModel, completes all 1,228 mappings, validates the TP8 distributed state, and persists iter_0000000. The checkpoint reloads at TP8 for strict GPU export; its complete round-trip audit is bitwise exact across all 1,436 tensors and 29,776,626,688 elements.
Export · CPU
✓ VerifiedExact command
./scripts/conversion/convert.sh export --executor slurm --device cpu --nodes 1 --hf-model meta-models/Muse-Glimmer-30B --hf-revision f84ecc3a0ea984a4c04542a84269e3d065350a6e --megatron-path work/model-verification/muse-glimmer-30b/cpu-megatron/iter_0000000 --hf-path work/model-verification/muse-glimmer-30b/cpu-hf-export --torch-dtype bfloat16 --overwrite
Expected result
Strict export exits successfully. All 1,436 BF16 tensors and 29,776,626,688 elements match the pinned source exactly in keys, shapes, dtypes, and values (maximum absolute difference 0). Transformers 5.15.0 reloads the result as MuseGlimmerForConditionalGeneration with no missing, unexpected, mismatched, or error keys and the same parameter count.
Export · GPU
✓ VerifiedExact command
./scripts/conversion/convert.sh export --executor slurm --device gpu --nodes 1 --gpus-per-node 8 --hf-model meta-models/Muse-Glimmer-30B --hf-revision f84ecc3a0ea984a4c04542a84269e3d065350a6e --megatron-path work/model-verification/muse-glimmer-30b/gpu-megatron/iter_0000000 --hf-path work/model-verification/muse-glimmer-30b/gpu-hf-export --torch-dtype bfloat16 --export-weight-dtype bfloat16 --distributed-save --tp 8 --overwrite
Expected result
Strict distributed export exits successfully. Every language, vision, adapter, and projection tensor matches the pinned source bitwise: 1,436 tensors and 29,776,626,688 elements with zero missing, unexpected, shape, dtype, or value mismatches. Transformers 5.15.0 strictly reloads the result with the same class and parameter count.
Pretrain · H100
✓ VerifiedRecorded metrics
- Initial loss
- 12.23977
- Final loss
- 1.4103
- Step time · last 10 avg
- 1,035.340 ms
- Model throughput · last 10 avg
- 42.490 TFLOP/s/GPU
- Token throughput · last 10 avg
- 247.262 tokens/s/GPU
Exact command
./scripts/training/train.sh --nodes 4 --gpus-per-node 8 --recipe muse_glimmer_30b_pretrain_32gpu_h100_bf16_multimodal_config --mode pretrain --max_steps 100 --save_dir work/model-verification/muse-glimmer-30b/pretrain-reference --save_interval 50 checkpoint.load=null checkpoint.finetune=false checkpoint.save_optim=true checkpoint.save_rng=true checkpoint.async_save=false ddp.check_for_large_grads=true rerun_state_machine.check_for_nan_in_loss=true logger.log_interval=1 logger.log_throughput=true logger.tensorboard_dir=null logger.save_config_filepath=work/model-verification/muse-glimmer-30b/pretrain-reference-resolved.yaml
Expected result
On 32 H100 GPUs at TP8/PP2/DP2, the command trains the complete randomly initialized HybridModel for exactly 100 optimizer steps on the revision-pinned public CORD-v2 multimodal dataset. All losses and gradient norms remain finite, with zero skipped and zero NaN iterations. Loss decreases from 12.23977 to 1.410300; steps 91-100 average 1,035.34 ms, 42.49 model TFLOP/s/GPU, and 247.262 token slots/s/GPU. The persisted post-setup config records the exact runtime. Complete step-50 and step-100 checkpoints each contain 32 nonempty distributed shards and 1,550 state entries, including model, optimizer, scheduler, and RNG state, plus metadata, train state, and the native HybridModel run config.
SFT · H100
✓ VerifiedRecorded metrics
- Initial loss
- 0.9772241
- Final loss
- 0.02695329
- Step time · last 10 avg
- 1,826.640 ms
- Model throughput · last 10 avg
- 94.030 TFLOP/s/GPU
- Token throughput · last 10 avg
- 560.592 tokens/s/GPU
Exact command
./scripts/training/train.sh --nodes 4 --gpus-per-node 8 --recipe muse_glimmer_30b_sft_32gpu_h100_bf16_config --mode sft --pretrained_checkpoint work/model-verification/muse-glimmer-30b/gpu-megatron/iter_0000000 --max_steps 100 --save_dir work/model-verification/muse-glimmer-30b/sft-checkpoints --save_interval 100 checkpoint.load=null checkpoint.save_optim=false checkpoint.save_rng=false checkpoint.async_save=false ddp.check_for_large_grads=true rerun_state_machine.check_for_nan_in_loss=true logger.log_interval=1 logger.log_throughput=true logger.tensorboard_dir=null logger.save_config_filepath=work/model-verification/muse-glimmer-30b/sft-resolved.yaml
Expected result
On 32 H100 GPUs at TP8/PP2/DP2, the command loads the complete imported HybridModel and performs exactly 100 full-model image-conditioned SFT optimizer steps on the revision-pinned public CORD-v2 dataset. All losses and gradient norms remain finite, with zero skipped and zero NaN iterations. Loss decreases from 0.9772241 to 0.02695329; steps 91-100 average 1,826.64 ms, 94.03 model TFLOP/s/GPU, and 560.592 token slots/s/GPU. The complete 32-shard step-100 full-model checkpoint, metadata, train state, and resolved config are saved and reloadable.
Long Context · H100
✓ VerifiedRecorded metrics
- Initial loss
- 1.034497
- Final loss
- 0.03382061
- Step time · last 10 avg
- 1,419.910 ms
- Model throughput · last 10 avg
- 16.070 TFLOP/s/GPU
- Token throughput · last 10 avg
- 1,442.345 tokens/s/GPU
Exact command
./scripts/training/train.sh --nodes 4 --gpus-per-node 8 --recipe muse_glimmer_30b_sft_32gpu_h100_bf16_long_context_config --mode sft --pretrained_checkpoint work/model-verification/muse-glimmer-30b/gpu-megatron/iter_0000000 --max_steps 100 --save_dir work/model-verification/muse-glimmer-30b/long-sft-checkpoints --save_interval 100 checkpoint.load=null checkpoint.save_optim=false checkpoint.save_rng=false checkpoint.async_save=false ddp.check_for_large_grads=true rerun_state_machine.check_for_nan_in_loss=true logger.log_interval=1 logger.log_throughput=true logger.tensorboard_dir=null logger.save_config_filepath=work/model-verification/muse-glimmer-30b/long-sft-resolved.yaml
Expected result
On 32 H100 GPUs, the command loads the complete imported HybridModel and completes exactly 100 full-model SFT optimizer steps at sequence length 8192 with in-batch packing, TP1/PP4/CP2, A2A context-parallel communication, and selective core-attention recompute. Every loss and gradient norm is finite, with zero skipped and zero NaN iterations. Loss decreases from 1.034497 to 0.03382061; steps 91-100 average 1,419.91 ms, 16.07 model TFLOP/s/GPU, and 1,442.345 token slots/s/GPU. The complete step-100 full-model checkpoint, metadata, train state, and resolved config are saved and reloadable.
LoRA · H100
✓ VerifiedRecorded metrics
- Initial loss
- 0.7263421
- Final loss
- 0.03298394
- Step time · last 10 avg
- 4,823.070 ms
- Model throughput · last 10 avg
- 283.360 TFLOP/s/GPU
- Token throughput · last 10 avg
- 1,698.503 tokens/s/GPU
Exact command
./scripts/training/train.sh --nodes 1 --gpus-per-node 8 --recipe muse_glimmer_30b_peft_8gpu_h100_bf16_config --mode lora --pretrained_checkpoint work/model-verification/muse-glimmer-30b/gpu-megatron/iter_0000000 --max_steps 100 --save_dir work/model-verification/muse-glimmer-30b/peft-checkpoints --save_interval 100 checkpoint.load=null checkpoint.async_save=false ddp.check_for_large_grads=true rerun_state_machine.check_for_nan_in_loss=true logger.log_interval=1 logger.log_throughput=true logger.tensorboard_dir=null logger.save_config_filepath=work/model-verification/muse-glimmer-30b/peft-resolved.yaml
Expected result
The command loads the complete imported native HybridModel checkpoint, freezes the base model, and inserts LoRA dim 8 / alpha 16 adapters only into the native linear_qkv and linear_proj attention projections. All 100 optimizer-step rows are present exactly once with finite loss, zero skipped iterations, and zero NaN iterations. Loss decreases from 0.7263421 to 0.03298394; steps 91-100 average 4,823.07 ms, 283.36 model TFLOP/s/GPU, and 1,698.503 token slots/s/GPU. The persisted post-setup config records the exact runtime. The complete eight-shard iter_0000100 checkpoint contains all 208 expected adapter tensors, its run config, metadata, and train state. A direct reload restores model and optimizer at step 100, completes finite step 101 with zero skipped or NaN iterations, and saves a distinct complete eight-shard checkpoint.
Benchmark · H100
✓ VerifiedRecorded metrics
- Initial loss
- 12.26882
- Final loss
- 2.442728
- Step time · last 10 avg
- 9,373.060 ms
- Model throughput · last 10 avg
- 428.240 TFLOP/s/GPU
- Token throughput · last 10 avg
- 2,621.983 tokens/s/GPU
Exact command
./scripts/training/train.sh --nodes 4 --gpus-per-node 8 --recipe muse_glimmer_30b_pretrain_32gpu_h100_fp8cs_config --mode pretrain --max_steps 50 checkpoint.async_save=false ddp.check_for_large_grads=true rerun_state_machine.check_for_nan_in_loss=true logger.log_interval=1 logger.log_throughput=true logger.tensorboard_dir=null logger.save_config_filepath=work/model-verification/muse-glimmer-30b/performance-resolved.yaml
Expected result
The exact 32-H100 command completes all 50 fixed-shape mock-token dense decoder steps at TP4/PP4/CP2, a 9|15|15|13 split of all 52 language layers, sequence length 4096, MBS6/GBS192, BF16 model and gradient state, and FP8 current-scaling Transformer GEMMs. The vision and projection towers are frozen and excluded from the fixed-shape decoder workload. Loss remains finite from 12.26882 to 2.442728 with zero skipped or NaN iterations. Steps 41-50 average 9,373.06 ms, 428.24 model TFLOP/s/GPU, and 2,621.983 token slots/s/GPU. The persisted post-setup config records the exact runtime, and the process exits successfully without checkpoint output.