BAGEL-7B-MoT#
Megatron Bridge supports BAGEL pretraining and fine-tuning with text-to-image, image editing, and vision-language objectives through Megatron MIMO. The H100 recipes use Megatron FSDP; the verification records below distinguish completed runs from unsupported or unverified capabilities.
BAGEL requires the tested Megatron-LM dev revision and the bagel extra.
The Megatron-LM revision pinned by Bridge main does not support its runtime.
Follow the BAGEL training guide
for dependency setup, checkpoint initialization, training, and direct inference.
See the BAGEL data tutorial
for dataset preparation.
Verified configurations#
Choose an exact recorded configuration to see its command and expected result. These selectors are generated from the authoritative verification cards and never synthesize combinations.
Run a configuration#
Choose a workflow, precision, and exact recorded combination. The command and expected result update below.
Import · CPU
× UnsupportedExact command
No runnable command is recorded for this status.
Expected result
BAGEL native-checkpoint initialization currently runs as part of GPU model construction and is not registered with the maintained public CPU conversion launcher.
Import · GPU
✓ VerifiedExact command
./scripts/conversion/convert.sh import --executor slurm --device gpu --nodes 1 --gpus-per-node 1 --hf-model ByteDance-Seed/BAGEL-7B-MoT --hf-revision 5019f57d168e5816e8f3f701b17cc816bb7cf24b --megatron-path work/model-verification/bagel/hf-to-megatron --torch-dtype bfloat16
Expected result
The public GPU importer consumed the pinned official EMA state, reported 1,223 source tensors consumed and 943 target tensors verified, and wrote a standalone torch_dist checkpoint. A clean one-GPU reload completed without missing or unexpected checkpoint keys; all 56 MoT language-layer input-norm tensors used as regression sentinels exactly matched the official EMA values.
Export · CPU
× UnsupportedExact command
No runnable command is recorded for this status.
Expected result
Exporting a BAGEL Megatron checkpoint to a complete reloadable Hugging Face checkpoint is not implemented by the public CPU converter.
Export · GPU
× UnsupportedExact command
No runnable command is recorded for this status.
Expected result
Exporting a BAGEL Megatron checkpoint to a complete reloadable Hugging Face checkpoint is not implemented by the public GPU converter.
Pretrain · H100
✓ VerifiedRecorded metrics
- Initial loss
- 0.9585528
- Final loss
- 1.242196
- Step time · last 10 avg
- 16,383.960 ms
- Model throughput · last 10 avg
- 195.540 TFLOP/s/GPU
- Token throughput · last 10 avg
- 2,250.005 tokens/s/GPU
- Peak allocated memory
- 57.060 GiB
- Peak reserved memory
- 70.154 GiB
Exact command
./scripts/training/train.sh --nodes 1 --gpus-per-node 8 --recipe bagel_7b_pretrain_8gpu_h100_bf16_config --mode pretrain --max_steps 10 model.bagel_repo=work/dependencies/Bagel model.model_path=work/model-verification/bagel/native model.vae_path=work/model-verification/bagel/native/ae.safetensors model.native_model_checkpoint=work/model-verification/bagel/native/ema.safetensors model.native_model_seed=336 model.native_world_size=8 model.validate_native_checkpoint_metadata=false model.reference_training_seed=336 model.reference_training_world_size=8 model.reset_reference_training_rng=true dataset.dataset_root=work/data/bagel-wds dataset.bagel_repo=work/dependencies/Bagel dataset.tokenizer_model=work/model-verification/bagel/native dataset.seed=42 dataset.data_seed=42 rng.seed=42 checkpoint.load=null checkpoint.load_optim=false checkpoint.load_rng=false dataset.dataloader_load=null checkpoint.save=work/model-verification/bagel/pretrain/checkpoints dataset.dataloader_save=work/model-verification/bagel/pretrain/checkpoints checkpoint.save_interval=10 checkpoint.async_save=false validation.eval_iters=0 validation.eval_interval=null logger.log_interval=1 logger.log_throughput=true logger.save_config_filepath=work/model-verification/bagel/pretrain/run_config.yaml
Expected result
The clean public-data run completes 10 BF16 optimizer steps on 8 H100s at TP1/PP1/CP1/DP8 and GBS/MBS 8/1. Every step has finite CE, MSE, and total loss with no skipped or NaN iterations. The final step reports CE 0.9156674, MSE 0.3265290, and total loss 1.242196. The final-10 averages include first-step compilation and are 16,383.960 ms and 195.540 model TFLOP/s/GPU; 2,250.005 token slots/GPU/s uses the fixed 36,864 maximum packed sequence slots rather than variable logical tokens. The post-setup configuration and iteration-10 model, optimizer, RNG, and per-DP dataloader state are persisted. A separate same-topology public-launcher load restored step 10 and exited without executing an additional optimizer step.
SFT · H100
â—‹ UnverifiedRecorded metrics
- Initial loss
- None
- Final loss
- None
- Step time · last 10 avg
- None ms
- Model throughput · last 10 avg
- None TFLOP/s/GPU
- Token throughput · last 10 avg
- None tokens/s/GPU
Exact command
No runnable command is recorded for this status.
Expected result
The BAGEL fine-tuning recipe must complete a bounded pinned-data run with finite losses, all required metrics, and a reloadable checkpoint.
Long Context · H100
× UnsupportedRecorded metrics
- Initial loss
- None
- Final loss
- None
- Step time · last 10 avg
- None ms
- Model throughput · last 10 avg
- None TFLOP/s/GPU
- Token throughput · last 10 avg
- None tokens/s/GPU
Exact command
No runnable command is recorded for this status.
Expected result
The current BAGEL integration does not expose a maintained packed long-context fine-tuning recipe with context parallelism.
LoRA · H100
× UnsupportedRecorded metrics
- Initial loss
- None
- Final loss
- None
- Step time · last 10 avg
- None ms
- Model throughput · last 10 avg
- None TFLOP/s/GPU
- Token throughput · last 10 avg
- None tokens/s/GPU
Exact command
No runnable command is recorded for this status.
Expected result
BAGEL LoRA and other parameter-efficient fine-tuning workflows are not implemented by the current integration.
Pretrain · FSDP · H100
✓ VerifiedRecorded metrics
- Initial loss
- 14.68331
- Final loss
- 10.809916
- Step time · last 10 avg
- 7,697.160 ms
- Model throughput · last 10 avg
- 234.268 TFLOP/s/GPU
- Token throughput · last 10 avg
- 4,789.299 tokens/s/GPU
- Peak allocated memory
- 58.594 GiB
- Peak reserved memory
- 62.919 GiB
Exact command
./scripts/training/train.sh --nodes 4 --gpus-per-node 8 --recipe bagel_7b_pretrain_32gpu_h100_bf16_config --mode pretrain --max_steps 30 model.bagel_repo=work/dependencies/Bagel model.model_path=work/model-verification/bagel/native model.vae_path=work/model-verification/bagel/native/ae.safetensors model.reference_training_seed=1344 model.reference_training_world_size=32 model.reset_reference_training_rng=true dataset.dataset_root=work/data/bagel-wds dataset.bagel_repo=work/dependencies/Bagel dataset.tokenizer_model=work/model-verification/bagel/native dataset.seed=42 dataset.data_seed=42 rng.seed=42 checkpoint.load=work/model-verification/bagel/seed42-mcore-init checkpoint.load_optim=true checkpoint.load_rng=true dataset.dataloader_load=work/model-verification/bagel/seed42-mcore-init ddp.overlap_grad_reduce=false ddp.overlap_param_gather=true ddp.fsdp_double_buffer=false checkpoint.save=null checkpoint.save_interval=0 validation.eval_iters=0 validation.eval_interval=0 logger.log_interval=1 logger.log_throughput=true
Expected result
The 32-GPU TP1/PP1/CP1/DP32 real-data run completes exactly 30 BF16 steps at GBS/MBS 32/1 with one microbatch, block-23 language-model recompute, full vision recompute, Megatron FSDP, finite CE/MSE/total losses, and no skipped or NaN iterations. Steps 21-30 average 7,697.160 ms, 234.268 model TFLOP/s/GPU, 23.687% MFU, and 4,789.299 token slots/GPU/s using 36,864 maximum packed sequence slots and GBS 32 across 32 GPUs. The separately measured logical throughput is 4,428.810 tokens/GPU/s from 1,090,856.3 logical tokens per step. Logical tokens include text, ViT, VAE-latent, and special tokens and exclude physical padding. The variable packed sequence length averages 34,070.15 tokens/GPU over steps 6-30. Peak allocated/reserved memory is 58.594/62.919 GiB. The same steps 6-30 workload reports standalone raw metrics of 10,167.941 ms, 176.936 model TFLOP/s/GPU, and 17.890% MFU under official BAGEL. Do not compute a relative speedup: that run uses full language-model recompute and performs an EMA update, while the Bridge run uses block-23 recompute without EMA.