BAGEL-7B-MoT#

Megatron Bridge supports BAGEL pretraining and fine-tuning with text-to-image, image editing, and vision-language objectives through Megatron MIMO. The H100 recipes use Megatron FSDP; the verification records below distinguish completed runs from unsupported or unverified capabilities.

BAGEL requires the tested Megatron-LM dev revision and the bagel extra. The Megatron-LM revision pinned by Bridge main does not support its runtime. Follow the BAGEL training guide for dependency setup, checkpoint initialization, training, and direct inference. See the BAGEL data tutorial for dataset preparation.

Verified configurations#

Choose an exact recorded configuration to see its command and expected result. These selectors are generated from the authoritative verification cards and never synthesize combinations.

Run a configuration#

Choose a workflow, precision, and exact recorded combination. The command and expected result update below.

Import · CPU

× Unsupported
Hardware
not specified
Precision
NOT SPECIFIED
Last verified
—
Exact command

No runnable command is recorded for this status.

Expected result

BAGEL native-checkpoint initialization currently runs as part of GPU model construction and is not registered with the maintained public CPU conversion launcher.

Import · GPU

✓ Verified
Hardware
not specified
Precision
BF16
Last verified
2026-09-01
Exact command
Command
./scripts/conversion/convert.sh import --executor slurm --device gpu --nodes 1 --gpus-per-node 1 --hf-model ByteDance-Seed/BAGEL-7B-MoT --hf-revision 5019f57d168e5816e8f3f701b17cc816bb7cf24b --megatron-path work/model-verification/bagel/hf-to-megatron --torch-dtype bfloat16
Expected result

The public GPU importer consumed the pinned official EMA state, reported 1,223 source tensors consumed and 943 target tensors verified, and wrote a standalone torch_dist checkpoint. A clean one-GPU reload completed without missing or unexpected checkpoint keys; all 56 MoT language-layer input-norm tensors used as regression sentinels exactly matched the official EMA values.

Export · CPU

× Unsupported
Hardware
not specified
Precision
NOT SPECIFIED
Last verified
—
Exact command

No runnable command is recorded for this status.

Expected result

Exporting a BAGEL Megatron checkpoint to a complete reloadable Hugging Face checkpoint is not implemented by the public CPU converter.

Export · GPU

× Unsupported
Hardware
not specified
Precision
NOT SPECIFIED
Last verified
—
Exact command

No runnable command is recorded for this status.

Expected result

Exporting a BAGEL Megatron checkpoint to a complete reloadable Hugging Face checkpoint is not implemented by the public GPU converter.

Pretrain · H100

✓ Verified
Hardware
H100
Precision
BF16
Last verified
2026-09-01
Recorded metrics
Initial loss
0.9585528
Final loss
1.242196
Step time · last 10 avg
16,383.960 ms
Model throughput · last 10 avg
195.540 TFLOP/s/GPU
Token throughput · last 10 avg
2,250.005 tokens/s/GPU
Peak allocated memory
57.060 GiB
Peak reserved memory
70.154 GiB
Exact command
Command
./scripts/training/train.sh --nodes 1 --gpus-per-node 8 --recipe bagel_7b_pretrain_8gpu_h100_bf16_config --mode pretrain --max_steps 10 model.bagel_repo=work/dependencies/Bagel model.model_path=work/model-verification/bagel/native model.vae_path=work/model-verification/bagel/native/ae.safetensors model.native_model_checkpoint=work/model-verification/bagel/native/ema.safetensors model.native_model_seed=336 model.native_world_size=8 model.validate_native_checkpoint_metadata=false model.reference_training_seed=336 model.reference_training_world_size=8 model.reset_reference_training_rng=true dataset.dataset_root=work/data/bagel-wds dataset.bagel_repo=work/dependencies/Bagel dataset.tokenizer_model=work/model-verification/bagel/native dataset.seed=42 dataset.data_seed=42 rng.seed=42 checkpoint.load=null checkpoint.load_optim=false checkpoint.load_rng=false dataset.dataloader_load=null checkpoint.save=work/model-verification/bagel/pretrain/checkpoints dataset.dataloader_save=work/model-verification/bagel/pretrain/checkpoints checkpoint.save_interval=10 checkpoint.async_save=false validation.eval_iters=0 validation.eval_interval=null logger.log_interval=1 logger.log_throughput=true logger.save_config_filepath=work/model-verification/bagel/pretrain/run_config.yaml
Expected result

The clean public-data run completes 10 BF16 optimizer steps on 8 H100s at TP1/PP1/CP1/DP8 and GBS/MBS 8/1. Every step has finite CE, MSE, and total loss with no skipped or NaN iterations. The final step reports CE 0.9156674, MSE 0.3265290, and total loss 1.242196. The final-10 averages include first-step compilation and are 16,383.960 ms and 195.540 model TFLOP/s/GPU; 2,250.005 token slots/GPU/s uses the fixed 36,864 maximum packed sequence slots rather than variable logical tokens. The post-setup configuration and iteration-10 model, optimizer, RNG, and per-DP dataloader state are persisted. A separate same-topology public-launcher load restored step 10 and exited without executing an additional optimizer step.

SFT · H100

â—‹ Unverified
Hardware
H100
Precision
BF16
Last verified
—
Recorded metrics
Initial loss
None
Final loss
None
Step time · last 10 avg
None ms
Model throughput · last 10 avg
None TFLOP/s/GPU
Token throughput · last 10 avg
None tokens/s/GPU
Exact command

No runnable command is recorded for this status.

Expected result

The BAGEL fine-tuning recipe must complete a bounded pinned-data run with finite losses, all required metrics, and a reloadable checkpoint.

Long Context · H100

× Unsupported
Hardware
H100
Precision
NOT SPECIFIED
Last verified
—
Recorded metrics
Initial loss
None
Final loss
None
Step time · last 10 avg
None ms
Model throughput · last 10 avg
None TFLOP/s/GPU
Token throughput · last 10 avg
None tokens/s/GPU
Exact command

No runnable command is recorded for this status.

Expected result

The current BAGEL integration does not expose a maintained packed long-context fine-tuning recipe with context parallelism.

LoRA · H100

× Unsupported
Hardware
H100
Precision
NOT SPECIFIED
Last verified
—
Recorded metrics
Initial loss
None
Final loss
None
Step time · last 10 avg
None ms
Model throughput · last 10 avg
None TFLOP/s/GPU
Token throughput · last 10 avg
None tokens/s/GPU
Exact command

No runnable command is recorded for this status.

Expected result

BAGEL LoRA and other parameter-efficient fine-tuning workflows are not implemented by the current integration.

Pretrain · FSDP · H100

✓ Verified
Hardware
H100
Precision
BF16
Last verified
2026-08-31
Recorded metrics
Initial loss
14.68331
Final loss
10.809916
Step time · last 10 avg
7,697.160 ms
Model throughput · last 10 avg
234.268 TFLOP/s/GPU
Token throughput · last 10 avg
4,789.299 tokens/s/GPU
Peak allocated memory
58.594 GiB
Peak reserved memory
62.919 GiB
Exact command
Command
./scripts/training/train.sh --nodes 4 --gpus-per-node 8 --recipe bagel_7b_pretrain_32gpu_h100_bf16_config --mode pretrain --max_steps 30 model.bagel_repo=work/dependencies/Bagel model.model_path=work/model-verification/bagel/native model.vae_path=work/model-verification/bagel/native/ae.safetensors model.reference_training_seed=1344 model.reference_training_world_size=32 model.reset_reference_training_rng=true dataset.dataset_root=work/data/bagel-wds dataset.bagel_repo=work/dependencies/Bagel dataset.tokenizer_model=work/model-verification/bagel/native dataset.seed=42 dataset.data_seed=42 rng.seed=42 checkpoint.load=work/model-verification/bagel/seed42-mcore-init checkpoint.load_optim=true checkpoint.load_rng=true dataset.dataloader_load=work/model-verification/bagel/seed42-mcore-init ddp.overlap_grad_reduce=false ddp.overlap_param_gather=true ddp.fsdp_double_buffer=false checkpoint.save=null checkpoint.save_interval=0 validation.eval_iters=0 validation.eval_interval=0 logger.log_interval=1 logger.log_throughput=true
Expected result

The 32-GPU TP1/PP1/CP1/DP32 real-data run completes exactly 30 BF16 steps at GBS/MBS 32/1 with one microbatch, block-23 language-model recompute, full vision recompute, Megatron FSDP, finite CE/MSE/total losses, and no skipped or NaN iterations. Steps 21-30 average 7,697.160 ms, 234.268 model TFLOP/s/GPU, 23.687% MFU, and 4,789.299 token slots/GPU/s using 36,864 maximum packed sequence slots and GBS 32 across 32 GPUs. The separately measured logical throughput is 4,428.810 tokens/GPU/s from 1,090,856.3 logical tokens per step. Logical tokens include text, ViT, VAE-latent, and special tokens and exclude physical padding. The variable packed sequence length averages 34,070.15 tokens/GPU over steps 6-30. Peak allocated/reserved memory is 58.594/62.919 GiB. The same steps 6-30 workload reports standalone raw metrics of 10,167.941 ms, 176.936 model TFLOP/s/GPU, and 17.890% MFU under official BAGEL. Do not compute a relative speedup: that run uses full language-model recompute and performs an EMA update, while the Bridge run uses block-23 recompute without EMA.