BAGEL-7B-MoT

View as Markdown

BAGEL-7B-MoT is a unified multimodal model from ByteDance Seed. It combines a Qwen2 language backbone, a SigLIP-NaViT vision encoder, and mixture-of-transformations layers for mixed understanding and visual-generation training.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Pretrain BAGEL-7B-MoT

From the repository root, run:

uv run automodel --nproc-per-node=8 examples/multimodal_pretrain/bagel/bagel_pretrain.yaml

Choose a Workflow

GoalStart Here
Joint text-understanding and image-generation pretrainingUse bagel_pretrain.yaml. Dataset: BAGEL-style packed multimodal data.
Joint understanding + generation fine-tuningUse bagel_sft.yaml. Dataset: BAGEL-style packed multimodal data.

Model Reference

Model Architecture

PropertyValue
TaskMultimodal Input/Output
ArchitectureBagelForUnifiedMultimodal, BagelForConditionalGeneration
Parameters14B (two 7B towers)
Hugging Face OrganizationByteDance-Seed

Available Models

ModelHF ID
BAGEL-7B-MoTByteDance-Seed/BAGEL-7B-MoT