diffusiongemma-26B-A4B-it

View as Markdown

DiffusionGemma is a block-diffusion language model from Google. Instead of generating tokens left-to-right, it denoises a fixed-length canvas of tokens in parallel: a causal encoder reads the prompt and a bidirectional decoder iteratively refines the response canvas. The released checkpoint is a Mixture-of-Experts model with 26B total parameters and ~4B active per token.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Fine-Tune diffusiongemma-26B-A4B-it

Install the repository’s locked environment:

uv sync --locked --all-groups --extra all

Generate the GSM8K chat dataset required by both DiffusionGemma recipes. Run this from the repository root; it writes gsm8k_chat_train.jsonl, the path used by both YAML files:

uv run python examples/dllm_sft/prep_gsm8k.py

This recipe was validated with Expert Parallelism (EP=8) on a single 8xH100 node. See the Launcher Guide for multi-node setup.

Then run either recipe:

# Full SFT
uv run automodel examples/dllm_sft/diffusion_gemma_sft.yaml --nproc-per-node 8
# LoRA SFT
uv run automodel examples/dllm_sft/diffusion_gemma_lora.yaml --nproc-per-node 8

Choose a Workflow

GoalStart Here
Full SFT - DiffusionGemma 26B-A4B with FSDP2 + Expert ParallelismUse diffusion_gemma_sft.yaml.
Low-rank adaptation (LoRA) SFT - DiffusionGemma 26B-A4B with FSDP2 + Expert ParallelismUse diffusion_gemma_lora.yaml.

Model Reference

Model Architecture

PropertyValue
TaskText Generation (Block Diffusion, MoE)
ArchitectureDiffusionGemmaForBlockDiffusion
Parameters26B total / ~4B active
Hugging Face Organizationgoogle
  • DiffusionGemmaForBlockDiffusion - block-diffusion MoE (causal prompt encoder + bidirectional canvas decoder).

Available Models

ModelHF ID
DiffusionGemma 26B-A4B-itgoogle/diffusiongemma-26B-A4B-it