diffusiongemma-26B-A4B-it
diffusiongemma-26B-A4B-it
DiffusionGemma is a block-diffusion language model from Google. Instead of generating tokens left-to-right, it denoises a fixed-length canvas of tokens in parallel: a causal encoder reads the prompt and a bidirectional decoder iteratively refines the response canvas. The released checkpoint is a Mixture-of-Experts model with 26B total parameters and ~4B active per token.
Set up NeMo AutoModel with the latest container or follow the installation instructions.
Fine-Tune diffusiongemma-26B-A4B-it
Install the repository’s locked environment:
Generate the GSM8K chat dataset required by both DiffusionGemma recipes. Run this from the repository root; it writes gsm8k_chat_train.jsonl, the path used by both YAML files:
This recipe was validated with Expert Parallelism (EP=8) on a single 8xH100 node. See the Launcher Guide for multi-node setup.
Then run either recipe:
Choose a Workflow
Model Reference
Model Architecture
DiffusionGemmaForBlockDiffusion- block-diffusion MoE (causal prompt encoder + bidirectional canvas decoder).