> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# Wan2.1-T2V-1.3B-Diffusers

> Reference Wan2.1-T2V-1.3B-Diffusers checkpoint and architecture details for NeMo AutoModel, with setup guidance and related training recipes.

[Wan 2.1](https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B-Diffusers) is a text-to-video diffusion model from Wan AI, trained with flow matching on a large-scale video dataset. It generates high-quality short video clips from text prompts.

Set up NeMo AutoModel with the [latest container](https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo-automodel) or follow the [installation instructions](/get-started/installation).

The 1.3B fine-tuning configuration uses 16 data-parallel ranks. Use the
[launcher guide](/job-launchers/slurm-cluster) to run it across the required ranks.

## Fine-Tune Wan2.1-T2V-1.3B-Diffusers

From the repository root, run:

```bash
uv run automodel examples/diffusion/finetune/wan2_1_t2v_flow_multinode.yaml --nproc-per-node 8
```

## Choose a Workflow

| Goal                                            | Start Here                                                                                                                                               |
| ----------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Fine-tune - Wan 2.1 T2V 1.3B with flow matching | Use [wan2\_1\_t2v\_flow\_multinode.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/diffusion/finetune/wan2_1_t2v_flow_multinode.yaml). |
| Pretrain - Wan 2.1 T2V with flow matching       | Use [wan2\_1\_t2v\_flow.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/diffusion/pretrain/wan2_1_t2v_flow.yaml).                      |

## Model Reference

### Model Architecture

| Property                  | Value                                   |
| ------------------------- | --------------------------------------- |
| Task                      | Text-to-Video                           |
| Architecture              | DiT (Flow Matching)                     |
| Parameters                | 1.3B                                    |
| Hugging Face Organization | [Wan-AI](https://huggingface.co/Wan-AI) |

### Task

* Text-to-Video (T2V)

### Available Models

| Model            | HF ID                                                                                         |
| ---------------- | --------------------------------------------------------------------------------------------- |
| Wan 2.1 T2V 1.3B | [`Wan-AI/Wan2.1-T2V-1.3B-Diffusers`](https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B-Diffusers) |

## Related Resources

* [Diffusion Fine-Tuning Guide](/recipes-e2e-examples/diffusion-fine-tuning)
* [Dataset Preparation](/datasets/diffusion-dataset)