Model CoverageDiffusionWan AIWan2.1-T2V-1.3B-Diffusers

Wan2.1-T2V-1.3B-Diffusers

View as Markdown

Wan 2.1 is a text-to-video diffusion model from Wan AI, trained with flow matching on a large-scale video dataset. It generates high-quality short video clips from text prompts.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

The 1.3B fine-tuning configuration uses 16 data-parallel ranks. Use the launcher guide to run it across the required ranks.

Fine-Tune Wan2.1-T2V-1.3B-Diffusers

From the repository root, run:

uv run automodel examples/diffusion/finetune/wan2_1_t2v_flow_multinode.yaml --nproc-per-node 8

Choose a Workflow

GoalStart Here
Fine-tune - Wan 2.1 T2V 1.3B with flow matchingUse wan2_1_t2v_flow_multinode.yaml.
Pretrain - Wan 2.1 T2V with flow matchingUse wan2_1_t2v_flow.yaml.

Model Reference

Model Architecture

PropertyValue
TaskText-to-Video
ArchitectureDiT (Flow Matching)
Parameters1.3B
Hugging Face OrganizationWan-AI

Task

  • Text-to-Video (T2V)

Available Models

ModelHF ID
Wan 2.1 T2V 1.3BWan-AI/Wan2.1-T2V-1.3B-Diffusers