Model CoverageDiffusionHunyuanVideo CommunityHunyuanVideo-1.5-Diffusers-720p_t2v

HunyuanVideo-1.5-Diffusers-720p_t2v

View as Markdown

HunyuanVideo 1.5 is a 13B parameter text-to-video diffusion model from the Hunyuan community, supporting 720p resolution video generation with flow matching training.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Fine-Tune HunyuanVideo-1.5-Diffusers-720p_t2v

From the repository root, run:

uv run torchrun --nproc-per-node=8 \
examples/diffusion/finetune/finetune.py \
-c examples/diffusion/finetune/hunyuan_t2v_flow.yaml

Choose a Workflow

GoalStart Here
Fine-tune - HunyuanVideo 1.5 with flow matchingUse hunyuan_t2v_flow.yaml.

Model Reference

Model Architecture

PropertyValue
TaskText-to-Video
ArchitectureDiT (Flow Matching)
Parameters13B
Hugging Face Organizationhunyuanvideo-community

Task

  • Text-to-Video (T2V)

Available Models