Text-to-Video
Text-to-Video
Generate videos from text prompts with vLLM-Omni, SGLang, TensorRT-LLM, or FastVideo
Choose a backend for text-to-video generation. See the Diffusion Overview for installation and shared configuration.
vLLM-Omni
SGLang
TensorRT-LLM
FastVideo
Text-to-video generation runs a vLLM-Omni worker with --output-modalities video.
Tested Models
To run a non-default model, pass --model to the launch script:
Launch
Launch using the provided script with Wan-AI/Wan2.1-T2V-1.3B-Diffusers:
Generate a Video
Generate a video via /v1/videos:
The response returns a video URL or base64 data depending on response_format (e.g. {"object": "video", "status": "completed", "data": [{"url": "file:///tmp/dynamo_media/videos/req-abc123.mp4"}]}).
MiniMax-H3
The initial MiniMax-H3 qualification serves text-to-video-and-audio (T2VA) with
one aggregated diffusion worker. The launcher passes --task-type fl2va, which
loads only H3’s FL2VA checkpoint partition. Dynamo’s standard image stays on
its VP9-only media stack, so build the opt-in video-audio overlay to mux H.264
video and AAC audio:
Start a shell with exactly four B200 GPUs visible. The commands pass the tested source revision and image ID into the qualification record:
Launch the profile inside the container. It uses four-way Ulysses and text-encoder tensor
parallelism while keeping VAE patch parallelism at one, so small supported
resolutions such as 448x256 still have at least one tile per participating
rank. Larger resolutions can opt into four-way VAE patch parallelism with
DYN_H3_VAE_PATCH_PARALLEL_SIZE=4:
When an offline cache contains only H3’s FL2VA/ partition, set
DYN_H3_MODEL_PATH to that partition directory. The launcher keeps
DYN_H3_MODEL as the public API model name and passes the local partition as
the engine path.
To use the four-step Dense FastH3 adapter, download and pin the startup adapter:
The launcher fuses this adapter into the checkpoint at load time. It is not a request-time LoRA and cannot be combined with another request-time adapter. The qualification script verifies the adapter SHA-256, requests four inference steps, omits request scheduler-shift overrides, and records request duration and adapter provenance. FastH3 v1 is T2VA-only and does not support CPU or layerwise offload.
To use FastH3 with Variable Sparse Attention (VSA), download a VSA adapter and identify the variant to the launcher and qualification script:
This path requires vLLM-Omni v0.29.0rc1 or later. The video-audio overlay
pins fastvideo-kernel==0.3.5. For a vsa-* adapter, the launcher selects
FASTVIDEO_VSA and uses the portable Triton route, retaining 64 key/value
blocks per query block. Override the top-k value with
DYN_H3_FASTVIDEO_VSA_TOPK. FastH3 VSA supports pure Ulysses sequence
parallelism; ring and all-gather degrees must remain 1. Check the first-forward
worker logs for FASTVIDEO_VSA H3 routing and verify that it does not fall back
to dense SDPA.
Run the supplied end-to-end qualification against the same worker. It requests a 10-second clip, then verifies a non-empty H.264 stream at 24 FPS and a non-silent 32-kHz stereo AAC stream:
The launcher pins model revision 42ed227ee7df40d41602854ae760620d6eb651fe.
Set DYN_H3_MODEL_REVISION to test another revision. Dynamo preserves
unrecognized top-level video fields and the vLLM-Omni adapter merges them into
the pipeline’s extra_args, where the upstream model validates them. Client
libraries that expose an extra_body option can use it to add these top-level
fields; do not send a literal nested extra_body object. Generated responses
report fps and audio_sample_rate for each MP4. The qualification output
records the exact Dynamo revision, container image ID, model revision, package
versions, visible GPU models, request, media checksum, and stream probe data.
First/last-frame FL2VA and reference-driven Ref2VA are not part of this initial qualification. They additionally require typed image, video, and audio input references.
Request Parameters (nvext)
The /v1/videos endpoint also accepts NVIDIA extensions via the nvext field for fine-grained control:
Frames per second.
Number of frames (overrides fps * seconds).
Negative prompt for guidance.
Number of denoising steps.
CFG guidance scale.
Random seed for reproducibility.
The nvext.boundary_ratio and nvext.guidance_scale_2 fields apply to the dual-expert MoE schedule used in image-to-video. See Image-to-Video with vLLM-Omni.
See Also
- Image-to-Video with vLLM-Omni — animate a source image with the same
/v1/videosendpoint - Text-to-Video with SGLang
- Text-to-Video with TensorRT-LLM
- Text-to-Video with FastVideo
- vLLM-Omni Configuration reference