nemo_automodel.components.models.hunyuan_image3.pipeline
nemo_automodel.components.models.hunyuan_image3.pipeline
HunyuanImage-3.0 text-to-image sampling with the native transformer.
Follows the release HunyuanImage3Text2ImagePipeline with its default settings: bf16 Gaussian latents, Euler
steps on the sigma schedule linspace(1, 0) shifted by flow_shift, the model fed sigma * 1000,
classifier-free guidance uncond + scale * (cond - uncond) over a [cond, uncond] batch, and the release VAE
decoding under fp16 autocast. The release caches the text keys and values across steps; recomputing them gives the
same result because text tokens never attend to the image.
Module Contents
Classes
Functions
Data
API
Text-to-image sampler around a HunyuanImage3ForCausalMM transformer.
Every rank of a sharded transformer must call it with the same prompt and seed: each denoising step is a collective forward pass.
Parameters:
The (possibly FSDP2 / expert-parallel sharded) HunyuanImage3ForCausalMM.
The release VAE (load_release_vae).
Token ids of the release prompt format.
Shift of the sigma schedule.
Device of the latents and token ids.
Generate one image.
Parameters:
Text prompt.
Seeds the initial latents (on self.device).
Number of Euler steps.
Classifier-free guidance scale; at most 1 disables guidance.
Requested image height, snapped to the release resolution group.
Requested image width, snapped to the release resolution group.
Returns: HunyuanImage3PipelineOutput
HunyuanImage3PipelineOutput holding one PIL image.
Decode [1, channels, h, w] scaled latents to a PIL image, as the release pipeline does.
Long [1 or 2, sequence]: the prompt row, then the <cfg> row when with_uncond.
Add the release VAE, prompt format and flow_shift from the checkpoint in model_dir.
Parameters:
The loaded HunyuanImage3ForCausalMM.
Local release checkpoint directory (remote code, VAE weights, generation config).
Generated images, one per prompt.
Sigma schedule of the release FlowMatchDiscreteScheduler (shift set, reverse=True).
Returns: torch.Tensor
fp32 tensor of shape [num_inference_steps + 1] running from 1 to 0.