nemo_automodel.components.models.hunyuan_image3.pipeline

View as Markdown

HunyuanImage-3.0 text-to-image sampling with the native transformer.

Follows the release HunyuanImage3Text2ImagePipeline with its default settings: bf16 Gaussian latents, Euler steps on the sigma schedule linspace(1, 0) shifted by flow_shift, the model fed sigma * 1000, classifier-free guidance uncond + scale * (cond - uncond) over a [cond, uncond] batch, and the release VAE decoding under fp16 autocast. The release caches the text keys and values across steps; recomputing them gives the same result because text tokens never attend to the image.

Module Contents

Classes

NameDescription
HunyuanImage3PipelineText-to-image sampler around a HunyuanImage3ForCausalMM transformer.
HunyuanImage3PipelineOutputGenerated images, one per prompt.

Functions

NameDescription
flow_sigmasSigma schedule of the release FlowMatchDiscreteScheduler (shift set, reverse=True).

Data

logger

API

class nemo_automodel.components.models.hunyuan_image3.pipeline.HunyuanImage3Pipeline(
transformer: torch.nn.Module,
vae: torch.nn.Module,
flow_shift: float,
device: torch.device
)

Text-to-image sampler around a HunyuanImage3ForCausalMM transformer.

Every rank of a sharded transformer must call it with the same prompt and seed: each denoising step is a collective forward pass.

Parameters:

transformer
torch.nn.Module

The (possibly FSDP2 / expert-parallel sharded) HunyuanImage3ForCausalMM.

vae
torch.nn.Module

The release VAE (load_release_vae).

prompt_tokenizer
HunyuanImage3PromptTokenizer

Token ids of the release prompt format.

flow_shift
float

Shift of the sigma schedule.

device
torch.device

Device of the latents and token ids.

nemo_automodel.components.models.hunyuan_image3.pipeline.HunyuanImage3Pipeline.__call__(
prompt: str,
generator: torch.Generator | None = None,
num_inference_steps: int = 50,
guidance_scale: float = 5.0,
height: int = 1024,
width: int = 1024

Generate one image.

Parameters:

prompt
str

Text prompt.

generator
torch.Generator | NoneDefaults to None

Seeds the initial latents (on self.device).

num_inference_steps
intDefaults to 50

Number of Euler steps.

guidance_scale
floatDefaults to 5.0

Classifier-free guidance scale; at most 1 disables guidance.

height
intDefaults to 1024

Requested image height, snapped to the release resolution group.

width
intDefaults to 1024

Requested image width, snapped to the release resolution group.

Returns: HunyuanImage3PipelineOutput

HunyuanImage3PipelineOutput holding one PIL image.

nemo_automodel.components.models.hunyuan_image3.pipeline.HunyuanImage3Pipeline._decode(
latents: torch.Tensor,
generator: torch.Generator | None
) -> PIL.Image.Image

Decode [1, channels, h, w] scaled latents to a PIL image, as the release pipeline does.

nemo_automodel.components.models.hunyuan_image3.pipeline.HunyuanImage3Pipeline._input_ids(
prompt: str,
height: int,
width: int,
with_uncond: bool
) -> torch.Tensor

Long [1 or 2, sequence]: the prompt row, then the <cfg> row when with_uncond.

classmethod

Add the release VAE, prompt format and flow_shift from the checkpoint in model_dir.

Parameters:

transformer
torch.nn.Module

The loaded HunyuanImage3ForCausalMM.

model_dir
str

Local release checkpoint directory (remote code, VAE weights, generation config).

class nemo_automodel.components.models.hunyuan_image3.pipeline.HunyuanImage3PipelineOutput(
images: list[PIL.Image.Image]
)
Dataclass

Generated images, one per prompt.

images
list[Image]
nemo_automodel.components.models.hunyuan_image3.pipeline.flow_sigmas(
num_inference_steps: int,
flow_shift: float
) -> torch.Tensor

Sigma schedule of the release FlowMatchDiscreteScheduler (shift set, reverse=True).

Returns: torch.Tensor

fp32 tensor of shape [num_inference_steps + 1] running from 1 to 0.

nemo_automodel.components.models.hunyuan_image3.pipeline.logger = logging.getLogger(__name__)