ReferenceFull Library ReferenceNemo AutomodelNemo AutomodelComponentsModelsHunyuan Image3nemo_automodel.components.models.hunyuan_image3.flow_adapter

nemo_automodel.components.models.hunyuan_image3.flow_adapter

View as Markdown

Flow-matching adapter for HunyuanImage-3.0 text-to-image training.

HunyuanImage-3.0 has no separate text encoder: the prompt is part of the transformer’s own token sequence. The preprocessing step stores, per sample, the token ids before the image (<bos> prompt <boi> <img_size_*> <img_ratio_*> <timestep>), the matching unconditional ids used for classifier-free guidance (the prompt replaced by <cfg> tokens of the same length), and the ids after the image (<eoi>). This adapter assembles prefix + <img> * (h*w) + suffix for every sample, right-pads the batch and calls the model.

The pipeline’s convention already matches the release: x_t = (1 - sigma) x_0 + sigma * noise, target noise - x_0 and model timestep sigma * 1000.

Module Contents

Classes

NameDescription
HunyuanImage3AdapterBuilds the joint token sequence around the noisy latents and returns the predicted velocity.

Data

PROMPT_IDS_KEY

PROMPT_SUFFIX_IDS_KEY

UNCOND_PROMPT_IDS_KEY

API

class nemo_automodel.components.models.hunyuan_image3.flow_adapter.HunyuanImage3Adapter(
image_token_id: int = 128006,
pad_token_id: int = 128009
)

Bases: ModelAdapter

Builds the joint token sequence around the noisy latents and returns the predicted velocity.

Parameters:

image_token_id
intDefaults to 128006

Id of the <img> placeholder token.

pad_token_id
intDefaults to 128009

Id used to right-pad sequences of different lengths.

nemo_automodel.components.models.hunyuan_image3.flow_adapter.HunyuanImage3Adapter.forward(
model: torch.nn.Module,
inputs: dict[str, typing.Any]
) -> torch.Tensor

Run the model on prepare_inputs output.

Returns: torch.Tensor

Tensor of shape [batch, channels, height, width]: the predicted velocity noise - x0.

nemo_automodel.components.models.hunyuan_image3.flow_adapter.HunyuanImage3Adapter.prepare_inputs(
) -> dict[str, typing.Any]

Assemble model inputs.

Parameters:

context
FlowMatchingContext

noisy_latents of shape [batch, channels, height, width], timesteps of shape [batch] (sigma * 1000) and a batch holding per-sample 1D long tensors under prompt_input_ids, uncond_prompt_input_ids and prompt_suffix_ids.

Returns: dict[str, Any]

input_ids long [batch, sequence] right-padded with pad_token_id, latents [batch, channels,

nemo_automodel.components.models.hunyuan_image3.flow_adapter.PROMPT_IDS_KEY = 'prompt_input_ids'
nemo_automodel.components.models.hunyuan_image3.flow_adapter.PROMPT_SUFFIX_IDS_KEY = 'prompt_suffix_ids'
nemo_automodel.components.models.hunyuan_image3.flow_adapter.UNCOND_PROMPT_IDS_KEY = 'uncond_prompt_input_ids'