nemo_automodel.components.models.hunyuan_image3.flow_adapter
nemo_automodel.components.models.hunyuan_image3.flow_adapter
Flow-matching adapter for HunyuanImage-3.0 text-to-image training.
HunyuanImage-3.0 has no separate text encoder: the prompt is part of the transformer’s own token sequence. The
preprocessing step stores, per sample, the token ids before the image (<bos> prompt <boi> <img_size_*> <img_ratio_*> <timestep>), the matching unconditional ids used for classifier-free guidance (the prompt replaced
by <cfg> tokens of the same length), and the ids after the image (<eoi>). This adapter assembles
prefix + <img> * (h*w) + suffix for every sample, right-pads the batch and calls the model.
The pipeline’s convention already matches the release: x_t = (1 - sigma) x_0 + sigma * noise, target
noise - x_0 and model timestep sigma * 1000.
Module Contents
Classes
Data
API
Bases: ModelAdapter
Builds the joint token sequence around the noisy latents and returns the predicted velocity.
Parameters:
Id of the <img> placeholder token.
Id used to right-pad sequences of different lengths.
Run the model on prepare_inputs output.
Returns: torch.Tensor
Tensor of shape [batch, channels, height, width]: the predicted velocity noise - x0.
Assemble model inputs.
Parameters:
noisy_latents of shape [batch, channels, height, width], timesteps of shape [batch]
(sigma * 1000) and a batch holding per-sample 1D long tensors under prompt_input_ids,
uncond_prompt_input_ids and prompt_suffix_ids.
Returns: dict[str, Any]
input_ids long [batch, sequence] right-padded with pad_token_id, latents [batch, channels,