nemo_automodel.components.models.hunyuan_image3.release

View as Markdown

The release’s VAE and prompt format, loaded from the checkpoint’s remote code.

The release ships its VAE, tokenizer wrapper and image processor as remote code inside the checkpoint (trust_remote_code); they are not vendored here. Preprocessing and sampling both go through this module, so the latents and token sequences a model is trained on are the ones it is sampled with.

Module Contents

Classes

NameDescription
HunyuanImage3PromptTokenizerToken ids of the release’s text-to-image sequence around the image span.

Functions

NameDescription
load_release_vaeBuild the release VAE (AutoencoderKLConv3D) and load the checkpoint’s vae.* weights.

API

class nemo_automodel.components.models.hunyuan_image3.release.HunyuanImage3PromptTokenizer(
wrapper: typing.Any,
image_processor: typing.Any,
image_base_size: int,
sequence_template: str
)

Token ids of the release’s text-to-image sequence around the image span.

Parameters:

wrapper
Any

The release TokenizerWrapper.

image_processor
Any

The release HunyuanImage3ImageProcessor (resolution group and image token grid).

image_base_size
int

image_base_size of the checkpoint config.

sequence_template
str

sequence_template of the checkpoint’s generation config.

nemo_automodel.components.models.hunyuan_image3.release.HunyuanImage3PromptTokenizer.__call__(
prompt: str,
height: int,
width: int
) -> dict[str, torch.Tensor]

Return the token ids around the image span for one prompt and image size.

The sequence is the release’s gen_image chat template with classifier-free guidance (bot_task auto, no system prompt), which is also what its generate_image samples with.

Returns: dict[str, torch.Tensor]

prompt_input_ids / uncond_prompt_input_ids: 1D long ids before the image span, ending in

classmethod

Load the release tokenizer wrapper and image processor from model_dir.

Parameters:

model_dir
str

Local checkpoint directory.

config
Any

The checkpoint config (needs image_base_size).

nemo_automodel.components.models.hunyuan_image3.release.HunyuanImage3PromptTokenizer.target_size(
width: int,
height: int
) -> tuple[int, int]

Snap width x height to the release’s resolution group (33 ratios around image_base_size).

Returns: tuple[int, int]

(width, height) of the closest size the release generates; its <img_ratio_*> token names it.

nemo_automodel.components.models.hunyuan_image3.release.load_release_vae(
model_dir: str,
vae_config: dict[str, typing.Any],
device: str | torch.device
) -> torch.nn.Module

Build the release VAE (AutoencoderKLConv3D) and load the checkpoint’s vae.* weights.

Parameters:

model_dir
str

Local checkpoint directory.

vae_config
dict[str, Any]

The vae section of the checkpoint config.

device
str | torch.device

Device for the returned VAE.

Returns: torch.nn.Module

The VAE in fp32 and eval mode.