nemo_automodel.components.models.hunyuan_image3.release
nemo_automodel.components.models.hunyuan_image3.release
The release’s VAE and prompt format, loaded from the checkpoint’s remote code.
The release ships its VAE, tokenizer wrapper and image processor as remote code inside the checkpoint
(trust_remote_code); they are not vendored here. Preprocessing and sampling both go through this module, so the
latents and token sequences a model is trained on are the ones it is sampled with.
Module Contents
Classes
Functions
API
Token ids of the release’s text-to-image sequence around the image span.
Parameters:
The release TokenizerWrapper.
The release HunyuanImage3ImageProcessor (resolution group and image token grid).
image_base_size of the checkpoint config.
sequence_template of the checkpoint’s generation config.
Return the token ids around the image span for one prompt and image size.
The sequence is the release’s gen_image chat template with classifier-free guidance (bot_task auto,
no system prompt), which is also what its generate_image samples with.
Returns: dict[str, torch.Tensor]
prompt_input_ids / uncond_prompt_input_ids: 1D long ids before the image span, ending in
Load the release tokenizer wrapper and image processor from model_dir.
Parameters:
Local checkpoint directory.
The checkpoint config (needs image_base_size).
Snap width x height to the release’s resolution group (33 ratios around image_base_size).
Returns: tuple[int, int]
(width, height) of the closest size the release generates; its <img_ratio_*> token names it.
Build the release VAE (AutoencoderKLConv3D) and load the checkpoint’s vae.* weights.
Parameters:
Local checkpoint directory.
The vae section of the checkpoint config.
Device for the returned VAE.
Returns: torch.nn.Module
The VAE in fp32 and eval mode.