nemo_automodel.components.models.hunyuan_image3.rope
nemo_automodel.components.models.hunyuan_image3.rope
2D rotary position embedding of HunyuanImage-3.0.
Every token gets a (y, x) position. Text tokens use their sequence index for both coordinates, so a text-only
sequence reduces to ordinary 1D RoPE. An image of h x w tokens starting at sequence index L is placed on a
grid centered inside the span it occupies: y = L + (h*w - h) / 2 + row and x = L + (h*w - w) / 2 + col.
Tokens after the image continue from L + h*w.
The rotary frequencies alternate between the two axes: of the head_dim / 2 frequencies, even ones rotate with
y and odd ones with x. The angles are laid out for the half-split rotate_half convention.
Module Contents
Functions
API
Rotate x of shape [batch, heads, seq, head_dim] with [batch, seq, head_dim] tables.
The result is fp32 (the tables are fp32), matching the reference, which rotates and then normalizes q / k in fp32 before casting back.
Per-token (y, x) positions of a batch of sequences that each hold one image.
Parameters:
Padded sequence length.
Long tensor of shape [batch], index of the first image token of every sample; every image
span must fit in seq_len.
Image height in tokens.
Image width in tokens.
Returns: torch.Tensor
Long tensor of shape [batch, seq_len, 2] with the (y, x) position of every token.
Compute fp32 rotary tables from (y, x) positions.
Parameters:
Integer tensor [..., seq, 2] of (y, x) positions.
Attention head dimension; must be divisible by 4.
RoPE base frequency.
Returns: tuple[torch.Tensor, torch.Tensor]
(cos, sin), each fp32 of shape [..., seq, head_dim].
Return [seq_len, 2] positions of a text-only sequence (y = x = index).