> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# nemo_automodel.components.flow_matching.adapters.qwen_image_21

Qwen-Image-2.1 model adapter for FlowMatching Pipeline.

Qwen-Image-2.1 is a single-stream, block-causal DiT. Unlike Qwen-Image it:

* consumes 64-channel latents unpatched (one token per 16x16 pixel tile),
* runs text and image tokens through one joint sequence whose layout is
  described by `img_mask` (one slot per 2x2 group of target latent tokens),
* predicts every joint-sequence token; only the trailing target-image tokens
  are supervised.

## Module Contents

### Classes

| Name                                                                                                       | Description                                            |
| ---------------------------------------------------------------------------------------------------------- | ------------------------------------------------------ |
| [`QwenImage21Adapter`](#nemo_automodel-components-flow_matching-adapters-qwen_image_21-QwenImage21Adapter) | Model adapter for Qwen-Image-2.1 text-to-image models. |

### Data

[`_IMG_TOKENS_PER_SLOT`](#nemo_automodel-components-flow_matching-adapters-qwen_image_21-_IMG_TOKENS_PER_SLOT)

### API

```python
class nemo_automodel.components.flow_matching.adapters.qwen_image_21.QwenImage21Adapter()
```

**Bases:** [ModelAdapter](/nemo-automodel/nemo_automodel/components/flow_matching/adapters/base#nemo_automodel-components-flow_matching-adapters-base-ModelAdapter)

Model adapter for Qwen-Image-2.1 text-to-image models.

Supports batch format from multiresolution dataloader:

* image\_latents: \[B, 64, H, W]
* text\_embeddings: Qwen3-VL embeddings \[B, seq\_len, 4096], right-padded
* text\_attention\_mask: optional \[B, seq\_len] bool marking valid text tokens

Qwen-Image-2.1 transformer forward interface:

* hidden\_states: Flattened latents \[B, H\*W, 64]
* encoder\_hidden\_states: Text embeddings \[B, text\_len, 4096]
* timestep: Normalized timesteps \[0, 1]
* img\_shapes: \[\[(1, H, W)]] per sample
* img\_mask: \[B, text\_len + H\*W/4] bool, True at target-image slots

The transformer lays out RoPE from row 0 of `img_mask` for the whole batch, and every text position
(padding included) advances the target image's frame index. Each sample is therefore run as its own
unpadded transformer call, so it sees exactly the positions of single-prompt inference.

```python
nemo_automodel.components.flow_matching.adapters.qwen_image_21.QwenImage21Adapter._build_img_mask(
    batch_size: int,
    text_len: int,
    image_tokens: int,
    device: torch.device
) -> torch.Tensor
```

staticmethod

Text positions first, then one slot per 2x2 group of target latent tokens.

```python
nemo_automodel.components.flow_matching.adapters.qwen_image_21.QwenImage21Adapter._pack_latents(
    latents: torch.Tensor
) -> torch.Tensor
```

staticmethod

Flatten latents from \[B, C, H, W] to \[B, H\*W, C] (no patching in 2.1).

```python
nemo_automodel.components.flow_matching.adapters.qwen_image_21.QwenImage21Adapter._unpack_latents(
    latents: torch.Tensor,
    height: int,
    width: int
) -> torch.Tensor
```

staticmethod

Restore \[B, H\*W, C] token predictions to \[B, C, H, W].

```python
nemo_automodel.components.flow_matching.adapters.qwen_image_21.QwenImage21Adapter.forward(
    model: torch.nn.Module,
    inputs: typing.Dict[str, typing.Any]
) -> torch.Tensor
```

Execute forward pass for Qwen-Image-2.1 model.

Runs one transformer call per sample, trimmed to that sample's prompt length, and returns the
target-image prediction in \[B, C, H, W] format.

The call count is always the local batch size, never the number of distinct prompt lengths: under FSDP
every call issues collectives, so all ranks must make the same number of calls, and the diffusion
sampler gives every rank the same local batch size.

```python
nemo_automodel.components.flow_matching.adapters.qwen_image_21.QwenImage21Adapter.prepare_inputs(
    context: nemo_automodel.components.flow_matching.adapters.base.FlowMatchingContext
) -> typing.Dict[str, typing.Any]
```

Prepare inputs for Qwen-Image-2.1 model from FlowMatchingContext.

Expects 4D image latents: \[B, C, H, W] with even H and W.

```python
nemo_automodel.components.flow_matching.adapters.qwen_image_21._IMG_TOKENS_PER_SLOT = 4
```