> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# nemo_automodel.components.models.qwen3_8_flash_next.backend

Runtime backend configuration owned by Qwen3.8-Flash-Next.

## Module Contents

### Classes

| Name                                                                                                                            | Description                                                       |
| ------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------- |
| [`Qwen3_8_FlashNextBackendConfig`](#nemo_automodel-components-models-qwen3_8_flash_next-backend-Qwen3_8_FlashNextBackendConfig) | Extend shared backend choices with this model's optional FA4 QSA. |

### API

```python
class nemo_automodel.components.models.qwen3_8_flash_next.backend.Qwen3_8_FlashNextBackendConfig(
    attn: typing.Literal['te', 'sdpa', 'flex', 'eager', 'tilelang', 'cudnn', 'cute'] = BackendConfig.attn,
    sparse_attn: typing.Literal['generic', 'msa'] = 'generic',
    linear: typing.Literal['torch', 'te', 'quack'] = 'te' if HAVE_TE and torch.c...,
    rms_norm: typing.Literal['torch', 'torch_fp32', 'te', 'quack'] = 'torch_fp32',
    rope: typing.Literal['torch', 'quack'] = 'torch',
    rope_fusion: bool = HAVE_TE and torch.cuda.is_a...,
    experts: typing.Literal['torch', 'te', 'gmm', 'torch_mm', 'torch_mm_mxfp8'] = 'torch_mm' if torch.cuda.is...,
    dispatcher: typing.Literal['torch', 'deepep', 'hybridep', 'uccl_ep', 'mok'] = 'deepep' if HAVE_DEEP_EP an...,
    dispatcher_num_sms: int = 20,
    dispatcher_share_token_dispatcher: bool = True,
    dispatcher_async_dispatch: bool = False,
    mok: nemo_automodel.components.models.common.utils.MoKBackendConfig = MoKBackendConfig(),
    enable_deepep: bool | None = None,
    fake_balanced_gate: bool = False,
    fake_gate_noise: float = 0.0,
    enable_hf_state_dict_adapter: bool = True,
    enable_fsdp_optimizations: bool = False,
    te_fp8: nemo_automodel.components.models.common.utils.TEFp8Config | None = None,
    gate_precision: str | torch.dtype | None = None,
    compile_attn: bool = False,
    compile_situ: bool = False,
    compile_norm: bool = False,
    shared_expert_overlap: bool = False,
    compile_router_weight: bool = False,
    benchmark_static_routing: bool = False,
    cuda_graph: nemo_automodel.components.models.common.utils.CudaGraphConfig = CudaGraphConfig()
)
```

Dataclass

**Bases:** [BackendConfig](/nemo-automodel/nemo_automodel/components/models/common/utils#nemo_automodel-components-models-common-utils-BackendConfig)

Extend shared backend choices with this model's optional FA4 QSA.

`attn="cute"` selects FA4 SM90 BF16 sparse GQA; `attn="flex"` selects
FlexAttention. CPU execution uses the numerical oracle with either choice.
Other values and the environment-dependent default are retained for
compatibility with existing BackendConfig callers; CUDA QSA requires
`"flex"` or `"cute"`. All other component settings are inherited.

**`attn`** `Literal['te', 'sdpa', 'flex', 'eager', 'tilelang', 'cudnn', 'cute'] = BackendConfig.attn`

---