nemo_automodel.components.models.qwen3_8_flash_next.backend

View as Markdown

Runtime backend configuration owned by Qwen3.8-Flash-Next.

Module Contents

Classes

NameDescription
Qwen3_8_FlashNextBackendConfigExtend shared backend choices with this model’s optional FA4 QSA.

API

class nemo_automodel.components.models.qwen3_8_flash_next.backend.Qwen3_8_FlashNextBackendConfig(
attn: typing.Literal['te', 'sdpa', 'flex', 'eager', 'tilelang', 'cudnn', 'cute'] = BackendConfig.attn,
sparse_attn: typing.Literal['generic', 'msa'] = 'generic',
linear: typing.Literal['torch', 'te', 'quack'] = 'te' if HAVE_TE and torch.c...,
rms_norm: typing.Literal['torch', 'torch_fp32', 'te', 'quack'] = 'torch_fp32',
rope: typing.Literal['torch', 'quack'] = 'torch',
rope_fusion: bool = HAVE_TE and torch.cuda.is_a...,
experts: typing.Literal['torch', 'te', 'gmm', 'torch_mm', 'torch_mm_mxfp8'] = 'torch_mm' if torch.cuda.is...,
dispatcher: typing.Literal['torch', 'deepep', 'hybridep', 'uccl_ep', 'mok'] = 'deepep' if HAVE_DEEP_EP an...,
dispatcher_num_sms: int = 20,
dispatcher_share_token_dispatcher: bool = True,
dispatcher_async_dispatch: bool = False,
enable_deepep: bool | None = None,
fake_balanced_gate: bool = False,
fake_gate_noise: float = 0.0,
enable_hf_state_dict_adapter: bool = True,
enable_fsdp_optimizations: bool = False,
gate_precision: str | torch.dtype | None = None,
compile_attn: bool = False,
compile_situ: bool = False,
compile_norm: bool = False,
shared_expert_overlap: bool = False,
compile_router_weight: bool = False,
benchmark_static_routing: bool = False,
)
Dataclass

Bases: BackendConfig

Extend shared backend choices with this model’s optional FA4 QSA.

attn="cute" selects FA4 SM90 BF16 sparse GQA; attn="flex" selects FlexAttention. CPU execution uses the numerical oracle with either choice. Other values and the environment-dependent default are retained for compatibility with existing BackendConfig callers; CUDA QSA requires "flex" or "cute". All other component settings are inherited.

attn
Literal['te', 'sdpa', 'flex', 'eager', 'tilelang', 'cudnn', 'cute'] = BackendConfig.attn