Model CoverageLarge Language ModelsQwenQwen3-Next-80B-A3B-Instruct

Qwen3-Next-80B-A3B-Instruct

View as Markdown

Qwen3-Next is an advanced MoE language model from Alibaba Cloud’s Qwen team designed for high-throughput inference with large total parameter counts and efficient per-token activation.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Fine-Tune Qwen3-Next-80B-A3B-Instruct

This recipe was validated on 4 nodes x 8 GPUs (32 H100s). See the Launcher Guide for multi-node setup.

From the repository root, run:

uv run automodel --nproc-per-node=8 examples/llm_finetune/qwen/qwen3_next_te_deepep.yaml

Choose a Workflow

GoalStart Here
Supervised fine-tuning (SFT) - Qwen3-Next with TE + DeepEPUse qwen3_next_te_deepep.yaml.

Model Reference

Model Architecture

PropertyValue
TaskText Generation (MoE)
ArchitectureQwen3NextForCausalLM
Parameters80B total / 3B active
Hugging Face OrganizationQwen

Available Models

ModelHF ID
Qwen3-Next 80B A3B InstructQwen/Qwen3-Next-80B-A3B-Instruct