Ling-mini-2.0
Ling-mini-2.0
Ling 2.0 is the Mixture-of-Experts LLM family from inclusionAI (Ant Group), released under the bailing_moe HF architecture (BailingMoeV2ForCausalLM). The line spans a 16 B mini through a 1 T flagship while sharing the same architecture.
Set up NeMo AutoModel with the latest container or follow the installation instructions.
Fine-Tune Ling-mini-2.0
From the repository root, run LoRA fine-tuning:
A single 80 GB H100 / A100 fits Ling-mini-2.0 in bf16 with the LoRA defaults in the example. Set distributed.ep_size > 1 for multi-GPU expert parallelism on the larger variants.
Choose a Workflow
Model Reference
Model Architecture
BailingMoeV2ForCausalLM(HFmodel_type: "bailing_moe")- GQA attention;
use_qk_norm: true - Half RoPE (
partial_rotary_factor=0.5) - DeepSeek-V3-style routing: sigmoid scoring, per-expert bias, grouped top-k (
n_group=8,topk_group=4) - 1 shared expert at
moe_intermediate_size first_k_dense_replacedense MLP layer(s) at the start of the stack