> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# Ling-mini-2.0

> Use Ling-mini-2.0 with NeMo AutoModel for language model fine-tuning, with documented checkpoints, runnable recipes, setup guidance, and model reference details.

[Ling 2.0](https://huggingface.co/collections/inclusionAI/ling-20) is the Mixture-of-Experts LLM family from inclusionAI (Ant Group), released under the `bailing_moe` HF architecture (`BailingMoeV2ForCausalLM`).  The line spans a 16 B mini through a 1 T flagship while sharing the same architecture.

Set up NeMo AutoModel with the [latest container](https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo-automodel) or follow the [installation instructions](/get-started/installation).

## Fine-Tune Ling-mini-2.0

From the repository root, run LoRA fine-tuning:

```bash
automodel examples/llm_finetune/ling/ling_mini_2_0_squad.yaml --nproc-per-node 1
```

A single 80 GB H100 / A100 fits Ling-mini-2.0 in bf16 with the LoRA defaults in the example.  Set `distributed.ep_size > 1` for multi-GPU expert parallelism on the larger variants.

## Choose a Workflow

| Goal                                                        | Start Here                                                                                                                                                                          |
| ----------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Low-rank adaptation (LoRA) SFT - Ling-mini-2.0 on SQuAD     | Use [ling\_mini\_2\_0\_squad.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/llm_finetune/ling/ling_mini_2_0_squad.yaml). Minimum hardware: 2x H100 80GB.         |
| Low-rank adaptation (LoRA) SFT - Ling-mini-2.0 on HellaSwag | Use [ling\_mini\_2\_0\_hellaswag.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/llm_finetune/ling/ling_mini_2_0_hellaswag.yaml). Minimum hardware: 2x H100 80GB. |
| Full SFT - Ling-mini-2.0 on HellaSwag, FSDP2 + EP=8         | Use [ling\_mini\_2\_0\_sft.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/llm_finetune/ling/ling_mini_2_0_sft.yaml). Minimum hardware: 8x H100 80GB.             |

## Model Reference

### Model Architecture

| Property                  | Value                                             |
| ------------------------- | ------------------------------------------------- |
| Task                      | Text Generation (MoE)                             |
| Architecture              | `BailingMoeV2ForCausalLM`                         |
| Parameters                | 16B total                                         |
| Hugging Face Organization | [inclusionAI](https://huggingface.co/inclusionAI) |

* `BailingMoeV2ForCausalLM` (HF `model_type: "bailing_moe"`)
* GQA attention; `use_qk_norm: true`
* Half RoPE (`partial_rotary_factor=0.5`)
* DeepSeek-V3-style routing: sigmoid scoring, per-expert bias, grouped top-k (`n_group=8`, `topk_group=4`)
* 1 shared expert at `moe_intermediate_size`
* `first_k_dense_replace` dense MLP layer(s) at the start of the stack

### Available Models

| Model         | HF ID                                                                           |
| ------------- | ------------------------------------------------------------------------------- |
| Ling-mini-2.0 | [`inclusionAI/Ling-mini-2.0`](https://huggingface.co/inclusionAI/Ling-mini-2.0) |

## Related Resources

* [LLM Fine-Tuning Guide](/recipes-e2e-examples/sft-peft)