> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# Kimi-Linear-48B-A3B-Instruct

> Use Kimi-Linear-48B-A3B-Instruct with NeMo AutoModel for language model fine-tuning, with documented checkpoints, runnable recipes, setup, and model reference details.

[Kimi Linear](https://huggingface.co/moonshotai/Kimi-Linear-48B-A3B-Instruct) is a hybrid-attention Mixture-of-Experts language model from Moonshot AI. Most layers use Kimi Delta Attention (KDA), a gated linear-attention variant with a recurrent state, and the remaining layers use full Multi-Head Latent Attention (MLA). NeMo AutoModel ships a native `KimiLinear48BForCausalLM` implementation with expert parallelism, packed sequences, and context parallelism.

Set up NeMo AutoModel with the [latest container](https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo-automodel) or follow the [installation instructions](/get-started/installation).

## Fine-Tune Kimi-Linear-48B-A3B-Instruct

From the repository root, run:

```bash
uv run automodel --nproc-per-node=8 examples/llm_finetune/kimi/kimi_linear_48b_a3b_hellaswag.yaml
```

## Choose a Workflow

| Goal                                                                     | Start Here                                                                                                                                                                   |
| ------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Supervised fine-tuning (SFT): Kimi Linear 48B A3B on HellaSwag with EP=8 | Use [kimi\_linear\_48b\_a3b\_hellaswag.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/llm_finetune/kimi/kimi_linear_48b_a3b_hellaswag.yaml).              |
| Supervised fine-tuning (SFT): 32k packed sequences with CP=8 and EP=8    | Use [kimi\_linear\_48b\_a3b\_longcontext\_cp8.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/llm_finetune/kimi/kimi_linear_48b_a3b_longcontext_cp8.yaml). |

## Model Reference

### Model Architecture

| Property                  | Value                                           |
| ------------------------- | ----------------------------------------------- |
| Task                      | Text Generation (hybrid linear attention, MoE)  |
| Architecture              | `KimiLinear48BForCausalLM`                      |
| Parameters                | 48B total / 3B active                           |
| Hugging Face Organization | [moonshotai](https://huggingface.co/moonshotai) |

* `KimiLinear48BForCausalLM` (`model_type: kimi_linear_48b_a3b`)
* Hybrid attention stack: KDA linear-attention layers interleaved with MLA layers
* MoE feed-forward blocks with sigmoid routing and grouped top-k selection

Moonshot publishes this model and the Kimi K3 text backbone under the same
`model_type: kimi_linear` and the same `architectures: ["KimiLinearForCausalLM"]`, so
neither field identifies the model on its own. NeMo AutoModel gives this implementation a
distinct identity, `kimi_linear_48b_a3b` / `KimiLinear48BForCausalLM`, and leaves
`kimi_linear` to the K3 text config. The example recipes name `KimiLinear48BConfig`
explicitly, which is what a published Moonshot checkpoint needs; checkpoints saved by
NeMo AutoModel already carry the distinct identity and load without that override.

### Supported Parallelism

| Feature                  | Notes                                                                                                                           |
| ------------------------ | ------------------------------------------------------------------------------------------------------------------------------- |
| FSDP2                    | Default sharding strategy for the recipes below                                                                                 |
| Expert parallelism (EP)  | Shards the MoE experts across ranks                                                                                             |
| Context parallelism (CP) | Contiguous per-rank sequence shards. KDA passes its recurrent state from rank to rank, and MLA gathers the compressed KV latent |
| Packed sequences         | Document boundaries from the THD collater are honored by both layer types, including under CP                                   |

### Available Models

| Model                        | HF ID                                                                                                       |
| ---------------------------- | ----------------------------------------------------------------------------------------------------------- |
| Kimi Linear 48B A3B Instruct | [`moonshotai/Kimi-Linear-48B-A3B-Instruct`](https://huggingface.co/moonshotai/Kimi-Linear-48B-A3B-Instruct) |

## Related Resources

* [LLM Fine-Tuning Guide](/recipes-e2e-examples/sft-peft)