Kimi-Linear-48B-A3B-Instruct

View as Markdown

Kimi Linear is a hybrid-attention Mixture-of-Experts language model from Moonshot AI. Most layers use Kimi Delta Attention (KDA), a gated linear-attention variant with a recurrent state, and the remaining layers use full Multi-Head Latent Attention (MLA). NeMo AutoModel ships a native KimiLinear48BForCausalLM implementation with expert parallelism, packed sequences, and context parallelism.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Fine-Tune Kimi-Linear-48B-A3B-Instruct

From the repository root, run:

uv run automodel --nproc-per-node=8 examples/llm_finetune/kimi/kimi_linear_48b_a3b_hellaswag.yaml

Choose a Workflow

GoalStart Here
Supervised fine-tuning (SFT): Kimi Linear 48B A3B on HellaSwag with EP=8Use kimi_linear_48b_a3b_hellaswag.yaml.
Supervised fine-tuning (SFT): 32k packed sequences with CP=8 and EP=8Use kimi_linear_48b_a3b_longcontext_cp8.yaml.

Model Reference

Model Architecture

PropertyValue
TaskText Generation (hybrid linear attention, MoE)
ArchitectureKimiLinear48BForCausalLM
Parameters48B total / 3B active
Hugging Face Organizationmoonshotai
  • KimiLinear48BForCausalLM (model_type: kimi_linear_48b_a3b)
  • Hybrid attention stack: KDA linear-attention layers interleaved with MLA layers
  • MoE feed-forward blocks with sigmoid routing and grouped top-k selection

Moonshot publishes this model and the Kimi K3 text backbone under the same model_type: kimi_linear and the same architectures: ["KimiLinearForCausalLM"], so neither field identifies the model on its own. NeMo AutoModel gives this implementation a distinct identity, kimi_linear_48b_a3b / KimiLinear48BForCausalLM, and leaves kimi_linear to the K3 text config. The example recipes name KimiLinear48BConfig explicitly, which is what a published Moonshot checkpoint needs; checkpoints saved by NeMo AutoModel already carry the distinct identity and load without that override.

Supported Parallelism

FeatureNotes
FSDP2Default sharding strategy for the recipes below
Expert parallelism (EP)Shards the MoE experts across ranks
Context parallelism (CP)Contiguous per-rank sequence shards. KDA passes its recurrent state from rank to rank, and MLA gathers the compressed KV latent
Packed sequencesDocument boundaries from the THD collater are honored by both layer types, including under CP

Available Models

ModelHF ID
Kimi Linear 48B A3B Instructmoonshotai/Kimi-Linear-48B-A3B-Instruct