Kimi-Linear-48B-A3B-Instruct
Kimi-Linear-48B-A3B-Instruct
Kimi Linear is a hybrid-attention Mixture-of-Experts language model from Moonshot AI. Most layers use Kimi Delta Attention (KDA), a gated linear-attention variant with a recurrent state, and the remaining layers use full Multi-Head Latent Attention (MLA). NeMo AutoModel ships a native KimiLinear48BForCausalLM implementation with expert parallelism, packed sequences, and context parallelism.
Set up NeMo AutoModel with the latest container or follow the installation instructions.
Fine-Tune Kimi-Linear-48B-A3B-Instruct
From the repository root, run:
Choose a Workflow
Model Reference
Model Architecture
KimiLinear48BForCausalLM(model_type: kimi_linear_48b_a3b)- Hybrid attention stack: KDA linear-attention layers interleaved with MLA layers
- MoE feed-forward blocks with sigmoid routing and grouped top-k selection
Moonshot publishes this model and the Kimi K3 text backbone under the same
model_type: kimi_linear and the same architectures: ["KimiLinearForCausalLM"], so
neither field identifies the model on its own. NeMo AutoModel gives this implementation a
distinct identity, kimi_linear_48b_a3b / KimiLinear48BForCausalLM, and leaves
kimi_linear to the K3 text config. The example recipes name KimiLinear48BConfig
explicitly, which is what a published Moonshot checkpoint needs; checkpoints saved by
NeMo AutoModel already carry the distinct identity and load without that override.