MiniMax-M2.1

View as Markdown

MiniMax-M2 is MiniMax’s large Mixture-of-Experts language model with linear attention for efficient long-context inference.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Fine-Tune MiniMax-M2.1

This recipe was validated on 8 nodes x 8 GPUs (64 H100s). See the Launcher Guide for multi-node setup.

From the repository root, run:

uv run automodel --nproc-per-node=8 examples/llm_finetune/minimax_m2/minimax_m2.1_hellaswag_pp.yaml

Choose a Workflow

GoalStart Here
Supervised fine-tuning (SFT) - MiniMax-M2.1 on HellaSwag with pipeline parallelismUse minimax_m2.1_hellaswag_pp.yaml.

Model Reference

Model Architecture

PropertyValue
TaskText Generation (MoE)
ArchitectureMiniMaxM2ForCausalLM
Parameters229B total
Hugging Face OrganizationMiniMaxAI

Available Models

ModelHF ID
MiniMax M2.1MiniMaxAI/MiniMax-M2.1