MiniMax-M2.1
MiniMax-M2.1
MiniMax-M2 is MiniMax’s large Mixture-of-Experts language model with linear attention for efficient long-context inference.
Set up NeMo AutoModel with the latest container or follow the installation instructions.
Fine-Tune MiniMax-M2.1
This recipe was validated on 8 nodes x 8 GPUs (64 H100s). See the Launcher Guide for multi-node setup.
From the repository root, run: