Ministral-3-3B-Base-2512
Ministral-3-3B-Base-2512
NeMo AutoModel loads Mistral AI’s Ministral3 with the stock Hugging Face Ministral3Model and configurable text attention for embedding and dense retrieval tasks. It defaults to is_causal: false, so each token can attend to both past and future tokens without requiring a custom model implementation.
The encoder can be loaded directly from text-only checkpoints (e.g., mistralai/Ministral-3B-Instruct) and also automatically extracts the language model from Ministral3 VLM checkpoints (e.g., mistralai/Ministral-3-3B-Base-2512 or mistralai/Ministral-3-3B-Instruct-2512). The custom Ministral3BidirectionalModel remains available through the legacy ministral3_bidirec model type for compatibility with existing checkpoints.
Set up NeMo AutoModel with the latest container or follow the installation instructions.
This page provides checkpoint and architecture details. Use the available checkpoint table below to choose the base or instruct variant.
Model Reference
Model Architecture
Embedding Models
The stock bi-encoder path is used for embedding generation and dense retrieval.
For cross-encoder scoring, stock ministral3 checkpoints use Hugging Face Ministral3ForSequenceClassification. The legacy ministral3_bidirec type does not provide a cross-encoder scoring class.
Attention Mode
Set model.is_causal to true or false in the bi-encoder recipe. An explicit value takes precedence over the
saved text-config value; when neither exists, the model defaults to bidirectional attention. The resolved value is saved
with the checkpoint. Changing the policy changes embeddings, so regenerate stored corpus embeddings before querying
with the updated model.
Pooling Strategies
The bi-encoder supports multiple pooling strategies to aggregate token representations into a single embedding vector: