Ministral-3-3B-Base-2512

View as Markdown

NeMo AutoModel loads Mistral AI’s Ministral3 with the stock Hugging Face Ministral3Model and configurable text attention for embedding and dense retrieval tasks. It defaults to is_causal: false, so each token can attend to both past and future tokens without requiring a custom model implementation.

The encoder can be loaded directly from text-only checkpoints (e.g., mistralai/Ministral-3B-Instruct) and also automatically extracts the language model from Ministral3 VLM checkpoints (e.g., mistralai/Ministral-3-3B-Base-2512 or mistralai/Ministral-3-3B-Instruct-2512). The custom Ministral3BidirectionalModel remains available through the legacy ministral3_bidirec model type for compatibility with existing checkpoints.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

This page provides checkpoint and architecture details. Use the available checkpoint table below to choose the base or instruct variant.

Model Reference

Model Architecture

PropertyValue
TasksEmbedding, Dense Retrieval
ArchitectureMinistral3Model with configurable attention (preferred); Ministral3BidirectionalModel (legacy)
Parameters3B
Hugging Face Organizationmistralai

Embedding Models

The stock bi-encoder path is used for embedding generation and dense retrieval.

ArchitectureTaskAuto ClassDescription
Ministral3ModelEmbeddingNeMoAutoModelBiEncoderPreferred stock Hugging Face path with configurable attention
Ministral3BidirectionalModelEmbeddingNeMoAutoModelBiEncoderLegacy compatibility path for custom retrieval checkpoints

For cross-encoder scoring, stock ministral3 checkpoints use Hugging Face Ministral3ForSequenceClassification. The legacy ministral3_bidirec type does not provide a cross-encoder scoring class.

Attention Mode

Set model.is_causal to true or false in the bi-encoder recipe. An explicit value takes precedence over the saved text-config value; when neither exists, the model defaults to bidirectional attention. The resolved value is saved with the checkpoint. Changing the policy changes embeddings, so regenerate stored corpus embeddings before querying with the updated model.

Pooling Strategies

The bi-encoder supports multiple pooling strategies to aggregate token representations into a single embedding vector:

StrategyDescription
avgAverage of all token hidden states (default)
clsFirst token hidden state
lastLast non-padding token hidden state
weighted_avgWeighted average of token hidden states

Available Models

ModelHF ID
Ministral-3 3B Basemistralai/Ministral-3-3B-Base-2512
Ministral-3 3B Instructmistralai/Ministral-3-3B-Instruct-2512