Ministral-3-3B-Base-2512
Ministral-3-3B-Base-2512
NeMo AutoModel loads Mistral AI’s Ministral3 with the stock Hugging Face Ministral3Model and configurable text attention for embedding and dense retrieval tasks. It defaults to is_causal: false, so each token can attend to both past and future tokens without requiring a custom model implementation.
The encoder loads text-only checkpoints directly. For Ministral3 vision-language model (VLM) checkpoints, the default retrieval path retains the vision tower. Set model.extract_submodel: language_model to train a text-only encoder instead. The custom Ministral3BidirectionalModel remains available through the legacy ministral3_bidirec model type for compatibility with existing checkpoints.
For multimodal retrieval training, the registered Mistral3BidirectionalModel combines the Pixtral vision tower with a bidirectional Ministral3 language tower. Mistral3VLBidirectionalForSequenceClassification adds the corresponding reranking head.
Set up NeMo AutoModel with the latest container or follow the installation instructions.
This page provides checkpoint and architecture details. Use the available checkpoint table below to choose the base or instruct variant.
Model Reference
Model Architecture
Embedding Models
The stock bi-encoder path is used for embedding generation and dense retrieval.
For multimodal cross-encoder scoring, mistral3 checkpoints with a Ministral3 text tower use Mistral3VLBidirectionalForSequenceClassification. Text-only ministral3 checkpoints continue to use Hugging Face Ministral3ForSequenceClassification.
Attention Mode
Set model.is_causal to true or false in the bi-encoder recipe. An explicit value takes precedence over the
saved text-config value; when neither exists, the model defaults to bidirectional attention. The resolved value is saved
with the checkpoint. Changing the policy changes embeddings, so regenerate stored corpus embeddings before querying
with the updated model.
Pooling Strategies
The bi-encoder supports multiple pooling strategies to aggregate token representations into a single embedding vector:
Available Models
Try with NeMo AutoModel
1. Clone and install from source (full instructions):
2. Prepare your retrieval dataset. Use the corpus-ID JSON format to associate each query with positive and negative document IDs. Include the referenced corpus and its merlin_metadata.json file, with a corpus class that matches your document format. These recipes do not require a particular dataset.
In the example you select, replace dataset.data_dir_list[0].path with the path to your training JSON file. The example path <DATASETS_PATH>/retrieval/train.json is a placeholder, not a bundled dataset.
The embedding and reranking examples set use_text_in_document: true, so image documents include companion text. Multimodal mining includes each document’s available image and text content and uses the checkpoint’s saved prompts, sequence limits, and image preprocessing. Text-only documents retain their text. Use the same document inputs when mining, training, and evaluating an embedding model.
3. Run one of the Mistral3 VL retrieval recipes from inside the repo:
See the Installation Guide.
The reranker’s scoring projection and model.temperature scaling run in FP32, including under BF16 autocast. The recipe also computes cross entropy in FP32. Keep the top-level temperature: 1.0 when using a non-unit model temperature; setup rejects two active temperature sources.
To freeze model towers, use the top-level freeze_config section. For example, setting freeze_vision_tower: true and freeze_language_model: true within that section leaves the projector and, for reranking, the scoring head trainable.