> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# Ministral-3-3B-Base-2512

> Reference Ministral-3-3B-Base-2512 checkpoint and architecture details for NeMo AutoModel, with model-family information, setup guidance, and upstream resources.

NeMo AutoModel loads [Mistral AI's Ministral3](https://mistral.ai/news/ministraux/) with the stock Hugging Face `Ministral3Model` and configurable text attention for embedding and dense retrieval tasks. It defaults to `is_causal: false`, so each token can attend to both past and future tokens without requiring a custom model implementation.

The encoder can be loaded directly from text-only checkpoints (e.g., `mistralai/Ministral-3B-Instruct`) and also automatically extracts the language model from Ministral3 VLM checkpoints (e.g., `mistralai/Ministral-3-3B-Base-2512` or `mistralai/Ministral-3-3B-Instruct-2512`). The custom `Ministral3BidirectionalModel` remains available through the legacy `ministral3_bidirec` model type for compatibility with existing checkpoints.

Set up NeMo AutoModel with the [latest container](https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo-automodel) or follow the [installation instructions](/get-started/installation).

This page provides checkpoint and architecture details. Use the available checkpoint table below to
choose the base or instruct variant.

## Model Reference

### Model Architecture

| Property                  | Value                                                                                              |
| ------------------------- | -------------------------------------------------------------------------------------------------- |
| Tasks                     | Embedding, Dense Retrieval                                                                         |
| Architecture              | `Ministral3Model` with configurable attention (preferred); `Ministral3BidirectionalModel` (legacy) |
| Parameters                | 3B                                                                                                 |
| Hugging Face Organization | [mistralai](https://huggingface.co/mistralai)                                                      |

### Embedding Models

The stock bi-encoder path is used for embedding generation and dense retrieval.

| Architecture                   | Task      | Auto Class                                                                                                                                                         | Description                                                   |
| ------------------------------ | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------- |
| `Ministral3Model`              | Embedding | [`NeMoAutoModelBiEncoder`](https://github.com/NVIDIA-NeMo/Automodel/blob/8dc00dcb4a35c2413c52c6e7eb7ac8f1c24836aa/nemo_automodel/_transformers/auto_model.py#L991) | Preferred stock Hugging Face path with configurable attention |
| `Ministral3BidirectionalModel` | Embedding | [`NeMoAutoModelBiEncoder`](https://github.com/NVIDIA-NeMo/Automodel/blob/8dc00dcb4a35c2413c52c6e7eb7ac8f1c24836aa/nemo_automodel/_transformers/auto_model.py#L991) | Legacy compatibility path for custom retrieval checkpoints    |

For cross-encoder scoring, stock `ministral3` checkpoints use Hugging Face `Ministral3ForSequenceClassification`. The legacy `ministral3_bidirec` type does not provide a cross-encoder scoring class.

### Attention Mode

Set `model.is_causal` to `true` or `false` in the bi-encoder recipe. An explicit value takes precedence over the
saved text-config value; when neither exists, the model defaults to bidirectional attention. The resolved value is saved
with the checkpoint. Changing the policy changes embeddings, so regenerate stored corpus embeddings before querying
with the updated model.

### Pooling Strategies

The bi-encoder supports multiple pooling strategies to aggregate token representations into a single embedding vector:

| Strategy       | Description                                  |
| -------------- | -------------------------------------------- |
| `avg`          | Average of all token hidden states (default) |
| `cls`          | First token hidden state                     |
| `last`         | Last non-padding token hidden state          |
| `weighted_avg` | Weighted average of token hidden states      |

### Available Models

| Model                   | HF ID                                                                                                     |
| ----------------------- | --------------------------------------------------------------------------------------------------------- |
| Ministral-3 3B Base     | [`mistralai/Ministral-3-3B-Base-2512`](https://huggingface.co/mistralai/Ministral-3-3B-Base-2512)         |
| Ministral-3 3B Instruct | [`mistralai/Ministral-3-3B-Instruct-2512`](https://huggingface.co/mistralai/Ministral-3-3B-Instruct-2512) |

## Related Resources

* [NeMo AutoModel Repository](https://github.com/NVIDIA-NeMo/Automodel)