Embedding Models

View as Markdown

Introduction

Text embedding models transform text into dense vector representations that power semantic search, dense retrieval, retrieval-augmented generation (RAG), and classification tasks. NeMo AutoModel includes custom bidirectional backbones and configures supported Hugging Face AutoModel backbones with non-causal attention for retrieval training.

For cross-encoder pairwise scoring, see Reranking Models.

Embedding models use bi-encoders to produce dense representations for queries and documents independently. They are the standard path for embedding generation and first-stage dense retrieval.

Optimized Backbones

OwnerModelArchitectureAuto ClassTasks
NVIDIALlama (Bidirectional)LlamaBidirectionalModelNeMoAutoModelBiEncoderEmbedding, Dense Retrieval
NVIDIALlama Nemotron VLLlamaNemotronVLModelNeMoAutoModelBiEncoderEmbedding, Dense Retrieval
Mistral AIMinistral3Ministral3Model with configurable attentionNeMoAutoModelBiEncoderEmbedding, Dense Retrieval

The stock ministral3 path is preferred. The ministral3_bidirec model type remains available for compatibility with legacy checkpoints and resolves to the custom Ministral3BidirectionalModel.

Hugging Face Auto Backbones

Unregistered text architectures load through Hugging Face AutoModel. The fallback path applies model.is_causal, restores a saved text-config value when the option is omitted, and otherwise defaults to bidirectional attention. Verify that an unregistered architecture honors the selected policy before using it for retrieval training.

Composite backbones must expose matching text configuration and decoder objects through get_text_config(decoder=True) and get_decoder(). Loading fails if these methods do not identify the same text tower. The attention policy changes only that tower, leaving vision attention unchanged.

Example Recipes

RecipeDescription
llama3_2_1b.yamlBi-encoder, Llama 3.2 1B embedding model
llama_embed_nemotron_8b.yamlBi-encoder, Llama-Embed-Nemotron-8B reproduction recipe
ministral3_3b_instruct.yamlBi-encoder Ministral3-3B recipe

Supported Models

This table combines recipe-backed checkpoints with documented model families. Dates show when the current checkpoint first appeared in a recipe, or when a documentation-only family page was added. See the combined model support log for recipe-backed checkpoints of every model type.

Supported Workflows

  • Fine-tuning (Bi-Encoder): Contrastive learning on query-document pairs to produce embedding models
  • LoRA/PEFT: Parameter-efficient fine-tuning for embedding backbones
  • ONNX Export: Export trained embedding models for deployment (case-by-case, model-dependent)

Dataset

Retrieval fine-tuning requires query-document pairs: each example is a query paired with one positive document and one or more negative documents. Both inline JSONL and corpus ID-based JSON formats are supported. See the Retrieval Dataset guide.

Train Embedding Models

For a complete walkthrough of training configuration, model-specific settings, and launch commands, see the Embedding and Reranking Fine-Tuning Guide.