llama-embed-nemotron-8b

View as Markdown

NeMo AutoModel provides a retrieval variant of Meta’s Llama for embedding and dense retrieval tasks. It defaults to bidirectional attention, so each token can attend to both past and future tokens in the sequence, and can also preserve or explicitly select standard causal attention.

For the cross-encoder variant, see Llama (Bidirectional) for Reranking.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Train Embeddings with llama-embed-nemotron-8b

From the repository root, run:

uv run automodel \
examples/retrieval/bi_encoder/llama_embed_nemotron_8b/llama_embed_nemotron_8b.yaml \
--nproc-per-node 8

Choose a Workflow

GoalStart Here
Bi-encoder - reproduction recipe for nvidia/llama-embed-nemotron-8b (uses nvidia/embed-nemotron-dataset-v1)Use llama_embed_nemotron_8b.yaml.

Model Reference

Model Architecture

PropertyValue
TasksEmbedding, Dense Retrieval
ArchitectureLlamaBidirectionalModel
Parameters8B
Hugging Face Organizationmeta-llama

Embedding Models

The bidirectional bi-encoder path is used for embedding generation and dense retrieval.

ArchitectureTaskAuto ClassDescription
LlamaBidirectionalModelEmbeddingNeMoAutoModelBiEncoderLlama with configurable attention and mean pooling for dense embeddings

Attention Mode

Set model.is_causal in a bi-encoder recipe:

model:
_target_: nemo_automodel.NeMoAutoModelBiEncoder.from_pretrained
is_causal: false

An explicit value takes precedence over the value saved in the checkpoint’s text config. When neither exists, the bi-encoder defaults to false. The resolved value is persisted when the model is saved. Cross-encoders expose their own is_causal setting and otherwise preserve their saved or native attention mode. For Llama Nemotron VL, the bi-encoder setting changes only the language tower; vision attention remains unchanged. Changing the policy changes embeddings, so regenerate any stored corpus embeddings before using the updated model.

Pooling Strategies

The bi-encoder supports multiple pooling strategies to aggregate token representations into a single embedding vector:

StrategyDescription
avgAverage of all token hidden states (default)
clsFirst token hidden state
lastLast non-padding token hidden state
weighted_avgWeighted average of token hidden states

Available Models

ModelHF ID
Llama 3.1 8Bmeta-llama/Llama-3.1-8B
llama-embed-nemotron-8bnvidia/llama-embed-nemotron-8b

NVIDIA trained and released the Llama Nemotron Embedding 1B model, which uses a bidirectional attention mechanism for multilingual and cross-lingual question-answer retrieval. The model supports long documents (up to 8,192 tokens) and dynamic embedding sizes using Matryoshka embeddings. For more details, see the model card on Hugging Face.