> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# llama-embed-nemotron-8b

> Use llama-embed-nemotron-8b with NeMo AutoModel for embedding training, with documented checkpoints, runnable recipes, setup guidance, and model reference details.

NeMo AutoModel provides a retrieval variant of [Meta's Llama](https://www.llama.com/) for embedding and dense retrieval tasks. It defaults to **bidirectional attention**, so each token can attend to both past and future tokens in the sequence, and can also preserve or explicitly select standard causal attention.

For the cross-encoder variant, see [Llama (Bidirectional) for Reranking](/model-coverage/reranking-models/nvidia/llama-nemotron-rerank-1b-v2).

Set up NeMo AutoModel with the [latest container](https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo-automodel) or follow the [installation instructions](/get-started/installation).

## Train Embeddings with llama-embed-nemotron-8b

From the repository root, run:

```bash
uv run automodel \
  examples/retrieval/bi_encoder/llama_embed_nemotron_8b/llama_embed_nemotron_8b.yaml \
  --nproc-per-node 8
```

## Choose a Workflow

| Goal                                                                                                                                                                                                                                         | Start Here                                                                                                                                                                    |
| -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Bi-encoder - reproduction recipe for [`nvidia/llama-embed-nemotron-8b`](https://huggingface.co/nvidia/llama-embed-nemotron-8b) (uses [`nvidia/embed-nemotron-dataset-v1`](https://huggingface.co/datasets/nvidia/embed-nemotron-dataset-v1)) | Use [llama\_embed\_nemotron\_8b.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/retrieval/bi_encoder/llama_embed_nemotron_8b/llama_embed_nemotron_8b.yaml). |

## Model Reference

### Model Architecture

| Property                  | Value                                           |
| ------------------------- | ----------------------------------------------- |
| Tasks                     | Embedding, Dense Retrieval                      |
| Architecture              | `LlamaBidirectionalModel`                       |
| Parameters                | 8B                                              |
| Hugging Face Organization | [meta-llama](https://huggingface.co/meta-llama) |

### Embedding Models

The bidirectional bi-encoder path is used for embedding generation and dense retrieval.

| Architecture              | Task      | Auto Class                                                                                                                                                         | Description                                                             |
| ------------------------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------- |
| `LlamaBidirectionalModel` | Embedding | [`NeMoAutoModelBiEncoder`](https://github.com/NVIDIA-NeMo/Automodel/blob/8dc00dcb4a35c2413c52c6e7eb7ac8f1c24836aa/nemo_automodel/_transformers/auto_model.py#L991) | Llama with configurable attention and mean pooling for dense embeddings |

### Attention Mode

Set `model.is_causal` in a bi-encoder recipe:

```yaml
model:
  _target_: nemo_automodel.NeMoAutoModelBiEncoder.from_pretrained
  is_causal: false
```

An explicit value takes precedence over the value saved in the checkpoint's text config. When neither exists, the
bi-encoder defaults to `false`. The resolved value is persisted when the model is saved. Cross-encoders expose their own
`is_causal` setting and otherwise preserve their saved or native attention mode. For Llama Nemotron VL, the bi-encoder
setting changes only the language tower; vision attention remains unchanged. Changing the policy changes embeddings, so
regenerate any stored corpus embeddings before using the updated model.

### Pooling Strategies

The bi-encoder supports multiple pooling strategies to aggregate token representations into a single embedding vector:

| Strategy       | Description                                  |
| -------------- | -------------------------------------------- |
| `avg`          | Average of all token hidden states (default) |
| `cls`          | First token hidden state                     |
| `last`         | Last non-padding token hidden state          |
| `weighted_avg` | Weighted average of token hidden states      |

### Available Models

| Model                   | HF ID                                                                                     |
| ----------------------- | ----------------------------------------------------------------------------------------- |
| Llama 3.1 8B            | [`meta-llama/Llama-3.1-8B`](https://huggingface.co/meta-llama/Llama-3.1-8B)               |
| llama-embed-nemotron-8b | [`nvidia/llama-embed-nemotron-8b`](https://huggingface.co/nvidia/llama-embed-nemotron-8b) |

## Related Resources

NVIDIA trained and released the `Llama Nemotron Embedding 1B` model, which uses a bidirectional attention mechanism for multilingual and cross-lingual question-answer retrieval. The model supports long documents (up to 8,192 tokens) and dynamic embedding sizes using Matryoshka embeddings. For more details, see the model card on Hugging Face.

* [NeMo AutoModel Repository](https://github.com/NVIDIA-NeMo/Automodel)
* [nvidia/llama-nemotron-embed-1b-v2](https://huggingface.co/nvidia/llama-nemotron-embed-1b-v2)