> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# llama-nemotron-rerank-1b-v2

> Reference llama-nemotron-rerank-1b-v2 checkpoint and architecture details for NeMo AutoModel, with model-family information, setup guidance, and upstream resources.

NeMo AutoModel provides a retrieval variant of [Meta's Llama](https://www.llama.com/) that defaults to bidirectional attention for reranking. This lets the query and document interact across the full sequence before a classification head produces a relevance score. Set `model.is_causal: true` to use causal attention. When the option is omitted, a saved text-config value takes precedence over the bidirectional default.

For the bi-encoder variant, see [Llama (Bidirectional) for Embedding](/model-coverage/embedding-models/nvidia/llama-embed-nemotron-8b).

Set up NeMo AutoModel with the [latest container](https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo-automodel) or follow the [installation instructions](/get-started/installation).

## Train a Reranker with llama-nemotron-rerank-1b-v2

From the repository root, run:

```bash
uv run automodel examples/retrieval/cross_encoder/llama3_2_1b.yaml --nproc-per-node 8
```

## Choose a Workflow

| Goal                               | Start Here                                                                                                                      |
| ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| Train a Llama 3.2 1B cross-encoder | Use [llama3\_2\_1b.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/retrieval/cross_encoder/llama3_2_1b.yaml). |

## Model Reference

### Model Architecture

| Property                  | Value                                           |
| ------------------------- | ----------------------------------------------- |
| Tasks                     | Reranking                                       |
| Architecture              | `LlamaBidirectionalForSequenceClassification`   |
| Parameters                | 1B                                              |
| Hugging Face Organization | [meta-llama](https://huggingface.co/meta-llama) |

### Reranking Models

The cross-encoder path is used for pairwise relevance scoring and reranking.

| Architecture                                  | Task      | Wrapper Class               | Description                                                        |
| --------------------------------------------- | --------- | --------------------------- | ------------------------------------------------------------------ |
| `LlamaBidirectionalForSequenceClassification` | Reranking | `NeMoAutoModelCrossEncoder` | Bidirectional Llama with classification head for relevance scoring |

### Available Models

| Model                       | HF ID                                                                                             |
| --------------------------- | ------------------------------------------------------------------------------------------------- |
| Llama 3.2 1B                | [`meta-llama/Llama-3.2-1B`](https://huggingface.co/meta-llama/Llama-3.2-1B)                       |
| llama-nemotron-rerank-1b-v2 | [`nvidia/llama-nemotron-rerank-1b-v2`](https://huggingface.co/nvidia/llama-nemotron-rerank-1b-v2) |

## Related Resources

NVIDIA trained and released the `Llama Nemotron Reranking 1B` model, optimized to produce a relevance logit score indicating how well a document matches a given query. The model was fine-tuned with a bidirectional attention mechanism for multilingual and cross-lingual question-answer retrieval, with support for long documents (up to 8,192 tokens).

* [nvidia/llama-nemotron-rerank-1b-v2](https://huggingface.co/nvidia/llama-nemotron-rerank-1b-v2)
* [NeMo AutoModel Repository](https://github.com/NVIDIA-NeMo/Automodel)