Model CoverageReranking ModelsNVIDIAllama-nemotron-rerank-1b-v2

llama-nemotron-rerank-1b-v2

View as Markdown

NeMo AutoModel provides a retrieval variant of Meta’s Llama that defaults to bidirectional attention for reranking. This lets the query and document interact across the full sequence before a classification head produces a relevance score. Set model.is_causal: true to use causal attention. When the option is omitted, a saved text-config value takes precedence over the bidirectional default.

For the bi-encoder variant, see Llama (Bidirectional) for Embedding.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Train a Reranker with llama-nemotron-rerank-1b-v2

From the repository root, run:

uv run automodel examples/retrieval/cross_encoder/llama3_2_1b.yaml --nproc-per-node 8

Choose a Workflow

GoalStart Here
Train a Llama 3.2 1B cross-encoderUse llama3_2_1b.yaml.

Model Reference

Model Architecture

PropertyValue
TasksReranking
ArchitectureLlamaBidirectionalForSequenceClassification
Parameters1B
Hugging Face Organizationmeta-llama

Reranking Models

The cross-encoder path is used for pairwise relevance scoring and reranking.

ArchitectureTaskWrapper ClassDescription
LlamaBidirectionalForSequenceClassificationRerankingNeMoAutoModelCrossEncoderBidirectional Llama with classification head for relevance scoring

Available Models

ModelHF ID
Llama 3.2 1Bmeta-llama/Llama-3.2-1B
llama-nemotron-rerank-1b-v2nvidia/llama-nemotron-rerank-1b-v2

NVIDIA trained and released the Llama Nemotron Reranking 1B model, optimized to produce a relevance logit score indicating how well a document matches a given query. The model was fine-tuned with a bidirectional attention mechanism for multilingual and cross-lingual question-answer retrieval, with support for long documents (up to 8,192 tokens).