> This page is for version Nightly (default).
> For other versions, use one of these documentation indexes:
> - Nightly (default): https://docs.nvidia.com/nemo/automodel/nightly/llms.txt
> - Latest: https://docs.nvidia.com/nemo/automodel/latest/llms.txt
> - 0.5.0 · 26.06: https://docs.nvidia.com/nemo/automodel/v0.5/llms.txt
> - 0.4.0 · 26.04: https://docs.nvidia.com/nemo/automodel/v0.4/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# Qwen3-Reranker-4B

> Train Qwen3-Reranker-4B with NeMo AutoModel using context-aware query-document scoring, compatible checkpoint loading, and the retrieval training recipe.

[Qwen3-Reranker-4B](https://huggingface.co/Qwen/Qwen3-Reranker-4B) scores query-document pairs with its pretrained language-model head. NeMo AutoModel loads it through `NeMoAutoModelCrossEncoder` using `Qwen3RerankerForCausalReranking`.

## Model Reference

### Model Architecture

| Property     | Value                             |
| ------------ | --------------------------------- |
| Architecture | `Qwen3RerankerForCausalReranking` |
| Task         | Reranking                         |
| Wrapper      | `NeMoAutoModelCrossEncoder`       |

## Scoring and Training

The model uses causal attention by default and returns `logit("yes") - logit("no")` at the final non-padding token. It adds no classification parameters. The [context-aware training example](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/retrieval/cross_encoder/qwen3_4b_reranker_agentir.yaml) includes optional reasoning and question context, per-field context dropout, and listwise training over positive and negative passages.

Prepare a local JSONL file using the schema documented in the example. Each row needs `query`, `pos_doc`, `neg_doc`, and a `global_qid` that groups all turns of the same originating question. `reasoning` and `global_query` are optional. Set both `dataset.data_dir_list` and `validation_dataset.data_dir_list` to this file; the example's dataset is not bundled.

From the repository root, launch the example on 8 GPUs:

```bash
uv run automodel examples/retrieval/cross_encoder/qwen3_4b_reranker_agentir.yaml --nproc-per-node 8
```

## Checkpoint Compatibility

Qwen3 causal language-model checkpoints loaded through `NeMoAutoModelCrossEncoder` use a single yes/no relevance score, including base Qwen3 checkpoints. They accept `num_labels: 1`; other label counts and `model.pooling` are unsupported. Set the training loss temperature at the recipe's top level, not under `model`.

Saved checkpoints identify as standard `Qwen3ForCausalLM` models for loading in stock Transformers or vLLM. The checkpoint's `tie_word_embeddings` setting is preserved; loading with a conflicting override raises an error.

Qwen3 sequence-classification checkpoints retain the Hugging Face classification loading path. Extracting a complete causal language model resolves missing yes/no token IDs from the source checkpoint's tokenizer. Extracting a bare decoder retains the classification fallback because it has no trained language-model head.

## Parallelism

The reranker does not declare support for tensor, context, pipeline, or expert parallelism. See the [example configuration](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/retrieval/cross_encoder/qwen3_4b_reranker_agentir.yaml) for its distributed training settings and the [retrieval fine-tuning guide](/recipes-e2e-examples/retrieval-finetune) for the training workflow.