> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# nemo_automodel.components.models.qwen3_reranker.model

Qwen3 reranker model for cross-encoder reranking tasks.

Unlike the bidirectional encoders in this package (e.g. `llama_bidirectional`),
the Qwen3 reranker keeps the standard **causal** attention of `Qwen3ForCausalLM`
and adds **no new classification head**. Instead, it reuses the language-model head
and turns reranking into a binary "yes"/"no" next-token prediction, exactly as the
official `Qwen/Qwen3-Reranker-*` models are trained and used.

Score convention

````
``forward`` returns a single **raw logit** per query-document pair at the final
(non-padding) token position::

    score(query, doc) = logit("yes") - logit("no")          # unbounded log-odds

This is the "raw logit difference" described in the official model card. Two
downstream consumers use it:

* **Inference** — apply a sigmoid to recover the exact probability produced by
  the official ``compute_logits`` reference implementation::

      p(yes) = sigmoid(score) = softmax([logit("no"), logit("yes")])[yes]

  (The serialized checkpoint is a plain ``Qwen3ForCausalLM``; running the
  official ``compute_logits`` on it reproduces this same ``p(yes)``.)

* **Training** — ``TrainCrossEncoderRecipe`` reshapes the raw scores with
  ``logits.view(-1, n_passages)``, divides by the recipe-level ``temperature``, and
  applies ``F.cross_entropy(..., labels=0)``: a single softmax over each query's
  candidate passages (a listwise contrastive loss). Temperature is a property of the
  training objective, not of the model, so it lives in the recipe config and never
  touches ``forward``; inference scores keep their full discriminative range.

The model is auto-discovered by ``ModelRegistry`` via the ``ModelClass`` export.

## Module Contents

### Classes

| Name | Description |
|------|-------------|
| [`Qwen3RerankerConfig`](#nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerConfig) | Configuration for `Qwen3RerankerForCausalReranking`. |
| [`Qwen3RerankerForCausalReranking`](#nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerForCausalReranking) | Qwen3 causal LM repurposed as a pointwise reranker. |

### Functions

| Name | Description |
|------|-------------|
| [`_last_token_indices`](#nemo_automodel-components-models-qwen3_reranker-model-_last_token_indices) | Return the index of the last attended token for each row. |
| [`_validate_reranker_options`](#nemo_automodel-components-models-qwen3_reranker-model-_validate_reranker_options) | Reject classifier options that cannot describe a single raw yes/no score. |

### Data

[`ModelClass`](#nemo_automodel-components-models-qwen3_reranker-model-ModelClass)

[`logger`](#nemo_automodel-components-models-qwen3_reranker-model-logger)

### API

<Anchor id="nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerConfig">

<CodeBlock showLineNumbers={false} wordWrap={true}>

```python
class nemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerConfig(
    yes_token_id: int | None = None,
    no_token_id: int | None = None,
    kwargs: typing.Any = {}
)
```

</CodeBlock>
</Anchor>

<Indent>

**Bases:** `Qwen3Config`

Configuration for `Qwen3RerankerForCausalReranking`.

Extends `Qwen3Config` with the token ids used for the binary
"yes"/"no" relevance decision.

Checkpoints are serialized as plain ``Qwen3ForCausalLM`` (``model_type:
"qwen3"``) so they load in vLLM and stock HF Transformers without custom
code. ``yes_token_id`` / ``no_token_id`` are preserved as extra JSON fields;
``PretrainedConfig.__init__`` stores unknown keys as instance attributes so
they survive the config round-trip.


<ParamField path="num_labels" type="= 1 if num_labels is None else num_labels">
</ParamField>
<Anchor id="nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerConfig-to_dict">

<CodeBlock showLineNumbers={false} wordWrap={true}>

```python
nemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerConfig.to_dict() -> dict
```

</CodeBlock>
</Anchor>

<Indent>

Serialize as a plain ``Qwen3ForCausalLM`` config.

Routing does not depend on this: the class is resolved through
``ModelRegistry`` by architecture name. Rewriting the serialized identity
is what lets saved checkpoints load in vLLM and plain HF Transformers
without ``trust_remote_code``, as standard ``Qwen3ForCausalLM`` weights.


</Indent>
</Indent>

<Anchor id="nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerForCausalReranking">

<CodeBlock links={{"nemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerConfig":"#nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerConfig"}} showLineNumbers={false} wordWrap={true}>

```python
class nemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerForCausalReranking(
    config: nemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerConfig
)
```

</CodeBlock>
</Anchor>

<Indent>

**Bases:** `Qwen3ForCausalLM`

Qwen3 causal LM repurposed as a pointwise reranker.

Scores each query-document pair with the raw logit difference
``logit("yes") - logit("no")`` at the final non-padding position. Apply a
sigmoid to recover ``p(yes)`` (the official ``compute_logits`` output);
``TrainCrossEncoderRecipe`` consumes the raw score directly. Attention remains
causal and the pretrained ``lm_head`` is reused (no new parameters).

**Why this subclasses ``Qwen3ForCausalLM`` rather than ``Qwen3PreTrainedModel``**

Qwen3-Reranker is a causal LM, not a sequence-classification model. Its
published ``config.json`` declares ``architectures: ["Qwen3ForCausalLM"]``, the
model card loads it with ``AutoModelForCausalLM``, and the official
``compute_logits`` scores by reading full-vocabulary next-token logits at the
final position (``model(**inputs).logits[:, -1, :]``) and indexing the "yes" /
"no" ids. Reranking here IS next-token prediction; there is no classification
head. Subclassing the causal LM keeps that identity, reuses upstream's
``model`` + ``lm_head`` construction and weight tying instead of duplicating
it, and keeps the state-dict keys (``model.*``, ``lm_head.*``) byte-identical
to the backbone.

Checkpoints saved from this class deserialize as plain ``Qwen3ForCausalLM``
(see `Qwen3RerankerConfig.to_dict`), so a finetuned model loads and
scores through the stock HuggingFace and vLLM paths exactly like the backbone,
with no custom code and no ``trust_remote_code``.

**Training-time forward contract**

`forward` is overridden for training and returns pooled scores of shape
``[batch, 1]`` rather than ``[batch, sequence, vocab]``, and ``labels`` are
binary relevance targets rather than next-token ids. Generation APIs are
therefore not usable on an instance of this class -- ``logits_to_keep`` is
rejected rather than silently ignored so that a mistaken ``generate()`` call
fails with a clear message. This affects the in-process training object only;
the saved checkpoint is a standard causal LM with full generation behavior.


<ParamField path="_tied_weights_keys" type="= &#123;'lm_head.weight': 'model.embed_tokens.weight'&#125;">
</ParamField>

<ParamField path="tie_word_embeddings_support" type="TieSupport = TieSupport.BOTH">
</ParamField>
<Anchor id="nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerForCausalReranking-forward">

<CodeBlock showLineNumbers={false} wordWrap={true}>

```python
nemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerForCausalReranking.forward(
    input_ids: torch.LongTensor | None = None,
    attention_mask: torch.Tensor | None = None,
    position_ids: torch.LongTensor | None = None,
    inputs_embeds: torch.FloatTensor | None = None,
    labels: torch.LongTensor | None = None,
    kwargs: typing.Any = {}
) -> transformers.modeling_outputs.SequenceClassifierOutputWithPast
```

</CodeBlock>
</Anchor>

<Indent>

Score each query-document pair at its last attended token.

**Parameters:**

<ParamField path="input_ids" type="torch.LongTensor | None" default="None">
Tensor of shape [batch, sequence]; omit when inputs_embeds is given.
</ParamField>

<ParamField path="attention_mask" type="torch.Tensor | None" default="None">
Tensor of shape [batch, sequence], with 1 for attended tokens
and 0 for padding. Left or right padding is supported; every row must
attend to at least one token.
</ParamField>

<ParamField path="position_ids" type="torch.LongTensor | None" default="None">
Optional tensor of shape [batch, sequence].
</ParamField>

<ParamField path="inputs_embeds" type="torch.FloatTensor | None" default="None">
Optional tensor of shape [batch, sequence, hidden].
</ParamField>

<ParamField path="labels" type="torch.LongTensor | None" default="None">
Optional tensor of shape [batch] containing binary relevance classes
(0 for no, 1 for yes), not next-token IDs. The training recipe instead
computes its own loss over each query's candidate passages.
</ParamField>

<ParamField path="**kwargs" type="Any" default="&#123;&#125;">
Hugging Face decoder options. Per-position vocabulary logits and
generation through this training class are unsupported.
</ParamField>

**Returns:** `SequenceClassifierOutputWithPast`

SequenceClassifierOutputWithPast with logits of shape [batch, 1], holding


</Indent>
<Anchor id="nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerForCausalReranking-from_pretrained">

<CodeBlock links={{"nemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerForCausalReranking":"#nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerForCausalReranking"}} showLineNumbers={false} wordWrap={true}>

```python
nemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerForCausalReranking.from_pretrained(
    pretrained_model_name_or_path: str | os.PathLike[str] | None,
    args: typing.Any = (),
    kwargs: typing.Any = {}
) -> nemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerForCausalReranking
```

</CodeBlock>
</Anchor>

<Indent>

<Badge>classmethod</Badge>

Load weights and resolve the "yes"/"no" token ids if not already set.

Explicit ``yes_token_id``/``no_token_id`` (e.g. from the recipe YAML or a
saved config) take precedence; otherwise they are resolved from the
tokenizer of ``pretrained_model_name_or_path``.

In-memory loads use ``config.name_or_path`` to locate the tokenizer.
A checkpoint's saved embedding ties cannot be overridden in either direction.


</Indent>
<Anchor id="nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerForCausalReranking-supports_config">

<CodeBlock showLineNumbers={false} wordWrap={true}>

```python
nemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerForCausalReranking.supports_config(
    config: transformers.PretrainedConfig
) -> bool
```

</CodeBlock>
</Anchor>

<Indent>

<Badge>classmethod</Badge>

Use causal reranking only for checkpoints declaring a compatible LM head.


</Indent>
<Anchor id="nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerForCausalReranking-tie_weights">

<CodeBlock showLineNumbers={false} wordWrap={true}>

```python
nemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerForCausalReranking.tie_weights(
    _args: object = (),
    _kwargs: object = {}
) -> None
```

</CodeBlock>
</Anchor>

<Indent>

Alias ``lm_head`` to the input embeddings when the config asks for it.

Declared model-locally rather than inherited: transformers v5 does not reliably tie a
custom model from the dict-shaped ``_tied_weights_keys`` alone, so the config flag is
honoured explicitly (mirroring ``Qwen2ForCausalLM``). This matters here because
scoring reads the yes/no rows of ``lm_head``; if the tie silently failed to apply, the
head would drift from the embeddings it is supposed to share and the published
checkpoint's scores would not reproduce.


</Indent>
</Indent>

<Anchor id="nemo_automodel-components-models-qwen3_reranker-model-_last_token_indices">

<CodeBlock showLineNumbers={false} wordWrap={true}>

```python
nemo_automodel.components.models.qwen3_reranker.model._last_token_indices(
    attention_mask: torch.Tensor
) -> torch.Tensor
```

</CodeBlock>
</Anchor>

<Indent>

Return the index of the last attended token for each row.

Padding-side agnostic: the result is the highest position index that the mask
attends to, so left padding resolves to the final position and right padding to
the last real token, with no assumption about which side a batch uses.

**Parameters:**

<ParamField path="attention_mask" type="torch.Tensor">
Tensor of shape [batch, sequence], 1 for attended tokens and
0 for padding. Each row is expected to attend to at least one token.
</ParamField>

**Returns:** `torch.Tensor`

Tensor of shape [batch] holding the last attended index per row.


</Indent>

<Anchor id="nemo_automodel-components-models-qwen3_reranker-model-_validate_reranker_options">

<CodeBlock showLineNumbers={false} wordWrap={true}>

```python
nemo_automodel.components.models.qwen3_reranker.model._validate_reranker_options(
    num_labels: int | None = None,
    pooling: str | None = None,
    temperature: float | None = None
) -> None
```

</CodeBlock>
</Anchor>

<Indent>

Reject classifier options that cannot describe a single raw yes/no score.


</Indent>

<Anchor id="nemo_automodel-components-models-qwen3_reranker-model-ModelClass">

<CodeBlock showLineNumbers={false} wordWrap={true}>

```python
nemo_automodel.components.models.qwen3_reranker.model.ModelClass = [Qwen3RerankerForCausalReranking]
```

</CodeBlock>
</Anchor>


<Anchor id="nemo_automodel-components-models-qwen3_reranker-model-logger">

<CodeBlock showLineNumbers={false} wordWrap={true}>

```python
nemo_automodel.components.models.qwen3_reranker.model.logger = logging.get_logger(__name__)
```

</CodeBlock>
</Anchor>


<style>{`
  .light .fern-code-block,
  .light .fern-prose code:not(.code-block) {
    background-color: var(--nv-color-bg-alt, #f7f7f7) !important;
  }

  .dark .fern-code-block,
  .dark .fern-prose code:not(.code-block) {
    background-color: #1f1f1f !important;
  }
`}</style>
````