ReferenceFull Library ReferenceNemo AutomodelNemo AutomodelComponentsModelsQwen3 Rerankernemo_automodel.components.models.qwen3_reranker.model
nemo_automodel.components.models.qwen3_reranker.model
nemo_automodel.components.models.qwen3_reranker.model
Qwen3 reranker model for cross-encoder reranking tasks.
Unlike the bidirectional encoders in this package (e.g. llama_bidirectional),
the Qwen3 reranker keeps the standard causal attention of Qwen3ForCausalLM
and adds no new classification head. Instead, it reuses the language-model head
and turns reranking into a binary “yes”/“no” next-token prediction, exactly as the
official Qwen/Qwen3-Reranker-* models are trained and used.
Score convention
``forward`` returns a single **raw logit** per query-document pair at the final(non-padding) token position::score(query, doc) = logit("yes") - logit("no") # unbounded log-oddsThis is the "raw logit difference" described in the official model card. Twodownstream consumers use it:* **Inference** — apply a sigmoid to recover the exact probability produced bythe official ``compute_logits`` reference implementation::p(yes) = sigmoid(score) = softmax([logit("no"), logit("yes")])[yes](The serialized checkpoint is a plain ``Qwen3ForCausalLM``; running theofficial ``compute_logits`` on it reproduces this same ``p(yes)``.)* **Training** — ``TrainCrossEncoderRecipe`` reshapes the raw scores with``logits.view(-1, n_passages)``, divides by the recipe-level ``temperature``, andapplies ``F.cross_entropy(..., labels=0)``: a single softmax over each query'scandidate passages (a listwise contrastive loss). Temperature is a property of thetraining objective, not of the model, so it lives in the recipe config and nevertouches ``forward``; inference scores keep their full discriminative range.The model is auto-discovered by ``ModelRegistry`` via the ``ModelClass`` export.## Module Contents### Classes| Name | Description ||------|-------------|| [`Qwen3RerankerConfig`](#nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerConfig) | Configuration for `Qwen3RerankerForCausalReranking`. || [`Qwen3RerankerForCausalReranking`](#nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerForCausalReranking) | Qwen3 causal LM repurposed as a pointwise reranker. |### Functions| Name | Description ||------|-------------|| [`_last_token_indices`](#nemo_automodel-components-models-qwen3_reranker-model-_last_token_indices) | Return the index of the last attended token for each row. || [`_validate_reranker_options`](#nemo_automodel-components-models-qwen3_reranker-model-_validate_reranker_options) | Reject classifier options that cannot describe a single raw yes/no score. |### Data[`ModelClass`](#nemo_automodel-components-models-qwen3_reranker-model-ModelClass)[`logger`](#nemo_automodel-components-models-qwen3_reranker-model-logger)### API<Anchor id="nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerConfig"><CodeBlock showLineNumbers={false} wordWrap={true}>```pythonclass nemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerConfig(yes_token_id: int | None = None,no_token_id: int | None = None,kwargs: typing.Any = {})```</CodeBlock></Anchor><Indent>**Bases:** `Qwen3Config`Configuration for `Qwen3RerankerForCausalReranking`.Extends `Qwen3Config` with the token ids used for the binary"yes"/"no" relevance decision.Checkpoints are serialized as plain ``Qwen3ForCausalLM`` (``model_type:"qwen3"``) so they load in vLLM and stock HF Transformers without customcode. ``yes_token_id`` / ``no_token_id`` are preserved as extra JSON fields;``PretrainedConfig.__init__`` stores unknown keys as instance attributes sothey survive the config round-trip.<ParamField path="num_labels" type="= 1 if num_labels is None else num_labels"></ParamField><Anchor id="nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerConfig-to_dict"><CodeBlock showLineNumbers={false} wordWrap={true}>```pythonnemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerConfig.to_dict() -> dict```</CodeBlock></Anchor><Indent>Serialize as a plain ``Qwen3ForCausalLM`` config.Routing does not depend on this: the class is resolved through``ModelRegistry`` by architecture name. Rewriting the serialized identityis what lets saved checkpoints load in vLLM and plain HF Transformerswithout ``trust_remote_code``, as standard ``Qwen3ForCausalLM`` weights.</Indent></Indent><Anchor id="nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerForCausalReranking"><CodeBlock links={{"nemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerConfig":"#nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerConfig"}} showLineNumbers={false} wordWrap={true}>```pythonclass nemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerForCausalReranking(config: nemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerConfig)```</CodeBlock></Anchor><Indent>**Bases:** `Qwen3ForCausalLM`Qwen3 causal LM repurposed as a pointwise reranker.Scores each query-document pair with the raw logit difference``logit("yes") - logit("no")`` at the final non-padding position. Apply asigmoid to recover ``p(yes)`` (the official ``compute_logits`` output);``TrainCrossEncoderRecipe`` consumes the raw score directly. Attention remainscausal and the pretrained ``lm_head`` is reused (no new parameters).**Why this subclasses ``Qwen3ForCausalLM`` rather than ``Qwen3PreTrainedModel``**Qwen3-Reranker is a causal LM, not a sequence-classification model. Itspublished ``config.json`` declares ``architectures: ["Qwen3ForCausalLM"]``, themodel card loads it with ``AutoModelForCausalLM``, and the official``compute_logits`` scores by reading full-vocabulary next-token logits at thefinal position (``model(**inputs).logits[:, -1, :]``) and indexing the "yes" /"no" ids. Reranking here IS next-token prediction; there is no classificationhead. Subclassing the causal LM keeps that identity, reuses upstream's``model`` + ``lm_head`` construction and weight tying instead of duplicatingit, and keeps the state-dict keys (``model.*``, ``lm_head.*``) byte-identicalto the backbone.Checkpoints saved from this class deserialize as plain ``Qwen3ForCausalLM``(see `Qwen3RerankerConfig.to_dict`), so a finetuned model loads andscores through the stock HuggingFace and vLLM paths exactly like the backbone,with no custom code and no ``trust_remote_code``.**Training-time forward contract**`forward` is overridden for training and returns pooled scores of shape``[batch, 1]`` rather than ``[batch, sequence, vocab]``, and ``labels`` arebinary relevance targets rather than next-token ids. Generation APIs aretherefore not usable on an instance of this class -- ``logits_to_keep`` isrejected rather than silently ignored so that a mistaken ``generate()`` callfails with a clear message. This affects the in-process training object only;the saved checkpoint is a standard causal LM with full generation behavior.<ParamField path="_tied_weights_keys" type="= {'lm_head.weight': 'model.embed_tokens.weight'}"></ParamField><ParamField path="tie_word_embeddings_support" type="TieSupport = TieSupport.BOTH"></ParamField><Anchor id="nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerForCausalReranking-forward"><CodeBlock showLineNumbers={false} wordWrap={true}>```pythonnemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerForCausalReranking.forward(input_ids: torch.LongTensor | None = None,attention_mask: torch.Tensor | None = None,position_ids: torch.LongTensor | None = None,inputs_embeds: torch.FloatTensor | None = None,labels: torch.LongTensor | None = None,kwargs: typing.Any = {}) -> transformers.modeling_outputs.SequenceClassifierOutputWithPast```</CodeBlock></Anchor><Indent>Score each query-document pair at its last attended token.**Parameters:**<ParamField path="input_ids" type="torch.LongTensor | None" default="None">Tensor of shape [batch, sequence]; omit when inputs_embeds is given.</ParamField><ParamField path="attention_mask" type="torch.Tensor | None" default="None">Tensor of shape [batch, sequence], with 1 for attended tokensand 0 for padding. Left or right padding is supported; every row mustattend to at least one token.</ParamField><ParamField path="position_ids" type="torch.LongTensor | None" default="None">Optional tensor of shape [batch, sequence].</ParamField><ParamField path="inputs_embeds" type="torch.FloatTensor | None" default="None">Optional tensor of shape [batch, sequence, hidden].</ParamField><ParamField path="labels" type="torch.LongTensor | None" default="None">Optional tensor of shape [batch] containing binary relevance classes(0 for no, 1 for yes), not next-token IDs. The training recipe insteadcomputes its own loss over each query's candidate passages.</ParamField><ParamField path="**kwargs" type="Any" default="{}">Hugging Face decoder options. Per-position vocabulary logits andgeneration through this training class are unsupported.</ParamField>**Returns:** `SequenceClassifierOutputWithPast`SequenceClassifierOutputWithPast with logits of shape [batch, 1], holding</Indent><Anchor id="nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerForCausalReranking-from_pretrained"><CodeBlock links={{"nemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerForCausalReranking":"#nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerForCausalReranking"}} showLineNumbers={false} wordWrap={true}>```pythonnemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerForCausalReranking.from_pretrained(pretrained_model_name_or_path: str | os.PathLike[str] | None,args: typing.Any = (),kwargs: typing.Any = {}) -> nemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerForCausalReranking```</CodeBlock></Anchor><Indent><Badge>classmethod</Badge>Load weights and resolve the "yes"/"no" token ids if not already set.Explicit ``yes_token_id``/``no_token_id`` (e.g. from the recipe YAML or asaved config) take precedence; otherwise they are resolved from thetokenizer of ``pretrained_model_name_or_path``.In-memory loads use ``config.name_or_path`` to locate the tokenizer.A checkpoint's saved embedding ties cannot be overridden in either direction.</Indent><Anchor id="nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerForCausalReranking-supports_config"><CodeBlock showLineNumbers={false} wordWrap={true}>```pythonnemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerForCausalReranking.supports_config(config: transformers.PretrainedConfig) -> bool```</CodeBlock></Anchor><Indent><Badge>classmethod</Badge>Use causal reranking only for checkpoints declaring a compatible LM head.</Indent><Anchor id="nemo_automodel-components-models-qwen3_reranker-model-Qwen3RerankerForCausalReranking-tie_weights"><CodeBlock showLineNumbers={false} wordWrap={true}>```pythonnemo_automodel.components.models.qwen3_reranker.model.Qwen3RerankerForCausalReranking.tie_weights(_args: object = (),_kwargs: object = {}) -> None```</CodeBlock></Anchor><Indent>Alias ``lm_head`` to the input embeddings when the config asks for it.Declared model-locally rather than inherited: transformers v5 does not reliably tie acustom model from the dict-shaped ``_tied_weights_keys`` alone, so the config flag ishonoured explicitly (mirroring ``Qwen2ForCausalLM``). This matters here becausescoring reads the yes/no rows of ``lm_head``; if the tie silently failed to apply, thehead would drift from the embeddings it is supposed to share and the publishedcheckpoint's scores would not reproduce.</Indent></Indent><Anchor id="nemo_automodel-components-models-qwen3_reranker-model-_last_token_indices"><CodeBlock showLineNumbers={false} wordWrap={true}>```pythonnemo_automodel.components.models.qwen3_reranker.model._last_token_indices(attention_mask: torch.Tensor) -> torch.Tensor```</CodeBlock></Anchor><Indent>Return the index of the last attended token for each row.Padding-side agnostic: the result is the highest position index that the maskattends to, so left padding resolves to the final position and right padding tothe last real token, with no assumption about which side a batch uses.**Parameters:**<ParamField path="attention_mask" type="torch.Tensor">Tensor of shape [batch, sequence], 1 for attended tokens and0 for padding. Each row is expected to attend to at least one token.</ParamField>**Returns:** `torch.Tensor`Tensor of shape [batch] holding the last attended index per row.</Indent><Anchor id="nemo_automodel-components-models-qwen3_reranker-model-_validate_reranker_options"><CodeBlock showLineNumbers={false} wordWrap={true}>```pythonnemo_automodel.components.models.qwen3_reranker.model._validate_reranker_options(num_labels: int | None = None,pooling: str | None = None,temperature: float | None = None) -> None```</CodeBlock></Anchor><Indent>Reject classifier options that cannot describe a single raw yes/no score.</Indent><Anchor id="nemo_automodel-components-models-qwen3_reranker-model-ModelClass"><CodeBlock showLineNumbers={false} wordWrap={true}>```pythonnemo_automodel.components.models.qwen3_reranker.model.ModelClass = [Qwen3RerankerForCausalReranking]```</CodeBlock></Anchor><Anchor id="nemo_automodel-components-models-qwen3_reranker-model-logger"><CodeBlock showLineNumbers={false} wordWrap={true}>```pythonnemo_automodel.components.models.qwen3_reranker.model.logger = logging.get_logger(__name__)```</CodeBlock></Anchor><style>{`.light .fern-code-block,.light .fern-prose code:not(.code-block) {background-color: var(--nv-color-bg-alt, #f7f7f7) !important;}.dark .fern-code-block,.dark .fern-prose code:not(.code-block) {background-color: #1f1f1f !important;}`}</style>