ReferenceFull Library ReferenceNemo AutomodelNemo AutomodelComponentsModelsMinistral Bidirectionalnemo_automodel.components.models.ministral_bidirectional.reranker_model

nemo_automodel.components.models.ministral_bidirectional.reranker_model

View as Markdown

Portable Transformers implementation of pooled Mistral3 sequence scoring.

Module Contents

Classes

NameDescription
Mistral3ForSequenceClassificationMistral3 backbone with masked pooling and an FP32 scoring projection.

API

class nemo_automodel.components.models.ministral_bidirectional.reranker_model.Mistral3ForSequenceClassification(
config: transformers.Mistral3Config
)

Bases: Mistral3PreTrainedModel

Mistral3 backbone with masked pooling and an FP32 scoring projection.

base_model_prefix
= 'model'
model
= Mistral3Model(config)
score
nemo_automodel.components.models.ministral_bidirectional.reranker_model.Mistral3ForSequenceClassification.forward(
input_ids: torch.Tensor | None = None,
attention_mask: torch.Tensor | None = None,
pixel_values: torch.Tensor | None = None,
image_sizes: torch.Tensor | None = None,
position_ids: torch.Tensor | None = None,
kwargs: typing.Any = {}
) -> transformers.modeling_outputs.SequenceClassifierOutputWithPast

Score tokenized text or multimodal pairs.

Parameters:

input_ids
torch.Tensor | NoneDefaults to None

Integer tensor of shape [batch, sequence].

attention_mask
torch.Tensor | NoneDefaults to None

Padding mask of shape [batch, sequence]; required for pooling.

pixel_values
torch.Tensor | NoneDefaults to None

Optional image tensor of shape [images, channels, height, width].

image_sizes
torch.Tensor | NoneDefaults to None

Optional integer tensor of shape [images, 2], height then width.

position_ids
torch.Tensor | NoneDefaults to None

Optional integer tensor of shape [batch, sequence].

**kwargs
AnyDefaults to {}

Additional Mistral3Model inputs, following its forward tensor contract.

Returns: SequenceClassifierOutputWithPast

Classifier output with raw, unscaled FP32 logits of shape [batch, num_labels].

classmethod

Load the shared classifier head without changing global Transformers conversion mappings.