ReferenceFull Library ReferenceNemo AutomodelNemo AutomodelComponentsModelsMinistral Bidirectionalnemo_automodel.components.models.ministral_bidirectional.processor

nemo_automodel.components.models.ministral_bidirectional.processor

View as Markdown

Pixtral processor extensions for Ministral3 retrieval models.

Module Contents

Classes

NameDescription
Mistral3BiEncoderProcessorPixtral processor with retrieval-specific query/document batching helpers.
PassageModalityDocument modality encoded in a multimodal retrieval batch.

Functions

NameDescription
_apply_mistral_retrieval_ownership_policyConfigure persistent tokenizer and retrieval-template ownership behavior.
load_imageLoad one retrieval image without fetching remote content.

Data

_CONTROL_TOKEN_ESCAPE

_CONTROL_TOKEN_ESCAPE_POLICY

_CONTROL_TOKEN_PREFIXES

_MISTRAL_RETRIEVAL_CHAT_TEMPLATE

API

class nemo_automodel.components.models.ministral_bidirectional.processor.Mistral3BiEncoderProcessor(
image_processor: typing.Any = None,
tokenizer: typing.Any = None,
patch_size: int = 16,
spatial_merge_size: int = 1,
chat_template: str | None = None,
image_token: str = '[IMG]',
image_break_token: str = '[IMG_BREAK]',
image_end_token: str = '[IMG_END]',
q_max_length: int | None = None,
p_max_length: int | None = None,
rerank_max_length: int | None = None,
q_max_len: int | None = None,
p_max_len: int | None = None,
pad_to_multiple_of: int | None = None,
query_prefix: str = 'query:',
passage_prefix: str = 'passage:',
padding: bool | str = True,
image_longest_edge: int | None = None,
use_prompt_template: bool = False,
export_as_stock_processor: bool = True,
kwargs: typing.Any = {}
)

Bases: PixtralProcessor

Pixtral processor with retrieval-specific query/document batching helpers.

The persistent ownership policy is applied when the processor is constructed: its tokenizer treats raw control markers as processor-owned, while its chat template marks externally supplied text for ordinary tokenization. Direct retrieval query/document helpers apply the same text marker because they do not render the chat template. Training helpers and the retrieval template normalize trailing ASCII spaces in query/document prefixes and add one separator before content. This class always installs its built-in retrieval template, including when loading a checkpoint, so training and export share that policy. Custom saved templates are not preserved by this training processor.

Saving defaults to the stock Pixtral class for portable bi-encoder checkpoints; cross-encoder recipes preserve this class because their evaluation path requires its custom batching helper.

chat_template
image_longest_edge
int | None

Return the canonical longest edge used by the image processor.

p_max_length
passage_prefix
= passage_prefix.rstrip(' ')
q_max_length
query_prefix
= query_prefix.rstrip(' ')
nemo_automodel.components.models.ministral_bidirectional.processor.Mistral3BiEncoderProcessor._extract_document_fields(
documents: dict[str, list[typing.Any]] | list[dict[str, typing.Any]]
) -> tuple[list[typing.Any], list[typing.Any]]
staticmethod
nemo_automodel.components.models.ministral_bidirectional.processor.Mistral3BiEncoderProcessor._load_tokenizer_from_pretrained(
sub_processor_type: str,
pretrained_model_name_or_path: str,
subfolder: str = '',
kwargs: typing.Any = {}
) -> transformers.tokenization_utils_tokenizers.TokenizersBackend
classmethod

Load the processor tokenizer directly, bypassing automatic backend selection.

nemo_automodel.components.models.ministral_bidirectional.processor.Mistral3BiEncoderProcessor._mark_user_text_ownership(
text: str
) -> str
staticmethod

Mark user-owned control-token prefixes for ordinary tokenization.

nemo_automodel.components.models.ministral_bidirectional.processor.Mistral3BiEncoderProcessor._process_text(
text: list[str],
max_length: int | None,
return_tensors: typing.Literal['pt', 'np'],
padding: bool | str | None,
truncation: bool,
images: list[PIL.Image.Image] | None = None,
kwargs: typing.Any = {}
) -> transformers.BatchEncoding

Tokenize text and optional images with the retrieval length policy.

Parameters:

text
list[str]

Text strings in batch order.

max_length
int | None

Maximum sequence length when truncation is enabled.

return_tensors
Literal['pt', 'np']

Output backend, "pt" or "np".

padding
bool | str | None

Padding strategy, or None to use the processor default.

truncation
bool

Whether to truncate text tokens.

images
list[Image.Image] | NoneDefaults to None

Optional PIL images aligned with image placeholders in text.

**kwargs
AnyDefaults to {}

Additional arguments forwarded to PixtralProcessor.

Returns: BatchEncoding

Mapping with integer input_ids and attention_mask of shape

nemo_automodel.components.models.ministral_bidirectional.processor.Mistral3BiEncoderProcessor._sync_reranker_defaults() -> None

Keep standard processor defaults independent of the training truncation limit.

nemo_automodel.components.models.ministral_bidirectional.processor.Mistral3BiEncoderProcessor.add_dummy_labels(
questions: list[str],
merged_batch_dict: dict[str, typing.Any],
return_tensors: typing.Literal['pt', 'np'] = 'pt'
) -> dict[str, typing.Any]

Add dummy labels expected by the retrieval training loop.

Parameters:

questions
list[str]

Query strings defining the batch size.

merged_batch_dict
dict[str, Any]

Merged processor output. Tensor or NumPy array values may have arbitrary shapes and are preserved without modification.

return_tensors
Literal['pt', 'np']Defaults to 'pt'

Backend for the labels, "pt" or "np".

Returns: dict[str, Any]

The input mapping, mutated in place with integer labels of shape [batch].

nemo_automodel.components.models.ministral_bidirectional.processor.Mistral3BiEncoderProcessor.check_argument_for_proper_class(
argument_name: str,
argument: typing.Any
) -> type | tuple[type, ...]

Validate the retrieval tokenizer without importing MistralCommonBackend.

Transformers validates processor tokenizers against every generally supported tokenizer class, which imports mistral-common even when the selected tokenizer is already a TokenizersBackend. This processor supports only the latter so its image markers and escaped user text retain the retrieval-specific semantics.

Parameters:

argument_name
str

Processor attribute being validated.

argument
Any

Runtime processor component.

Returns: type | tuple[type, ...]

The accepted component class, or the superclass validation result.

Raises:

  • TypeError: If a non-TokenizersBackend tokenizer is supplied.
nemo_automodel.components.models.ministral_bidirectional.processor.Mistral3BiEncoderProcessor.from_pretrained(
pretrained_model_name_or_path: str | os.PathLike,
kwargs: typing.Any = {}
) -> 'Mistral3BiEncoderProcessor'
classmethod

Restore saved retrieval prompts unless explicitly overridden.

Parameters:

pretrained_model_name_or_path
str | os.PathLike

Local checkpoint or Hugging Face model identifier.

**kwargs
AnyDefaults to {}

Standard processor loading arguments and explicit runtime overrides.

Returns: 'Mistral3BiEncoderProcessor'

Processor with Sentence Transformers query/document prompts and saved image settings.

Raises:

  • ValueError: If saved retrieval prompt metadata is malformed or a prompt-policy override conflicts with the saved retrieval chat template.
nemo_automodel.components.models.ministral_bidirectional.processor.Mistral3BiEncoderProcessor.get_hf_export_processor() -> transformers.PixtralProcessor

Return the behaviorally equivalent stock processor for portable export.

Subclasses must explicitly re-assert _export_as_stock_processor = True. This prevents inherited export from silently discarding custom image or text preprocessing that stock PixtralProcessor cannot reproduce.

Returns: PixtralProcessor

Stock processor configured with the persistent retrieval behavior.

Raises:

  • TypeError: If a subclass has not explicitly opted into stock export.
nemo_automodel.components.models.ministral_bidirectional.processor.Mistral3BiEncoderProcessor.merge_batch_dict(
query_batch_dict: dict[str, typing.Any],
doc_batch_dict: dict[str, typing.Any]
) -> dict[str, typing.Any]

Prefix and merge query and document processor outputs.

Parameters:

query_batch_dict
dict[str, Any]

Query mapping, typically containing input_ids and attention_mask of shape [batch, query_sequence]. Tensor or NumPy values of arbitrary rank and non-array values are accepted.

doc_batch_dict
dict[str, Any]

Document mapping, typically containing input_ids and attention_mask of shape [documents, document_sequence], pixel_values of shape [images, channels, height, width], and image_sizes of shape [images, 2] in height-width order. Image fields may be None; other arbitrary-rank values are also accepted.

Returns: dict[str, Any]

One mapping with query keys prefixed by q_ and document keys by d_.

nemo_automodel.components.models.ministral_bidirectional.processor.Mistral3BiEncoderProcessor.process_documents(
documents: dict[str, list[typing.Any]] | list[dict[str, typing.Any]],
return_tensors: typing.Literal['pt', 'np'] = 'pt',
padding: bool | str | None = None,
truncation: bool = True,
kwargs: typing.Any = {}
) -> dict[str, typing.Any]

Process text and image documents into model inputs.

Parameters:

documents
dict[str, list[Any]] | list[dict[str, Any]]

Either a dict with images and texts lists, or a list of dicts with image and text keys.

return_tensors
Literal['pt', 'np']Defaults to 'pt'

Output format, "pt" or "np".

padding
bool | str | NoneDefaults to None

Padding strategy for tokenization.

truncation
boolDefaults to True

Whether to truncate document tokens.

**kwargs
AnyDefaults to {}

Extra keyword arguments forwarded to PixtralProcessor.

Returns: dict[str, Any]

A mapping with integer input_ids and attention_mask of shape

nemo_automodel.components.models.ministral_bidirectional.processor.Mistral3BiEncoderProcessor.process_queries(
queries: list[str],
return_tensors: typing.Literal['pt', 'np'] = 'pt',
padding: bool | str | None = None,
truncation: bool = True,
kwargs: typing.Any = {}
) -> transformers.BatchEncoding

Process query strings into tokenized model inputs.

Parameters:

queries
list[str]

Query texts for a batch of size batch.

return_tensors
Literal['pt', 'np']Defaults to 'pt'

Output format, "pt" or "np".

padding
bool | str | NoneDefaults to None

Padding strategy for tokenization.

truncation
boolDefaults to True

Whether to truncate query tokens.

**kwargs
AnyDefaults to {}

Extra keyword arguments forwarded to PixtralProcessor.

Returns: BatchEncoding

A mapping with integer input_ids and attention_mask of shape

nemo_automodel.components.models.ministral_bidirectional.processor.Mistral3BiEncoderProcessor.process_queries_documents_biencoder(
features: list[dict[str, typing.Any]],
return_tensors: typing.Literal['pt', 'np'] = 'pt',
kwargs: typing.Any = {}
) -> dict[str, typing.Any]

Process grouped retrieval examples into the bi-encoder batch format.

Parameters:

features
list[dict[str, Any]]

Query examples with aligned doc_text and doc_image candidate lists. Optional doc_id lists must contain one string per candidate. An all-empty list means IDs are absent; populated IDs must be present for every candidate in the batch. Distributed callers must use the same ID policy on every rank.

return_tensors
Literal['pt', 'np']Defaults to 'pt'

Output format, "pt" or "np".

**kwargs
AnyDefaults to {}

Extra keyword arguments forwarded to the query and document processors.

Returns: dict[str, Any]

A mapping with q_input_ids and q_attention_mask of shape

Raises:

  • ValueError: Supplied document IDs are partially populated or misaligned.
nemo_automodel.components.models.ministral_bidirectional.processor.Mistral3BiEncoderProcessor.process_queries_documents_crossencoder(
features: list[dict[str, typing.Any]],
return_tensors: typing.Literal['pt', 'np'] = 'pt',
padding: bool | str | None = None,
truncation: bool = True,
kwargs: typing.Any = {}
) -> dict[str, typing.Any]

Process flattened query-document pairs for vision cross-encoder training.

Each feature contains one question, doc_text, and optional doc_image. Image, text, and image-plus-text documents are supported.

Parameters:

features
list[dict[str, Any]]

Flattened query-document examples. num_labels contains the number of original query groups when labels are requested.

return_tensors
Literal['pt', 'np']Defaults to 'pt'

Output format, "pt" or "np".

padding
bool | str | NoneDefaults to None

Padding strategy for tokenization.

truncation
boolDefaults to True

Whether to truncate the combined sequences.

**kwargs
AnyDefaults to {}

Extra keyword arguments forwarded to PixtralProcessor.

Returns: dict[str, Any]

A mapping with integer input_ids and attention_mask of shape

nemo_automodel.components.models.ministral_bidirectional.processor.Mistral3BiEncoderProcessor.prompt_template_question_passage(
question: str,
text: str
) -> str

Format one question-passage pair for cross-encoder scoring.

Parameters:

question
str

Query text.

text
str

Passage text, including an image placeholder when applicable.

Returns: str

The formatted question-passage text.

nemo_automodel.components.models.ministral_bidirectional.processor.Mistral3BiEncoderProcessor.save_pretrained(
save_directory: str | os.PathLike,
push_to_hub: bool = False,
kwargs: typing.Any = {}
) -> list[str]

Save either a stock Pixtral processor or this custom processor.

Bi-encoder processors default to a portable stock export. Cross-encoder recipes disable stock export because evaluation requires process_queries_documents_crossencoder from this class.

Parameters:

save_directory
str | os.PathLike

Directory for the standard Transformers processor assets.

push_to_hub
boolDefaults to False

Whether to upload the saved assets to the Hugging Face Hub.

**kwargs
AnyDefaults to {}

Additional arguments forwarded to the standard processor save.

Returns: list[str]

Paths written by the selected processor save operation.

class nemo_automodel.components.models.ministral_bidirectional.processor.PassageModality

Bases: enum.Enum

Document modality encoded in a multimodal retrieval batch.

IMAGE_ONLY
= 1
IMAGE_TEXT
= 2
TEXT_ONLY
= 0
nemo_automodel.components.models.ministral_bidirectional.processor._apply_mistral_retrieval_ownership_policy(
tokenizer: transformers.tokenization_utils_tokenizers.TokenizersBackend
) -> str

Configure persistent tokenizer and retrieval-template ownership behavior.

nemo_automodel.components.models.ministral_bidirectional.processor.load_image(
image: typing.Any
) -> PIL.Image.Image

Load one retrieval image without fetching remote content.

Parameters:

image
Any

PIL image, local path, or a mapping containing disk_path, base64, or bytes. Encoded payloads require the mapping form.

Returns: Image.Image

The decoded PIL image.

Raises:

  • ValueError: If the input is unsupported or contains a remote URL.
nemo_automodel.components.models.ministral_bidirectional.processor._CONTROL_TOKEN_ESCAPE = '\u200c'
nemo_automodel.components.models.ministral_bidirectional.processor._CONTROL_TOKEN_ESCAPE_POLICY = 'mistral_retrieval_zwnj_prefix_v1'
nemo_automodel.components.models.ministral_bidirectional.processor._CONTROL_TOKEN_PREFIXES = ('[', '<')
nemo_automodel.components.models.ministral_bidirectional.processor._MISTRAL_RETRIEVAL_CHAT_TEMPLATE = '\n{%- if messages | length == 2 and messages[0][\'role\'] == \'query\' and mess...