nemo_automodel.components.models.ministral_bidirectional.processor
nemo_automodel.components.models.ministral_bidirectional.processor
Pixtral processor extensions for Ministral3 retrieval models.
Module Contents
Classes
Functions
Data
_MISTRAL_RETRIEVAL_CHAT_TEMPLATE
API
Bases: PixtralProcessor
Pixtral processor with retrieval-specific query/document batching helpers.
The persistent ownership policy is applied when the processor is constructed: its tokenizer treats raw control markers as processor-owned, while its chat template marks externally supplied text for ordinary tokenization. Direct retrieval query/document helpers apply the same text marker because they do not render the chat template. Training helpers and the retrieval template normalize trailing ASCII spaces in query/document prefixes and add one separator before content. This class always installs its built-in retrieval template, including when loading a checkpoint, so training and export share that policy. Custom saved templates are not preserved by this training processor.
Saving defaults to the stock Pixtral class for portable bi-encoder checkpoints; cross-encoder recipes preserve this class because their evaluation path requires its custom batching helper.
Return the canonical longest edge used by the image processor.
Load the processor tokenizer directly, bypassing automatic backend selection.
Mark user-owned control-token prefixes for ordinary tokenization.
Tokenize text and optional images with the retrieval length policy.
Parameters:
Text strings in batch order.
Maximum sequence length when truncation is enabled.
Output backend, "pt" or "np".
Padding strategy, or None to use the processor default.
Whether to truncate text tokens.
Optional PIL images aligned with image placeholders in text.
Additional arguments forwarded to PixtralProcessor.
Returns: BatchEncoding
Mapping with integer input_ids and attention_mask of shape
Keep standard processor defaults independent of the training truncation limit.
Add dummy labels expected by the retrieval training loop.
Parameters:
Query strings defining the batch size.
Merged processor output. Tensor or NumPy array values may have arbitrary shapes and are preserved without modification.
Backend for the labels, "pt" or "np".
Returns: dict[str, Any]
The input mapping, mutated in place with integer labels of shape [batch].
Validate the retrieval tokenizer without importing MistralCommonBackend.
Transformers validates processor tokenizers against every generally supported
tokenizer class, which imports mistral-common even when the selected tokenizer
is already a TokenizersBackend. This processor supports only the latter so its
image markers and escaped user text retain the retrieval-specific semantics.
Parameters:
Processor attribute being validated.
Runtime processor component.
Returns: type | tuple[type, ...]
The accepted component class, or the superclass validation result.
Raises:
TypeError: If a non-TokenizersBackend tokenizer is supplied.
Restore saved retrieval prompts unless explicitly overridden.
Parameters:
Local checkpoint or Hugging Face model identifier.
Standard processor loading arguments and explicit runtime overrides.
Returns: 'Mistral3BiEncoderProcessor'
Processor with Sentence Transformers query/document prompts and saved image settings.
Raises:
ValueError: If saved retrieval prompt metadata is malformed or a prompt-policy override conflicts with the saved retrieval chat template.
Return the behaviorally equivalent stock processor for portable export.
Subclasses must explicitly re-assert _export_as_stock_processor = True.
This prevents inherited export from silently discarding custom image or
text preprocessing that stock PixtralProcessor cannot reproduce.
Returns: PixtralProcessor
Stock processor configured with the persistent retrieval behavior.
Raises:
TypeError: If a subclass has not explicitly opted into stock export.
Prefix and merge query and document processor outputs.
Parameters:
Query mapping, typically containing input_ids
and attention_mask of shape [batch, query_sequence]. Tensor
or NumPy values of arbitrary rank and non-array values are accepted.
Document mapping, typically containing input_ids
and attention_mask of shape [documents, document_sequence],
pixel_values of shape [images, channels, height, width], and
image_sizes of shape [images, 2] in height-width order. Image
fields may be None; other arbitrary-rank values are also accepted.
Returns: dict[str, Any]
One mapping with query keys prefixed by q_ and document keys by d_.
Process text and image documents into model inputs.
Parameters:
Either a dict with images and texts lists, or a list
of dicts with image and text keys.
Output format, "pt" or "np".
Padding strategy for tokenization.
Whether to truncate document tokens.
Extra keyword arguments forwarded to PixtralProcessor.
Returns: dict[str, Any]
A mapping with integer input_ids and attention_mask of shape
Process query strings into tokenized model inputs.
Parameters:
Query texts for a batch of size batch.
Output format, "pt" or "np".
Padding strategy for tokenization.
Whether to truncate query tokens.
Extra keyword arguments forwarded to PixtralProcessor.
Returns: BatchEncoding
A mapping with integer input_ids and attention_mask of shape
Process grouped retrieval examples into the bi-encoder batch format.
Parameters:
Query examples with aligned doc_text and doc_image
candidate lists. Optional doc_id lists must contain one
string per candidate. An all-empty list means IDs are absent;
populated IDs must be present for every candidate in the batch.
Distributed callers must use the same ID policy on every rank.
Output format, "pt" or "np".
Extra keyword arguments forwarded to the query and document processors.
Returns: dict[str, Any]
A mapping with q_input_ids and q_attention_mask of shape
Raises:
ValueError: Supplied document IDs are partially populated or misaligned.
Process flattened query-document pairs for vision cross-encoder training.
Each feature contains one question, doc_text, and optional
doc_image. Image, text, and image-plus-text documents are supported.
Parameters:
Flattened query-document examples. num_labels contains
the number of original query groups when labels are requested.
Output format, "pt" or "np".
Padding strategy for tokenization.
Whether to truncate the combined sequences.
Extra keyword arguments forwarded to PixtralProcessor.
Returns: dict[str, Any]
A mapping with integer input_ids and attention_mask of shape
Format one question-passage pair for cross-encoder scoring.
Parameters:
Query text.
Passage text, including an image placeholder when applicable.
Returns: str
The formatted question-passage text.
Save either a stock Pixtral processor or this custom processor.
Bi-encoder processors default to a portable stock export. Cross-encoder
recipes disable stock export because evaluation requires
process_queries_documents_crossencoder from this class.
Parameters:
Directory for the standard Transformers processor assets.
Whether to upload the saved assets to the Hugging Face Hub.
Additional arguments forwarded to the standard processor save.
Returns: list[str]
Paths written by the selected processor save operation.
Bases: enum.Enum
Document modality encoded in a multimodal retrieval batch.
Configure persistent tokenizer and retrieval-template ownership behavior.
Load one retrieval image without fetching remote content.
Parameters:
PIL image, local path, or a mapping containing disk_path,
base64, or bytes. Encoded payloads require the mapping form.
Returns: Image.Image
The decoded PIL image.
Raises:
ValueError: If the input is unsupported or contains a remote URL.