nemo_automodel.components.models.llama_nemotron_vl
nemo_automodel.components.models.llama_nemotron_vl
Llama Nemotron VL model for multimodal embedding and retrieval tasks.
Submodules
nemo_automodel.components.models.llama_nemotron_vl.modelnemo_automodel.components.models.llama_nemotron_vl.processor
Package Contents
Classes
API
Bases: PretrainedConfig
Base configuration for vision-language models combining vision and language components. This serves as the foundation for LlamaNemotronVL configurations.
Return the language decoder config using the Transformers composite-config contract.
Bases: PreTrainedModel
LlamaNemotron VL model for vision-language reranking. Combines a vision encoder (SigLIP) with a bidirectional language model (LLaMA) for cross-modal reranking tasks.
Encodes the inputs into a tensor of embeddings. Args: inputs: A dictionary of inputs to the model. You can prepare the inputs using the processor.process_queries and processor.process_documents methods. pool_type: The type of pooling to use. If None, the pooling type is set to the pooling type configured in the model. Returns: A tensor of embeddings.
Encodes the input document images and texts into a tensor of embeddings. Args: images: A list of PIL.Image of document pages images. texts: A list of document page texts. Returns: A tensor of embeddings.
Encodes the input queries into a tensor of embeddings. Args: queries: A list of queries. Returns: A tensor of embeddings.
Extract and project vision features to language model space.
Return the language tower used as the retrieval text decoder.
Bases: ProcessorMixin
Processor for LlamaNemotronVL model.
Process text and/or image inputs into model-ready features. This method provides compatibility with the standard HuggingFace processor interface used by Sentence Transformers. For image inputs, it delegates to process_documents. For text-only inputs, it tokenizes directly (assuming any task prefix has already been applied by the caller). Args: text: List of text strings. For text-only inputs, these should already include any task prefix (e.g. “query: ” or “passage: ”). images: List of PIL Images for document encoding. text_kwargs: Keyword arguments for text processing (e.g. padding, truncation). images_kwargs: Keyword arguments for image processing (unused, for API compat). common_kwargs: Common keyword arguments (e.g. return_tensors). **kwargs: Additional keyword arguments (ignored). Returns: Dict with “input_ids”, “attention_mask”, and optionally “pixel_values”.
Process documents into model inputs with tokenized text and pixel values. Args: documents: Either a dict with “images” and “texts” lists, or a list of dicts each with “image” and “text” keys. Images can be PIL Images, file paths, or None/empty string for text-only documents. return_tensors: Output format — “pt” for PyTorch tensors, “np” for numpy arrays. padding: Padding strategy passed to the tokenizer. Defaults to the value set in the processor constructor. truncation: Whether to truncate sequences to p_max_length. pixel_values_layout: How to structure the pixel values output:
- “flat_tiles”: All image tiles concatenated into a single tensor of shape (total_tiles, C, H, W). Different images may contribute different numbers of tiles. None if no images are present. This is the format expected by the model’s forward() method.
- “per_image”: A list aligned with the input documents, where each entry is either a tensor of shape (num_tiles, C, H, W) or None. Returns: Dict with “input_ids”, “attention_mask”, and “pixel_values”.
Process queries into model inputs with tokenized text. Args: queries: List of query strings. return_tensors: Output format — “pt” for PyTorch tensors, “np” for numpy arrays. padding: Padding strategy passed to the tokenizer. Defaults to the value set in the processor constructor. truncation: Whether to truncate sequences to q_max_length. Returns: Dict with “input_ids” and “attention_mask”.
(Pdb) features [{‘image’: [<PIL.Image.Image image mode=RGB size=1275x1650 at 0x155059A5C3A0>, <PIL.Image.Image image mode=RGB size=1275x1650 at 0x155059A5C580>, <PIL.Image.Image image mode=RGB size=1275x1650 at 0x155059A5C940>], ‘text’: [‘passage: ’, ‘passage: ’, ‘passage: ’], ‘question’: “query: What change did Carl Rey suggest for the Strategic Plan’s website objective deadline?”}, {‘image’: [<PIL.Image.Image image mode=RGB size=1275x1650 at 0x155059A5C0D0>, <PIL.Image.Image image mode=RGB size=1275x1650 at 0x155059A5DC00>, <PIL.Image.Image image mode=RGB size=1275x1650 at 0x155059A5EBF0>], ‘text’: [‘passage: ’, ‘passage: ’, ‘passage: ’], ‘question’: ‘query: What are the name and TIN requirements for individuals with real estate transactions?’}, {‘image’: [<PIL.Image.Image image mode=RGB size=1275x1650 at 0x155059A5D390>, <PIL.Image.Image image mode=RGB size=1275x1650 at 0x155059A5C850>, <PIL.Image.Image image mode=RGB size=1275x1650 at 0x155059A5C070>], ‘text’: [‘passage: ’, ‘passage: ’, ‘passage: ’], ‘question’: ‘query: How does Richard Hooker view human inclinations?’}]