ReferenceFull Library ReferenceNemo AutomodelNemo AutomodelComponentsModelsnemo_automodel.components.models.llama_nemotron_vl

nemo_automodel.components.models.llama_nemotron_vl

View as Markdown

Llama Nemotron VL model for multimodal embedding and retrieval tasks.

Submodules

Package Contents

Classes

NameDescription
LlamaNemotronVLConfigBase configuration for vision-language models combining vision and language components.
LlamaNemotronVLModelLlamaNemotron VL model for vision-language reranking.
LlamaNemotronVLProcessorProcessor for LlamaNemotronVL model.

API

class nemo_automodel.components.models.llama_nemotron_vl.model.LlamaNemotronVLConfig(
vision_config = None,
llm_config = None,
use_backbone_lora = 0,
use_llm_lora = 0,
select_layer = -1,
force_image_size = None,
downsample_ratio = 0.5,
template = None,
dynamic_image_size = False,
use_thumbnail = False,
min_dynamic_patch = 1,
max_dynamic_patch = 6,
mlp_checkpoint = True,
pre_feature_reduction = False,
keep_aspect_ratio = False,
vocab_size = -1,
q_max_length: int | None = 512,
p_max_length: int | None = 10240,
query_prefix: str = 'query:',
passage_prefix: str = 'passage:',
pooling: str = 'last',
bidirectional_attention: bool = False,
max_input_tiles: int = 2,
img_context_token_id: int = 128258,
kwargs = {}
)

Bases: PretrainedConfig

Base configuration for vision-language models combining vision and language components. This serves as the foundation for LlamaNemotronVL configurations.

llm_config
= LlamaBidirectionalConfig(**llm_config)
model_type
= 'llama_nemotron_vl'
sub_configs
vision_config
= SiglipVisionConfig(**vision_config)
vocab_size
= self.llm_config.vocab_size
nemo_automodel.components.models.llama_nemotron_vl.model.LlamaNemotronVLConfig.get_text_config(
decoder: bool | None = None,
encoder: bool | None = None
) -> transformers.configuration_utils.PretrainedConfig

Return the language decoder config using the Transformers composite-config contract.

class nemo_automodel.components.models.llama_nemotron_vl.model.LlamaNemotronVLModel(
vision_model: transformers.PreTrainedModel | None = None,
language_model: transformers.PreTrainedModel | None = None
)

Bases: PreTrainedModel

LlamaNemotron VL model for vision-language reranking. Combines a vision encoder (SigLIP) with a bidirectional language model (LLaMA) for cross-modal reranking tasks.

_no_split_modules
= ['LlamaDecoderLayer']
downsample_ratio
= config.downsample_ratio
main_input_name
= 'pixel_values'
mlp1
num_image_token
= int((grid_size * config.downsample_ratio) ** 2)
patch_size
= 14
processor
select_layer
= config.select_layer
template
= config.template
nemo_automodel.components.models.llama_nemotron_vl.model.LlamaNemotronVLModel._embed_batch(
inputs: typing.Dict[str, typing.Any],
pool_type: str | None = None
)

Encodes the inputs into a tensor of embeddings. Args: inputs: A dictionary of inputs to the model. You can prepare the inputs using the processor.process_queries and processor.process_documents methods. pool_type: The type of pooling to use. If None, the pooling type is set to the pooling type configured in the model. Returns: A tensor of embeddings.

nemo_automodel.components.models.llama_nemotron_vl.model.LlamaNemotronVLModel.build_collator(
processor = None,
kwargs = {}
)
nemo_automodel.components.models.llama_nemotron_vl.model.LlamaNemotronVLModel.encode_documents(
images: typing.List[typing.Any] | None = None,
texts: typing.List[str] | None = None,
kwargs = {}
)

Encodes the input document images and texts into a tensor of embeddings. Args: images: A list of PIL.Image of document pages images. texts: A list of document page texts. Returns: A tensor of embeddings.

nemo_automodel.components.models.llama_nemotron_vl.model.LlamaNemotronVLModel.encode_queries(
queries: typing.List[str],
kwargs = {}
)

Encodes the input queries into a tensor of embeddings. Args: queries: A list of queries. Returns: A tensor of embeddings.

nemo_automodel.components.models.llama_nemotron_vl.model.LlamaNemotronVLModel.extract_feature(
pixel_values
)

Extract and project vision features to language model space.

nemo_automodel.components.models.llama_nemotron_vl.model.LlamaNemotronVLModel.forward(
pixel_values: torch.FloatTensor = None,
input_ids: torch.LongTensor = None,
attention_mask: torch.Tensor | None = None,
position_ids: torch.LongTensor | None = None,
image_flags: torch.LongTensor | None = None,
past_key_values: typing.List[torch.FloatTensor] | None = None,
labels: torch.LongTensor | None = None,
use_cache: bool | None = None,
output_attentions: bool | None = None,
output_hidden_states: bool | None = None,
return_dict: bool | None = None,
num_patches_list: typing.List[torch.Tensor] | None = None,
run_dummy_vision: bool | None = None
) -> typing.Union[typing.Tuple, transformers.modeling_outputs.CausalLMOutputWithPast]
nemo_automodel.components.models.llama_nemotron_vl.model.LlamaNemotronVLModel.get_decoder() -> transformers.PreTrainedModel

Return the language tower used as the retrieval text decoder.

nemo_automodel.components.models.llama_nemotron_vl.model.LlamaNemotronVLModel.get_input_embeddings()
nemo_automodel.components.models.llama_nemotron_vl.model.LlamaNemotronVLModel.get_output_embeddings()
nemo_automodel.components.models.llama_nemotron_vl.model.LlamaNemotronVLModel.pixel_shuffle(
x,
scale_factor = 0.5
)
nemo_automodel.components.models.llama_nemotron_vl.model.LlamaNemotronVLModel.post_loss(
loss,
inputs
)
class nemo_automodel.components.models.llama_nemotron_vl.processor.LlamaNemotronVLProcessor(
tokenizer: typing.Any,
q_max_length: int | None = None,
p_max_length: int | None = None,
pad_to_multiple_of: int | None = None,
query_prefix: str = 'query:',
passage_prefix: str = 'passage:',
max_input_tiles: int = 6,
num_image_token: int = 256,
dynamic_image_size: bool = True,
image_size: int = 512,
use_thumbnail: bool = True,
template: str = 'bidirectional-llama-retrie...,
num_channels: int = 3,
norm_type: str = 'siglip',
system_message: str = '',
padding: typing.Union[bool, str] = True,
kwargs = {}
)

Bases: ProcessorMixin

Processor for LlamaNemotronVL model.

attributes
= ['tokenizer']
image_processor
tokenizer_class
= 'AutoTokenizer'
nemo_automodel.components.models.llama_nemotron_vl.processor.LlamaNemotronVLProcessor.__call__(
text: typing.List[str] | None = None,
images: typing.List[typing.Any] | None = None,
text_kwargs: typing.Dict[str, typing.Any] | None = None,
images_kwargs: typing.Dict[str, typing.Any] | None = None,
common_kwargs: typing.Dict[str, typing.Any] | None = None,
kwargs = {}
) -> typing.Dict[str, typing.Any]

Process text and/or image inputs into model-ready features. This method provides compatibility with the standard HuggingFace processor interface used by Sentence Transformers. For image inputs, it delegates to process_documents. For text-only inputs, it tokenizes directly (assuming any task prefix has already been applied by the caller). Args: text: List of text strings. For text-only inputs, these should already include any task prefix (e.g. “query: ” or “passage: ”). images: List of PIL Images for document encoding. text_kwargs: Keyword arguments for text processing (e.g. padding, truncation). images_kwargs: Keyword arguments for image processing (unused, for API compat). common_kwargs: Common keyword arguments (e.g. return_tensors). **kwargs: Additional keyword arguments (ignored). Returns: Dict with “input_ids”, “attention_mask”, and optionally “pixel_values”.

nemo_automodel.components.models.llama_nemotron_vl.processor.LlamaNemotronVLProcessor.add_dummy_labels(
questions,
merged_batch_dict
)
nemo_automodel.components.models.llama_nemotron_vl.processor.LlamaNemotronVLProcessor.merge_batch_dict(
query_batch_dict,
doc_batch_dict
)
nemo_automodel.components.models.llama_nemotron_vl.processor.LlamaNemotronVLProcessor.process_documents(
documents: typing.Union[typing.Dict, typing.List[typing.Dict]],
return_tensors: typing.Literal['pt', 'np'] = 'pt',
padding: bool | str | None = None,
truncation: bool = True,
pixel_values_layout: typing.Literal['per_image', 'flat_tiles'] = 'flat_tiles',
kwargs = {}
) -> typing.Dict[str, typing.Any]

Process documents into model inputs with tokenized text and pixel values. Args: documents: Either a dict with “images” and “texts” lists, or a list of dicts each with “image” and “text” keys. Images can be PIL Images, file paths, or None/empty string for text-only documents. return_tensors: Output format — “pt” for PyTorch tensors, “np” for numpy arrays. padding: Padding strategy passed to the tokenizer. Defaults to the value set in the processor constructor. truncation: Whether to truncate sequences to p_max_length. pixel_values_layout: How to structure the pixel values output:

  • “flat_tiles”: All image tiles concatenated into a single tensor of shape (total_tiles, C, H, W). Different images may contribute different numbers of tiles. None if no images are present. This is the format expected by the model’s forward() method.
  • “per_image”: A list aligned with the input documents, where each entry is either a tensor of shape (num_tiles, C, H, W) or None. Returns: Dict with “input_ids”, “attention_mask”, and “pixel_values”.
nemo_automodel.components.models.llama_nemotron_vl.processor.LlamaNemotronVLProcessor.process_queries(
queries: typing.List[str],
return_tensors: typing.Literal['pt', 'np'] = 'pt',
padding: bool | str | None = None,
truncation: bool = True,
kwargs = {}
) -> transformers.BatchEncoding

Process queries into model inputs with tokenized text. Args: queries: List of query strings. return_tensors: Output format — “pt” for PyTorch tensors, “np” for numpy arrays. padding: Padding strategy passed to the tokenizer. Defaults to the value set in the processor constructor. truncation: Whether to truncate sequences to q_max_length. Returns: Dict with “input_ids” and “attention_mask”.

nemo_automodel.components.models.llama_nemotron_vl.processor.LlamaNemotronVLProcessor.process_queries_documents_biencoder(
features: typing.Dict,
kwargs = {}
) -> typing.Dict[str, typing.Any]

(Pdb) features [{‘image’: [<PIL.Image.Image image mode=RGB size=1275x1650 at 0x155059A5C3A0>, <PIL.Image.Image image mode=RGB size=1275x1650 at 0x155059A5C580>, <PIL.Image.Image image mode=RGB size=1275x1650 at 0x155059A5C940>], ‘text’: [‘passage: ’, ‘passage: ’, ‘passage: ’], ‘question’: “query: What change did Carl Rey suggest for the Strategic Plan’s website objective deadline?”}, {‘image’: [<PIL.Image.Image image mode=RGB size=1275x1650 at 0x155059A5C0D0>, <PIL.Image.Image image mode=RGB size=1275x1650 at 0x155059A5DC00>, <PIL.Image.Image image mode=RGB size=1275x1650 at 0x155059A5EBF0>], ‘text’: [‘passage: ’, ‘passage: ’, ‘passage: ’], ‘question’: ‘query: What are the name and TIN requirements for individuals with real estate transactions?’}, {‘image’: [<PIL.Image.Image image mode=RGB size=1275x1650 at 0x155059A5D390>, <PIL.Image.Image image mode=RGB size=1275x1650 at 0x155059A5C850>, <PIL.Image.Image image mode=RGB size=1275x1650 at 0x155059A5C070>], ‘text’: [‘passage: ’, ‘passage: ’, ‘passage: ’], ‘question’: ‘query: How does Richard Hooker view human inclinations?’}]