nemo_automodel.components.models.llama_nemotron_vl.model
nemo_automodel.components.models.llama_nemotron_vl.model
Module Contents
Classes
Functions
Data
_HAS_NATIVE_BIDIRECTIONAL_MASK
API
Bases: LlamaConfig
Configuration for bidirectional (non-causal) LLaMA model.
Bases: LlamaModel
Legacy Llama retrieval model supporting causal and bidirectional attention.
Supports Transformers 4.44 through 5.x with a compatibility forward path.
Encode the language sequence using the configured attention mode.
Parameters:
Integer token IDs [batch, sequence], exclusive with inputs_embeds.
Optional padding mask [batch, sequence], with zero for padding.
Integer positions [batch, sequence] or [1, sequence].
Optional Transformers cache in its native layout.
Floating-point embeddings [batch, sequence, hidden], which may include projected vision embeddings at image-token positions.
Integer cache positions [sequence].
Whether to populate the returned cache.
Whether to return intermediate hidden states.
Additional Hugging Face forward options.
Returns: BaseModelOutputWithPast
Output with floating-point last_hidden_state [batch, sequence, hidden],
Bases: PretrainedConfig
Base configuration for vision-language models combining vision and language components. This serves as the foundation for LlamaNemotronVL configurations.
Return the language decoder config using the Transformers composite-config contract.
Bases: PreTrainedModel
LlamaNemotron VL model for vision-language reranking. Combines a vision encoder (SigLIP) with a bidirectional language model (LLaMA) for cross-modal reranking tasks.
Encodes the inputs into a tensor of embeddings. Args: inputs: A dictionary of inputs to the model. You can prepare the inputs using the processor.process_queries and processor.process_documents methods. pool_type: The type of pooling to use. If None, the pooling type is set to the pooling type configured in the model. Returns: A tensor of embeddings.
Encodes the input document images and texts into a tensor of embeddings. Args: images: A list of PIL.Image of document pages images. texts: A list of document page texts. Returns: A tensor of embeddings.
Encodes the input queries into a tensor of embeddings. Args: queries: A list of queries. Returns: A tensor of embeddings.
Extract and project vision features to language model space.
Return the language tower used as the retrieval text decoder.
Keep only vision embeddings marked as real images.
Register bidirectional models with HuggingFace Auto classes.
This is needed so that AutoModel.from_config(LlamaBidirectionalConfig) works inside LlamaForSequenceClassification.init.
Replace image placeholder token embeddings with vision embeddings.