nemo_automodel.components.models.laguna
nemo_automodel.components.models.laguna
Submodules
nemo_automodel.components.models.laguna.confignemo_automodel.components.models.laguna.modelnemo_automodel.components.models.laguna.state_dict_adapter
Package Contents
Classes
Data
API
Bases: PretrainedConfig
Configuration for Poolside Laguna causal language models.
Bases: HFCheckpointingMixin, Module, MoEFSDPSyncMixin
Causal LM wrapper for Laguna with Automodel checkpoint adapters.
Run the Laguna causal language model.
Parameters:
Optional token IDs of shape [batch, sequence].
Optional embeddings of shape [batch, sequence, hidden].
Optional position IDs of shape [batch, sequence].
Optional 2D bool/int mask of shape [batch, sequence], a 4D additive mask, or a mapping with per-attention-type masks keyed by “full_attention” and “sliding_attention”.
Optional bool tensor of shape [batch, sequence], where True marks tokens excluded from MoE routing.
Unsupported KV-cache state.
Unsupported cache flag.
If 0, compute logits for all sequence positions. If an int or tensor, compute logits only for the selected trailing positions.
When true, include final hidden states in the output.
Additional attention backend arguments.
Returns: CausalLMOutputWithPast
Causal LM output with logits and optional hidden states.
Bases: Module
Backbone model for Laguna SFT.
Build additive attention masks for full and sliding Laguna layers.
Parameters:
Input embedding tensor of shape [batch, sequence, hidden].
Optional sequence mask of shape [batch, sequence], 4D bool/additive mask, or mapping with masks keyed by “full_attention” and “sliding_attention”.
Position tensor of shape [batch, sequence].
Returns: dict[str, torch.Tensor]
Mapping from attention type to additive masks of shape
Run the Laguna decoder stack.
Parameters:
Optional token IDs of shape [batch, sequence].
Optional embeddings of shape [batch, sequence, hidden].
Optional position IDs of shape [batch, sequence].
Optional 2D bool/int mask of shape [batch, sequence], a 4D additive mask, or a mapping with per-attention-type masks keyed by “full_attention” and “sliding_attention”.
Optional bool tensor of shape [batch, sequence], where True marks tokens excluded from MoE routing.
Additional attention backend arguments.
Returns: torch.Tensor
Final hidden states of shape [batch, sequence, hidden].