nemo_automodel.components.models.qwen3
nemo_automodel.components.models.qwen3
Dense Qwen3 model support.
Submodules
Package Contents
Classes
API
Bases: HFCheckpointingMixin, Qwen3PreTrainedModel, GenerationMixin
Dense Qwen3 causal LM with packed THD context parallelism.
BSHD and NEAT use HF dispatch; TE handles only pre-packed THD.
Run causal LM projection for BSHD or packed THD inputs.
Parameters:
Token IDs [B, S] or packed local IDs [T].
Optional padded mask. THD uses cu_seqlens.
Position IDs [B, S] or packed local IDs [T].
Optional BSHD KV cache; unsupported for THD.
Optional hidden inputs [B, S, H] or [T, H].
Optional labels [B, S] or packed [T].
Whether to update the BSHD KV cache.
Whether to request attention outputs.
Whether to return per-layer hidden states.
Whether to return CausalLMOutputWithPast.
Optional BSHD cache positions [S].
Positions to project from hidden size H to
vocabulary size V.
THD metadata. cu_seqlens is [N + 1] and CP adds
cp_size and cp_rank.
Returns: CausalLMOutputWithPast
Causal LM output with BSHD logits [B, S, V]. Packed local