Llama-3.2-1B
Llama-3.2-1B
Meta’s Llama is a family of open-weight autoregressive language models built on the transformer decoder architecture. Key design choices include pre-normalization with RMSNorm, SwiGLU activations, and Rotary Positional Embeddings (RoPE). Llama 3+ models add Grouped Query Attention (GQA) for memory-efficient inference at larger scales.
Set up NeMo AutoModel with the latest container or follow the installation instructions.
Fine-Tune Llama-3.2-1B
From the repository root, run: