Nemotron-Flash-1B

View as Markdown

NVIDIA Nemotron-Flash is a compact, fast language model designed for low-latency inference workloads.

This model requires trust_remote_code: true in your recipe YAML.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Fine-Tune Nemotron-Flash-1B

From the repository root, run:

uv run automodel --nproc-per-node=8 examples/llm_finetune/nemotron_flash/nemotron_flash_1b_squad.yaml

Choose a Workflow

GoalStart Here
Supervised fine-tuning (SFT) - Nemotron-Flash 1B on SQuADUse nemotron_flash_1b_squad.yaml.
Low-rank adaptation (LoRA) - Nemotron-Flash 1B on SQuADUse nemotron_flash_1b_squad_peft.yaml.

Model Reference

Model Architecture

PropertyValue
TaskText Generation
ArchitectureNemotronFlashForCausalLM
Parameters1B
Hugging Face Organizationnvidia

Available Models

ModelHF ID
Nemotron-Flash 1Bnvidia/Nemotron-Flash-1B