Nemotron-Flash-1B
Nemotron-Flash-1B
NVIDIA Nemotron-Flash is a compact, fast language model designed for low-latency inference workloads.
This model requires trust_remote_code: true in your recipe YAML.
Set up NeMo AutoModel with the latest container or follow the installation instructions.
Fine-Tune Nemotron-Flash-1B
From the repository root, run: