gpt-oss-20b

View as Markdown

GPT-OSS is OpenAI’s open-weight model family featuring QuickGELU activations and activation clamping for training stability.

The documented workflows cover full-parameter fine-tuning, LoRA, packed sequences, and a DGX Spark LoRA configuration.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Fine-Tune gpt-oss-20b

From the repository root, run:

uv run automodel --nproc-per-node=8 examples/llm_finetune/gpt_oss/gpt_oss_20b.yaml

Choose a Workflow

GoalStart Here
Full-parameter fine-tuningUse the full-parameter recipe.
LoRA fine-tuningUse the LoRA recipe.
Fine-tuning with 1,024-token packed sequencesUse the packed-sequence recipe.
LoRA fine-tuning on DGX SparkUse the DGX Spark recipe.

Model Reference

Model Architecture

PropertyValue
TaskText Generation
ArchitectureGptOssForCausalLM
Parameters21B total / 3.6B active
Hugging Face Organizationopenai

Available Models

ModelHF ID
GPT-OSS 20Bopenai/gpt-oss-20b