Step-3.5-Flash

View as Markdown

Step-3.5-Flash is a Mixture-of-Experts language model from Stepfun AI, designed for efficient inference.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Fine-Tune Step-3.5-Flash

This recipe was validated on 16 nodes x 8 GPUs (128 H100s). See the Launcher Guide for multi-node setup.

From the repository root, run:

uv run automodel --nproc-per-node=8 examples/llm_finetune/stepfun/step_3.5_flash_hellaswag_pp.yaml

Choose a Workflow

GoalStart Here
Supervised fine-tuning (SFT) - Step-3.5-Flash on HellaSwag with pipeline parallelismUse step_3.5_flash_hellaswag_pp.yaml.

Model Reference

Model Architecture

PropertyValue
TaskText Generation (MoE)
ArchitectureStep3p5ForCausalLM
Parameters196B total / 11B active
Hugging Face Organizationstepfun-ai

Available Models

ModelHF ID
Step-3.5-Flashstepfun-ai/Step-3.5-Flash