DeepSeek-V4-Flash-Vision-Exp

View as Markdown

DeepSeek-V4-Flash-Vision-Exp has a checked-in NeMo AutoModel recipe for image-text-to-text. The Hugging Face configuration declares the DeepseekV4ForCausalLM architecture.

Fine-Tune DeepSeek-V4-Flash-Vision-Exp

Follow the installation instructions, then review the checked-in recipe and its declared topology below before choosing a local or cluster launcher.

Use the Slurm launcher guide for the multi-node run.

Choose a Workflow

GoalStart Here
Run the primary recipe for this modelUse the recipe configuration.
Prepare your environmentFollow the installation instructions.

Configuration

SettingConfiguration
Hardware-
StrategyFSDP2; tp_size=1, pp_size=4, cp_size=1, ep_size=32
Nodes16
FeaturesActivation checkpointing (activation_checkpointing=true), HybridEP (dispatcher=hybridep)
Advancedbfloat16, attention=tilelang, linear=torch, experts=torch_mm, dispatcher=hybridep

Model Reference

Model Architecture

PropertyValue
TaskImage-text-to-text
ArchitectureDeepseekV4ForCausalLM

Available Models