For AI agents: a documentation index is available at the root level at /llms.txt. Append /llms.txt to any URL for a page-level index, or .md for the markdown version of any page.
LogoLogoNeMo AutoModel
    • Home
  • Get Started
    • About
    • Key Features
    • Install NeMo AutoModel
    • Configuration
    • 🤗 HF Compatibility
    • Repo Structure
  • What's New
    • Announcements
    • Release Notes
    • Model Support Log
  • Performance
    • Performance Summary
  • Model Coverage
    • Overview
    • Model Support Log
      • Overview
      • Meta
        • Llama-3.2-1B
      • Mistral AI
      • NVIDIA
    • Troubleshooting
  • Recipes & E2E Examples
    • Recipes and End-to-End Examples
    • SFT & PEFT
    • Function Calling with FunctionGemma
    • Multi-Turn Agent (Tool-Calling) SFT
    • Knowledge Distillation
    • Fine-Tune Large MoE LLMs
    • DeepSeek-V4 Flash
    • Hy3-preview
    • Nemotron-3-Ultra-550B
    • Pretraining
    • NanoGPT Pretraining
    • Sequence Classification (SFT/PEFT)
    • Retrieval Fine-Tuning
    • Gemma 3 / 3n
    • Gemma 4 31B
    • Qwen3.5-VL
    • Fine-Tune Qwen3.8-27B
    • Nemotron-Omni
    • Nemotron 3.5 Super VL
    • Nemotron 3.5 Super VL: Text-to-SQL with LoRA
    • Nemotron 3.5 Super VL: Continued Pretraining on FineWeb
    • Mistral Medium 3.5 VL
    • MiniMax-M3
    • Fine-Tune Step-3.7-Flash
    • ASR with Qwen3-Omni
    • Wan2.1-T2V Fine-Tuning
    • Fine-Tuning DiffusionGemma
    • Train an EAGLE Drafter for Speculative Decoding
    • Train a DFlash Drafter for Speculative Decoding
    • Train a DSpark Drafter for Speculative Decoding
    • dLLM Fine-Tuning
    • Quantization-Aware Training (QAT)
    • Model Training on Databricks
  • Data
    • Overview
    • Text Dataset
    • Retrieval Dataset
    • ColumnMapped Dataset
    • ColumnMapped Iterable
    • Multi-Modal Dataset
    • Diffusion Dataset
  • Run Jobs
    • Overview
    • Local Workstation
    • SLURM Cluster
    • NeMo Run
    • SkyPilot
    • k8s with SkyPilot
  • Advanced Training
    • Checkpointing
    • Gradient Checkpointing
    • Distributed Setup (Python API)
    • Pipeline Parallelism
    • Context-Parallel Vision Frame Sharding
    • FP8 Training
    • Mixed-Precision Training
    • MLflow Logging
    • Breaking Changes
  • Reference
    • API Reference
    • Model Parallelizer API
  • Home
  • About
  • Key Features
  • Install NeMo AutoModel
  • Configuration
  • 🤗 HF Compatibility
  • Repo Structure
  • Announcements
  • Release Notes
  • Model Support Log
  • Performance Summary
  • Overview
  • Model Support Log
  • Overview
  • Allen AI
  • OLMo-1B-hf
  • OLMo2-7B-1124
  • OLMoE-1B-7B-0924
  • BAAI
  • Aquila-7B
  • Baichuan Inc.
  • Baichuan2-13B-Chat
  • Baichuan2-7B-Chat
  • Baidu
  • ERNIE-4.5-0.3B-PT
  • ERNIE-4.5-21B-A3B-PT
  • BigCode
  • starcoder
  • starcoder2-3b
  • ByteDance Seed
  • Seed-Coder-8B-Instruct
  • Seed-OSS-36B-Instruct
  • Cohere
  • c4ai-command-r-v01
  • DeepSeek AI
  • deepseek-llm-7b-chat
  • DeepSeek-V3
  • DeepSeek-V3.2
  • DeepSeek-V4-Flash
  • DeepSeek-V4-Pro
  • DeepSeek-V4.1-Flash
  • EleutherAI
  • gpt-j-6b
  • gpt-neox-20b
  • Google
  • functiongemma-270m-it
  • gemma-2-9b-it
  • gemma-3-270m
  • gemma-4-12B
  • gemma-4-31B
  • gemma-7b
  • Google T5
  • t5-small
  • IBM
  • granite-3.0-1b-a400m-base
  • granite-3.0-2b-base
  • IBM AI Platform
  • Bamba-9B
  • Inception AI
  • jais-13b
  • Inclusion AI
  • Ling-1T
  • Ling-flash-2.0
  • Ling-mini-2.0
  • InternLM
  • internlm3-8b-instruct
  • LG AI EXAONE
  • EXAONE-3.0-7.8B-Instruct
  • Meta
  • Llama-3.1-70B
  • Llama-3.1-8B
  • Llama-3.2-1B
  • Llama-3.2-3B-Instruct
  • Llama-3.3-70B-Instruct
  • Meta Models
  • Muse-Glimmer-30B
  • Microsoft
  • phi-2
  • Phi-3-mini-4k-instruct
  • Phi-3-small-8k-instruct
  • Phi-4
  • MiniMax AI
  • MiniMax-M2.1
  • MiniMax-M2.5
  • MiniMax-M2.7
  • Mistral AI
  • Devstral-2-123B-Instruct-2512
  • Devstral-Small-2-24B-Instruct-2512
  • Ministral-3-3B-Instruct-2512
  • Mistral-7B-v0.1
  • Mistral-Nemo-Base-2407
  • Mistral-Small-4-119B-2603
  • Mixtral-8x7B-Instruct-v0.1
  • Mixtral-8x7B-v0.1
  • Moonshot AI
  • Kimi-K2-Base
  • Kimi-K3
  • Kimi-Linear-48B-A3B-Instruct
  • Moonlight-16B-A3B
  • NVIDIA
  • Llama-3.1-Nemotron-Nano-8B-v1
  • Llama-3_3-Nemotron-Super-49B-v1
  • Minitron-8B-Base
  • Nemotron-Flash-1B
  • NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
  • NVIDIA-Nemotron-3-Nano-4B-BF16
  • NVIDIA-Nemotron-3-Super-120B-A12B-BF16
  • NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
  • NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
  • NVIDIA-Nemotron-Nano-9B-v2
  • OpenAI
  • gpt-oss-120b
  • gpt-oss-20b
  • OpenAI Community
  • gpt2
  • OpenBMB
  • MiniCPM5-1B
  • OrionStar AI
  • Orion-14B-Base
  • ParaSail AI
  • GritLM-7B-vllm
  • Poolside
  • Laguna-S-2.1
  • Laguna-XS-2.1
  • Qwen
  • Qwen1.5-MoE-A2.7B
  • Qwen2.5-0.5B
  • Qwen2.5-32B-Instruct
  • Qwen2.5-3B
  • Qwen2.5-7B-Instruct
  • Qwen3-0.6B
  • Qwen3-235B-A22B
  • Qwen3-30B-A3B
  • Qwen3-32B
  • Qwen3-4B-Base
  • Qwen3-8B
  • Qwen3-Next-80B-A3B-Instruct
  • Qwen3.5-35B-A3B
  • Qwen3.8-2.4T-A95B
  • Qwen3.8-Flash-Next
  • QwQ-32B
  • Stability AI
  • stablelm-3b-4e1t
  • StepFun AI
  • Step-3.5-Flash
  • Tencent
  • Hy-MT2-30B-A3B
  • Hy3-preview
  • TII
  • falcon-7b
  • Falcon-H1-0.5B-Instruct
  • Falcon-H1-1.5B-Deep-Instruct
  • Falcon-H1-34B-Instruct
  • Falcon3-7B-Instruct
  • Upstage
  • solar-pro-preview-instruct
  • XiaomiMiMo
  • MiMo-V2-Flash
  • MiMo-V2.5-Pro
  • Z.ai
  • chatglm3-6b
  • glm-4-9b-chat-hf
  • GLM-4.5-Air
  • GLM-4.7
  • GLM-5
  • GLM-5.1
  • GLM-5.2
  • GLM-5.3
  • Overview
  • Cohere Labs
  • North-Micro-Vision-Instruct
  • DeepSeek AI
  • DeepSeek-V4-Flash-Vision-Exp
  • DeepSeek-V4.1-Flash
  • Google
  • gemma-3-4b-it
  • gemma-3n-E4B-it
  • gemma-4-26B-A4B-it
  • gemma-4-31B-it
  • gemma-4-E2B-it
  • gemma-4-E4B-it
  • HuggingFaceTB
  • SmolVLM-Instruct
  • LLaVA HF
  • llava-1.5-7b-hf
  • LMMS Lab
  • LLaVA-OneVision-1.5-4B-Instruct
  • LLaVA-OneVision-1.5-8B-Instruct
  • Meta
  • Llama-4-Scout-17B-16E-Instruct
  • Meta Models
  • Muse-Glimmer-30B
  • MiniMax AI
  • MiniMax-M3
  • Mistral AI
  • Devstral-2-123B-Instruct-2512
  • Ministral-3-14B-Reasoning-2512
  • Ministral-3-3B-Instruct-2512
  • Ministral-3-8B-Reasoning-2512
  • Mistral-Medium-3.5-128B
  • Mistral-Small-4-119B-2603
  • Moonshot AI
  • Kimi-K2.5
  • Kimi-K3
  • Kimi-VL-A3B-Instruct
  • NVIDIA
  • NVIDIA-Nemotron-Parse-v1.1
  • OpenGVLab
  • InternVL3_5-4B
  • Qwen
  • Qwen2.5-VL-3B-Instruct
  • Qwen3-VL-235B-A22B-Instruct
  • Qwen3-VL-30B-A3B-Instruct
  • Qwen3-VL-4B-Instruct
  • Qwen3-VL-8B-Instruct
  • Qwen3.5-122B-A10B
  • Qwen3.5-27B
  • Qwen3.5-35B-A3B
  • Qwen3.5-397B-A17B
  • Qwen3.5-4B
  • Qwen3.5-9B
  • Qwen3.6-27B
  • Qwen3.6-35B-A3B
  • Qwen3.8-27B
  • StepFun AI
  • Step-3.7-Flash
  • Thinking Machines
  • Inkling
  • Inkling-Small
  • Z.ai
  • GLM-5.3-Flash
  • Overview
  • ByteDance Seed
  • BAGEL-7B-MoT
  • Overview
  • Microsoft
  • Phi-4-multimodal-instruct
  • NVIDIA
  • Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16
  • NVIDIA-Nemotron-3.5-Super-VL-120B-A12B-BF16
  • Qwen
  • Qwen2.5-Omni-3B
  • Qwen2.5-Omni-7B
  • Qwen3-Omni-30B-A3B-Instruct
  • Overview
  • Google
  • diffusiongemma-26B-A4B-it
  • GSAI ML
  • LLaDA-8B-Base
  • Inclusion AI
  • LLaDA2.1-mini
  • NVIDIA
  • Nemotron-Labs-Diffusion-8B-Base
  • Qwen
  • Qwen3-8B
  • Z Lab
  • Qwen3-4B-DFlash-b16
  • Overview
  • Black Forest Labs
  • FLUX.1-dev
  • FLUX.2-dev
  • Diffusers
  • LTX-2.3-Diffusers
  • HunyuanVideo Community
  • HunyuanVideo-1.5-Diffusers-720p_t2v
  • Lightricks
  • LTX-2.3
  • Qwen
  • Qwen-Image
  • Qwen-Image-2.1
  • Qwen-Image-Edit-2511
  • Wan AI
  • Wan2.1-T2V-1.3B-Diffusers
  • Wan2.1-T2V-14B-Diffusers
  • Wan2.2-T2V-A14B-Diffusers
  • Overview
  • Meta
  • Llama-3.2-1B
  • Mistral AI
  • Ministral-3-3B-Base-2512
  • NVIDIA
  • llama-embed-nemotron-8b
  • llama-nemotron-embed-vl-1b-v2
  • Overview
  • NVIDIA
  • llama-nemotron-rerank-1b-v2
  • Troubleshooting
  • Recipes and End-to-End Examples
  • SFT & PEFT
  • Function Calling with FunctionGemma
  • Multi-Turn Agent (Tool-Calling) SFT
  • Knowledge Distillation
  • Fine-Tune Large MoE LLMs
  • DeepSeek-V4 Flash
  • Hy3-preview
  • Nemotron-3-Ultra-550B
  • Pretraining
  • NanoGPT Pretraining
  • Sequence Classification (SFT/PEFT)
  • Retrieval Fine-Tuning
  • Gemma 3 / 3n
  • Gemma 4 31B
  • Qwen3.5-VL
  • Fine-Tune Qwen3.8-27B
  • Nemotron-Omni
  • Nemotron 3.5 Super VL
  • Nemotron 3.5 Super VL: Text-to-SQL with LoRA
  • Nemotron 3.5 Super VL: Continued Pretraining on FineWeb
  • Mistral Medium 3.5 VL
  • MiniMax-M3
  • Fine-Tune Step-3.7-Flash
  • ASR with Qwen3-Omni
  • Wan2.1-T2V Fine-Tuning
  • Fine-Tuning DiffusionGemma
  • Train an EAGLE Drafter for Speculative Decoding
  • Train a DFlash Drafter for Speculative Decoding
  • Train a DSpark Drafter for Speculative Decoding
  • dLLM Fine-Tuning
  • Quantization-Aware Training (QAT)
  • Model Training on Databricks
  • Overview
  • Text Dataset
  • Retrieval Dataset
  • ColumnMapped Dataset
  • ColumnMapped Iterable
  • Multi-Modal Dataset
  • Diffusion Dataset
  • Overview
  • Local Workstation
  • SLURM Cluster
  • NeMo Run
  • SkyPilot
  • k8s with SkyPilot
  • Checkpointing
  • Gradient Checkpointing
  • Distributed Setup (Python API)
  • Pipeline Parallelism
  • Context-Parallel Vision Frame Sharding
  • FP8 Training
  • Mixed-Precision Training
  • MLflow Logging
  • Breaking Changes
  • API Reference
  • Model Parallelizer API
  • Nemo Automodel
  • Autonvtx
  • Cli
  • App
  • Query Capabilities
  • Utils
  • Components
  • Attention
  • Dflash Mask
  • Ffpa Attention
  • Flex Attention
  • Idlm Mask
  • Utils
  • Checkpoint
  • Addons
  • Checkpointing
  • Config
  • Conversion Mapping
  • Lifecycle
  • State Dict Adapter
  • Stateful Wrappers
  • Utils
  • Config
  • Loader
  • Cuda Graphs
  • Partial
  • Datasets
  • Audio
  • Collate Fns
  • Datasets
  • Multi En
  • Datum
  • Diffusion
  • Base Dataset
  • Collate Fns
  • Image Edit Dataset
  • Loader
  • Meta Files Dataset
  • Mock Dataloader
  • Multi Tier Bucketing
  • Sampler
  • Text To Image Dataset
  • Text To Video Dataset
  • Dllm
  • Collate
  • Corruption
  • Lazy Mapped Dataset
  • Llm
  • Agent Chat
  • Chat Dataset
  • Column Mapped Text Instruction Dataset
  • Column Mapped Text Instruction Iterable Dataset
  • Delta Lake Dataset
  • Dspark Cache
  • Eagle3
  • Eagle3 Cache
  • Formatting Utils
  • Hellaswag
  • Length Grouped Sampler
  • Megatron
  • Builder
  • Gpt Dataset
  • Helpers
  • Indexed Dataset
  • Megatron Utils
  • Sampler
  • Megatron Dataset
  • Mock
  • Mock Iterable Dataset
  • Mock Packed
  • Mock Prefix Tree
  • Mock Seq Cls
  • Nanogpt Dataset
  • Neat Packing
  • Offline Cache
  • Packed Sequence
  • Prefix Tree
  • Retrieval Collator
  • Retrieval Dataset
  • Retrieval Dataset Inline
  • Retrieval Dataset Normalized
  • Retrieval Distill Collator
  • Seq Cls
  • Seq2seq
  • Squad
  • Xlam
  • Loader
  • Multimodal
  • Collate Fns
  • Datasets
  • Distributed Iterable
  • Interleave
  • Loader
  • Packing
  • Parquet Utils
  • Transforms
  • Utils
  • Video
  • Reservoir Sampler
  • Utils
  • Vlm
  • Collate Fns
  • Datasets
  • Dspark Collate
  • Fake Image
  • Loader
  • Mock
  • Neat Packing Vlm
  • Pp Media
  • Samplers
  • Utils
  • Distributed
  • Activation Checkpointing
  • Blockdiag Cp
  • Batch
  • Exchange
  • Kernels
  • Packed
  • Runtime
  • State
  • Config
  • Context Parallel
  • Magi
  • Mamba
  • Sharder
  • Utils
  • Cp Vision Frame Shard
  • Ddp
  • Fsdp Patches
  • Fsdp2
  • Grad Utils
  • Init Utils
  • Megatron Fsdp
  • Mesh
  • Mesh Utils
  • Model Parallelizer
  • Multimodal Fsdp
  • Optimized Tp Plans
  • Parallel Styles
  • Parallelizer
  • Parallelizer Utils
  • Pipelining
  • Autopipeline
  • Config
  • Functional
  • Hf Utils
  • Recv Buffer Pool
  • Runtime
  • Tensor Utils
  • Thd Utils
  • Tp Replicas
  • Utils
  • Eval
  • Tool Call Evaluator
  • Tool Call Parser
  • Flow Matching
  • Adapters
  • Base
  • Flux
  • Flux2
  • Hunyuan
  • Ltx2
  • Qwen Image
  • Qwen Image 21
  • Simple
  • Pipeline
  • Time Shift Utils
  • Launcher
  • Base
  • Interactive
  • Nemo Run
  • Config
  • Launcher
  • Utils
  • Skypilot
  • Config
  • Launcher
  • Utils
  • Loggers
  • Comet Utils
  • Log Utils
  • Loggers
  • Metric Logger
  • Mlflow Utils
  • Wandb Utils
  • Loss
  • Chunked Ce
  • Dist Utils
  • Dllm Loss
  • Embedding Distill
  • Infonce
  • Intermediate Distill
  • Kd Loss
  • Linear Ce
  • Linear Ce Base
  • Listmle
  • Loss
  • Masked Ce
  • Mtp
  • Soft Ce
  • Te Parallel Ce
  • Utils
  • Models
  • Bagel
  • Attention Masks
  • Autoencoder
  • Configuration
  • Connector
  • Embeddings
  • Hf Backbone Loader
  • Model
  • Modeling Qwen2 Packed
  • Modeling Siglip Navit
  • Parallelization
  • State Dict Adapter
  • Baichuan
  • Configuration
  • Model
  • Common
  • Bidirectional
  • Cudnn Sparse Attention
  • Gated Delta Net Fp32
  • Hf Checkpointing Mixin
  • Inbatch Neg Utils
  • Mtp
  • Mtp
  • Packing
  • Tie Word Embeddings
  • Utils
  • Deepseek V3
  • Layers
  • Model
  • Rope Utils
  • State Dict Adapter
  • Deepseek V32
  • Config
  • Layers
  • Model
  • State Dict Adapter
  • Deepseek V4
  • Config
  • Cp
  • Fsdp
  • Kernels
  • Sparse Attention
  • Tilelang Indexer
  • Tilelang Indexer Bwd
  • Tilelang Indexer Fwd
  • Tilelang Sparse Mla Bwd
  • Tilelang Sparse Mla Fwd
  • Layers
  • Model
  • Mtp
  • Optimized Kernels
  • Parallelization
  • Processing
  • State Dict Adapter
  • Vision
  • Deepseek V41
  • Attention
  • Config
  • Cp
  • Dspark
  • Engram
  • Fsdp
  • Indexer
  • Layers
  • Model
  • Packing
  • Processing
  • Quantization
  • State Dict Adapter
  • Vision
  • Deprecation
  • Diffusion Gemma
  • Attention Mask
  • Fsdp
  • Layers
  • Model
  • State Dict Adapter
  • Ernie4 5
  • Model
  • Rope Utils
  • State Dict Adapter
  • Gemma4 Drafter
  • Composite
  • Model
  • Gemma4 Moe
  • Cp Attention
  • Cp Batch
  • Cp Local Ring
  • Loss
  • Model
  • Parallelization
  • Sdpa Fp32
  • State Dict Adapter
  • Gemma4 Unified
  • Model
  • State Dict Adapter
  • Glm Moe Dsa
  • Config
  • Cp
  • Kernels
  • Cudnn Dsa
  • Indexer
  • Sparse Mla
  • Tilelang Indexer Bwd
  • Tilelang Indexer Fwd
  • Tilelang Sparse Mla Bwd
  • Tilelang Sparse Mla Fwd
  • Layers
  • Model
  • Optimized Kernels
  • Rope Utils
  • State Dict Adapter
  • Glm4 Moe
  • Layers
  • Model
  • State Dict Adapter
  • Glm4 Moe Lite
  • Model
  • Glm5 Next
  • Config
  • Cp
  • Image Processing
  • Layers
  • Model
  • Processing
  • State Dict Adapter
  • Vision
  • Gpt Oss
  • Layers
  • Model
  • Rope Utils
  • State Dict Adapter
  • Gpt2
  • Hy Mt2
  • Config
  • Dispatch
  • Layers
  • Model
  • State Dict Adapter
  • Hy V3
  • Config
  • Layers
  • Model
  • State Dict Adapter
  • Inkling
  • Configuration
  • Feature Extraction
  • Image Processing
  • Layers
  • Model
  • Multimodal
  • Processing
  • State Dict Adapter
  • Text
  • Kimi K2
  • Config
  • Kimi K25 Vl
  • Model
  • State Dict Adapter
  • Kimi K3
  • Attn Res Triton
  • Config
  • Cp
  • Encoding
  • Kda Fused
  • Chunk Kda Bwd Triton
  • Chunk Kda Fwd Cuda
  • Model
  • Multimodal
  • Situ
  • Situ Triton
  • State Dict Adapter
  • Tokenization
  • Vision
  • Kimi Linear
  • Config
  • Cp
  • Model
  • State Dict Adapter
  • Kimivl
  • Model
  • Laguna
  • Config
  • Model
  • State Dict Adapter
  • Ling V2
  • Config
  • Layers
  • Model
  • State Dict Adapter
  • Llama
  • Model
  • Rope Utils
  • Llama Bidirectional
  • Export Onnx
  • Model
  • Llama Nemotron Vl
  • Model
  • Processor
  • Llava Onevision
  • Model
  • Rice Vit
  • State Dict Adapter
  • Mimo V2 Flash
  • Config
  • Cp
  • Model
  • Parallelization
  • State Dict Adapter
  • Vision
  • Mimo V25
  • Config
  • Model
  • State Dict Adapter
  • Minimax M2
  • Layers
  • Model
  • State Dict Adapter
  • Minimax M3 Vl
  • Config
  • Cp Sparse Attn
  • Kernels
  • Msa Backward Postprocess Sm100
  • Msa Backward Preprocess Sm100
  • Msa Backward Sm100
  • Msa Schedule
  • Msa Task Build Sm100
  • Layers
  • Model
  • Msa
  • Msa Bindings
  • Mtp
  • Processing
  • State Dict Adapter
  • Vision Encoder
  • Ministral Bidirectional
  • Model
  • Mistral3
  • Model
  • State Dict Adapter
  • Mistral3 Vlm
  • Model
  • State Dict Adapter
  • Mistral4
  • Configuration
  • Model
  • State Dict Adapter
  • Muse Glimmer
  • Config
  • Model
  • Parallelization
  • State Dict Adapter
  • Vision
  • Nemotron Omni
  • Model
  • State Dict Adapter
  • Nemotron Parse
  • Model
  • Nemotron Parse Loss
  • Nemotron V3
  • Cache
  • Layers
  • Model
  • Mtp
  • Parallelization
  • State Dict Adapter
  • Qwen Image Edit
  • Adapter
  • Preprocessing
  • Qwen2
  • Model
  • Qwen2 5 Omni
  • Model
  • State Dict Adapter
  • Qwen3
  • Model
  • Qwen3 5
  • Model
  • Packing
  • Parallelization
  • State Dict Adapter
  • Qwen3 5 Moe
  • Cp Linear Attn
  • Model
  • State Dict Adapter
  • Qwen3 8 Flash Next
  • Backend
  • Config
  • Cp
  • Engram
  • Fa4 Qsa
  • Flex Qsa
  • Layers
  • Model
  • Qsa
  • State Dict Adapter
  • Qwen3 Moe
  • Layers
  • Model
  • State Dict Adapter
  • Qwen3 Next
  • Layers
  • Model
  • State Dict Adapter
  • Qwen3 Omni Moe
  • Model
  • State Dict Adapter
  • Qwen3 Vl
  • Model
  • Parallelization
  • Qwen3 Vl Moe
  • Model
  • State Dict Adapter
  • Step3p5
  • Layers
  • Model
  • State Dict Adapter
  • Step3p7
  • Configuration Step3p7
  • Model
  • Mtp
  • Processing Step3
  • State Dict Adapter
  • Vision Encoder
  • Moe
  • Config
  • Experts
  • Fsdp Mixin
  • Layers
  • Load Balance Metrics
  • Megatron
  • Fused A2a
  • Fused Indices Converter
  • Moe Utils
  • Token Dispatcher
  • Mok Experts
  • Mxfp8
  • Optimized Ops
  • Parallelizer
  • Router Replay
  • State Dict Mixin
  • State Dict Utils
  • Tp Plan Validation
  • Uccl Ep
  • Buffer
  • Optim
  • Dion
  • Optimizer
  • Precision Warnings
  • Scheduler
  • Quantization
  • Fp8
  • Qat
  • Qlora
  • Speculative
  • Bench Common
  • Bench Sglang
  • Bench Sweep
  • Bench Vllm
  • Decode Eval
  • Dflash
  • Core
  • Dflash2 Core
  • Domino Core
  • Draft Kimi K3
  • Draft Qwen3
  • Draft Qwen3 Dflash2
  • Jetspec Core
  • Registry
  • Target
  • Dspark
  • Common
  • Config
  • Core
  • Draft Deepseek V4
  • Draft Gemma4
  • Draft Glm 5 2
  • Draft Kimi K3
  • Draft Minimax M3
  • Draft Qwen3
  • Loss
  • Markov Head
  • Registry
  • Target
  • Target Utils
  • Eagle
  • Backend
  • Core
  • Core V12
  • Draft Deepseek
  • Draft Gemma
  • Draft Gpt Oss
  • Draft Kimi K3
  • Draft Llama
  • Draft Llama V12
  • Msd
  • Msd Curriculum
  • Msd Decode
  • Msd Target
  • Peagle Attention
  • Peagle Data
  • Peagle Draft
  • Peagle Trainer
  • Registry
  • Remote
  • Client
  • Protocol
  • Server
  • Transport
  • Wire
  • Ring Attention
  • Sglang Runner
  • Sglang Target
  • Target
  • Target Runner
  • Target V12
  • Ulysses Attention
  • Vispec Core
  • Vispec Draft
  • Vispec Target
  • Vllm Runner
  • Vllm Target
  • Zigzag Ring Attention
  • Precompute Dspark
  • Precompute Eagle3
  • Regen Loop
  • Regenerate
  • Regenerate Vlm
  • Serve Sglang
  • Serve Target
  • Serve Vllm
  • Streaming
  • Async Pipeline
  • Eagle3
  • Loader
  • Producer
  • Queue
  • Refs
  • Store
  • Stores
  • Local
  • Shared Dir
  • Target Cp
  • Training
  • Domain Mixture
  • Ema
  • Embedding Row Repair
  • Garbage Collection
  • Model Output Utils
  • Neftune
  • Prewarm
  • Rng
  • Signal Handler
  • Step Scheduler
  • Timers
  • Utils
  • Utils
  • Compile Utils
  • Flops Utils
  • Model Utils
  • Yaml Utils
  • Package Info
  • Shared
  • Embedding Padding
  • Import Utils
  • Model Utils
  • Parameter Names
  • Te Patches
  • Tied Weights
  • Torch Patches
  • Tp Linear
  • Transformers Patches
  • Utils
Model CoverageEmbedding ModelsMeta

Meta

||View as Markdown|
  • Llama-3.2-1B
Previous
Embedding Models
Next

Llama-3.2-1B

NVIDIANVIDIA
Developer-friendly docs for your API
Privacy Policy | Your Privacy Choices | Terms of Service | Accessibility | Corporate Policies | Product Security | Contact

Copyright © 2026, NVIDIA Corporation.