SmolVLM-Instruct

View as Markdown

SmolVLM is HuggingFace’s compact vision language model designed for on-device and memory-constrained deployment, featuring an efficient image token compression strategy.

Use this page as a checkpoint and architecture reference. Set up NeMo AutoModel with the latest container or follow the installation instructions.

Model Reference

Model Architecture

PropertyValue
TaskImage-Text-to-Text
ArchitectureSmolVLMForConditionalGeneration
Parameters2.25B
Hugging Face OrganizationHuggingFaceTB

Available Models

ModelHF ID
SmolVLM InstructHuggingFaceTB/SmolVLM-Instruct