SmolVLM-Instruct
SmolVLM-Instruct
SmolVLM is HuggingFace’s compact vision language model designed for on-device and memory-constrained deployment, featuring an efficient image token compression strategy.
Use this page as a checkpoint and architecture reference. Set up NeMo AutoModel with the latest container or follow the installation instructions.