Qwen2.5-VL-3B-Instruct
Qwen2.5-VL-3B-Instruct
Qwen2.5-VL is Alibaba Cloud’s vision language model series supporting image and video understanding. It features dynamic resolution processing and integrates with the Qwen2.5 language backbone.
Set up NeMo AutoModel with the latest container or follow the installation instructions.
Fine-Tune Qwen2.5-VL-3B-Instruct
From the repository root, run:
Choose a Workflow
Model Reference
Model Architecture
Qwen2_5VLForConditionalGeneration- Qwen2.5-VLQwen2VLForConditionalGeneration- Qwen2-VL