Kimi-VL-A3B-Instruct

View as Markdown

Kimi-VL and Kimi-K25-VL are vision language models from Moonshot AI. Kimi-VL-A3B uses a MoE language backbone (3B active parameters) with a vision encoder, supporting image understanding and multimodal reasoning.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Fine-Tune Kimi-VL-A3B-Instruct

From the repository root, run:

uv run automodel --nproc-per-node=8 examples/vlm_finetune/kimi/kimi2vl_cordv2.yaml

Choose a Workflow

GoalStart Here
Supervised fine-tuning (SFT) - Kimi-VL on CORD-v2Use kimi2vl_cordv2.yaml. Dataset: cord-v2.

Model Reference

Model Architecture

PropertyValue
TaskImage-Text-to-Text
ArchitectureKimiVLForConditionalGeneration
Parameters~3B active (MoE)
Hugging Face Organizationmoonshotai

Available Models

ModelHF ID
Kimi-VL-A3B-Instructmoonshotai/Kimi-VL-A3B-Instruct