gemma-3-4b-it

View as Markdown

gemma-3-4b-it is a multimodal Gemma 3 checkpoint for image-text inputs. NeMo AutoModel provides full-parameter and Low-Rank Adaptation (LoRA) recipes for supervised fine-tuning.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Fine-Tune gemma-3-4b-it

From the repository root, run:

uv run automodel examples/vlm_finetune/gemma3/gemma3_vl_4b_cord_v2.yaml --nproc-per-node 8

Choose a Workflow

GoalStart Here
Fine-tune on CORD-v2Use gemma3_vl_4b_cord_v2.yaml.
Fine-tune with LoRAUse gemma3_vl_4b_cord_v2_peft.yaml.
Fine-tune with Megatron FSDPUse gemma3_vl_4b_cord_v2_megatron_fsdp.yaml.
Fine-tune on MedPix-VQAUse gemma3_vl_4b_medpix.yaml.

Model Reference

Model Architecture

PropertyValue
TaskImage-text-to-text
Hugging Face ArchitectureGemma3ForConditionalGeneration
Checkpointgoogle/gemma-3-4b-it