> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# Phi-4-multimodal-instruct

> Use Phi-4-multimodal-instruct with NeMo AutoModel for multimodal fine-tuning, with documented checkpoints, runnable recipes, setup guidance, and model reference details.

[Phi-4-multimodal-instruct](https://huggingface.co/microsoft/Phi-4-multimodal-instruct) is Microsoft's multimodal extension of Phi-4, supporting text, image, and audio inputs - making it suitable for speech, vision, and combined multimodal tasks.

Set up NeMo AutoModel with the [latest container](https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo-automodel) or follow the [installation instructions](/get-started/installation).

## Fine-Tune Phi-4-multimodal-instruct

From the repository root, run:

```bash
uv run automodel --nproc-per-node=8 examples/vlm_finetune/phi4/phi4_mm_cv17.yaml
```

## Choose a Workflow

| Goal                                                                        | Start Here                                                                                                                                           |
| --------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
| Supervised fine-tuning (SFT) - Phi-4-multimodal on CommonVoice (audio-text) | Use [phi4\_mm\_cv17.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/vlm_finetune/phi4/phi4_mm_cv17.yaml). Dataset: CommonVoice 17. |

## Model Reference

### Model Architecture

| Property                  | Value                                         |
| ------------------------- | --------------------------------------------- |
| Task                      | Omnimodal (Text/Image/Audio)                  |
| Architecture              | `Phi4MultimodalForCausalLM`                   |
| Parameters                | 5.6B                                          |
| Hugging Face Organization | [microsoft](https://huggingface.co/microsoft) |

### Available Models

| Model                     | HF ID                                                                                               |
| ------------------------- | --------------------------------------------------------------------------------------------------- |
| Phi-4-multimodal-instruct | [`microsoft/Phi-4-multimodal-instruct`](https://huggingface.co/microsoft/Phi-4-multimodal-instruct) |

## Related Resources

* [VLM Fine-Tuning Guide](/recipes-e2e-examples/gemma-3-3n)