country_code
Skip to main content
Ctrl+K
📢 Notice: NIM LLM and VLM 3.0 is now available alongside NIM LLM and VLM 2.0. Consider 3.0 for workloads that depend on Dynamo-based distributed inference with independently scaled workers.
NVIDIA NIM for Large Language Models (LLM) and Vision Language Models (VLM) - Home NVIDIA NIM for Large Language Models (LLM) and Vision Language Models (VLM) - Home

NVIDIA NIM for Large Language Models (LLM) and Vision Language Models (VLM)

  • Documentation Home
NVIDIA NIM for Large Language Models (LLM) and Vision Language Models (VLM) - Home NVIDIA NIM for Large Language Models (LLM) and Vision Language Models (VLM) - Home

NVIDIA NIM for Large Language Models (LLM) and Vision Language Models (VLM)

  • Documentation Home

Table of Contents

About NIM LLM and VLM

  • Overview
  • NIM Offerings
  • Enterprise-Grade Inference Software Stack
  • Release Notes

Get Started

  • About Get Started
  • Prerequisites
  • Configuration
  • Installation
  • Quickstart
  • Advanced
    • Get Started with DeepSeek-V4-Pro-0813

Deployment

  • Model Profiles and Selection
  • Model Download
  • Model-Free NIM
  • Kubernetes Deployment
    • Helm and Kubernetes
    • KServe
    • OpenShift
    • Run:ai
    • NIM Operator Deployment
  • Cloud Service Provider (CSP) Deployment
    • Google Cloud
    • AWS
    • Azure
    • Oracle
  • Air-Gap Deployment
  • Multi-Node Deployment
  • vGPU Deployment

Advanced Use Cases

  • Fine-Tuning with LoRA
  • Speculative Decoding
  • Payload Capture
  • Tool Calling and MCP Integration
  • Custom Parsers and Chat Templates
  • Custom Logits Processing
  • Prompt Embeddings
  • Image, Audio, and Video Input

AI Assistant Integrations

  • Use Claude Code with NIM
  • Use Codex CLI with NIM

Reference

  • Benchmarking
  • Architecture
  • Environment Variables
  • API Reference
  • CLI Reference
  • Advanced Configuration
  • Logging and Observability
  • Model Signature Verification
  • 1.x Migration Guide
  • Support Matrix for NIMs
  • Archived Versions

Troubleshooting

  • GPU Memory (OOM) Errors
  • CUDA Driver Initialization
  • OpenShift SCC Admission Failures

Resources

  • Support and FAQ
  • Related Products
  • Legal
  • Related Software
Is this page helpful?

Related Software#

NVIDIA NIM for Large Language Models (LLM) and Vision Language Models (VLM) fits into a broader inference and platform ecosystem. The following software products are highly relevant when you are deploying, operating, or extending LLM and VLM workloads.

Product

Relation to NIM LLM and VLM

vLLM

NVIDIA NIM for Large Language Models (LLM) and Vision Language Models (VLM) packages vLLM as its inference backend, so many request semantics and tuning concepts come directly from vLLM.

NIM Operator

The operator manages NIM deployments by using Kubernetes custom resources and is especially useful for repeatable, production-scale rollouts.

NVIDIA NIM for Vision Language Models

Earlier VLM-specific documentation remains useful for deployments that still follow that path. Use these NIM LLM and VLM docs for the combined LLM and VLM workflow.

Usage Guidance#

To help you choose the right tool for your specific use case, consider the following recommendations:

  • Use NIM LLM and VLM when your workload uses text, image, audio, or video input and you want a curated, enterprise-ready container for production inference.

  • Use vLLM documentation alongside NIM documentation when you need deeper, backend-specific context for passthrough arguments or upstream model-serving behavior.

  • Use the NIM Operator when your primary deployment target is Kubernetes and you want lifecycle automation around NIM services.

previous

Support and FAQ

next

Legal

On this page
  • Usage Guidance
NVIDIA NVIDIA
Privacy Policy | Your Privacy Choices | Terms of Service | Accessibility | Corporate Policies | Product Security | Contact

Copyright © 2024-2026, NVIDIA Corporation.

Last updated on Sep 11, 2026.