country_code
Skip to main content
Ctrl+K
NVIDIA NIM for LLM and VLM - Home NVIDIA NIM for LLM and VLM - Home

NVIDIA NIM for LLM and VLM

  • Documentation Home
NVIDIA NIM for LLM and VLM - Home NVIDIA NIM for LLM and VLM - Home

NVIDIA NIM for LLM and VLM

  • Documentation Home

Table of Contents

About NIM LLM and VLM

  • Overview
  • NIM Offerings
  • Enterprise-Grade Inference Software Stack
  • Release Notes

Get Started

  • About Get Started
  • Prerequisites
  • Configuration
  • Installation
  • Quickstart
  • Advanced
    • Model-Specific Deployment Considerations
    • Get Started with DeepSeek-V4-Pro-0813

Deployment

  • Model Profiles and Selection
  • Model Download
  • Model-Free NIM
  • Kubernetes Deployment
    • Helm and Kubernetes
    • KServe
    • OpenShift
    • Run:ai
    • NIM Operator Deployment
  • Cloud Service Provider (CSP) Deployment
    • Google Cloud
    • AWS
    • Azure
    • Oracle
  • Air-Gap Deployment
  • Multi-Node Deployment
  • vGPU Deployment

Advanced Use Cases

  • Fine-Tuning with LoRA
  • Speculative Decoding
  • Payload Capture
  • Tool Calling and MCP Integration
  • Custom Parsers and Chat Templates
  • Custom Logits Processing
  • Prompt Embeddings
  • Image, Audio, and Video Input

AI Assistant Integrations

  • Use Claude Code with NIM
  • Use Codex CLI with NIM

Reference

  • Benchmarking
  • Architecture
  • Environment Variables
  • API Reference
  • CLI Reference
  • Advanced Configuration
  • Logging and Observability
  • Model Signature Verification
  • 1.x Migration Guide
  • Support Matrix for NIMs
  • Archived Versions

Troubleshooting

  • GPU Memory (OOM) Errors
  • Request Timeouts and Responses
  • CUDA Driver Initialization
  • OpenShift SCC Admission Failures

Resources

  • Support and FAQ
  • Related Products
  • Legal and Acknowledgements
  • Legal and Acknowledgements
Is this page helpful?

Legal and Acknowledgements#

This page contains the primary legal references and third-party acknowledgements for NVIDIA NIM for Large Language Models (LLM) and Vision Language Models (VLM).

NVIDIA AI Product Agreement#

By using this NIM, you acknowledge that you have read and agreed to the NVIDIA AI Product Agreement.

NOTICE AND DISCLAIMER: This software automatically retrieves, accesses or interacts with external materials. Those retrieved materials are not distributed with this software and are governed solely by separate terms, conditions and licenses. You are solely responsible for finding, reviewing and complying with all applicable terms, conditions, and licenses, and for verifying the security, integrity and suitability of any retrieved materials for your specific use case. This software is provided “AS IS”, without warranty of any kind. The author makes no representations or warranties regarding any retrieved materials, and assumes no liability for any losses, damages, liabilities or legal consequences from your use or inability to use this software or any retrieved materials. Use this software and the retrieved materials at your own risk.

Open Source Software License Acknowledgements#

NVIDIA NIM for Large Language Models (LLM) and Vision Language Models (VLM) (NIM LLM and VLM) is built on the work of open source communities whose ongoing innovation drives the LLM inference ecosystem forward. NVIDIA is grateful for the collaboration, contributions, and shared commitment to advancing AI infrastructure that these projects represent.

vLLM#

vLLM is the inference engine at the core of NIM LLM and VLM. The project provides a high-throughput, memory-efficient serving engine for large language models, and NIM LLM and VLM packages vLLM directly as its inference backend.

NVIDIA recognizes the vLLM community for:

  • Foundational inference technology. vLLM introduced PagedAttention, a novel memory management algorithm for KV cache, and has subsequently implemented cutting-edge theory such as continuous batching, automatic prefix caching, and chunked prefill into a production-grade, open source engine. These capabilities form the backbone of every NIM LLM and VLM deployment.

  • Sustained collaboration. NVIDIA engineers contribute upstream to vLLM, and NVIDIA is dedicated to continuing this investment. Improvements flow in both directions - NVIDIA contributes optimizations to the vLLM-project and community-driven innovations are available to NIM LLM and VLM users.

  • An open development model. The vLLM project maintains a transparent development process that welcomes contributors across organizations. This openness accelerates the pace of innovation for the entire ecosystem.

For vLLM documentation, refer to the vLLM project documentation. For details on how NIM LLM and VLM integrates with vLLM, refer to Architecture.

Open Source Projects#

NVIDIA NIM for Large Language Models (LLM) and Vision Language Models (VLM) includes open source software components. For the acknowledgements that apply to a specific container, refer to the NVIDIA OSS archive.

previous

Related Software

On this page
  • NVIDIA AI Product Agreement
  • Open Source Software License Acknowledgements
    • vLLM
    • Open Source Projects
NVIDIA NVIDIA
Privacy Policy | Your Privacy Choices | Terms of Service | Accessibility | Corporate Policies | Product Security | Contact

Copyright © 2024-2026, NVIDIA Corporation.

Last updated on Oct 02, 2026.