Troubleshooting GPU Memory Out-of-Memory Errors#

GPU out-of-memory (OOM) errors occur when the model needs more VRAM than the GPU provides. The usual causes are a model profile that does not match the hardware, or a misconfigured memory setting.

Startup allocates GPU memory in phases, and the fix depends on which phase failed. Check the container logs, then follow the matching issue below. If the default logs are not enough, refer to Enable Detailed Logging.

Start with a GPU/profile combination listed in the Support Matrix for NIMs and check the Release Notes for model-specific startup failures. The steps below diagnose memory pressure; they do not establish a workaround for a known backend defect or make an unsupported profile supported. An illegal-memory-access error or worker crash during warm-up is not, by itself, evidence of an OOM.

How GPU Memory Is Used#

During startup, NIM and the vLLM backend allocate GPU memory in stages. The largest single consumer is the model weights. You can estimate weight memory from the model’s parameter count (typically listed on the model card, or available in the model.safetensors.index.json file under the metadata.total_size field) and the precision of the profile.

A good heuristic formula for per-GPU weight memory is:

weight_memory_per_gpu = total_parameters x bytes_per_parameter / TP

Where bytes_per_parameter depends on precision:

Precision

Bytes Per Parameter

BF16

2

FP16

2

FP8

1

INT4

0.5

NVFP4

0.5

Examples

Model

Precision

Tensor Parallelism
(TP)

Calculation

Memory per GPU

Notes

Llama 3.1 8B

BF16

1

8 billion x 2 bytes

16 GB

Fits on a single 24 GB GPU (for example, A10G or RTX 4090) with room for KV cache and overhead

Llama 3.3 70B

BF16

4

70 billion x 2 bytes / 4 GPUs

35 GB

Fits on four A100-40GB or four A100-80GB GPUs, with varying amounts of room remaining for KV cache

Llama 3.3 70B

FP8

2

70 billion x 1 byte / 2 GPUs

35 GB

Using FP8 halves the weight memory, allowing the same model to run on two GPUs instead of four

Beyond weights, GPU memory is needed for KV cache, activations, communication buffers, CUDA graphs, and any LoRA adapters, multimodal reservations, or hybrid-model state. The vLLM --gpu-memory-utilization parameter controls the GPU-memory budget used to size the KV cache after other measured allocations are accounted for. NIM_KVCACHE_PERCENT configures this same parameter, despite its name. Check the effective startup configuration: the image, profile, and user overrides can change the value. Refer to Advanced Configuration for configuration precedence.

GPU Total Memory
├── Requested Budget = total x gpu_memory_utilization
│   ├── Model Weights           (loaded from checkpoint files)
│   ├── Non-Torch Overhead      (NCCL communication buffers, CUDA context)
│   ├── Peak Activations        (intermediate computation tensors)
│   └── KV Cache                (fills all remaining budget)
│
└── Remaining Headroom = total x (1 - gpu_memory_utilization)
    └── Allocations not accounted for during memory profiling

The backend uses the remaining budget to size the KV cache. Allocation order and memory accounting vary by backend version and model. CUDA graphs, adapter buffers, multimodal buffers, and hybrid-model state can also consume memory before or during profiling; use the logs to identify the actual failing allocation.

Out of Memory During Weight Loading#

Symptoms

The OOM error appears early in startup, during model loading, before any messages about KV cache or graph compilation. Typical log patterns include the following:

torch.OutOfMemoryError: CUDA out of memory.

Cause

The GPU does not have enough memory to hold the model weights at the selected precision and tensor parallelism (TP) degree. For example, a 70-billion-parameter model in BF16 requires approximately 140 GB of weight memory, which does not fit on a single 80 GB GPU.

Resolution

Run list-model-profiles to find profiles that distribute the model across more GPUs or use a lower-precision quantization:

docker run --rm --gpus=all \
  -p 8000:8000 \
  ${NIM_LLM_IMAGE} \
  list-model-profiles

Look for profiles with higher tensor parallelism (TP) or pipeline parallelism (PP), or profiles that use FP8 or NVFP4 precision – preferably one with native hardware support on your GPU to avoid a performance penalty. Refer to the weight estimation formula in How GPU Memory Is Used for a quick check of whether the weights fit. For details on profile selection, refer to Model Profiles and Selection.

Out of Memory During LoRA Adapter Allocation#

LoRA adapter weights require memory in addition to the base model. Increasing --max-lora-rank or --max-loras can exhaust that memory even when the non-LoRA profile fits. If a failure begins after increasing these settings, compare with the profile’s defaults and use a supported profile with enough memory for the required adapters. Reducing the context length does not reduce adapter weight memory.

A kernel assertion about the maximum supported LoRA rank is a separate limit. Adding GPU memory or reducing the number of adapter slots does not remove that limit. Use an adapter rank supported by the model and backend; refer to the Release Notes for model-specific restrictions.

Out of Memory During KV Cache Allocation#

KV cache allocation failures fall into two categories: insufficient memory (the KV cache genuinely does not fit) and memory fragmentation (the KV cache fits in total but the allocator cannot find contiguous blocks). Check the error message to determine which applies.

Insufficient Memory#

Symptoms

The error occurs after weights are loaded, during memory profiling or KV cache block allocation. Look for log messages mentioning KV cache, determine_available_memory, or block allocation failures. NIM may also print an advisory warning before the crash similar to the following:

WARNING: Estimated VRAM (45.2 GB) exceeds available GPU memory (39.6 GB).
Consider reducing context length with --max-model-len=4096 (estimated 30.1 GB).

Cause

The model’s default context length (max_position_embeddings from config.json) requires more KV cache memory than fits in the GPU budget after weights, activations, and overhead are subtracted. This is common when running a model with a long native context (for example, 128K tokens) on a GPU with limited free memory.

Resolution

Reduce the context length with --max-model-len. Use the value suggested in the NIM warning, or choose a value appropriate for your workload:

docker run --gpus=all \
  -p 8000:8000 \
  -e NGC_API_KEY \
  ${NIM_LLM_IMAGE} \
  --max-model-len 4096

Note

Reducing --max-model-len limits the maximum sequence length (input + output tokens) per request. Choose a value that fits your use case.

The equivalent NIM setting is NIM_MAX_MODEL_LEN. Lowering --gpu-memory-utilization instead reduces the KV-cache budget and can make a KV-capacity failure worse. If model, adapter, or multimodal allocations already exhaust memory before KV-cache sizing, reducing the context length alone might not make the deployment fit. Check the known issue for any validated workaround.

Memory Fragmentation#

Symptoms

The error occurs during KV cache allocation and reports that the free memory is less than the requested allocation size. The error also mentions a large amount of memory “reserved by PyTorch but unallocated” and suggests setting expandable_segments:True similar to the following:

torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 1.09 GiB.
GPU 0 has a total capacity of 31.36 GiB of which 1.02 GiB is free.
Of the allocated memory 23.18 GiB is allocated by PyTorch, and
6.48 GiB is reserved by PyTorch but unallocated. If reserved but
unallocated memory is large try setting
PYTORCH_ALLOC_CONF=expandable_segments:True to avoid fragmentation.

Cause

The PyTorch CUDA memory allocator requires contiguous memory blocks for each tensor allocation. During model loading and torch.compile, many temporary allocations fragment the GPU address space into small, non-contiguous pieces. When KV cache allocation begins, the total free memory may be sufficient, but no single contiguous block is large enough for individual per-layer KV cache tensors (typically 1+ GiB each).

This is particularly common with hybrid architectures (such as NemotronH models that combine Mamba and attention layers) where torch.compile creates extensive temporary allocations.

Resolution

Set PYTORCH_ALLOC_CONF=expandable_segments:True to instruct PyTorch to use CUDA virtual memory APIs, which allow allocations to grow without requiring physically contiguous pages:

docker run --gpus=all \
  -p 8000:8000 \
  -e NGC_API_KEY \
  -e PYTORCH_ALLOC_CONF=expandable_segments:True \
  ${NIM_LLM_IMAGE}

This setting changes allocator behavior and can reduce fragmentation; it does not add GPU capacity. Check CUDA IPC compatibility limitations before using it with deployments that share CUDA allocations between processes.

Out of Memory During CUDA Graph Compilation or Warmup#

Symptoms

The OOM error appears during CUDA graph capture, memory profiling, or sampler warm-up. Depending on the backend and model, this can occur before or after KV-cache allocation. Typical log patterns include the following:

torch.OutOfMemoryError: CUDA out of memory.

preceded by messages such as:

Graph capturing finished in ...

or:

compile_or_warm_up_model

Cause

Graph capture or warm-up needs more memory than remains available. The amount depends on the model, capture configuration, and other allocations; there is no single amount of headroom that works for every profile.

Resolution

If the failure occurs after KV-cache allocation, reducing --gpu-memory-utilization can leave more headroom by shrinking the KV cache. For example, changing a configured value of 0.9 to 0.85 reduces the requested budget by 5% of total GPU memory. This is an illustration, not a validated value for every profile:

docker run --gpus=all \
  -p 8000:8000 \
  -e NGC_API_KEY \
  ${NIM_LLM_IMAGE} \
  --gpu-memory-utilization 0.85

For vLLM, setting NIM_DISABLE_CUDA_GRAPH=1 (or passing --enforce-eager) disables CUDA graphs and can help determine whether graph capture causes the failure. This can reduce inference throughput and does not remove other model or adapter allocations:

docker run --gpus=all \
  -p 8000:8000 \
  -e NGC_API_KEY \
  -e NIM_DISABLE_CUDA_GRAPH=1 \
  ${NIM_LLM_IMAGE}

Enable Detailed Logging#

NIM prints a GPU Memory Report at startup when NIM_LOG_LEVEL is set to INFO (the default) or DEBUG. At the INFO or DEBUG log levels, NIM also prints GPU Diagnostics (GPU summary, NVLink status, NVLink capabilities, and GPU topology from nvidia-smi) before model loading begins, and the startup banner includes estimated memory per GPU and the CPU core count required by the model configuration.

To enable detailed logging, set the NIM_LOG_LEVEL environment variable:

docker run --gpus=all \
  -p 8000:8000 \
  -e NGC_API_KEY \
  -e NIM_LOG_LEVEL=INFO \
  ${NIM_LLM_IMAGE}