Troubleshooting GPU Memory Out-of-Memory Errors#
GPU out-of-memory (OOM) errors occur when the model needs more VRAM than the GPU provides. The usual causes are a model profile that does not match the hardware, or a misconfigured memory setting.
Startup allocates GPU memory in phases, and the fix depends on which phase failed. Check the container logs, then follow the matching issue below. If the default logs are not enough, refer to Enable Detailed Logging.
Start with a GPU/profile combination listed in the Support Matrix for NIMs and check the Release Notes for model-specific startup failures. The steps below diagnose memory pressure; they do not establish a workaround for a known backend defect or make an unsupported profile supported. An illegal-memory-access error or worker crash during warm-up is not, by itself, evidence of an OOM.
How GPU Memory Is Used#
During startup, NIM and the vLLM backend allocate GPU memory in stages. The largest single consumer is the model weights. You can estimate weight memory from the model’s parameter count (typically listed on the model card, or available in the model.safetensors.index.json file under the metadata.total_size field) and the precision of the profile.
A good heuristic formula for per-GPU weight memory is:
weight_memory_per_gpu = total_parameters x bytes_per_parameter / TP
Where bytes_per_parameter depends on precision:
Precision |
Bytes Per Parameter |
|---|---|
BF16 |
2 |
FP16 |
2 |
FP8 |
1 |
INT4 |
0.5 |
NVFP4 |
0.5 |
Examples
Model |
Precision |
Tensor Parallelism (TP) |
Calculation |
Memory per GPU |
Notes |
|---|---|---|---|---|---|
Llama 3.1 8B |
BF16 |
1 |
8 billion x 2 bytes |
16 GB |
Fits on a single 24 GB GPU (for example, A10G or RTX 4090) with room for KV cache and overhead |
Llama 3.3 70B |
BF16 |
4 |
70 billion x 2 bytes / 4 GPUs |
35 GB |
Fits on four A100-40GB or four A100-80GB GPUs, with varying amounts of room remaining for KV cache |
Llama 3.3 70B |
FP8 |
2 |
70 billion x 1 byte / 2 GPUs |
35 GB |
Using FP8 halves the weight memory, allowing the same model to run on two GPUs instead of four |
Beyond weights, GPU memory is needed for KV cache, activations, communication
buffers, CUDA graphs, and any LoRA adapters, multimodal reservations, or hybrid-model
state. The vLLM --gpu-memory-utilization parameter controls the GPU-memory budget
used to size the KV cache after other measured allocations are accounted for.
NIM_KVCACHE_PERCENT configures this same parameter, despite its name. Check the
effective startup configuration: the image, profile, and user overrides can change
the value. Refer to Advanced Configuration for configuration precedence.
GPU Total Memory
├── Requested Budget = total x gpu_memory_utilization
│ ├── Model Weights (loaded from checkpoint files)
│ ├── Non-Torch Overhead (NCCL communication buffers, CUDA context)
│ ├── Peak Activations (intermediate computation tensors)
│ └── KV Cache (fills all remaining budget)
│
└── Remaining Headroom = total x (1 - gpu_memory_utilization)
└── Allocations not accounted for during memory profiling
The backend uses the remaining budget to size the KV cache. Allocation order and memory accounting vary by backend version and model. CUDA graphs, adapter buffers, multimodal buffers, and hybrid-model state can also consume memory before or during profiling; use the logs to identify the actual failing allocation.
Out of Memory During Weight Loading#
Symptoms
The OOM error appears early in startup, during model loading, before any messages about KV cache or graph compilation. Typical log patterns include the following:
torch.OutOfMemoryError: CUDA out of memory.
Cause
The GPU does not have enough memory to hold the model weights at the selected precision and tensor parallelism (TP) degree. For example, a 70-billion-parameter model in BF16 requires approximately 140 GB of weight memory, which does not fit on a single 80 GB GPU.
Resolution
Run list-model-profiles to find profiles that distribute the model across more GPUs or use a lower-precision quantization:
docker run --rm --gpus=all \
-p 8000:8000 \
${NIM_LLM_IMAGE} \
list-model-profiles
Look for profiles with higher tensor parallelism (TP) or pipeline parallelism (PP), or profiles that use FP8 or NVFP4 precision – preferably one with native hardware support on your GPU to avoid a performance penalty. Refer to the weight estimation formula in How GPU Memory Is Used for a quick check of whether the weights fit. For details on profile selection, refer to Model Profiles and Selection.
Out of Memory During LoRA Adapter Allocation#
LoRA adapter weights require memory in addition to the base model. Increasing
--max-lora-rank or --max-loras can exhaust that memory even when the non-LoRA
profile fits. If a failure begins after increasing these settings, compare with
the profile’s defaults and use a supported profile with enough memory for the
required adapters. Reducing the context length does not reduce adapter weight memory.
A kernel assertion about the maximum supported LoRA rank is a separate limit. Adding GPU memory or reducing the number of adapter slots does not remove that limit. Use an adapter rank supported by the model and backend; refer to the Release Notes for model-specific restrictions.
Out of Memory During KV Cache Allocation#
KV cache allocation failures fall into two categories: insufficient memory (the KV cache genuinely does not fit) and memory fragmentation (the KV cache fits in total but the allocator cannot find contiguous blocks). Check the error message to determine which applies.
Insufficient Memory#
Symptoms
The error occurs after weights are loaded, during memory profiling or KV cache block allocation. Look for log messages mentioning KV cache, determine_available_memory, or block allocation failures. NIM may also print an advisory warning before the crash similar to the following:
WARNING: Estimated VRAM (45.2 GB) exceeds available GPU memory (39.6 GB).
Consider reducing context length with --max-model-len=4096 (estimated 30.1 GB).
Cause
The model’s default context length (max_position_embeddings from config.json) requires more KV cache memory than fits in the GPU budget after weights, activations, and overhead are subtracted. This is common when running a model with a long native context (for example, 128K tokens) on a GPU with limited free memory.
Resolution
Reduce the context length with --max-model-len. Use the value suggested in the NIM warning, or choose a value appropriate for your workload:
docker run --gpus=all \
-p 8000:8000 \
-e NGC_API_KEY \
${NIM_LLM_IMAGE} \
--max-model-len 4096
Note
Reducing --max-model-len limits the maximum sequence length (input + output tokens) per request. Choose a value that fits your use case.
The equivalent NIM setting is NIM_MAX_MODEL_LEN. Lowering
--gpu-memory-utilization instead reduces the KV-cache budget and can make a
KV-capacity failure worse. If model, adapter, or multimodal allocations already
exhaust memory before KV-cache sizing, reducing the context length alone might
not make the deployment fit. Check the known issue for any validated workaround.
Memory Fragmentation#
Symptoms
The error occurs during KV cache allocation and reports that the free memory is less than the requested allocation size. The error also mentions a large amount of memory “reserved by PyTorch but unallocated” and suggests setting expandable_segments:True similar to the following:
torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 1.09 GiB.
GPU 0 has a total capacity of 31.36 GiB of which 1.02 GiB is free.
Of the allocated memory 23.18 GiB is allocated by PyTorch, and
6.48 GiB is reserved by PyTorch but unallocated. If reserved but
unallocated memory is large try setting
PYTORCH_ALLOC_CONF=expandable_segments:True to avoid fragmentation.
Cause
The PyTorch CUDA memory allocator requires contiguous memory blocks for each tensor allocation. During model loading and torch.compile, many temporary allocations fragment the GPU address space into small, non-contiguous pieces. When KV cache allocation begins, the total free memory may be sufficient, but no single contiguous block is large enough for individual per-layer KV cache tensors (typically 1+ GiB each).
This is particularly common with hybrid architectures (such as NemotronH models that combine Mamba and attention layers) where torch.compile creates extensive temporary allocations.
Resolution
Set PYTORCH_ALLOC_CONF=expandable_segments:True to instruct PyTorch to use CUDA virtual memory APIs, which allow allocations to grow without requiring physically contiguous pages:
docker run --gpus=all \
-p 8000:8000 \
-e NGC_API_KEY \
-e PYTORCH_ALLOC_CONF=expandable_segments:True \
${NIM_LLM_IMAGE}
This setting changes allocator behavior and can reduce fragmentation; it does not add GPU capacity. Check CUDA IPC compatibility limitations before using it with deployments that share CUDA allocations between processes.
Out of Memory During CUDA Graph Compilation or Warmup#
Symptoms
The OOM error appears during CUDA graph capture, memory profiling, or sampler warm-up. Depending on the backend and model, this can occur before or after KV-cache allocation. Typical log patterns include the following:
torch.OutOfMemoryError: CUDA out of memory.
preceded by messages such as:
Graph capturing finished in ...
or:
compile_or_warm_up_model
Cause
Graph capture or warm-up needs more memory than remains available. The amount depends on the model, capture configuration, and other allocations; there is no single amount of headroom that works for every profile.
Resolution
If the failure occurs after KV-cache allocation, reducing
--gpu-memory-utilization can leave more headroom by shrinking the KV cache.
For example, changing a configured value of 0.9 to 0.85 reduces the requested
budget by 5% of total GPU memory. This is an illustration, not a validated value
for every profile:
docker run --gpus=all \
-p 8000:8000 \
-e NGC_API_KEY \
${NIM_LLM_IMAGE} \
--gpu-memory-utilization 0.85
For vLLM, setting NIM_DISABLE_CUDA_GRAPH=1 (or passing --enforce-eager)
disables CUDA graphs and can help determine whether graph capture causes the
failure. This can reduce inference throughput and does not remove other model
or adapter allocations:
docker run --gpus=all \
-p 8000:8000 \
-e NGC_API_KEY \
-e NIM_DISABLE_CUDA_GRAPH=1 \
${NIM_LLM_IMAGE}
Enable Detailed Logging#
NIM prints a GPU Memory Report at startup when NIM_LOG_LEVEL is set to INFO (the default) or DEBUG. At the INFO or DEBUG log levels, NIM also prints GPU Diagnostics (GPU summary, NVLink status, NVLink capabilities, and GPU topology from nvidia-smi) before model loading begins, and the startup banner includes estimated memory per GPU and the CPU core count required by the model configuration.
To enable detailed logging, set the NIM_LOG_LEVEL environment variable:
docker run --gpus=all \
-p 8000:8000 \
-e NGC_API_KEY \
-e NIM_LOG_LEVEL=INFO \
${NIM_LLM_IMAGE}