Configuration
This section describes the various ways to configure a NIM container.
GPU Selection
Passing --gpus all to docker run is acceptable in homogeneous environments with one or more of the same GPU.
In heterogeneous environments with a combination of GPUs, such as an A6000 + a GeForce display GPU, workloads should only run on compute-capable GPUs. Expose specific GPUs inside the container using either:
- the
--gpusflag (ex:--gpus="device=1") - the environment variable
NVIDIA_VISIBLE_DEVICES(ex:-e NVIDIA_VISIBLE_DEVICES=1)
The device ID(s) to use as input(s) are listed in the output of nvidia-smi -L:
Refer to the NVIDIA Container Toolkit documentation for more instructions.
Shared memory flag
Tokenization uses Triton’s Python backend capabilities that scales with the number of CPU cores available. You may need to increase the available shared memory given to the microservice container.
Example providing 1g of shared memory:
Environment Variables
The following table describes the environment variables that can be passed into a NIM, as a -e argument added to a docker run command:
Volumes
The following table describes the paths inside the container into which local paths can be mounted.