Getting Started#
This page shows how to authenticate with NGC, launch the LipSync NIM container, and configure its runtime.
Prerequisites#
To ensure that you have the supported hardware and software stack, check the Support Matrix.
NGC Authentication#
An NGC API key is required to pull the container image and download models. Complete the following steps to create the key, export it, and log in to the NVIDIA Container Registry.
Generate an API Key#
LipSync NIM is available only through the AI for Media Private Access Program. Joining the Private Access Program gives you an NGC API key with the permissions required for this NIM.
You can generate a key at https://org.ngc.nvidia.com/setup/api-keys after joining the Private Access Program.
When creating an NGC API Personal key, ensure that at least NGC Catalog is selected from the Services Included dropdown. You can include more services if this key is to be reused for other purposes.
Note
Personal keys allow you to configure an expiration date, revoke or delete the key using an action button, and rotate the key as needed. For more information about key types, refer to NGC API Keys in the NGC User Guide.
Export the NGC API Key#
Pass the value of the API key to the docker run command in the next section as the NGC_API_KEY environment variable to download the appropriate models and resources when starting the NIM.
If you are not familiar with how to create the NGC_API_KEY environment variable, the simplest way is to export it in your terminal:
export NGC_API_KEY=<value>
Run one of the following commands to make the key available at startup:
# If using bash
echo "export NGC_API_KEY=<value>" >> ~/.bashrc
# If using zsh
echo "export NGC_API_KEY=<value>" >> ~/.zshrc
Note
Other, more secure options include saving the value in a file, so that you can retrieve with cat $NGC_API_KEY_FILE, or using a password manager.
Docker Login to NGC#
To pull the NIM container image from NGC, first authenticate with the NVIDIA Container Registry with the following command:
echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin
Use $oauthtoken as the username and NGC_API_KEY as the password. The $oauthtoken username is a special name that indicates that you will authenticate with an API key and not a user name and password.
Note
Before pulling the container image, you must accept the Governing Terms on the LipSync NIM container page in NGC. This page is accessible only after enrollment in the AI for Media Private Access Program.
Launching the NIM Container#
The following command launches the LipSync NIM container with the gRPC service. The NIM detects the GPU architecture of the host and downloads the matching model automatically. (For a list of parameters, refer to Runtime Parameters for the Container.)
docker run -it --rm --name=lipsync-nim \
--runtime=nvidia \
--gpus all \
--shm-size=8GB \
-e NGC_API_KEY=$NGC_API_KEY \
-e NV_AI4M_MAX_CONCURRENCY_PER_GPU=1 \
-e NIM_HTTP_API_PORT=8000 \
-e NIM_GRPC_API_PORT=8001 \
-p 8000:8000 \
-p 8001:8001 \
-p 9002:9002 \
nvcr.io/nim/nvidia/lipsync:latest
Note
Port 9002 publishes the Triton Inference Server metrics endpoint, which
Observability uses. It is not required for inference; omit it if you do
not intend to scrape metrics.
Mount a host directory at NIM_CACHE_PATH (for example
-v $HOME/.cache/nim:/opt/nim/.cache) to persist model artifacts across runs and avoid
re-downloading them on every launch.
Note
The flag --gpus all is used to assign all available GPUs to the NIM container.
To assign specific GPUs to the NIM container (if your machine has multiple GPUs), use --gpus '"device=0,1,2..."'.
On bare-metal Linux systems with PCIe topology and multiple GPUs, we recommend setting IOMMU to passthrough. Refer to PCIe Multi-GPU Systems.
If the NIM launch is successful, you get a response similar to the following.
I1027 22:31:44.952125 123 grpc_server.cc:2560] "Started GRPCInferenceService at 127.0.0.1:9001"
I1027 22:31:44.952247 123 http_server.cc:4755] "Started HTTPService at 127.0.0.1:9000"
I1027 22:31:44.993329 123 http_server.cc:358] "Started Metrics Service at 127.0.0.1:9002"
Triton server is ready
[INFO AI4M BASE LOGGER 2025-10-27 22:31:46.097 PID:207] Using threading mode for gRPC service
[INFO AI4M BASE LOGGER 2025-10-27 22:31:46.097 PID:207] Starting threading gRPC service with 1 threads
[INFO AI4M BASE LOGGER 2025-10-27 22:31:46.105 PID:207] Using Insecure Server Credentials
[INFO AI4M BASE LOGGER 2025-10-27 22:31:46.107 PID:207] Listening to 0.0.0.0:8001
Note
By default, the LipSync NIM gRPC service is hosted on port 8001. You must use this port for inferencing requests. The port is configurable via the NIM_GRPC_API_PORT environment variable.
Selecting a Language-Specific Model#
The LipSync NIM includes a generic, language-agnostic model (used by default) and fine-tuned models for specific languages. You select a language-specific model at NIM container launch time by using the NIM_TAGS_SELECTOR environment variable. Each container loads exactly one model.
|
Model used |
|---|---|
(not set) |
Generic, language-agnostic model (default) |
|
German |
|
Spanish |
|
French |
Only de, es, and fr are supported. Any other value causes the container to fail at startup; omit the variable to use the generic model.
Profile selection stays automatic when you request a language: the NIM selects the profile that matches both your GPU architecture and the requested language.
The following example launches the NIM with the German model:
docker run -it --rm --name=lipsync-nim \
--runtime=nvidia \
--gpus all \
--shm-size=8GB \
-e NGC_API_KEY=$NGC_API_KEY \
-e NIM_TAGS_SELECTOR="language=de" \
-e NIM_HTTP_API_PORT=8000 \
-e NIM_GRPC_API_PORT=8001 \
-p 8000:8000 \
-p 8001:8001 \
nvcr.io/nim/nvidia/lipsync:latest
Environment Variables#
The following table describes the environment variables that can be passed into a NIM as a -e argument added to a docker run command. The NV_AI4M_LS_* variables that govern buffering and timeouts for the input streams are grouped and explained in Input Stream Handling.
ENV |
Required? |
Default |
Notes |
|---|---|---|---|
|
Yes |
None |
You must set this variable to the value of your personal NGC API key. |
|
Optional |
|
Location (in container) where the container caches model artifacts. |
|
No |
|
Publish the gRPC |
|
No |
|
Publish the NIM HTTP service to the prescribed port inside the container. Supported endpoints are |
|
Optional |
None |
Pins the NIM to one specific model profile. Leave it unset to let the NIM select the profile that matches the detected GPU architecture. The earlier name |
|
Optional |
None |
Selects the language-specific model to load, in the form |
|
No |
disabled |
Set SSL security on the endpoints to |
|
No |
None |
Set the path to CA root certificate inside the NIM. This is required only when |
|
No |
None |
Set the path to the server’s public SSL certificate inside the NIM. This is required only when an SSL mode is enabled. For example, if the SSL certificates are mounted at |
|
No |
None |
Set the path to the server’s private key inside the NIM. This is required only when an SSL mode is enabled. For example, if the SSL certificates are mounted at |
|
No |
|
Set to |
|
No |
|
Number of concurrent inference requests the NIM server supports per GPU. Higher values consume more GPU memory and can cause out-of-memory errors. Buffer caps and coverage timeouts apply per request, so review them when you raise this value. Refer to Input Stream Handling. |
|
No |
|
Seconds to wait for the beginning of the video, which selects streaming or transactional mode. A client that opens a request and then sends no video fails with |
|
No |
|
Seconds to wait for the video needed to produce one output frame. Refer to Coverage Timeouts. |
|
No |
|
Seconds to wait for the speech audio needed to produce one output frame. Refer to Coverage Timeouts. |
|
No |
|
Seconds to wait for the speaker information needed to produce one output frame. Applies only when you supply speaker data. Refer to Coverage Timeouts. |
|
No |
|
Seconds to wait for the background audio needed to produce one output frame. On expiry the background audio drops to silence instead of failing the request. Refer to Coverage Timeouts. |
|
No |
Sum of the four values above |
Overall cap on the time spent assembling one output frame across every input. Set to |
|
No |
|
Per-frame budget used in transactional mode, replacing the four per-stream values. Refer to Coverage Timeouts. |
|
No |
|
Cap, in MB, on encoded video held while the video decoder is behind. Refer to Buffer Caps. |
|
No |
|
Cap, in MB, on encoded audio held while the audio decoder is behind. Speech and background audio each get this allowance. Refer to Buffer Caps. |
|
No |
|
Cap, in |
|
No |
|
Bound on the queue of decoded video frames waiting for inference. Refer to Buffer Caps. |
Runtime Parameters for the Container#
The following table describes the docker run flags used to launch the LipSync NIM container.
Flags |
Description |
|---|---|
|
|
|
Delete the container after it stops (see docker container run). |
|
Give a name to the NIM container. Use any preferred value. |
|
Ensure NVIDIA drivers are accessible in the container. |
|
Expose NVIDIA GPUs inside the container. If you are running on a host with multiple GPUs, you need to specify which GPU to use. You can also specify multiple GPUs. For more information about mounting specific GPUs, see GPU Enumeration. |
|
Allocate host memory for multi-process communication. |
|
Number of concurrent inference requests to be supported by the NIM server per GPU (default: 1). Higher values consume more GPU memory and can cause out-of-memory errors. |
|
Provide the container with the token necessary to download adequate models and resources from NGC. See NGC Authentication. |
|
Ports published by the container are directly accessible on the host port. |
|
Environment variable to enable debug mode that overlays frame number, lipsync effect status and bounding boxes for each frame. Enable by setting it to 1. (default: 0). |
HTTP API Endpoints#
In addition to the gRPC Lipsync service on NIM_GRPC_API_PORT, the NIM serves three
read-only HTTP endpoints on NIM_HTTP_API_PORT (default 8000). Publish that port with
-p 8000:8000 to reach them from the host.
Endpoint |
Method |
Purpose |
|---|---|---|
|
GET |
Returns the license text shipped in the container. |
|
GET |
Returns asset information, license information, model information, and version. |
|
GET |
Exposes Prometheus metrics via an ASGI app endpoint. |
The following example queries the /v1/metadata endpoint from the host:
curl -s http://localhost:8000/v1/metadata
An abbreviated response:
{
"assetInfo": [""],
"licenseInfo": {
"name": "LICENSE",
"path": "/opt/nim/LICENSE",
"type": "file"
},
"modelInfo": [
{
"modelUrl": "ngc://nim/nvidia/lipsync:r22-sm89-en"
}
]
}
Note
/v1/metrics on the HTTP port is distinct from the Triton metrics endpoint on port 9002
described in Observability. They are different services on different
ports.
Stopping the Container#
The following command can be used to stop the container.
docker stop lipsync-nim