Advanced Usage#

This page describes deployment options, request parameters, and runtime behavior for the LipSync NIM.

Model Caching#

When the container launches for the first time, it downloads the required models from NGC. To avoid downloading the models on subsequent runs, you can cache them locally by using a cache directory:

# Create the cache directory on the host machine
export LOCAL_NIM_CACHE=~/.cache/nim
mkdir -p "$LOCAL_NIM_CACHE"
chmod 777 $LOCAL_NIM_CACHE

# Run the container with the cache directory mounted in the appropriate location
docker run -it --rm --name=lipsync-nim \
  --runtime=nvidia \
  --gpus all \
  --shm-size=8GB \
  -e NGC_API_KEY=$NGC_API_KEY \
  -e NIM_HTTP_API_PORT=8000 \
  -p 8000:8000 \
  -p 8001:8001 \
  -v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
  nvcr.io/nim/nvidia/lipsync:latest

The NIM selects the model profile automatically from the detected GPU architecture. To cache a specific profile instead, add -e NIM_MODEL_PROFILE=<model_profile_id>; refer to Model Manifest Profile.

Model Manifest Profile#

NIM_MODEL_PROFILE is optional. Set it only when you need to pin the NIM to one specific profile instead of the profile it selects automatically. The model manifest holds one profile for every combination of GPU architecture and language, listed in the following table.

Warning

Pin a profile only when you have a specific reason to, and make sure the ID you pass matches both your GPU architecture and the language you intend to load. Automatic selection is the supported path and is correct for almost every deployment.

A profile ID that does not exist in the manifest is reported at startup as NIMProfileIDNotFound. A profile ID that exists but was built for a different GPU architecture can leave the container in a degraded state: it starts normally, reports the model as READY, and begins listening, but is not able to serve inference requests on that GPU. Verify the ID against the table below before using it.

GPU Architecture (compute capability)

Language

NIM_TAGS_SELECTOR

Manifest Profile ID

Blackwell (cc 12.0)

Generic (language-agnostic)

(not set)

cd156ba5006903499651d2853e077e7ce06c65b9eb675266355351d91778860f

Blackwell (cc 12.0)

German

language=de

c383edbc670ce48fce1440fc72ecae6a21dbff2dc62654e24c4b24f6b65c7c61

Blackwell (cc 12.0)

Spanish

language=es

783e2890c3ddc4779608697556b789cce5e3ee2cf3ee726e76bd342dac85ba31

Blackwell (cc 12.0)

French

language=fr

615d3ea18b11fa8242e489079025b2e10192bfd3c84695723f9330f661c13380

Ada (cc 8.9)

Generic (language-agnostic)

(not set)

94714c8a037124df18c71cf1eed83a282547d1dea9a30a691a2d7e337c1d64b8

Ada (cc 8.9)

German

language=de

cbddaf2240ab082ccdec4aa22c98c23b2e6ea1bda31dd3e12113a827e95c3ca6

Ada (cc 8.9)

Spanish

language=es

6a371d465d795693e251de7e4515723d095ecf1e390884ec1cd5f2ed997cb028

Ada (cc 8.9)

French

language=fr

c1e415cab2d2c3c4015f93395667894b5bee830ef87eaf6223d5e647b4b1d86b

Ampere (cc 8.6)

Generic (language-agnostic)

(not set)

fedf140a2ac1bf0a149ee0acd81cb23d5dc2cc1036e2ec2762489b1457506bc8

Ampere (cc 8.6)

German

language=de

84740f0aa874fe3e765f6e6bb80523d38ba42fa71f44b696e59d42d8b3903adb

Ampere (cc 8.6)

Spanish

language=es

f34aaf8b97d20e476a7a908940054e757dddcf36a322398ef661bd9f1b0fdfc9

Ampere (cc 8.6)

French

language=fr

56d8891fdd8b00438ae0b78fea32a2ae544d1b982da13be238c8bc3e7c355dd6

Turing (cc 7.5)

Generic (language-agnostic)

(not set)

682fdf66c18c462977c0fd8d98afae03d27f6ed4657d86eeb4268d581ab96a66

Turing (cc 7.5)

German

language=de

b1544cba4251da166a18047113a10bf7f01010abb0c96a8d6ca11be65c317754

Turing (cc 7.5)

Spanish

language=es

4708dbd582b564f0835a5011c4f52c3b48250f08594a5cceb1829bf377bd53da

Turing (cc 7.5)

French

language=fr

511b4df5f555f76fe73f5245d4f2554c32ba4c57d18b1650fbc45e7755673169

To pin a profile, pass its ID to the launch command:

# Choose the model profile ID that matches BOTH the target GPU architecture and the language.
export MODEL_PROFILE_ID=<enter_valid_model_profile_id>

docker run -it --rm --name=lipsync-nim \
  --runtime=nvidia \
  --gpus all \
  --shm-size=8GB \
  -e NGC_API_KEY=$NGC_API_KEY \
  -e NIM_MODEL_PROFILE=$MODEL_PROFILE_ID \
  -e NV_AI4M_MAX_CONCURRENCY_PER_GPU=1 \
  -e NIM_HTTP_API_PORT=8000 \
  -e NIM_GRPC_API_PORT=8001 \
  -p 8000:8000 \
  -p 8001:8001 \
  nvcr.io/nim/nvidia/lipsync:latest

Ensure that the GPU architecture associated with the manifest profile ID is compatible with the target hardware. If an incorrect manifest profile ID is used, a deserialization error occurs on inference.

Note

The manifest profile ID determines only which model files are downloaded. The model that the service loads is resolved from NIM_TAGS_SELECTOR, so the two must agree: when you pin a language-specific profile ID, also set NIM_TAGS_SELECTOR to the language in the same row (refer to Selecting a Language-Specific Model).

SSL enablement#

LipSync NIM provides an SSL mode to ensure secure communication between clients and the server by encrypting data in transit. To enable SSL, you must provide the path to the SSL certificate and key files in the container. The following example shows how to do this:

export NGC_API_KEY=<add-your-api-key>
SSL_CERT=path/to/ssl_key

docker run -it --rm --name=lipsync-nim \
  --runtime=nvidia \
  --gpus all \
  --shm-size=8GB \
  -v $SSL_CERT:/opt/nim/crt/:ro \
  -e NGC_API_KEY=$NGC_API_KEY \
  -p 8000:8000 \
  -p 8001:8001 \
  -e NIM_SSL_MODE="mtls" \
  -e NIM_SSL_CA_CERTS_PATH="/opt/nim/crt/ssl_ca_cert.pem" \
  -e NIM_SSL_CERT_PATH="/opt/nim/crt/ssl_cert_server.pem" \
  -e NIM_SSL_KEY_PATH="/opt/nim/crt/ssl_key_server.pem" \
  nvcr.io/nim/nvidia/lipsync:latest

NIM_SSL_MODE can be set to mtls, tls, or disabled. If set to mtls, the container uses mutual TLS authentication. If set to tls, the container uses TLS authentication. For more information, refer to Environment Variables.

Be sure to verify the permissions of the SSL certificate and key files on the host machine. The container cannot access the files if they are not readable by the user running the container.

NIM Service Configuration Parameters via Client#

The following options can be configured by each request to the Lipsync NIM:

Video-Audio Alignment Options#

The following parameters control how the NIM handles mismatched audio and video durations.

  • extend_audio: Controls how to handle cases where video is longer than audio. This parameter is useful when working with content where the video track extends beyond the audio duration, such as when processing silent video segments, incomplete audio recordings, or when synchronizing content with varying durations.

    • EXTEND_AUDIO_UNSPECIFIED (default): Truncates video to match audio length.

    • EXTEND_AUDIO_SILENCE: Adds silent audio padding to keep audio and video synchronized.

    This option can be configured by setting the extend_audio parameter to EXTEND_AUDIO_SILENCE in the configuration message when making requests to the NIM.

    The following example uses the sample Python client to extend audio with custom options for handling mismatched input video and audio durations:

    python3 lipsync.py --target 127.0.0.1:8001 --video-input /path/to/video.mp4 --audio-input /path/to/audio.wav --extend-audio silence
    
  • extend_video: Controls how to handle cases where audio is longer than video. This parameter is useful when working with content where the audio track extends beyond the video duration, such as when adding voiceovers, dubbing content etc.

    • EXTEND_VIDEO_UNSPECIFIED (default): Keeps original video length and truncates longer audio.

    • EXTEND_VIDEO_FORWARD: Extends video duration by repeating the final 5 seconds of frames in forward order until matching audio length.

    • EXTEND_VIDEO_REVERSE: Extends video duration by repeating the final 5 seconds of frames in reverse order until matching audio length.

    This option can be configured by setting the extend_video to EXTEND_VIDEO_FORWARD or EXTEND_VIDEO_REVERSE in the configuration message when making requests to the NIM.

    The following example uses the sample Python client to extend video with custom options for handling mismatched input video and audio durations:

    python3 lipsync.py --target 127.0.0.1:8001 --video-input /path/to/video.mp4 --audio-input /path/to/audio.wav --extend-video reverse
    

Encoding Options#

The output video encoding quality can be configured through the output_video_encoding parameter in the following ways:

  • Lossless encoding: Provides maximum quality with no compression artifacts, preserving the original video quality pixel-for-pixel. This mode maintains the highest fidelity but results in significantly larger file sizes. Use this mode when quality is the top priority.

    • To configure this option in the configuration message when making requests to the NIM, set output_video_encoding to VideoEncoding(lossless=True).

    • To run LipSync with lossless encoding (which overrides the bitrate setting) via the sample client, use the --lossless option:

      python3 lipsync.py --target 127.0.0.1:8001 --video-input /path/to/video.mp4 --audio-input /path/to/audio.wav --lossless
      
  • Bitrate control: Allows balancing quality and file size by specifying the video bitrate in megabits per second (Mbps). The default is 30 Mbps. Higher bitrates produce better quality but larger file sizes; lower bitrates reduce file size at the cost of quality.

    • To configure this option in the configuration message when making requests to the NIM, set output_video_encoding to VideoEncoding(lossy=LossyEncoding(bitrate_mbps=<desired-bitrate-in-mbps>)).

    • To run LipSync with a desired output bitrate via the sample client, use the --bitrate option:

      python3 lipsync.py --target 127.0.0.1:8001 --video-input /path/to/video.mp4 --audio-input /path/to/audio.wav --bitrate 20
      
  • IDR interval control: Sets the interval between Instantaneous Decoder Refresh (IDR) frames (default: 8) for seeking and random-access capabilities. Lower values improve seeking accuracy, random access, and overall encoding quality but increase file size; higher values reduce file size but can impact seeking performance and quality.

    • To configure this option in the configuration message when making requests to the NIM, set output_video_encoding to VideoEncoding(lossy=LossyEncoding(idr_interval=<desired_idr_interval>)).

    • To run LipSync with a desired IDR interval via the sample client, use the --idr-interval option:

      python3 lipsync.py --target 127.0.0.1:8001 --video-input /path/to/video.mp4 --audio-input /path/to/audio.wav --idr-interval 10
      
  • Custom encoding parameters: Provides fine-grained control for expert users using JSON configuration. These parameters configure properties of the DeepStream H264 encoder.

    To specify custom encoding parameters, set output_video_encoding to custom_encoding in the configuration message sent to the NIM in every request:

    Custom encoding parameters are passed as a JSON object whose keys are nvv4l2h264enc GStreamer property names. Note that these are the encoder’s own property spellings, which differ from the proto field names — the IDR interval property is idrinterval, not idr_interval.

    {
      "bitrate": 5000000,
      "idrinterval": 16,
      "maxbitrate": 6000000
    }
    

    Via the sample client:

    python3 lipsync.py --target 127.0.0.1:8001 --video-input /path/to/video.mp4 --audio-input /path/to/audio.wav \
      --custom-encoding-params '{"bitrate": 5000000, "idrinterval": 16, "maxbitrate": 6000000}'
    

    Warning

    Passing a property name the encoder does not recognize currently fails the request with INTERNAL and the message Server Error; the underlying cause appears only in the container log. Verify property names against the DeepStream encoder documentation before use.

    Note

    Custom encoding parameters override standard bitrate and IDR interval settings. Use with caution because incorrect values can affect output quality or encoding stability.

Speaker Data Option#

The speaker_data_input option specifies the path to a JSON file that defines per-frame speaker bounding box information for targeted lip synchronization. This option is specifically designed to support multi-speaker scenarios by letting you specify the active speaker to lip-sync on a per-frame basis. The file contains bounding-box coordinates (x, y, width, height), speaker identity, and speaking status for each frame. If not provided, the LipSync NIM automatically detects faces in each frame using its built-in face detection capabilities.

When providing speaker data, set is_speaker_info_provided to True in the configuration message when making requests to the NIM.

The SpeakerInfo structure contains the following fields:

Field

Type

Description

speaker_bbox

BoundingBox

Defines the bounding box coordinates for the speaker’s face region.

speaker_id

int32

Unique identifier for this speaker across frames.

is_speaking

bool

Flag indicating whether the speaker is currently speaking.

The BoundingBox structure contains the following fields:

Field

Type

Description

x

float

X-coordinate of the top-left corner of the bounding box.

y

float

Y-coordinate of the top-left corner of the bounding box.

width

float

Width of the bounding box in pixels.

height

float

Height of the bounding box in pixels.

The SpeakerInfoPerFrame structure wraps per-frame data:

Field

Type

Description

frame_id

uint32

Frame index in the video.

speaker_infos

repeated SpeakerInfo

List of all speakers in this frame.

bypass

bool (optional)

If set to true, LipSync processing is bypassed for this frame and the original frame is returned unchanged. If false or unset, the frame is processed normally. This field is useful for cutscenes or frames that contain no subject.

When providing speaker data, you must send a SpeakerInfoPerFrame object for every frame in the input video. For multi-speaker scenarios, each frame can contain multiple speakers.

JSON File Format for Speaker Data

When creating a JSON file for speaker data, use the following format. The file must contain a top-level frames array with one entry per frame, as in the following example:

{
  "frames": [
    {
      "speakers": [
        {
          "bbox": [186, 191, 175, 254],
          "speaker_id": 1,
          "is_speaking": false
        },
        {
          "bbox": [815, 188, 263, 357],
          "speaker_id": 2,
          "is_speaking": true
        }
      ]
    },
    {
      "speakers": [
        {
          "bbox": [188, 191, 174, 254],
          "speaker_id": 1,
          "is_speaking": false
        }
      ],
      "bypass": false
    }
  ]
}

Each frame entry accepts the following fields:

  • speakers: Array of speaker entries for that frame. If empty or absent, the server auto-detects faces for that frame.

  • bypass: (Optional) Boolean indicating whether to skip lip sync processing for this frame. When true, the original video frame is passed through unchanged. When false or absent, lip sync is applied normally. Use it to leave segments without a speaking subject untouched.

Each entry in the speakers array contains:

  • bbox: Array of four values [x, y, width, height] defining the bounding box of the speaker’s face in pixel coordinates.

  • speaker_id: Integer identifier for tracking this face across frames.

  • is_speaking: Boolean indicating whether the speaker is currently speaking.

Obtaining Bounding Box Coordinates

To create the speaker data JSON file, you need to obtain bounding-box coordinates for each frame. You can use an external face-detection system to generate these coordinates.

Note

The quality and accuracy of your face detection directly affects the LipSync results. For optimal lip-synchronization performance, ensure that your bounding boxes accurately encompass the facial regions.

To run LipSync NIM with a speaker data file via the sample client, use the --speaker-data-input option:

python3 lipsync.py --target 127.0.0.1:8001 --video-input /path/to/video.mp4 --audio-input /path/to/audio.wav --speaker-data-input /path/to/speaker_data.json

Head Movement Speed#

The head_movement_speed parameter controls the expected speed of head movement in the input video:

  • 0: Static or slow-moving head (default).

  • 1: Fast-moving head.

This parameter helps the model optimize lip synchronization based on the dynamics of the subject’s head movement.

To run LipSync with head movement speed via the sample client, use the --head-movement-speed option:

python3 lipsync.py --target 127.0.0.1:8001 --video-input /path/to/video.mp4 --audio-input /path/to/audio.wav --head-movement-speed 1

Output Audio Codec#

The output_audio_codec parameter specifies the audio codec used in the output video file:

  • opus: Opus audio codec (default).

  • mp3: MP3 audio codec.

To run LipSync with a specific output audio codec via the sample client, use the --output-audio-codec option:

python3 lipsync.py --target 127.0.0.1:8001 --video-input /path/to/video.mp4 --audio-input /path/to/audio.wav --output-audio-codec mp3

Background Audio Mixing#

The LipSync NIM supports mixing background audio with the output to preserve ambient sounds. This is useful when the original video contains background music or environmental sounds that should be retained in the output.

The background audio configuration includes the following options:

  • --mix-background-audio: Flag to enable background audio mixing.

  • --background-audio-input: Path to the background audio file (WAV or MP3).

  • --background-audio-volume: Volume level for the background audio (0.0 to 1.0; default: 0.5).

    • 0.0: Background audio is not included.

    • 1.0: Background audio is included at its full, original volume.

    • Values between 0.0 and 1.0: Background volume is reduced proportionally.

Note

The background audio and the speech audio must have the same sample rate.

To run LipSync with background audio mixing via the sample client:

python3 lipsync.py --target 127.0.0.1:8001 --video-input /path/to/video.mp4 --audio-input /path/to/audio.wav --mix-background-audio --background-audio-input /path/to/background.wav --background-audio-volume 0.3

Python Configuration Example#

The following example shows how to set configuration parameters while sending an inference request from a Python client to the LipSync NIM:

import nvidia.ai4m.lipsync.v1.lipsync_pb2 as lipsync_pb2
import nvidia.ai4m.video.v1.video_pb2 as video_pb2
import nvidia.ai4m.audio.v1.audio_pb2 as audio_pb2

params = {
    "input_audio_codec": audio_pb2.AudioCodec.AUDIO_CODEC_WAV,
    "extend_audio": lipsync_pb2.ExtendAudio.EXTEND_AUDIO_SILENCE,
    "extend_video": lipsync_pb2.ExtendVideo.EXTEND_VIDEO_REVERSE,
    "output_video_encoding": video_pb2.VideoEncoding(lossless=True),
    "is_speaker_info_provided": True,
    "output_audio_codec": audio_pb2.AudioCodec.AUDIO_CODEC_OPUS,
    "head_movement_speed": 0,
}

yield lipsync_pb2.LipsyncRequest(config=lipsync_pb2.LipsyncConfig(**params))

Service Information in the Response#

Each response stream opens with a one-shot ServiceInfo message on the service_info field of LipsyncResponse, before any response data. It reports feature_name, feature_version, model_info, server_request_id, and client_session_id. Quote server_request_id when reporting an issue — it locates the request in the server logs.

The following snippet prints the server_request_id from the ServiceInfo banner and writes the video payload to disk:

for response in response_iterator:
    if response.HasField("service_info"):
        print(response.service_info.server_request_id)
    elif response.HasField("video_file_data"):
        output_file.write(response.video_file_data)

Client Session ID#

client_session_id is whatever the client sent in the client-session-id gRPC metadata header, echoed back verbatim. It is empty if the client sends nothing.

The sample client mints a random UUID for each run by default, so the banner and the server logs can always be tied together. Pass --client-session-id to supply your own identifier, for example to group a batch of runs under one job ID:

python3 lipsync.py --target 127.0.0.1:8001 --client-session-id my-session-001

The client prints the following banner when it receives the stream:

ServiceInfo: feature=lipsync version=<version> model=Lipsync server_request_id=<guid> client_session_id=my-session-001

To send no session ID at all, pass an empty string: --client-session-id "".

If you are writing your own client rather than using the sample, set the metadata key directly when invoking the RPC:

metadata = (("client-session-id", "my-session-001"),)
response_iterator = stub.Lipsync(request_iterator, metadata=metadata)

Note

The session ID is opaque to the NIM — it is never interpreted, only logged and echoed back. Do not put sensitive data in it.

Debug Mode#

The LipSync NIM includes a debug mode that provides visual feedback during processing. When enabled, diagnostic overlays are rendered directly onto each output video frame, making it easier to verify effect behavior and troubleshoot issues.

To enable debug mode, set the environment variable NV_AI4M_LS_DEBUG_MODE=1 when launching the NIM container:

docker run -it --rm --name=lipsync-nim \
  --runtime=nvidia \
  --gpus all \
  --shm-size=8GB \
  -e NGC_API_KEY=$NGC_API_KEY \
  -e NV_AI4M_LS_DEBUG_MODE=1 \
  -p 8000:8000 \
  -p 8001:8001 \
  nvcr.io/nim/nvidia/lipsync:latest

When debug mode is enabled, the following overlays appear on each frame:

Overlay

Description

Frame number

Displayed at the top-center of each frame with a white background.

LipSync effect status

A LIPSYNC ON or LIPSYNC OFF indicator below the frame number. The background is green when the effect is active and red when it is bypassed.

Activation bounding box

A square bounding box showing where the LipSync effect is being applied. The box is green when the effect is active (strength > 0) and red when bypassed (strength = 0). Box coordinates and dimensions are labeled above the box.

Speaker bounding boxes

White bounding boxes for each speaker, shown only when speaker data is provided. Each box is labeled with a speaker identifier (e.g., S0, S1) and shows [speaking] when the speaker is actively speaking.

Debug mode is useful for:

  • Verifying that speaker bounding boxes are correctly positioned.

  • Confirming that the LipSync effect is being applied to the intended regions.

  • Troubleshooting cases where the effect appears inactive or misaligned.

Input Stream Handling#

A LipSync request requires up to four input streams on a single gRPC connection — video, speech audio, background audio, and per-frame speaker information — and each input stream is decoded by its own pipeline inside the NIM. The NIM buffers each input stream independently. You can send the entire video before any audio, or interleave the two.

Two families of environment variables bound that buffering:

  • Coverage timeouts limit how long the NIM waits for an input stream before it gives up on the request. Tune these first: they are the intended control for how patient the service is with a slow or stalled client.

  • Buffer caps limit how much of an input stream the NIM holds in memory while the rest of the request catches up. They are an out-of-memory backstop rather than a tuning knob, and most deployments should leave them at their defaults.

The defaults suit a single request per GPU from a client on a reasonably fast link, and most deployments never need to change them. Tune them when you run at high concurrency, when your clients upload over slow or unreliable links, or when you want a stalled client to fail faster than the defaults allow.

Streaming and Transactional Modes#

Which timeouts apply depends on how the NIM ingests the request, which it decides from the first bytes of the input video:

  • Streaming mode is used when the video is a fast-start MP4 (its moov atom is at the front of the file). The NIM decodes the video as it arrives, and the per-stream coverage timeouts below apply.

  • Transactional mode is used for any other input. The NIM receives the complete input before it starts processing, and a single, more generous timeout replaces the per-stream ones.

NV_AI4M_LS_FIRST_BYTES_TIMEOUT_S (default 30) bounds the wait for the bytes that decide the mode. A client that opens a request and then sends no video fails with a DEADLINE_EXCEEDED gRPC error once this elapses. This wait occupies a request slot, so keep it short: at the default concurrency of 1, a single stalled client would otherwise hold the whole service until it disconnects.

Coverage Timeouts#

A coverage timeout is how long the NIM waits for one input stream to supply what it needs to produce one output frame. The budget applies per frame rather than once per request, so a client that keeps sending data never trips it, however long the asset is.

Environment Variable

Default

Description

NV_AI4M_LS_VIDEO_COVERAGE_TIMEOUT_S

300

Wait for the video needed to cover the next output frame. Deliberately generous: it is sized against how fast a client can upload, not how fast the GPU decodes.

NV_AI4M_LS_AUDIO_COVERAGE_TIMEOUT_S

30

Wait for the speech audio needed to cover the next output frame.

NV_AI4M_LS_SPEAKER_INFO_COVERAGE_TIMEOUT_S

30

Wait for the speaker information needed to cover the next output frame. Applies only when you supply speaker data.

NV_AI4M_LS_BACKGROUND_COVERAGE_TIMEOUT_S

30

Wait for the background audio needed to cover the next output frame. Applies only when you enable background audio mixing.

NV_AI4M_LS_UNIT_COVERAGE_TIMEOUT_S

Sum of the four values above

Overall cap on the time spent assembling one output frame across every input stream. Set it to 0 to disable the overall cap.

NV_AI4M_LS_TRANSACTIONAL_COVERAGE_TIMEOUT_S

600

Single per-frame budget used in transactional mode, replacing the four per-stream values. Larger because the input is already on local storage, so the wait covers decoding rather than upload.

When video, speech audio, or speaker information exceeds its budget, the request fails with an INVALID_ARGUMENT gRPC error that names the input stream that stalled. Background audio is the exception: on expiry it drops to silence for the remainder of the request and processing continues, so a missing background track degrades the output instead of failing it.

Note

Apart from 0, which disables the overall cap, NV_AI4M_LS_UNIT_COVERAGE_TIMEOUT_S must never be set lower than the largest per-stream value. Each input stream is capped at whichever deadline expires first, so an overall cap below the video budget silently shrinks the video wait to that cap.

Lower the audio and speaker-information budgets when you want a client that stops mid-upload to fail fast and release its GPU slot. Raise the video budget when clients legitimately go quiet for long stretches, such as a large asset over a thin link. If you raise any per-stream value, raise NV_AI4M_LS_UNIT_COVERAGE_TIMEOUT_S to match, or set it to 0.

Buffer Caps#

Each input stream accumulates in its own buffer whenever its decoder falls behind — most commonly video piling up while the NIM is still waiting for the audio that pairs with it. These caps are an out-of-memory backstop, not flow control. Exceeding one fails the request with a RESOURCE_EXHAUSTED gRPC error rather than blocking, because blocking the request thread would stall every other input stream arriving on the same gRPC connection.

Environment Variable

Default

Description

NV_AI4M_LS_MAX_VIDEO_INPUT_BUFFER_MB

2048

Cap, in MB, on encoded video held while the video decoder is behind. Video gets the largest share because it is orders of magnitude larger per second than audio.

NV_AI4M_LS_MAX_AUDIO_INPUT_BUFFER_MB

512

Cap, in MB, on encoded audio held while the audio decoder is behind. Speech audio and background audio each get this allowance separately.

NV_AI4M_LS_MAX_SPEAKER_INFO_INPUT_BUFFER_FRAMES

1000000

Cap, in SpeakerInfoPerFrame entries, on buffered speaker information. Applies only when you supply speaker data.

NV_AI4M_LS_MAX_VIDEO_QUEUE_SIZE

60

Bound on the queue of decoded video frames waiting for inference. Raising it lets decoding run further ahead of inference, at the cost of memory per request.

Raise a cap only when a legitimate client is rejected with RESOURCE_EXHAUSTED — for example, when uploading a large 4K asset ahead of its audio track exceeds the 2048 MB video allowance.

Note

Exercise caution when raising concurrency together with these buffer caps. Caps apply per request, so total host memory scales with NV_AI4M_MAX_CONCURRENCY_PER_GPU and the number of GPUs. At high concurrency, lower the caps so a stalled reader cannot exhaust host memory, and keep timeout budgets consistent with the reduced buffers. Raising both concurrency and cap values without reviewing memory headroom is not recommended.

Tuning Example#

The following example shortens the coverage timeouts and lowers the buffer caps for a deployment that serves several concurrent clients on a fast local network, where a stalled client should be dropped quickly:

docker run -it --rm --name=lipsync-nim \
  --runtime=nvidia \
  --gpus all \
  --shm-size=8GB \
  -e NGC_API_KEY=$NGC_API_KEY \
  -e NV_AI4M_MAX_CONCURRENCY_PER_GPU=4 \
  -e NV_AI4M_LS_FIRST_BYTES_TIMEOUT_S=10 \
  -e NV_AI4M_LS_VIDEO_COVERAGE_TIMEOUT_S=60 \
  -e NV_AI4M_LS_AUDIO_COVERAGE_TIMEOUT_S=15 \
  -e NV_AI4M_LS_UNIT_COVERAGE_TIMEOUT_S=120 \
  -e NV_AI4M_LS_MAX_VIDEO_INPUT_BUFFER_MB=512 \
  -e NV_AI4M_LS_MAX_AUDIO_INPUT_BUFFER_MB=128 \
  -e NV_AI4M_LS_MAX_VIDEO_QUEUE_SIZE=30 \
  -e NIM_HTTP_API_PORT=8000 \
  -e NIM_GRPC_API_PORT=8001 \
  -p 8000:8000 \
  -p 8001:8001 \
  nvcr.io/nim/nvidia/lipsync:latest

For more information, refer to Runtime Parameters for the Container.

For more information about AI for Media NIM clients, refer to the GitHub repository NVIDIA-Maxine/nim-clients.