Speculative Decoding#

Speculative decoding speeds up generation in NIM LLM and VLM by having a small, fast draft predict several tokens ahead, which the large target model then verifies in a single forward pass.

When the draft is right, you get multiple tokens for roughly the cost of one; when it is wrong, the target falls back to normal decoding, so output is never degraded. For a conceptual primer, refer to NVIDIA’s blog post An Introduction to Speculative Decoding for Reducing Latency in AI Inference.

Note

This guide documents speculative decoding on the vLLM backend.

Key Terms#

The following terms are used throughout this page:

  • Target model – the large, accurate model you are serving.

  • Draft – the small, fast predictor. It can be a separate model (EAGLE3), extra heads built into the target (MTP), or a model-free prompt-lookup histogram (n-gram).

  • Acceptance rate – the fraction of drafted tokens the target accepts. Higher is better; it is highly dependent on the workload (coding, math, writing, and so on).

  • Acceptance length – the average number of tokens accepted per step.

Speculative decoding trades memory and compute for latency. It helps most when a GPU has spare compute (for example, low-to-moderate concurrency). Under saturated load with no spare compute, it can fail to improve throughput, and can reduce it. Always measure on your own traffic. Refer to Benchmark it Yourself.

How a NIM Chooses a Method#

Each NIM ships with a default speculative-decoding method, chosen in this order of preference:

  1. MTP (multi-token prediction) – used when the base model includes MTP heads. No extra model to download.

  2. EAGLE3 – used when NVIDIA has published an EAGLE3 draft for the model with a useful acceptance rate. The draft downloads through the same path as the checkpoint. Refer to NVIDIA’s published drafts: Speculative Decoding Modules collection.

  3. n-gram – the model-free fallback when neither MTP nor a published EAGLE3 draft is available. No extra model to download.

Important

The shipped method is a supported, validated option, not necessarily the optimal one for your workload. Speculative-decoding performance is workload-specific, and an EAGLE3 draft tuned to your traffic will usually beat the default. You can override the built-in configuration or bring your own draft model at any time. Refer to Bring Your Own Draft Model. Memory footprint can also drive that choice: speculative decoding always costs extra GPU memory, and the amount differs sharply by method (refer to the note below the table).

The methods differ in what they load and how much extra GPU memory they need. The following table compares extra model and GPU memory cost:

Method

Extra Model

Extra GPU Memory

n-gram

No

Negligible – model-free (a prompt-lookup histogram).

MTP

No

Low – a few extra prediction heads on the target model.

EAGLE3

Yes (a small draft)

Higher – a separate draft model with its own weights and KV cache.

Important

Speculative decoding needs spare GPU memory. Every method drafts and verifies several tokens per step, so enabling it costs more memory than serving the target alone, and the extra footprint grows EAGLE3 >> MTP > n-gram (refer to the table above). Confirm the GPU has headroom before turning it on, especially for EAGLE3, whose separate draft model dominates the cost. If you hit out-of-memory, reduce the KV-cache footprint (a shorter NIM_MAX_MODEL_LEN, lower concurrency, or vLLM’s --gpu-memory-utilization) or use a lighter method. Refer to GPU memory troubleshooting.

Enable or Disable the Built-In Method#

Speculative decoding is a runtime toggle, not a separate profile. The same profile serves both modes, so you do not need to pick a different profile to turn it on or off.

The profile decides the default. A profile that ships a built-in speculative-decoding config serves with speculative decoding; a profile without one serves without it. NIM_SPECDEC_ENABLE is a global override on top of that default.

The following table describes the effect of each variable:

Variable

Values

Effect

NIM_SPECDEC_ENABLE

unset (default)

The selected profile decides: its built-in speculative-decoding config, if any, is applied.

1

Force speculative decoding on wherever a config exists (built-in or through NIM_SPECDEC_ARGS). A profile with no config logs a warning and serves without it; it never fails to start.

0

Force speculative decoding off for every profile.

NIM_SPECDEC_ARGS

JSON object

Override the built-in config (refer to Override the Built-In Configuration). Consumed only when speculative decoding is active; on its own it does not activate it.

To enable or disable the built-in method, complete the following steps:

  1. Optional: Leave NIM_SPECDEC_ENABLE unset to use the selected profile’s default.

  2. To force speculative decoding off for a profile that ships it on by default, set NIM_SPECDEC_ENABLE=0.

    export IMG=nvcr.io/nim/<org>/<nim>:<tag>   # your NIM image
    export LOCAL_NIM_CACHE=~/.cache/nim
    mkdir -p "$LOCAL_NIM_CACHE"
    docker run --gpus all --shm-size=16GB \
      -e NGC_API_KEY \
      -e NIM_SPECDEC_ENABLE=0 \
      -v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
      -p 8000:8000 \
      $IMG
    
  3. Set NIM_SPECDEC_ENABLE=1 to force the built-in method on regardless of the profile’s default.

    A profile that ships an EAGLE3 draft stages the draft weights alongside the checkpoint, so they are cached with the model and present whether or not speculative decoding is active (the toggle only gates whether the draft is used). A draft supplied only through NIM_DRAFT_MODEL_PATH downloads at launch, and only when speculative decoding is active.

Note

A profile that serves speculative decoding by default ships the EAGLE3 draft weights as part of the model: download-to-cache caches them alongside the checkpoint and the draft is materialized locally at launch, so it is air-gap ready with no startup fetch. A draft supplied only through NIM_DRAFT_MODEL_PATH is fetched at startup instead; for air-gapped deployments pre-cache it alongside the model (or point NIM_DRAFT_MODEL_PATH at a local directory). Set NIM_SPECDEC_ENABLE=0 to serve without speculative decoding.

Override the Built-In Configuration#

NIM_SPECDEC_ARGS is a JSON object of vLLM runtime-config keys that overrides the model’s default speculative config. Use it to tune parameters or switch methods (for example, force n-gram on a model that ships MTP).

The following examples reuse IMG and LOCAL_NIM_CACHE from Enable or Disable the Built-In Method. If you run them on their own, export those values first.

To override the built-in configuration, complete the following steps:

  1. Set NIM_SPECDEC_ENABLE=1 so speculative decoding is active.

  2. Pass the override JSON in NIM_SPECDEC_ARGS.

    docker run --gpus all --shm-size=16GB \
      -e NGC_API_KEY \
      -e NIM_SPECDEC_ENABLE=1 \
      -e NIM_SPECDEC_ARGS='{"speculative_config": "{\"method\": \"ngram\", \"num_speculative_tokens\": 3, \"prompt_lookup_max\": 4, \"prompt_lookup_min\": 1}"}' \
      -v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
      -p 8000:8000 \
      $IMG
    

Configure Speculative Decoding for Model-Free Deployments#

In model-free mode (NIM_MODEL_PATH) the NIM has no built-in speculative config, so you supply one yourself. The simplest path is to pass the vLLM speculative config directly as a backend argument.

Use n-gram (No Draft Model)#

The following command enables n-gram speculative decoding without a draft model:

docker run --gpus all --shm-size=16GB \
  -e NGC_API_KEY \
  -v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
  -p 8000:8000 \
  -e NIM_MODEL_PATH=hf://meta-llama/Llama-3.1-8B-Instruct \
  $IMG \
  --speculative-config '{"method": "ngram", "num_speculative_tokens": 3, "prompt_lookup_max": 4, "prompt_lookup_min": 1}'

Use MTP (Model with Built-In MTP Heads)#

The following command enables MTP when the model includes MTP heads:

docker run --gpus all --shm-size=16GB \
  -e NGC_API_KEY \
  -v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
  -p 8000:8000 \
  -e NIM_MODEL_PATH=ngc://<org>/<model>:<tag> \
  $IMG \
  --speculative-config '{"method": "mtp", "num_speculative_tokens": 1}'

Bring Your Own Draft Model#

EAGLE3 (and other external-draft methods) need a separate draft model. Point NIM_DRAFT_MODEL_PATH at it and reference the draft by a bare name in the speculative config; the NIM downloads the draft into that subdirectory and rewrites the reference to the local path before launch.

NIM_DRAFT_MODEL_PATH mirrors NIM_MODEL_PATH: it downloads from the same locations (ngc://, hf://, s3://, gs://, or an absolute local path), is cached, and works in air-gapped deployments.

To bring your own draft model, complete the following steps:

  1. Set NIM_MODEL_PATH to the target model and NIM_DRAFT_MODEL_PATH to the draft source.

  2. Pass a speculative config that names the draft by the bare subdirectory ("model": "eagle3_draft" in this example).

    docker run --gpus all --shm-size=16GB \
      -e NGC_API_KEY \
      -v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
      -p 8000:8000 \
      -e NIM_MODEL_PATH=hf://meta-llama/Llama-3.1-8B-Instruct \
      -e NIM_DRAFT_MODEL_PATH=hf://<org>/<your-eagle3-draft> \
      $IMG \
      --speculative-config '{"method": "eagle3", "model": "eagle3_draft", "num_speculative_tokens": 3}'
    

    Here "model": "eagle3_draft" is the bare subdirectory name the draft is materialized into; it does not need to match the remote repository name. To override the draft on a model-specific NIM instead, set NIM_DRAFT_MODEL_PATH while speculative decoding is active (the profile’s default, or forced with NIM_SPECDEC_ENABLE=1). An explicit draft source takes precedence over the one shipped with the NIM.

Create Your Own Draft Model#

The best draft for your workload is usually an EAGLE3 model trained on it. NVIDIA publishes ready-to-use EAGLE3 drafts for common models in the Speculative Decoding Modules collection, or you can train your own. Refer to the guide to training a speculative-decoding model. Export the trained draft and serve it through NIM_DRAFT_MODEL_PATH as shown above.

Use a Local Draft Model (Air-Gapped)#

For air-gapped deployments, point NIM_DRAFT_MODEL_PATH at an absolute local directory holding the draft (its config.json and weights). The draft is linked into the workspace with no network access, the same way a local target model is served. Mount the directory into the container and use its in-container path:

docker run --gpus all --shm-size=16GB \
  -v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
  -v /data/models/llama-3.1-8b:/models/target \
  -v /data/models/eagle3-draft:/models/draft \
  -p 8000:8000 \
  -e NIM_MODEL_PATH=/models/target \
  -e NIM_DRAFT_MODEL_PATH=/models/draft \
  $IMG \
  --speculative-config '{"method": "eagle3", "model": "eagle3_draft", "num_speculative_tokens": 3}'

No NGC_API_KEY or HF_TOKEN is required: the target loads from its local path (or a pre-populated /opt/nim/.cache) and the draft links from /models/draft. Keep the bare "model": "eagle3_draft" reference; NIM materializes the local draft into that workspace subdirectory and rewrites the reference to the absolute path at launch.

Benchmark it Yourself#

Because speculative-decoding gains are workload-specific, measure on representative traffic before committing. NVIDIA AIPerf collects acceptance rate, acceptance length, and throughput directly from the server’s metrics endpoint. NIM exposes Prometheus metrics at /v1/metrics, so pass --server-metrics http://localhost:8000/v1/metrics.

Benchmark with SPEED-Bench (Per-Category Acceptance)#

SPEED-Bench is NVIDIA’s dataset for evaluating speculative decoding across 11 semantic domains (coding, math, writing, and so on). Prepare the data, then profile per category and assemble a matrix report:

SPEED_BENCH_DIR="./datasets/speed-bench"
curl -LsSf https://raw.githubusercontent.com/NVIDIA-NeMo/Skills/refs/heads/main/nemo_skills/dataset/speed-bench/prepare.py \
  | python3 - --output_dir $SPEED_BENCH_DIR

CATEGORIES="coding humanities math multilingual qa rag reasoning roleplay stem summarization writing"
for cat in $CATEGORIES; do
  aiperf profile \
    --model <served-model-name> \
    --endpoint-type chat --streaming \
    --url localhost:8000 \
    --custom-dataset-type speed_bench_${cat} \
    --input-file ${SPEED_BENCH_DIR}/qualitative.jsonl \
    --server-metrics http://localhost:8000/v1/metrics \
    --osl 4096 --extra-inputs temperature:0 --concurrency 16 \
    --output-artifact-dir ./artifacts/speed_bench_${cat}
done

aiperf speed-bench-report ./artifacts/ --format both          # acceptance length matrix
aiperf speed-bench-report ./artifacts/ --metric accept_rate   # acceptance rate matrix
aiperf speed-bench-report ./artifacts/ --metric throughput    # throughput matrix

To quantify speedup, run the same matrix with NIM_SPECDEC_ENABLE=0 (acceptance is then 0) and compare the throughput tables.

Replay Your Own Traffic#

To benchmark against your production traffic rather than a public dataset, capture it with NIM’s built-in payload capture and replay it with AIPerf.

To replay captured traffic, complete the following steps:

  1. Capture real requests by enabling payload capture (off by default; it records exact prompts; refer to the payload capture reference). With format: mooncake_payload, NIM writes a Mooncake payload-mode trace ({"timestamp": <ms>, "payload": <request>} per line) of the actual requests:

    docker run --gpus all --shm-size=16GB \
      -e NGC_API_KEY \
      -e NIM_SPECDEC_ENABLE=1 \
      -e NIM_CAPTURE_ENABLE=1 \
      -e NIM_CAPTURE_ARGS='{"path": "/captures/traffic.jsonl", "format": "mooncake_payload", "endpoint_pattern": "^/v1/chat/completions$"}' \
      -v /data/captures:/captures \
      -p 8000:8000 \
      $IMG
    

    The endpoint_pattern scopes capture to one endpoint. AIPerf replays each payload verbatim to the single --endpoint-type you pass below, so a trace must contain one endpoint’s payloads; refer to the payload capture reference for replaying completion traffic.

  2. Replay the captured trace with AIPerf, collecting server-side spec metrics:

    aiperf profile \
      --model <served-model-name> \
      --endpoint-type chat --streaming \
      --url localhost:8000 \
      --custom-dataset-type mooncake_trace \
      --input-file /data/captures/traffic.jsonl \
      --server-metrics http://localhost:8000/v1/metrics
    

    AIPerf loads a tokenizer for the model to compute token metrics; if it cannot be fetched (private model name or air-gapped host), pass a local tokenizer directory with --tokenizer <dir> (or --tokenizer builtin).

  3. Read the results. AIPerf’s console report covers latency and throughput (time to first token, request latency, output token throughput); it does not report speculative-decoding acceptance. Those come from the server-side counters that --server-metrics records (written to server_metrics_export.csv/.json, and also served live on /v1/metrics). On vLLM:

    The following table describes the vLLM speculative-decoding counters:

    Counter

    Meaning

    vllm:spec_decode_num_draft_tokens_total

    draft tokens proposed

    vllm:spec_decode_num_accepted_tokens_total

    draft tokens accepted

    vllm:spec_decode_num_drafts_total

    draft steps

    Acceptance rate (AR) = accepted / draft tokens; acceptance length (AL) = 1 + accepted / drafts (mean tokens emitted per step). SGLang exposes equivalent acceptance counters under its own metric names. Re-run with NIM_SPECDEC_ENABLE=0 to compare throughput and latency on the same traffic.

Because capture stores the real payloads (not hashed or synthetic prompts), the replay reflects your true workload. For raw_payload capture use --custom-dataset-type raw_payload instead. Refer to the payload capture reference and the AIPerf SPEED-Bench tutorial.

Best Practices#

Use the following practices when you enable speculative decoding:

  • Measure acceptance and throughput on representative traffic before you commit to a method. Refer to Benchmark it Yourself.

  • Confirm GPU memory headroom before you enable speculative decoding, especially for EAGLE3.

  • Prefer an EAGLE3 draft trained on your workload when the shipped default is not optimal.

  • For air-gapped deployments, pre-cache profile-shipped drafts with download-to-cache, or point NIM_DRAFT_MODEL_PATH at a local directory.

  • Scope payload capture to one endpoint when you plan a single AIPerf replay run.

Next Steps#

After you enable a speculative-decoding method, capture production traffic and replay it to measure speedup. For more information, refer to Payload Capture and the Environment Variables reference for NIM_SPECDEC_ENABLE, NIM_SPECDEC_ARGS, and NIM_DRAFT_MODEL_PATH.