Speculative Decoding#
Speculative decoding speeds up generation in NIM LLM and VLM by having a small, fast draft predict several tokens ahead, which the large target model then verifies in a single forward pass.
When the draft is right, you get multiple tokens for roughly the cost of one; when it is wrong, the target falls back to normal decoding, so output is never degraded. For a conceptual primer, refer to NVIDIA’s blog post An Introduction to Speculative Decoding for Reducing Latency in AI Inference.
Note
This guide documents speculative decoding on the vLLM backend.
Key Terms#
The following terms are used throughout this page:
Target model – the large, accurate model you are serving.
Draft – the small, fast predictor. It can be a separate model (EAGLE3), extra heads built into the target (MTP), or a model-free prompt-lookup histogram (n-gram).
Acceptance rate – the fraction of drafted tokens the target accepts. Higher is better; it is highly dependent on the workload (coding, math, writing, and so on).
Acceptance length – the average number of tokens accepted per step.
Speculative decoding trades memory and compute for latency. It helps most when a GPU has spare compute (for example, low-to-moderate concurrency). Under saturated load with no spare compute, it can fail to improve throughput, and can reduce it. Always measure on your own traffic. Refer to Benchmark it Yourself.
How a NIM Chooses a Method#
Each NIM ships with a default speculative-decoding method, chosen in this order of preference:
MTP (multi-token prediction) – used when the base model includes MTP heads. No extra model to download.
EAGLE3 – used when NVIDIA has published an EAGLE3 draft for the model with a useful acceptance rate. The draft downloads through the same path as the checkpoint. Refer to NVIDIA’s published drafts: Speculative Decoding Modules collection.
n-gram – the model-free fallback when neither MTP nor a published EAGLE3 draft is available. No extra model to download.
Important
The shipped method is a supported, validated option, not necessarily the optimal one for your workload. Speculative-decoding performance is workload-specific, and an EAGLE3 draft tuned to your traffic will usually beat the default. You can override the built-in configuration or bring your own draft model at any time. Refer to Bring Your Own Draft Model. Memory footprint can also drive that choice: speculative decoding always costs extra GPU memory, and the amount differs sharply by method (refer to the note below the table).
The methods differ in what they load and how much extra GPU memory they need. The following table compares extra model and GPU memory cost:
Method |
Extra Model |
Extra GPU Memory |
|---|---|---|
n-gram |
No |
Negligible – model-free (a prompt-lookup histogram). |
MTP |
No |
Low – a few extra prediction heads on the target model. |
EAGLE3 |
Yes (a small draft) |
Higher – a separate draft model with its own weights and KV cache. |
Important
Speculative decoding needs spare GPU memory. Every method drafts and verifies
several tokens per step, so enabling it costs more memory than serving the target
alone, and the extra footprint grows EAGLE3 >> MTP > n-gram (refer to the table above).
Confirm the GPU has headroom before turning it on, especially for EAGLE3, whose
separate draft model dominates the cost. If you hit out-of-memory, reduce the KV-cache
footprint (a shorter NIM_MAX_MODEL_LEN, lower concurrency, or vLLM’s
--gpu-memory-utilization) or use a lighter method. Refer to
GPU memory troubleshooting.
Enable or Disable the Built-In Method#
Speculative decoding is a runtime toggle, not a separate profile. The same profile serves both modes, so you do not need to pick a different profile to turn it on or off.
The profile decides the default. A profile that ships a built-in
speculative-decoding config serves with speculative decoding; a profile without
one serves without it. NIM_SPECDEC_ENABLE is a global override on top of that
default.
The following table describes the effect of each variable:
Variable |
Values |
Effect |
|---|---|---|
|
unset (default) |
The selected profile decides: its built-in speculative-decoding config, if any, is applied. |
|
Force speculative decoding on wherever a config exists (built-in or through |
|
|
Force speculative decoding off for every profile. |
|
|
JSON object |
Override the built-in config (refer to Override the Built-In Configuration). Consumed only when speculative decoding is active; on its own it does not activate it. |
To enable or disable the built-in method, complete the following steps:
Optional: Leave
NIM_SPECDEC_ENABLEunset to use the selected profile’s default.To force speculative decoding off for a profile that ships it on by default, set
NIM_SPECDEC_ENABLE=0.export IMG=nvcr.io/nim/<org>/<nim>:<tag> # your NIM image export LOCAL_NIM_CACHE=~/.cache/nim mkdir -p "$LOCAL_NIM_CACHE" docker run --gpus all --shm-size=16GB \ -e NGC_API_KEY \ -e NIM_SPECDEC_ENABLE=0 \ -v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \ -p 8000:8000 \ $IMG
Set
NIM_SPECDEC_ENABLE=1to force the built-in method on regardless of the profile’s default.A profile that ships an EAGLE3 draft stages the draft weights alongside the checkpoint, so they are cached with the model and present whether or not speculative decoding is active (the toggle only gates whether the draft is used). A draft supplied only through
NIM_DRAFT_MODEL_PATHdownloads at launch, and only when speculative decoding is active.
Note
A profile that serves speculative decoding by default ships the EAGLE3 draft
weights as part of the model: download-to-cache caches them alongside the
checkpoint and the draft is materialized locally at launch, so it is air-gap ready
with no startup fetch. A draft supplied only through NIM_DRAFT_MODEL_PATH is fetched
at startup instead; for air-gapped deployments pre-cache it alongside the model (or
point NIM_DRAFT_MODEL_PATH at a local directory). Set NIM_SPECDEC_ENABLE=0 to
serve without speculative decoding.
Override the Built-In Configuration#
NIM_SPECDEC_ARGS is a JSON object of vLLM runtime-config keys that overrides the
model’s default speculative config. Use it to tune parameters or switch methods (for
example, force n-gram on a model that ships MTP).
The following examples reuse IMG and LOCAL_NIM_CACHE from
Enable or Disable the Built-In Method.
If you run them on their own, export those values first.
To override the built-in configuration, complete the following steps:
Set
NIM_SPECDEC_ENABLE=1so speculative decoding is active.Pass the override JSON in
NIM_SPECDEC_ARGS.docker run --gpus all --shm-size=16GB \ -e NGC_API_KEY \ -e NIM_SPECDEC_ENABLE=1 \ -e NIM_SPECDEC_ARGS='{"speculative_config": "{\"method\": \"ngram\", \"num_speculative_tokens\": 3, \"prompt_lookup_max\": 4, \"prompt_lookup_min\": 1}"}' \ -v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \ -p 8000:8000 \ $IMG
Configure Speculative Decoding for Model-Free Deployments#
In model-free mode (NIM_MODEL_PATH) the NIM has no built-in speculative config, so you
supply one yourself. The simplest path is to pass the vLLM speculative config directly as
a backend argument.
Use n-gram (No Draft Model)#
The following command enables n-gram speculative decoding without a draft model:
docker run --gpus all --shm-size=16GB \
-e NGC_API_KEY \
-v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
-p 8000:8000 \
-e NIM_MODEL_PATH=hf://meta-llama/Llama-3.1-8B-Instruct \
$IMG \
--speculative-config '{"method": "ngram", "num_speculative_tokens": 3, "prompt_lookup_max": 4, "prompt_lookup_min": 1}'
Use MTP (Model with Built-In MTP Heads)#
The following command enables MTP when the model includes MTP heads:
docker run --gpus all --shm-size=16GB \
-e NGC_API_KEY \
-v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
-p 8000:8000 \
-e NIM_MODEL_PATH=ngc://<org>/<model>:<tag> \
$IMG \
--speculative-config '{"method": "mtp", "num_speculative_tokens": 1}'
Bring Your Own Draft Model#
EAGLE3 (and other external-draft methods) need a separate draft model. Point
NIM_DRAFT_MODEL_PATH at it and reference the draft by a bare name in the
speculative config; the NIM downloads the draft into that subdirectory and rewrites the
reference to the local path before launch.
NIM_DRAFT_MODEL_PATH mirrors NIM_MODEL_PATH: it downloads from the same locations
(ngc://, hf://, s3://, gs://, or an absolute local path), is cached, and works in
air-gapped deployments.
To bring your own draft model, complete the following steps:
Set
NIM_MODEL_PATHto the target model andNIM_DRAFT_MODEL_PATHto the draft source.Pass a speculative config that names the draft by the bare subdirectory (
"model": "eagle3_draft"in this example).docker run --gpus all --shm-size=16GB \ -e NGC_API_KEY \ -v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \ -p 8000:8000 \ -e NIM_MODEL_PATH=hf://meta-llama/Llama-3.1-8B-Instruct \ -e NIM_DRAFT_MODEL_PATH=hf://<org>/<your-eagle3-draft> \ $IMG \ --speculative-config '{"method": "eagle3", "model": "eagle3_draft", "num_speculative_tokens": 3}'
Here
"model": "eagle3_draft"is the bare subdirectory name the draft is materialized into; it does not need to match the remote repository name. To override the draft on a model-specific NIM instead, setNIM_DRAFT_MODEL_PATHwhile speculative decoding is active (the profile’s default, or forced withNIM_SPECDEC_ENABLE=1). An explicit draft source takes precedence over the one shipped with the NIM.
Create Your Own Draft Model#
The best draft for your workload is usually an EAGLE3 model trained on it. NVIDIA
publishes ready-to-use EAGLE3 drafts for common models in the
Speculative Decoding Modules
collection, or you can train your own. Refer to the
guide to training a speculative-decoding model.
Export the trained draft and serve it through NIM_DRAFT_MODEL_PATH as shown above.
Use a Local Draft Model (Air-Gapped)#
For air-gapped deployments, point NIM_DRAFT_MODEL_PATH at an absolute local
directory holding the draft (its config.json and weights). The draft is linked into
the workspace with no network access, the same way a local target model is served.
Mount the directory into the container and use its in-container path:
docker run --gpus all --shm-size=16GB \
-v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
-v /data/models/llama-3.1-8b:/models/target \
-v /data/models/eagle3-draft:/models/draft \
-p 8000:8000 \
-e NIM_MODEL_PATH=/models/target \
-e NIM_DRAFT_MODEL_PATH=/models/draft \
$IMG \
--speculative-config '{"method": "eagle3", "model": "eagle3_draft", "num_speculative_tokens": 3}'
No NGC_API_KEY or HF_TOKEN is required: the target loads from its local path (or a
pre-populated /opt/nim/.cache) and the draft links from /models/draft. Keep the bare
"model": "eagle3_draft" reference; NIM materializes the local draft into that
workspace subdirectory and rewrites the reference to the absolute path at launch.
Benchmark it Yourself#
Because speculative-decoding gains are workload-specific, measure on representative
traffic before committing. NVIDIA AIPerf collects
acceptance rate, acceptance length, and throughput directly from the server’s metrics
endpoint. NIM exposes Prometheus metrics at /v1/metrics, so pass
--server-metrics http://localhost:8000/v1/metrics.
Benchmark with SPEED-Bench (Per-Category Acceptance)#
SPEED-Bench is NVIDIA’s dataset for evaluating speculative decoding across 11 semantic domains (coding, math, writing, and so on). Prepare the data, then profile per category and assemble a matrix report:
SPEED_BENCH_DIR="./datasets/speed-bench"
curl -LsSf https://raw.githubusercontent.com/NVIDIA-NeMo/Skills/refs/heads/main/nemo_skills/dataset/speed-bench/prepare.py \
| python3 - --output_dir $SPEED_BENCH_DIR
CATEGORIES="coding humanities math multilingual qa rag reasoning roleplay stem summarization writing"
for cat in $CATEGORIES; do
aiperf profile \
--model <served-model-name> \
--endpoint-type chat --streaming \
--url localhost:8000 \
--custom-dataset-type speed_bench_${cat} \
--input-file ${SPEED_BENCH_DIR}/qualitative.jsonl \
--server-metrics http://localhost:8000/v1/metrics \
--osl 4096 --extra-inputs temperature:0 --concurrency 16 \
--output-artifact-dir ./artifacts/speed_bench_${cat}
done
aiperf speed-bench-report ./artifacts/ --format both # acceptance length matrix
aiperf speed-bench-report ./artifacts/ --metric accept_rate # acceptance rate matrix
aiperf speed-bench-report ./artifacts/ --metric throughput # throughput matrix
To quantify speedup, run the same matrix with NIM_SPECDEC_ENABLE=0 (acceptance is then
0) and compare the throughput tables.
Replay Your Own Traffic#
To benchmark against your production traffic rather than a public dataset, capture it with NIM’s built-in payload capture and replay it with AIPerf.
To replay captured traffic, complete the following steps:
Capture real requests by enabling payload capture (off by default; it records exact prompts; refer to the payload capture reference). With
format: mooncake_payload, NIM writes a Mooncake payload-mode trace ({"timestamp": <ms>, "payload": <request>}per line) of the actual requests:docker run --gpus all --shm-size=16GB \ -e NGC_API_KEY \ -e NIM_SPECDEC_ENABLE=1 \ -e NIM_CAPTURE_ENABLE=1 \ -e NIM_CAPTURE_ARGS='{"path": "/captures/traffic.jsonl", "format": "mooncake_payload", "endpoint_pattern": "^/v1/chat/completions$"}' \ -v /data/captures:/captures \ -p 8000:8000 \ $IMG
The
endpoint_patternscopes capture to one endpoint. AIPerf replays each payload verbatim to the single--endpoint-typeyou pass below, so a trace must contain one endpoint’s payloads; refer to the payload capture reference for replaying completion traffic.Replay the captured trace with AIPerf, collecting server-side spec metrics:
aiperf profile \ --model <served-model-name> \ --endpoint-type chat --streaming \ --url localhost:8000 \ --custom-dataset-type mooncake_trace \ --input-file /data/captures/traffic.jsonl \ --server-metrics http://localhost:8000/v1/metrics
AIPerf loads a tokenizer for the model to compute token metrics; if it cannot be fetched (private model name or air-gapped host), pass a local tokenizer directory with
--tokenizer <dir>(or--tokenizer builtin).Read the results. AIPerf’s console report covers latency and throughput (time to first token, request latency, output token throughput); it does not report speculative-decoding acceptance. Those come from the server-side counters that
--server-metricsrecords (written toserver_metrics_export.csv/.json, and also served live on/v1/metrics). On vLLM:The following table describes the vLLM speculative-decoding counters:
Counter
Meaning
vllm:spec_decode_num_draft_tokens_totaldraft tokens proposed
vllm:spec_decode_num_accepted_tokens_totaldraft tokens accepted
vllm:spec_decode_num_drafts_totaldraft steps
Acceptance rate (AR) = accepted / draft tokens; acceptance length (AL) = 1 + accepted / drafts (mean tokens emitted per step). SGLang exposes equivalent acceptance counters under its own metric names. Re-run with
NIM_SPECDEC_ENABLE=0to compare throughput and latency on the same traffic.
Because capture stores the real payloads (not hashed or synthetic prompts), the replay
reflects your true workload. For raw_payload capture use
--custom-dataset-type raw_payload instead. Refer to the
payload capture reference and the AIPerf
SPEED-Bench tutorial.
Best Practices#
Use the following practices when you enable speculative decoding:
Measure acceptance and throughput on representative traffic before you commit to a method. Refer to Benchmark it Yourself.
Confirm GPU memory headroom before you enable speculative decoding, especially for EAGLE3.
Prefer an EAGLE3 draft trained on your workload when the shipped default is not optimal.
For air-gapped deployments, pre-cache profile-shipped drafts with
download-to-cache, or pointNIM_DRAFT_MODEL_PATHat a local directory.Scope payload capture to one endpoint when you plan a single AIPerf replay run.
Next Steps#
After you enable a speculative-decoding method, capture production traffic and replay it to measure speedup.
For more information, refer to Payload Capture and the
Environment Variables reference for NIM_SPECDEC_ENABLE, NIM_SPECDEC_ARGS, and NIM_DRAFT_MODEL_PATH.