Motif-3 NVFP4

Serve Motif-3 NVFP4 with Dynamo and the Motif vLLM runtime on NVIDIA B200 GPUs.

以 Markdown 格式查看

This experimental recipe serves Motif-Technologies/Motif-3-NVFP4 on two NVIDIA B200 GPUs with vLLM tensor parallelism, expert parallelism, Dynamo MTP2 speculative decoding, and a 262K-token context. Only an aggregated target is provided; there is no disaggregated target for Motif-3 yet.

Deployment target

Topology
Hardware 2x NVIDIA B200Runtime Motif vLLM image, Dynamo 1.5.0Serving TP2, expert parallelism, Dynamo MTP2Cache Shared model cache, FP8 KV cache

Prerequisites

  • A Kubernetes cluster with the Dynamo platform installed and 2xB200 GPUs available.
  • Create a namespace and an hf-token-secret containing access to the model. The token is used only by the model-download Job.
export NAMESPACE=your-namespace
kubectl create namespace "${NAMESPACE}"
kubectl create secret generic hf-token-secret \
--from-literal=HF_TOKEN="$HF_TOKEN" \
-n "${NAMESPACE}"

Deploy

Edit storageClassName in the model-cache manifest, then create the cache and download the checkpoint:

kubectl apply -f recipes/motif-3/model-cache/model-cache.yaml -n "${NAMESPACE}"
kubectl apply -f recipes/motif-3/model-cache/model-download.yaml -n "${NAMESPACE}"
kubectl wait --for=condition=Complete job/model-download -n "${NAMESPACE}" --timeout=7200s

Apply the aggregated chat manifest:

kubectl apply -f recipes/motif-3/vllm/agg-b200-chat/base/deploy.yaml -n "${NAMESPACE}"
kubectl get dgd motif3-agg-b200 -n "${NAMESPACE}" -w

The base manifest has no cluster-specific scheduling. To add node selectors or tolerations for your cluster, use the Kustomization described in the recipe README.

Smoke Test

First, forward the frontend port for your target:

kubectl port-forward svc/motif3-agg-b200-frontend 8000:8000 -n "${NAMESPACE}"

Send a request:

curl http://localhost:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "Motif-Technologies/Motif-3-NVFP4",
"messages": [{"role": "user", "content": "Write a one-sentence readiness check."}],
"max_tokens": 64,
"temperature": 0
}'

Benchmark

The AIPerf manifest replays the 8k_1k_70kv_chat_new_noschedule_short_15perc.jsonl trace at concurrency 11 against the Dynamo MTP2, TP2 deployment. Stage that trace at /shared-model-cache/traces/8k_1k_70kv_chat_new_noschedule_short_15perc.jsonl, then apply the Job:

kubectl apply -f recipes/motif-3/perf/perf.yaml -n "${NAMESPACE}"
kubectl logs -f job/motif3-aiperf-job -n "${NAMESPACE}"

The deployment uses actual MTP verification by default. For benchmark-only synthetic acceptance-length trials, change the SPECULATIVE_CONFIG ConfigMap key in the deployment manifest from speculative-config to speculative-config-synthetic. These acceptance lengths were calculated by running SPEED-Bench on the coding domain:

MTP speculative tokensAcceptance length
11.75
22.13
32.21

Expected Performance

This is a benchmark-only synthetic proxy, not the expected performance of the shipped manifests. It was measured with speculative-config-synthetic (MTP2, acceptance length 2.13 from SPEED-Bench coding) on the chat trace. The deployment manifest uses real MTP verification, and the AIPerf Job does not select the synthetic configuration, so applying them as shipped will not reproduce these figures. A real-MTP chat-trace result is not yet available.

Under the synthetic proxy, the run meets user output throughput P50 ≥ 50 tokens/second/user and TTFT P50 < 5 seconds. Thirty-five trace requests exceeded the configured 262,144-token context limit.

WorkloadRecipeGPUConcurrencySystem output tok/s/GPUUser output tok/s (P50)TTFT P50 (ms)
Chat trace (synthetic AL 2.13 proxy)Dynamo MTP2, TP2B20011293.9666.27388.18

Notes

  • The recipe uses the public nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.5.0-motif-3-dev.1 image; no image-pull secret is required.
  • The model image carries Motif-specific vLLM compatibility code. Do not replace its vLLM package with an unrelated nightly wheel.

Source