> For clean Markdown content of this page, append .md to this URL. For the complete documentation index, see https://docs.nvidia.com/dynamo/latest/llms.txt. For section-specific indexes, append /llms.txt to any section URL.

# Inkling NVFP4

Day-0 aggregated SGLang deployment for `thinkingmachines/Inkling-NVFP4` — the NVFP4-quantized checkpoint of Thinking Machines Lab's first open-weights model, a Mixture-of-Experts that reasons natively over text, image, and audio input with controllable reasoning effort. One TP8 worker on 8x B200 with EAGLE speculative decoding, FA4 attention, and the Inkling reasoning and tool-call parsers wired end to end.

<p>
  Deployment target
</p>

<b>Checkpoint</b> thinkingmachines/Inkling-NVFP4 (\~592 GB)

<b>Precision</b> NVFP4 (modelopt\_fp4)

<b>GPUs</b> 8x B200, one TP8 worker

<b>Techniques</b> Aggregated serving, EAGLE speculative decoding, FA4 attention, Mamba radix cache

<b>Modalities</b> Text + image + audio input, text output

## Prerequisites

* A Kubernetes cluster with the Dynamo platform installed and **8x B200** available on one node — see the [Kubernetes Deployment Guide](/dynamo/dev/kubernetes-deployment/start-here/kubernetes-quickstart).
* An image pull secret with access to the runtime image registry.
* A Hugging Face token (the model is public, but a token avoids rate limits).

Create the namespace and secrets:

```bash
export NAMESPACE=your-namespace
kubectl create namespace ${NAMESPACE}

kubectl create secret docker-registry nvcr-imagepullsecret \
  --docker-server=nvcr.io \
  --docker-username='$oauthtoken' \
  --docker-password="your-ngc-api-key" \
  -n ${NAMESPACE}

kubectl create secret generic hf-token-secret \
  --from-literal=HF_TOKEN="your-token" \
  -n ${NAMESPACE}
```

Update `storageClassName` in `model-cache/model-cache.yaml` to match your cluster (`kubectl get storageclass`) before applying. The checkpoint is \~592 GB — the download job can take a while. On clusters that already provide a shared RWX PVC, skip the cache step and replace the `model-cache` claim name in the manifests with the shared PVC name.

## Deploy

Prepare the model cache and download the checkpoint:

```bash
# Update storageClassName in model-cache.yaml first!
kubectl apply -f recipes/inkling/model-cache/model-cache.yaml -n ${NAMESPACE}
kubectl apply -f recipes/inkling/model-cache/model-download.yaml -n ${NAMESPACE}
kubectl wait --for=condition=Complete job/inkling-model-download -n ${NAMESPACE} --timeout=7200s
```

Then deploy:

```bash
kubectl apply -f recipes/inkling/sglang/agg-b200/deploy.yaml -n ${NAMESPACE}
```

## Smoke Test

Inkling accepts text, image, and audio input in a single OpenAI-compatible API. First, forward the frontend port:

```bash
kubectl port-forward svc/tml-inkling-sglang-agg-frontend 8000:8000 -n ${NAMESPACE} &
```

### Text

```bash
curl -s http://localhost:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "thinkingmachines/Inkling-NVFP4",
    "messages": [{"role": "user", "content": "Hello, who are you?"}],
    "max_tokens": 64
  }'
```

### Image

Images ride as an `image_url` content part — a public HTTP(S) URL (the worker pod fetches it, so it needs egress) or a base64 `data:` URI:

```bash
curl -s http://localhost:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "thinkingmachines/Inkling-NVFP4",
    "messages": [{
      "role": "user",
      "content": [
        {"type": "text", "text": "Describe this image, including its text and visual elements."},
        {"type": "image_url", "image_url": {"url": "https://raw.githubusercontent.com/ai-dynamo/dynamo/main/docs/assets/img/frontpage-banner.png"}}
      ]
    }],
    "max_tokens": 256
  }'
```

### Audio

Audio rides as an `audio_url` content part, same URL-or-data-URI rule. Inkling expects 16 kHz audio; WAV is ideal, and MP3/FLAC/OGG also decode in this image. This sample is 16 kHz mono WAV, matching the model card spec:

```bash
curl -s http://localhost:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "thinkingmachines/Inkling-NVFP4",
    "messages": [{
      "role": "user",
      "content": [
        {"type": "audio_url", "audio_url": {"url": "https://huggingface.co/datasets/Xenova/transformers.js-docs/resolve/main/mlk.wav"}},
        {"type": "text", "text": "Transcribe the following speech to text."}
      ]
    }],
    "max_tokens": 256
  }'
```

For air-gapped clusters, send local files as `data:` URIs instead. Encode to a single line first (`base64` wraps output by default, which breaks the JSON payload): `AUDIO_B64=$(base64 < sample.wav | tr -d '\n')`, then `{"type": "audio_url", "audio_url": {"url": "data:audio/wav;base64,'"${AUDIO_B64}"'"}}`. Multiple media parts can ride in one message as the context budget allows.

### Reasoning effort

Inkling's controllable thinking is exposed per request: pass `reasoning_effort` as a named level (`none` / `minimal` / `low` / `medium` / `high` / `max`) or a float in `[0.0, 0.99]`; omitted requests default to `0.9` (high). These are the values in this checkpoint's chat template — `xhigh` from the launch blog is not in its map and is rejected:

```bash
curl -s http://localhost:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "thinkingmachines/Inkling-NVFP4",
    "messages": [{"role": "user", "content": "What is 17 times 24?"}],
    "chat_template_kwargs": {"reasoning_effort": "low"},
    "max_tokens": 128
  }'
```

## Notes

* This is a Day-0 recipe on a dedicated dev runtime image (`sglang-runtime:1.4.0-inkling-dev.1`, a custom SGLang build carrying the Inkling model support); it is functional but not yet promoted to a release runtime image, and no benchmark results are published yet.
* Audio input: the model card specifies WAV sampled at 16 kHz, ideally under 20 minutes. The runtime image ships without royalty-bearing media codecs — AAC/`.m4a` decode is not included; other common formats (WAV, FLAC, OGG/Opus, MP3) decode via `libsndfile`. Image input is unaffected.
* Video input is not supported in this recipe: the model card notes video handling exists for downstream fine-tuning but is unevaluated out of the box.
* Reasoning and tool calling use the model's dedicated parsers, wired in both the frontend and the worker (`--reasoning-parser inkling --tool-call-parser inkling`).
* EAGLE speculative decoding is enabled by default in the deploy manifest (multi-layer, 8 steps, rejection sampling); remove the `--speculative-*` flags to run without it.
* The manifest sets `DYN_FORWARDPASS_METRIC_PORT` to an empty value — a deliberate workaround that keeps forward-pass metrics off, because this rc image crashes on EAGLE batches when they are enabled. Keep it (and don't set that variable) until a fixed image ships.
* The worker enables the Mamba radix cache (`extra_buffer` strategy) and FlashInfer TRT-LLM GEMM/MoE backends; these flags are part of the validated configuration — change them only deliberately.

## Source

* Source README: [recipes/inkling/README.md](https://github.com/ai-dynamo/dynamo/blob/main/recipes/inkling/README.md)
* SGLang aggregated B200: [deploy.yaml](https://github.com/ai-dynamo/dynamo/blob/main/recipes/inkling/sglang/agg-b200/deploy.yaml)
* Setup assets: [model-cache.yaml](https://github.com/ai-dynamo/dynamo/blob/main/recipes/inkling/model-cache/model-cache.yaml) and [model-download.yaml](https://github.com/ai-dynamo/dynamo/blob/main/recipes/inkling/model-cache/model-download.yaml)
* Model announcement: [Inkling — Thinking Machines Lab](https://thinkingmachines.ai/news/introducing-inkling/)