Reasoning Mode

View as Markdown

Reasoning (thinking) mode is off by default in NeMo Labs Voice Agent. A reasoning model emits a reasoning block before its spoken answer. In a voice pipeline, no content reaches text-to-speech (TTS) until that block closes. default.yaml therefore ships llm.enable_reasoning: false to minimize latency, and the default large language model (LLM) sub-configuration (server_configs/llm_configs/nemotron_nano_v3.yaml) sends enable_thinking: False to vLLM.

Enable reasoning when answer quality on multi-step or tool-heavy tasks is more important than time to first audio.

Reasoning Components

The following settings and configuration files control whether reasoning runs and whether it reaches audio output.

PieceLocationFunction
llm.enable_reasoningserver_configs/default.yaml (and default_nvidia.yaml)Boolean switch. Drives the _think.yaml swap and, in the NVIDIA configuration, is interpolated into the request body.
*_think.yaml sub-configurationserver_configs/llm_configs/Sibling of each LLM configuration that changes the model’s thinking flag and raises the token budget.
tts.think_tokensserver_configs/tts_configs/*.yamlDelimiter pair the local TTS services use to suppress the reasoning span so the user never hears it.
--reasoning-parserinside llm.vllm_server_paramsMakes vLLM split reasoning into a separate response field, so it never enters the text stream at all.

Enable Reasoning

llm.enable_reasoning: true on its own does not always change the model’s behavior. The swap in nemo_voice_agent/utils/config_manager.py (_configure_llm) occurs only when all three conditions hold:

  1. server.use_model_registry: true is set.
  2. llm.model_config is not set. An explicit model_config short-circuits the registry lookup and marks the configuration as non-registry.
  3. The model entry in server/model_registry.yaml has reasoning_supported: true.

Only then is the resolved path rewritten from <name>.yaml to <name>_think.yaml. Today Qwen/Qwen3-8B is the sole registry entry with reasoning_supported: true.

Because default.yaml pins llm.model_config explicitly, the swap does not occur for the shipped default model. Point model_config at the think variant explicitly.

Route 1 — Select the Think Configuration for the Shipped Default

This route uses model_config to select a _think.yaml file explicitly.

Edit examples/generic_voice_agent/server/server_configs/default.yaml:

1llm:
2 model: "nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4"
3 model_config: "./server_configs/llm_configs/nemotron_nano_v3_think.yaml"
4 enable_reasoning: true # documentary here; the think config is what actually flips the model

The sub-YAML overrides default.yaml, so settings in nemotron_nano_v3_think.yaml, including type: vllm and max_new_tokens, take precedence. For details, refer to Server Configuration. To avoid editing the shipped file, copy it and select the copy at launch:

$cd examples/generic_voice_agent/server
$SERVER_CONFIG_PATH=./server_configs/my_think.yaml python server.py

Route 2 — Use a Registry-Driven Swap

Remove llm.model_config and let the registry resolve the file:

1server:
2 use_model_registry: true
3llm:
4 model: "Qwen/Qwen3-8B"
5 enable_reasoning: true # resolves qwen3-8B.yaml -> qwen3-8B_think.yaml

To use this route for a model you added, give it a reasoning_supported: true entry and a _think.yaml sibling. For details, refer to Model Registry.

Route 3 — Interpolate the Flag for Hosted NVIDIA NIM Endpoints

default_nvidia.yaml wires the switch directly into the request body with OmegaConf interpolation, so no file swap is needed — flipping llm.enable_reasoning is enough:

1llm:
2 type: nvidia
3 enable_reasoning: false
4 nvidia_generation_params:
5 extra:
6 extra_body:
7 chat_template_kwargs:
8 enable_thinking: ${llm.enable_reasoning}
9 thinking_token_budget: 3000

For endpoint configuration, refer to NVIDIA NIM Endpoints.

How Think Configuration Variants Work

Each _think.yaml variant changes only the settings listed below.

Comparing each pair shows the complete difference. The rest of each configuration is identical.

PairDifference
nemotron_nano_v3.yaml to nemotron_nano_v3_think.yamlChanges enable_thinking to True, raises max_new_tokens from 1024 to 4096, and adds thinking_budget: 2048, passed on as thinking_token_budget.
nemotron_nano_v3_omni.yaml to nemotron_nano_v3_omni_think.yamlApplies the same three changes and moves sampling from near-greedy (temperature: 0.2, top_k: 1) to temperature: 0.6 and top_p: 0.95.
qwen3-8B.yaml to qwen3-8B_think.yamlChanges system_prompt_suffix from /no_think to /think and drops the extra_body block that forced thinking off.

Raising max_new_tokens matters: the reasoning block and the spoken answer share one completion budget, so a think configuration left at 1024 tokens can truncate mid-thought and produce no audio.

Keep Reasoning Out of the Audio

tts.think_tokens is a two-element list of delimiters. The local NeMo TTS services safely strip the delimited content from a stream before synthesis. Text before the opening tag is spoken, chunks inside the block are dropped, and speech resumes after the closing tag. The logic lives in _handle_think_tokens in nemo_voice_agent/pipecat/services/nemo/tts.py.

All three shipped TTS sub-configurations (kokoro_82M.yaml, nemo_fastpitch-hifigan.yaml, magpie_tts_multilingual_357m.yaml) already set it:

1think_tokens: ["<think>", "</think>"]

Set it to null if you want the model to think out loud. This setting is useful for debugging but unsuitable for production use.

The value must be a list of exactly two strings (asserted at construction), and only the local NeMo TTS services honor it. tts.type: nvidia (Riva or NVIDIA Cloud Functions (NVCF) Magpie) is built without the think_tokens argument, so with hosted TTS, rely on the vLLM-side parser. For details, refer to Text-to-Speech.

Filter Reasoning with vLLM

The preferred option is to let vLLM separate the reasoning. Add --reasoning-parser to llm.vllm_server_params. vLLM then routes the thinking block to a separate response field that the pipeline never reads, so no <think> delimiter ever reaches TTS and think_tokens becomes a secondary safeguard.

The shipped configurations use the following reasoning parsers:

ConfigurationReasoning Parser
llm_configs/nemotron_nano_v3*.yaml (including omni and think)nemotron_v3 (vLLM built-in)
evaluation/server_configs/agent.yaml, user.yamldeepseek_r1
All other shipped LLM configurationsNone

Because nemotron_nano_v3.yaml sets start_vllm_on_init: false, you launch vLLM yourself with the same flags the configuration expects:

$vllm serve nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 \
> --trust-remote-code --tensor-parallel-size 1 --enable-prefix-caching \
> --max-num-seqs 1 --gpu-memory-utilization 0.8 \
> --enable-auto-tool-choice --tool-call-parser qwen3_coder \
> --reasoning-parser nemotron_v3

For details, refer to Serving with vLLM.

Bound Reasoning Time

Two independent mechanisms bound reasoning time:

  • llm.thinking_budget — used by the think configurations, forwarded to the server as thinking_token_budget inside vllm_generation_params.extra.extra_body. Models or servers that implement that field, including Nemotron-3 Nano and hosted NIM, honor the setting.
  • ReasoningBudgetLogitsProcessor — a vLLM plugin shipped in this repository (nemo_voice_agent/vllm/v1/sample/logits_processor/) that counts tokens inside the thinking block and forces the closing sequence when the budget is hit. It is loaded with --logits-processors and driven for each request through vllm_xargs, so it works for models that ignore thinking_token_budget. No shipped configuration enables it. For details, refer to vLLM Plugins.

Verify Reasoning Behavior

Reasoning can be misconfigured without a startup error, so check the log. With server.log_level: DEBUG (the default), bot_server.log shows:

  • Loading LLM config from: ... — confirms whether the _think.yaml path was chosen.
  • Final LLM config: ... — the merged config, including the resolved enable_thinking value.
  • LLM starts thinking: and LLM is done thinking: — emitted by the TTS think-token handler, so their presence proves both that the model is reasoning and that the span is being suppressed.

Use the following symptom-to-cause mapping:

SymptomLikely Cause
Log still loads the non-think config after setting enable_reasoning: truellm.model_config is set (short-circuits the registry) or the registry entry lacks reasoning_supported: true.
The bot speaks its chain of thoughttts.think_tokens is null, the model uses different delimiters, or tts.type: nvidia is in use with no --reasoning-parser.
Long silence, then a truncated or empty replymax_new_tokens too low for reasoning plus answer, or the thinking budget is unbounded.

For more diagnostic guidance, refer to Troubleshooting and the Configuration Schema.

Use these pages to configure the surrounding model-serving and runtime behavior.