vLLM
vLLM is a popular LLM inference engine. The NeMo Gym VLLMModel server wraps vLLM’s Chat Completions endpoint and converts requests and responses to NeMo Gym’s native format, the OpenAI Responses API schema.
Most open-source models use Chat Completions format, while NeMo Gym uses the Responses API natively. VLLMModel bridges this gap by converting between the two formats automatically. For background on why NeMo Gym chose the Responses API and how the two schemas differ, see responses-api-evolution.
VLLMModel provides a Responses API to Chat Completions mapping middleware layer via responses_api_models/vllm_model. It assumes you are pointing to a vLLM instance since it relies on vLLM-specific endpoints like /tokenize and vLLM-specific arguments like return_tokens_as_token_ids.
Two upstream backends
VLLMModel can drive either of vLLM’s OpenAI-compatible endpoints, selected by the use_completions_api config flag. Both backends keep the same external Gym surface — /v1/responses and /v1/chat/completions continue to work identically; only the call to vLLM swaps.
To use VLLMModel, just change the responses_api_models/openai_model/configs/openai_model.yaml in your config paths to responses_api_models/vllm_model/configs/vllm_model.yaml!
VLLMModel connects NeMo Gym to a vLLM server that you start and manage yourself. If you would prefer NeMo Gym to launch and manage vLLM itself, use LocalVLLMModel instead. See LocalVLLMModel to learn more.
Use VLLMModel
Below is an e2e example of how to spin up a NeMo Gym compatible vLLM Chat Completions OpenAI server and run rollout collection with it.
This section walks through starting a vLLM server manually and connecting NeMo Gym to it through responses_api_models/vllm_model.
If you want NeMo Gym to manage the vLLM server lifecycle for you instead, see LocalVLLMModel.
Install vLLM
Please run the steps below in a separate terminal than your NeMo Gym terminal! The installation will take a few minutes.
Recommended on workstations without a system CUDA toolkit (i.e. no /usr/local/cuda). Install FlashInfer’s pre-built kernel packages so vLLM’s sampler and allreduce paths do not try to JIT-compile via nvcc at startup. Both companion packages enforce strict version equality with the flashinfer-python that vLLM pulled in, so pin them to its exact version:
Without these, the first sampling step inside vllm serve may fail with:
because PyPI’s PyTorch wheels ship the CUDA runtime libraries but not the nvcc compiler that FlashInfer’s JIT path falls back to.
If you install the companions without pinning their versions (e.g. uv pip install flashinfer-jit-cache --index-url https://flashinfer.ai/whl/cu130), the resolver picks the newest wheel on the index, and you will hit a different error at first sampling:
That’s the version-equality check in flashinfer/jit/env.py — fix by pinning both companions to flashinfer.__version__ as shown above. See the FlashInfer installation docs for details.
Download the model
This download will take a few minutes.
If you get errors relating to HuggingFace rate limits, please provide your HF token to command above.
If you do not have a HuggingFace token, please follow the instructions here to create one!
Spin up a vLLM server
vLLM server configuration
- If you want to use tools, find the appropriate vLLM arguments regarding the tool call parser to use. In this example, we use
Qwen/Qwen3-4B-Thinking-2507, which is suggested to use thehermestool call parser. - If you are using a reasoning model, find the appropriate vLLM arguments regarding reasoning parser to use. In this example, we use
Qwen/Qwen3-4B-Thinking-2507, which is suggested to use thedeepseek_r1reasoning parser. - The example below uses
--tensor-parallel-size 1which requires 1 GPU.
The spinup step will take a few minutes.
Configure NeMo Gym to use the local vLLM server
In a second terminal on the same GPU node that was used to spin up the vLLM server, enter the NeMo Gym Python environment, and start the NeMo Gym servers.
If you want to run NeMo Gym on a separate machine from the one used to spin up the vLLM server, please get the hostname of the machine used to run the vLLM server.
Then replace the policy_base_url=http://0.0.0.0:10240/v1 to point to the hostname policy_base_url=http://{hostname}:10240/v1.
Run rollout collection
In a third terminal on the same GPU node that was used to spin up the vLLM server, enter the NeMo Gym Python environment, and run rollout collection.
The /v1/completions backend
Set use_completions_api: true to drive vLLM’s text-completions endpoint instead of chat completions. A ready-to-use config ships at responses_api_models/vllm_model/configs/vllm_model_completions.yaml:
Render modes
A second flag, render_chat_template, picks how the caller’s messages list becomes the prompt string:
render_chat_template: false(default) — raw render. Forwards bytes verbatim. Required input: a single user message, optionally preceded by a single system message; their content is joined with\n\nand sent asprompt.tools, multi-turn turns, and non-text content blocks are rejected. This is the cheapest path and the most common base-model setup.render_chat_template: true— chat-template render. Renders the messages list to a prompt string client-side via HFAutoTokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, tools=..., **chat_template_kwargs). Multi-turn assistant / tool turns andtoolsare allowed.
For Gym’s /v1/responses endpoint, a string input is forwarded as the prompt directly under raw mode; under chat-template mode the converter still wraps it as a single user message before rendering.
tokenizer config (chat-template mode)
When render_chat_template: true, the HF tokenizer is loaded once at server startup. By default it’s loaded from the same identifier as model. Override with tokenizer: for base-model setups where the model checkpoint has no chat template in its tokenizer config — point tokenizer: at a different model whose template you want to inherit:
If the loaded tokenizer has no chat_template, the server fails at startup — silent fallback to raw mode would mask a config bug. The transformers package is required at runtime for chat-template mode (it’s a Gym dep, but if it’s missing the startup error message points at it explicitly).
Constraint matrix
Tools in chat-template mode. The chat template renders the tool definitions into the prompt (Hermes / Qwen / etc. all do this), but /v1/completions doesn’t run vLLM’s tool-call parser regardless of how the prompt is built — so any tool-call output text the model emits is not parsed by Gym. The caller is responsible for parsing tool calls out of the assistant text. If you need first-class tool-call routing, use use_completions_api: false.
<think>...</think> blocks emitted inline by the model are extracted downstream by VLLMConverter._extract_reasoning_from_content when results are converted back to a Response, exactly as on the chat path.
Token-ID extraction (RL training)
When return_token_id_information: true, the completions backend automatically adds logprobs: 0, return_token_ids: true, and return_tokens_as_token_ids: true to outbound requests. It reads prompt and generation token IDs directly from each choice. For older vLLM versions that omit inline IDs, it falls back to /tokenize for the prompt and the "token_id:<int>" entries in logprobs.tokens for the generation.
Use it
The end-to-end flow is identical to the Use VLLMModel walkthrough above; just swap the model-server config:
Make sure your input JSONL respects the selected mode in the Constraint matrix above.
VLLMModel configuration reference
Advanced: chat_template_kwargs
Override chat template behavior for specific models:
Advanced: extra_body
Pass vLLM-specific parameters not in the standard OpenAI API:
Use VLLMModel with multiple replicas of a model endpoint
The vLLM model server supports multiple endpoints for horizontal scaling:
How it works:
- Initial assignment: New sessions are assigned to endpoints using round-robin (session 1 → endpoint 1, session 2 → endpoint 2, etc.)
- Session affinity: Once assigned, a session always uses the same endpoint (tracked via HTTP session cookies)
- Why affinity? Multi-turn conversations and agentic workflows that call the model multiple times in one trajectory need to hit the same model endpoint in order to hit the prefix cache, which significantly speeds up the prefill phase of model inference.
Context Length Exceeded Handling
When a conversation exceeds the vLLM model’s maximum context length (max_seq_length), VLLMModel handles the error gracefully instead of crashing the entire rollout collection.
How it works
- vLLM rejects the request: vLLM returns an HTTP 400 error with a message like
"This model's maximum context length is 32768 tokens. However, you requested 32818 tokens...". - VLLMModel catches the error: Instead of propagating the exception, VLLMModel returns an empty response with
finish_reason: "length". - Responses API mapping: The
finish_reason: "length"is converted toincomplete_details: { reason: "max_output_tokens" }in the Responses API response returned to the agent.
This is particularly important for multi-turn agentic rollouts where conversation length can grow unpredictably across tool-call turns.
How to detect truncated responses
Downstream consumers (agents, RL training frameworks) can check the incomplete_details field on the response:
When incomplete_details.reason == "max_output_tokens", the response output is empty because vLLM rejected the request before generation began. This differs from a normal max_output_tokens truncation where the model generates up to the token limit — in this case, the input itself was too long.
Implications for training
When using NeMo Gym with NeMo RL or another training framework, responses with incomplete_details.reason == "max_output_tokens" indicate that the full conversation (prompt + prior generations) exceeded max_seq_length. Training frameworks should filter or handle these responses appropriately since they contain no generated tokens.
Training vs Offline Inference
By default, VLLMModel will not track any token IDs explicitly. However, token IDs are necessary when using NeMo Gym in conjunction with a training framework in order to train a model. For training workflows, use the training-dedicated config which enables token ID tracking:
This enables:
prompt_token_ids: Token IDs for the input promptgeneration_token_ids: Token IDs for generated textgeneration_log_probs: Log probabilities for each generated token