Tool Calling and MCP Integration#

NIM LLM and VLM supports OpenAI-compatible tool calling through vLLM’s tool calling engine. Tool calling lets models invoke external functions by returning structured tool calls instead of text responses, enabling integration with Model Context Protocol (MCP) servers and any client library that uses the OpenAI tools format.

Prerequisites#

Before you start, complete the following prerequisites:

  1. Deploy a chat-capable NIM LLM and VLM container. For deployment instructions, refer to the Quickstart.

  2. Identify the built-in --tool-call-parser name that matches your model, or plan to supply a custom plugin. For the built-in parser list and plugin path rules, refer to Custom Parsers and Chat Templates.

Enable Tool Calling#

Tool calling requires two vLLM engine arguments: --enable-auto-tool-choice and a --tool-call-parser that matches your model.

To enable tool calling, complete the following steps:

  1. Pass both arguments as CLI arguments to nim-serve.

    docker run --gpus all \
      -e NIM_MODEL_PATH=hf://meta-llama/Llama-3.1-8B-Instruct \
      -p 8000:8000 \
      ${NIM_LLM_MODEL_FREE_IMAGE}:2.0.13 \
      nim-serve --enable-auto-tool-choice --tool-call-parser llama3_json
    
  2. Optional: For a Kubernetes workload, pass the same arguments through the container args field.

    containers:
      - name: nim
        image: <nim_llm_image>
        args:
          - "--enable-auto-tool-choice"
          - "--tool-call-parser"
          - "llama3_json"
    

    The following table describes the tool-calling arguments:

    Argument

    Description

    --enable-auto-tool-choice

    Allows the model to choose between generating text or calling a tool. Required for tool calling.

    --tool-call-parser <parser>

    Built-in parser name for extracting tool calls from model output. Must match the model’s tool calling format (for example, llama3_json for Llama 3.1 and 3.3).

    --tool-parser-plugin <path>

    Path to a custom tool-call parser file. Accepts an absolute path or a path relative to the model checkpoint directory (the path is used as-is, with no basename stripping). Required only when no built-in parser fits your model.

    Both --enable-auto-tool-choice and --tool-call-parser are required together. For the full list of built-in parsers, the supported models, and how to provide a custom plugin file, refer to Custom Parsers and Chat Templates. For request and response format and schema details, refer to the vLLM tool calling documentation. For more information on container arguments, the NIM_PASSTHROUGH_ARGS fallback, and configuration precedence, refer to Advanced Configuration.

  3. After you enable tool calling, send tool definitions on the /v1/chat/completions endpoint with the tools parameter and receive tool calls.

    For request and response format, examples, and the tool result loop, refer to the vLLM tool calling documentation.

Use MCP Tools with NIM#

NIM LLM and VLM does not connect to MCP servers directly. Your client application connects to MCP servers, converts tool definitions to the OpenAI tools format, and sends them in the tools array of a Chat Completions request.

To use MCP tools with NIM, complete the following steps:

  1. Connect your client application to the MCP servers that expose the tools you need.

  2. Convert those MCP tool definitions to the OpenAI tools format.

  3. Send the converted definitions in the tools array of a Chat Completions request.

  4. Execute the tool calls from the model response against the MCP server, then return the results to the model.

MCP tool schemas converted to OpenAI format can include extra fields such as strict or additionalProperties that are not part of the core function definition schema. vLLM accepts these fields without error. For details on schema handling, refer to the vLLM tool calling documentation.

Integrate LangChain and LangGraph#

When building agentic applications with LangChain and LangGraph, use create_agent from langchain.agents to implement tool calling with NIM LLM and VLM. This agent correctly executes the full tool-calling loop: generating a tool call, executing the tool, feeding the result back to the model, and producing a final response.

To integrate LangChain and LangGraph, complete the following steps:

  1. Launch the container with --enable-auto-tool-choice and --tool-call-parser <applicable-parser> as container arguments.

    Important

    Without these flags the server rejects "tool_choice": "auto" with a 400 error, and create_agent cannot invoke the tool.

    Important

    Do not pass a ProviderStrategy to create_agent for tool calling. That pattern bypasses the tool execution loop and causes the model to describe intended tool calls instead of executing them. Constructing create_agent with (llm, tools=[...]) and no strategy uses the correct execution loop.

  2. Install the required packages inside an isolated Python environment.

    On Ubuntu 23.04+, Debian 12+, and other distributions that ship with a PEP 668 “externally managed” system Python, running pip install directly against the system interpreter fails with an error: externally-managed-environment message. Create a virtual environment first, then install:

    python3 -m venv .venv
    source .venv/bin/activate
    
    pip install --upgrade pip
    pip install "langchain-nvidia-ai-endpoints>=0.3.0" "langchain>=1.0.0" "langchain-core>=1.0.0"
    

    create_react_agent from langgraph.prebuilt is deprecated in LangGraph 1.0 and scheduled for removal in LangGraph 2.0. Use create_agent from langchain.agents for new code. Existing code using from langgraph.prebuilt import create_react_agent still works with a LangGraphDeprecatedSinceV10 warning; migrate by swapping the import.

  3. Define a tool with the @tool decorator so create_agent can bind it, then create the agent and invoke it end-to-end.

    from langchain_core.tools import tool
    from langchain_nvidia_ai_endpoints import ChatNVIDIA
    from langchain.agents import create_agent
    
    
    @tool
    def get_weather(location: str) -> str:
        """Return the current weather for a location."""
        return f"The weather in {location} is 72F and sunny."
    
    
    llm = ChatNVIDIA(base_url="http://localhost:8000/v1", model="your-model-name")
    agent = create_agent(llm, tools=[get_weather])
    
    response = agent.invoke(
        {"messages": [{"role": "user", "content": "What is the weather in Santa Clara, CA?"}]}
    )
    print(response["messages"][-1].content)
    

    The snippet above pins the LangChain package versions to the API surface it uses: create_agent accepts the model as its first positional argument, tools are @tool-decorated callables, and the response is a dict with a messages list where the final entry is the agent’s answer. For more information on defining tools and building agents, refer to the LangChain agents documentation.

Troubleshooting#

Use the following topics when the model does not return tool calls or the server rejects "tool_choice": "auto".

Tool Calls Not Generated#

If the model returns text instead of a tool call, do the following:

  • Verify that --enable-auto-tool-choice is set.

  • Verify that --tool-call-parser is set to the correct parser for your model.

  • Check that the tools array is included in the request.

Error: “auto” Tool Choice Requires Configuration#

If you see "auto" tool choice requires --enable-auto-tool-choice and --tool-call-parser to be set, both arguments must be provided when launching the container.

Next Steps#

After tool calling is enabled, continue with the following topics: