This page is for version v0.9.0.
For other versions, use one of these documentation indexes:
- Latest (v1.5.1) (default): https://docs.nvidia.com/dynamo/latest/llms.txt
- dev: https://docs.nvidia.com/dynamo/dev/llms.txt
- v1.5.1: https://docs.nvidia.com/dynamo/v1.5.1/llms.txt
- v1.5.0: https://docs.nvidia.com/dynamo/v1.5.0/llms.txt
- v1.4.2: https://docs.nvidia.com/dynamo/v1.4.2/llms.txt
- v1.4.1: https://docs.nvidia.com/dynamo/v1.4.1/llms.txt
- v1.4.0: https://docs.nvidia.com/dynamo/v1.4.0/llms.txt
- v1.3.0: https://docs.nvidia.com/dynamo/v1.3.0/llms.txt
- v1.2.1: https://docs.nvidia.com/dynamo/v1.2.1/llms.txt
- v1.2.0: https://docs.nvidia.com/dynamo/v1.2.0/llms.txt
- v1.1.1: https://docs.nvidia.com/dynamo/v1.1.1/llms.txt
- v1.1.0: https://docs.nvidia.com/dynamo/v1.1.0/llms.txt
- v1.0.2: https://docs.nvidia.com/dynamo/v1.0.2/llms.txt
- v1.0.1: https://docs.nvidia.com/dynamo/v1.0.1/llms.txt
- v1.0.0: https://docs.nvidia.com/dynamo/v1.0.0/llms.txt
- v0.9.1: https://docs.nvidia.com/dynamo/v-0-9-1/llms.txt
- v0.9.0: https://docs.nvidia.com/dynamo/v-0-9-0/llms.txt
- v0.8.1: https://docs.nvidia.com/dynamo/v-0-8-1/llms.txt
- v0.8.0: https://docs.nvidia.com/dynamo/v-0-8-0/llms.txt
- v0.7.1: https://docs.nvidia.com/dynamo/v-0-7-1/llms.txt
- v0.7.0: https://docs.nvidia.com/dynamo/v-0-7-0/llms.txt
The Dynamo Frontend is the API gateway for serving LLM inference requests. It provides OpenAI-compatible HTTP endpoints and KServe gRPC endpoints, handling request preprocessing, routing, and response formatting.
Feature Matrix
Quick Start
Prerequisites
- Dynamo platform installed
etcd and nats-server -js running
- At least one backend worker registered
HTTP Frontend
This starts an OpenAI-compatible HTTP server with integrated preprocessing and routing. Backends are auto-discovered when they call register_llm.
KServe gRPC Frontend
See the Frontend Guide for KServe-specific configuration and message formats.
Kubernetes
Configuration
See the Frontend Guide for full configuration options.
Next Steps