NVIDIA NIM LLM with Run:ai and Vanilla Kubernetes for Enterprise RA# Deployment, Scale and Sizing Guide Abstract Introduction Scope Intended Audience Systems Overview Run:ai Overview Enterprise Reference Architecture Overview Hardware Enterprise Reference Architecture Software Reference Stack NVIDIA Inference Microservice (NIM) System Configuration Pre-requisites for installing Run:ai and NIM LLM Pre-reqs for RunAI Deploy and Configure NIM LLM on Run:ai Create Data Source for NIM Create a Secret Deploy a Specific NIM LLM Alternate YAML method Query the Inference Server Performance and Scale Methodology Benchmarking NIM LLM Scale Methodology for NIM LLM with Run:ai Installing and Configuring Gen-AI Perf Inference Performance and Scale Results with Run:ai Benchmarking without Run:ai Benchmarking with Run:ai using a full GPU Benchmarking with Run:ai using a Fractional GPU Performance of NIM LLM at Scale with Run:ai Simultaneous Multiple NIMs with Run:ai Sizing Guidelines Summary Appendix Appendix A Installing Pre-reqs for RunAI Get the Run.ai SaaS Login Installing Nginx Installing Prometheus Create Certs and Private Keys Update BCM Ingress with CA Certificates Expose Ingress Controller to use Public IP from the Metal LB pool Update DNS Server Configure Run:ai Configure the addition of a cluster to Run.ai SaaS Install Run:ai Cluster Install Knative for Inference Workloads Configure Knative to use with Run:ai Configure HPA for Autoscaling by Knative Update Knative timeout Create a Project in Run:ai Change the Placement Strategy Add Users in Run:ai Notices Notices Notice Trademarks Copyright