Troubleshooting#

Start with the affected job ID, the backend /health response, and the relevant container or pod logs. Check the symptom below, then rerun the failed validation step from Configure and validate.

Cursor Key Is Rejected or Generation Does Not Start#

The default auto contract-first path needs a valid Cursor credential. Open Settings and use Validate & Save Cursor Key if user credentials are allowed. Otherwise, ask the deployment operator to check the administrator key and generator policy. CODE_GENERATOR_TOOL=llm is a separate, intentionally selected inference-backed full-project mode.

Launch Preflight Reports Missing Values#

Put NIM_NGC_ORG, database credentials, and ENCRYPTION_KEY in the deployment .env. Exporting them only in your shell does not populate the preflight container’s environment file. Compare the current .env with the matching release template while retaining existing keys and database values.

Startup Reports MODEL_CURSOR Is Required#

An administrator CURSOR_API_KEY requires MODEL_CURSOR. Set MODEL_CURSOR=auto for Cursor’s automatic routing or a supported explicit model identifier. Changing the inference provider key does not replace the Cursor credential.

Docker Compose Rejects env_file or required#

Check docker compose version. The release bundle uses syntax requiring Compose plugin 2.24.0 or later. Use the Compose plugin, not the legacy docker-compose executable.

Download Fails Although the POC Appears on Dashboard#

Compare the POC database record with the generated artifact path and mounted storage. Restore the matching artifact from the coordinated backup or use the authorized synchronization or cleanup workflow. Do not delete a record or volume before confirming the retained copy.

Frontend Returns 502 or 504#

Check the frontend proxy target, backend Service or container health, and port 8000. On Kubernetes, also check ingress routing and timeouts. Use the browser-facing frontend URL for user traffic; the backend service is internal to the deployment.

Image Pull Reports ImagePullBackOff#

Check the registry host, NGC product access or mirror pull Secret, nvidia namespace, selected image tag, and target CPU architecture. Inspect docker compose --profile mcp config --images or the rendered Helm manifests. The backend, frontend, and proxy must come from the same release set.

Provider Reports Quota Exhaustion or Repeated Retries#

Permanent quota exhaustion needs account quota or billing action; repeatedly retrying the same job will not resolve it. For transient provider errors, inspect the job and backend logs before creating a duplicate. See Release notes for changes to provider error classification.

Inference Profile Is Required or Blueprint Matching Fails#

Check Settings for a complete administrator or user profile. The profile needs a base URL, API key, and five model roles. Validate chat roles through /chat/completions and the embedding role through /embeddings; a working chat route does not prove blueprint matching can embed.

Login Loops or Settings Returns 401#

For shared deployments, confirm HTTPS, secure cookies, exact ALLOWED_ORIGINS and ALLOWED_HOSTS values, and an OAuth callback matching https://<host>/api/auth/callback. For private localhost HTTP evaluation, use development mode with secure cookies disabled. Clear stale browser cookies after changing cookie settings.

Progress Stops Updating#

Check the durable job on Dashboard and the backend logs. For Kubernetes ingress or a reverse proxy, disable response buffering for Server-Sent Events and allow sufficiently long read and send timeouts. Reopen the existing job rather than submitting a duplicate request.

Saved Credentials Disappear or Cannot Be Decrypted#

Check that ENCRYPTION_KEY is present, Fernet-compatible, and identical across restarts and replicas. Restore the original key with the corresponding database backup. If the key has been rotated without re-encrypting stored values, users must save their credentials again.

Storage Reports Permission Denied#

Check ownership and permissions on the generated-POC and job-store mounts. The Kubernetes production example runs application containers as UID/GID 1000; PVCs must be writable by that identity. Check the mounted path and storage class before restarting the workload.

Fine-Tuning Has No Available Base Models#

Validate all three NeMo endpoints from the training-target panel. The URLs must be reachable from the backend, and the intended model needs both a NIM model entry and a matching Customizer configuration. Read-only fields indicate deployment-managed endpoints; ask the operator to check the connection settings.

Fine-Tuning Waits for Training or Model Readiness#

Inspect the existing run and backend logs. Give the job ID to the NeMo platform owner to check scheduling, the selected customization configuration, and deployment readiness. Training completion does not mean the adapter is serving; wait for readiness before chat or evaluation. Use the NeMo collection’s platform guidance for cluster-side failures.

Fine-Tuning Evaluation Fails or LLM Judge Is Unavailable#

Check model readiness and Evaluator availability for an evaluation failure. For a missing LLM-judge score, validate the active inference profile’s balanced model role and inspect backend provider errors. Missing evidence requires human review when the accepted plan requires that metric; an unavailable score is not a passing result.

Fine-Tuned Model Cleanup Reports Skipped Actions#

Review the Clean Up Model result and backend logs for the existing job. Ask the platform owner to resolve remote resource failures before assuming the deployment or adapter was removed. Preserve needed dataset and config downloads before cleanup; Fine-tune a model explains the application actions.