Operate and Restore#
Treat POC Factory as a stateful service. Keep the release image set, configuration revision, database, generated artifacts, job state, and encryption material aligned through operations and recovery.
Prerequisites#
Assign owners and record recovery targets before the first operating cycle.
Record the deployment handoff from Configure and validate.
Assign owners for the database, PVC or volume backups, secrets, certificates, and incident response.
Set retention and recovery objectives for POCs and any fine-tuning data.
Check the Running Service#
Use the health endpoint and deployment logs to establish current service state.
Check the frontend URL and backend
/healthresponse. Review failed jobs and provider errors with their job IDs.Inspect container or pod status, events, and logs:
docker compose --profile mcp ps docker logs --tail 200 poc-factory-mcp curl -fsS http://localhost:8000/health
For Kubernetes, use:
kubectl get pods,svc,ingress,pvc -n poc-factory kubectl get events -n poc-factory --sort-by=.lastTimestamp kubectl logs -n poc-factory <backend-pod-name> --tail=200
Check storage usage, certificate and Secret expiry, database growth, and the running image digests on your normal operations cadence.
At every release, render the chart or inspect Compose image resolution, perform the acceptance test, and confirm restore readiness.
The health endpoint reports database and state-manager status. For an incident, collect the job ID, status response, relevant event stream, logs, image digest, and configuration revision without secret values.
Back Up and Restore State#
Test the complete restore path, including credentials and generated artifacts.
Define recovery point and recovery time objectives for PostgreSQL, generated POCs, NAT job state, fine-tuning artifacts, and fine-tuning datasets.
Take database-native PostgreSQL backups and consistent volume or PVC snapshots at a coordinated point in time.
Protect the stable
ENCRYPTION_KEY, JWT signing material, OAuth secrets, and the image and configuration revisions needed to recreate the service.Restore into an isolated environment on the same release. Check sign-in, POC inventory, archive download, job lineage, and decryption of saved Settings credentials.
Record the restore-test result and correct any missing or mismatched state before relying on the backup set.
Do not treat docker compose down as data deletion: it preserves named volumes and host-mounted data. Avoid --volumes unless permanent deletion is intended and verified backups exist.
Apply Retention#
Set POC_RETENTION_DAYS to the approved period and verify that database records, archives, job state, fine-tuning data, and remote adapters follow the intended policy. The release bundle’s local template uses seven days; the Helm production example sets 30 days. Choose the value for your deployment rather than copying either without review.
Next Steps#
Use these references during incident review or a release upgrade.
Use Troubleshooting to investigate a failed health or generation check.
Use the deployment reference when comparing ports, images, and state stores.