Accept a Kubernetes deployment for production
Verify an OpsKnight Kubernetes installation across security, migration, runtime roles, disruption, recovery, and incident delivery.
Prerequisites
Complete the selected Helm/Kustomize and integrated/split path. Configure production secrets, database, ingress/TLS, NetworkPolicy, monitoring, backup, and restore ownership.
Prepare the acceptance record
Record cluster/namespace, package revision, values/overlay revision, image digest, topology, database class, secret version identifiers, ingress origin, test operator, and rollback owner. Never record secret values.
Run production acceptance
- Render manifests and confirm all images are immutable and expected.
- Confirm no placeholder Secret and no unnecessary runtime RBAC/API token.
- Require the migration Job to complete once over a direct database path.
- Confirm exactly one integrated or split ownership model.
- Confirm all selected Deployments, probes, PDBs, and spread rules match the availability plan.
- Verify public TLS/readiness, forwarded headers, realtime streams, and signed webhooks.
- Confirm DNS host = certificate host = Ingress host =
NEXTAUTH_URL= normallyNEXT_PUBLIC_APP_URL= saved Application URL; no internal Service host appears in redirects or links. - Confirm
/setupwas completed through public HTTPS, login and provider callbacks remain on that hostname, and an unrelated host returns 421. - Verify allowed NetworkPolicy paths and an intended denied path.
- Confirm database TLS, connection headroom, backup, and an isolated restore.
- Verify role heartbeats, queue age/throughput, provider metrics, and status projection.
- Run a synthetic alert through notification, acknowledgement, escalation/assignment where configured, resolution, and status projection.
- Evict or restart one application Pod and confirm continued or timely restored service.
- Review dashboards and alerts with the operational on-call.
- Record evidence and make an explicit go/no-go decision.
Verify acceptance
Every applicable step requires objective evidence. Running Pods alone do not prove migration, queue progress, notifications, recovery, or public routing. Block go-live on unresolved security, data-recovery, migration, notification, or incident-processing failures.
Operate it in production
Repeat affected checks after cluster, ingress, database, topology, identity, provider, NetworkPolicy, or release changes. Schedule restore and disruption drills, secret/certificate rotation, capacity reviews, and upgrade rehearsal.
Troubleshooting failed acceptance
Use Kubernetes troubleshooting. Preserve events/logs before restarting. Correct the actual layer and rerun the failed check plus every dependent check. Never remove policy, disable TLS/signatures, bypass migration, or delete PVCs just to obtain a green result.
Change or undo the release decision
If acceptance fails after traffic begins, stop new traffic where safe, preserve data/evidence, and invoke Rollback or Backup and restore. Application rollback never implies schema rollback.
Next steps
Last updated for v2.0.0
Edit this page on GitHub