Operate2 min read

Accept a Kubernetes deployment for production

Verify an OpsKnight Kubernetes installation across security, migration, runtime roles, disruption, recovery, and incident delivery.

Prerequisites

Complete the selected Helm/Kustomize and integrated/split path. Configure production secrets, database, ingress/TLS, NetworkPolicy, monitoring, backup, and restore ownership.

Prepare the acceptance record

Record cluster/namespace, package revision, values/overlay revision, image digest, topology, database class, secret version identifiers, ingress origin, test operator, and rollback owner. Never record secret values.

Run production acceptance

  1. Render manifests and confirm all images are immutable and expected.
  2. Confirm no placeholder Secret and no unnecessary runtime RBAC/API token.
  3. Require the migration Job to complete once over a direct database path.
  4. Confirm exactly one integrated or split ownership model.
  5. Confirm all selected Deployments, probes, PDBs, and spread rules match the availability plan.
  6. Verify public TLS/readiness, forwarded headers, realtime streams, and signed webhooks.
  7. Confirm DNS host = certificate host = Ingress host = NEXTAUTH_URL = normally NEXT_PUBLIC_APP_URL = saved Application URL; no internal Service host appears in redirects or links.
  8. Confirm /setup was completed through public HTTPS, login and provider callbacks remain on that hostname, and an unrelated host returns 421.
  9. Verify allowed NetworkPolicy paths and an intended denied path.
  10. Confirm database TLS, connection headroom, backup, and an isolated restore.
  11. Verify role heartbeats, queue age/throughput, provider metrics, and status projection.
  12. Run a synthetic alert through notification, acknowledgement, escalation/assignment where configured, resolution, and status projection.
  13. Evict or restart one application Pod and confirm continued or timely restored service.
  14. Review dashboards and alerts with the operational on-call.
  15. Record evidence and make an explicit go/no-go decision.

Verify acceptance

Every applicable step requires objective evidence. Running Pods alone do not prove migration, queue progress, notifications, recovery, or public routing. Block go-live on unresolved security, data-recovery, migration, notification, or incident-processing failures.

Operate it in production

Repeat affected checks after cluster, ingress, database, topology, identity, provider, NetworkPolicy, or release changes. Schedule restore and disruption drills, secret/certificate rotation, capacity reviews, and upgrade rehearsal.

Troubleshooting failed acceptance

Use Kubernetes troubleshooting. Preserve events/logs before restarting. Correct the actual layer and rerun the failed check plus every dependent check. Never remove policy, disable TLS/signatures, bypass migration, or delete PVCs just to obtain a green result.

Change or undo the release decision

If acceptance fails after traffic begins, stop new traffic where safe, preserve data/evidence, and invoke Rollback or Backup and restore. Application rollback never implies schema rollback.

Next steps

Last updated for v2.0.0

Edit this page on GitHub