Choose and deploy an OpsKnight topology
Select the supported OpsKnight deployment path and continue to a complete installation and production acceptance workflow.
Start here when installing OpsKnight. Choose one packaging path and one runtime topology, then keep that choice consistent for installation, upgrades, troubleshooting, and recovery.
Start with the closest workload shape
| Current planning shape | Starting architecture to evaluate |
|---|---|
| Small — about 40 users and 12 services | Integrated Compose when one host and no HA are acceptable |
| Medium — about 120 users and 32 services | Split Compose, or Helm Split when Kubernetes or HA is required |
| Large — about 400 users and 80 services | Helm/Kustomize Split with PgBouncer and external PostgreSQL |
| Storm — about 1,000 users and 200 services | Multi-replica Split, PgBouncer, external HA PostgreSQL, and workload-specific certification |
These are starting architectures derived from current test-data shapes, not certified user limits. User count alone cannot size OpsKnight: alert bursts, notification fanout, simultaneous responders, SSE sessions, status-page subscribers, provider quotas, availability requirements, and the PostgreSQL connection budget can change the answer.
Use the full Deployment Planner to compare all dimensions, then review the historical benchmark results without treating an observed peak as supported capacity.
Follow the complete installation journey
Every supported packaging path has the same control points. Do not skip ahead when a Pod or container becomes healthy:
- Choose the package and integrated or split topology.
- Provision PostgreSQL, calculate connections, and establish backup/restore ownership.
- Generate and protect stable secrets; pin the tested image digest.
- Create public DNS and TLS, then align the ingress/proxy host,
NEXTAUTH_URL, andNEXT_PUBLIC_APP_URL. - Implement the reverse-proxy contract.
- Run migration once through a direct database connection, then deploy workloads.
- Require readiness through the public HTTPS origin.
- Open public
/setup, verify the detected Application URL, and create the first administrator. - Sign in through the same hostname and confirm Settings → System → App URL.
- Run domain, authentication, webhook, realtime, incident, notification, backup, and restore acceptance tests.
Read Application URL and host routing before exposing any production installation. A wrong canonical host can cause HTTP 421 after bootstrap.
Choose how to run OpsKnight
| Requirement | Recommended path |
|---|---|
| Evaluation or one small server | Docker Compose: integrated |
| One server with isolated workers | Docker Compose: split |
| Kubernetes with packaged, schema-validated configuration | Helm |
| Kubernetes with GitOps or owned overlays | Kustomize |
| Docker across multiple manager/worker nodes | Swarm |
| Operator-managed production database | Use the selected path with external PostgreSQL |
| High web connection count | Use split runtime with PgBouncer |
Compose is a single-host orchestrator. Swarm and Kubernetes can reschedule workloads after a host failure, but availability still depends on database, ingress, storage, replica, and disruption design.
Choose integrated or split runtime
- Integrated runs the web application and background responsibilities together. It is the least complex path for evaluation and smaller installations.
- Split runs Web, Scheduler, General Worker, Critical Worker, Bulk Worker, and Status Projector separately. Choose it when you need role-specific scaling, failure isolation, or Web-only pooling.
Read Integrated versus split before choosing. Do not run integrated and split ownership at the same time against one database.
Choose the database path
- Bundled PostgreSQL is convenient, but the supplied deployment is not a highly available database service.
- External PostgreSQL is the normal production choice when another team or managed service owns availability, backups, upgrades, and failover.
- PgBouncer is supported for Web in split mode. Migration, scheduler, and worker roles retain direct PostgreSQL connections.
Read Database connections and calculate the connection budget before setting replicas or pool sizes.
Production acceptance applies to every path
Do not declare an installation ready because its process or Pod is running. Before accepting production traffic:
- Pin the exact tested image digest.
- Back up stable secrets independently from the database.
- Complete database migration with exactly one owner.
- Verify readiness through the public HTTPS origin.
- Complete
/setupthrough that origin and verify the saved Application URL. - Verify every selected runtime role and its heartbeat or queue progress.
- Trigger a synthetic alert and complete acknowledgement and resolution.
- Verify at least one real notification provider and any configured ChatOps destination.
- Test a logical backup and isolated restore.
- Record the deployment files, values, overlays, image digest, database endpoint class, and rollback decision.
The packaging-specific production checklist gives exact commands and expected results.
Related guides
Last updated for v2.0.0
Edit this page on GitHub