DEPLOY / v2.0.0

Run OpsKnight
on infrastructure you operate.

Start with the maintained Compose quickstart, then choose Split, Swarm, Helm, or Kustomize when isolation, availability, platform standards, or connection pressure require it.

DEPLOYMENT CHOOSER

Choose from operational requirements, not an invented traffic threshold.

Split mode exists for independent scaling, failure isolation and notification-lane protection. Multi-node availability requires the surrounding database, proxy, storage and recovery design to match.

Where are you running it?
What are you planning?

YOUR STARTING POINT

Docker Compose · integrated

Start with the integrated runtime and PostgreSQL. Split the runtime when independent scaling is useful.

View guide

FASTEST START

Bring up an isolated Compose evaluation.

The maintained quickstart runs the integrated application and PostgreSQL on one Docker host. It is an evaluation or small-test path, not a multi-node HA claim.

01 / CONFIGURE

Pin the release and set the minimum contract.

Use an explicit 2.0.0 image tag or immutable digest. Configure the database password, public URL values, NEXTAUTH_SECRET, API_KEY_SECRET, and the stable 64-hex-character ENCRYPTION_KEY. Other provider or API secrets are conditional on the features you enable.

OPSKNIGHT_IMAGE=ghcr.io/opsknight-labs/opsknight:2.0.0
POSTGRES_PASSWORD=<unique-password>
NEXTAUTH_URL=http://localhost:3000
NEXT_PUBLIC_APP_URL=http://localhost:3000
NEXTAUTH_SECRET=<random-base64-secret>
API_KEY_SECRET=<separate-base64-secret>
ENCRYPTION_KEY=<64-hex-character-key>
02 / START & VERIFY

Use the maintained Compose path.

Pull the pinned image, wait for health, inspect the services, and require readiness to return HTTP 200 before setup.

docker compose -f deploy/compose/docker-compose.yml pull
docker compose -f deploy/compose/docker-compose.yml up -d --wait
docker compose -f deploy/compose/docker-compose.yml ps
curl --fail --show-error 'http://localhost:3000/api/health?mode=readiness'

PRODUCTION FOUNDATIONS

The topology is only one part of production readiness.

Public routing, database durability, secrets, migrations and recovery remain operator responsibilities in a self-hosted system.

Public HTTPS & Application URL

DNS, TLS, proxy or ingress, NEXTAUTH_URL, normally NEXT_PUBLIC_APP_URL, and the saved Application URL must resolve to the same browser-facing origin.

Host-routing contract

PostgreSQL & recovery

Budget aggregate connections, keep migrations on the direct database route where required, automate backups, and test a restore with the matching encryption secrets.

Backup and restore

Secrets that survive restarts

Keep NEXTAUTH_SECRET and ENCRYPTION_KEY stable and protected. Add provider, metrics, SCIM, voice, or API secrets only when the corresponding capability requires them.

Configuration reference

Migrations & readiness gates

A successful container start is not acceptance. Require migration success, readiness, canonical-host behavior, and an end-to-end synthetic incident before production cutover.

Production acceptance

SUPPORTED TOPOLOGIES

Pick the operating model that matches your platform.

The v2.0.0 deployment guide does not publish a certified requests-per-second threshold for choosing between these paths.

TopologyBest fitMulti-node HAIndependent scalingPgBouncerCapacity status
Integrated ComposeGuideSimplest single-host operationNoNoNoNot certified
Split ComposeGuideSingle-host workload isolationNoYesOptionalNot certified
Split + PgBouncerGuideIsolation plus web connection poolingNoYesEnabledNot certified
Docker SwarmGuideDocker-native multi-node application operationApplication tierYesOptional in SplitNot certified
Kubernetes HelmGuidePackaged, schema-validated KubernetesWhen configuredYesOptional in SplitNot certified
Kubernetes KustomizeGuideGitOps and environment-owned overlaysWhen configuredYesOptional in SplitNot certified

Docker Swarm can replace application tasks across hosts, but bundled PostgreSQL is not highly available. A Swarm HA objective requires external HA PostgreSQL, an external TLS load balancer, durable backups, and at least three managers.

MAINTAINED INSTALL PATHS

Use the gates each orchestrator already provides.

DOCKER SWARM

Use the maintained deploy orchestrator.

The script validates manager/capacity prerequisites, creates secrets, owns migration, deploys the stack, waits for convergence, and checks readiness. Routine raw docker stack deploy bypasses those gates.

# Split mode (default; requires explicit release image)
export OPSKNIGHT_IMAGE="ghcr.io/opsknight-labs/opsknight:2.0.0"
./deploy/swarm/scripts/deploy.sh

# Integrated mode (defaults to 2.0.0 image)
SWARM_RUNTIME_MODE=integrated ./deploy/swarm/scripts/deploy.sh
Swarm installation guide
KUBERNETES / HELM

Render and validate the maintained chart.

Create the namespace and externally managed Secret, render the local chart, validate it, then install the reviewed production values. The published v2 docs do not require a hosted chart repository.

helm lint deploy/kubernetes/helm/opsknight -f values.production.yaml
helm template opsknight deploy/kubernetes/helm/opsknight \
  --namespace opsknight -f values.production.yaml > rendered.yaml
kubectl apply --dry-run=server -f rendered.yaml

helm upgrade --install opsknight deploy/kubernetes/helm/opsknight \
  --namespace opsknight --create-namespace \
  --values values.production.yaml --wait --timeout 15m
Helm installation guide

CAPACITY

Build a budget from workload shape and measured saturation.

Until a matching benchmark is certified, do not turn a measured peak into a production promise.

Describe the workload first.

  • Peak and sustained alert ingestion
  • Incident deduplication ratio
  • Notifications and escalations per incident
  • Concurrent users and SSE streams
  • Status subscribers and fan-out
  • External provider rate limits

Budget PostgreSQL connections.

Sum each runtime role's replicas × pool size plus migration, administration, monitoring and safety headroom. PgBouncer can pool supported web traffic in Split mode, but direct-role pools still count.

Build a capacity budget
SIGNALLIKELY CONSTRAINTFIRST RESPONSE
Critical notification age risesCritical delivery lane

Inspect provider/DB pressure, then scale critical workers if that lane is constrained.

Bulk queue growsBulk worker lane

Scale bulk workers without consuming critical-delivery capacity.

General job age risesGeneral background work

Scale general workers after checking downstream dependencies.

API latency risesWeb tier

Scale web replicas only after checking PostgreSQL latency and connections.

DB connections approach budgetPostgreSQL / pools

Reduce pools, add supported PgBouncer, or scale PostgreSQL.

Provider 429s riseProvider quota

Reduce provider concurrency and honor retry timing.

Status projection lag risesStatus projector

Scale the projector and inspect its direct DB pool.

SSE latency/disconnects riseWeb / realtime path

Inspect proxy timeouts, web capacity, and database pressure.

Read scaling signals

GO-LIVE ACCEPTANCE

Prove the deployment before production alerting depends on it.

Pin the exact image tag or digest and preserve the reviewed deployment configuration.

Require migration success and public readiness through the intended HTTPS origin.

Verify DNS, TLS/proxy host, authentication URLs, and saved Application URL agree.

Verify database backups and an isolated restore with the required stable secrets.

Run a synthetic alert through notification, acknowledgement, resolution, and status projection.

Observe queue/worker/database/provider signals and repeat the affected checks after topology or release changes.

VERIFICATION GATES

First-deployment verification checklist

Run these five operational smoke tests before redirecting production monitoring signals to your new instance.

01

Bootstrap & Health Probe

Verify API gateway readiness, database pool connectivity, and Redis cache health.

PROBE COMMAND
curl -fsSL https://opsknight.internal/api/v1/health | jq .
EXPECTED RESULT

{"status":"healthy","database":"connected","redis":"connected","version":"2.0.0"}

02

Inbound Webhook Verification

Simulate a signed Prometheus Alertmanager or Datadog alert payload.

PROBE COMMAND
curl -X POST https://opsknight.internal/api/v1/webhooks/raw-synthetic \
  -H "Content-Type: application/json" \
  -H "X-OpsKnight-Signature: $SYNTHETIC_HMAC" \
  -d '{"event":"ping","service":"checkout","severity":"sev1"}'
EXPECTED RESULT

HTTP/2 202 Accepted · Event acknowledged and routed into Command Center triage stream.

03

On-Call Paging Carrier Test

Dispatch a high-priority paging test through configured Twilio/carrier routes to primary responder.

PROBE COMMAND
curl -X POST https://opsknight.internal/api/v1/schedules/primary/test-page \
  -H "Authorization: Bearer $ADMIN_TOKEN"
EXPECTED RESULT

Dispatched carrier notification token. Primary handset rings in <5s with acknowledgment prompt.

04

ChatOps War Room Handshake

Confirm Slack or Microsoft Teams webhook handshake and slash command interactivity.

PROBE COMMAND
/opsknight ack INC-2026-0042
EXPECTED RESULT

Bidirectional sync: thread updated with "Acknowledged by @responder. Auto-escalation halted."

05

Operations Queue Diagnostics

Probe background worker pools, BullMQ queue depths, and database connection overhead.

PROBE COMMAND
curl -fsSL https://opsknight.internal/api/v1/operations/queues \
  -H "Authorization: Bearer $ADMIN_TOKEN"
EXPECTED RESULT

{"critical_queue":{"waiting":0,"active":1},"bulk_queue":{"waiting":0},"delayed":0}

OPTIONAL PROFESSIONAL HELP

Commercial support & implementation services

Deployment assistance, architecture review, upgrades, troubleshooting and implementation can be scoped separately. No 24×7 coverage or response-time SLA is advertised on this website; any service commitment must be agreed in writing.

Discuss support & services

Your incidents should belong to you.

Run OpsKnight on infrastructure you control.

v2.0.0 · AGPL-3.0-only · Self-hosted · 28 inbound integrations