Deploy with Docker Swarm
Deploy, verify, operate, upgrade, and recover a multi-node OpsKnight installation on Docker Swarm.
Use this deployment when Docker Swarm is already your production scheduler and you want OpsKnight spread across multiple Docker hosts. The supplied deployment orchestrator validates capacity, creates encrypted Swarm secrets, runs database migrations, waits for every service to converge, and checks HTTP readiness.
Swarm makes application roles replaceable; it does not make every dependency highly available. For production, use an external highly available PostgreSQL service, an external TLS load balancer, durable backups, and at least three Swarm manager nodes. The bundled PostgreSQL service is one replica on one labeled node. Its volume survives task replacement on that node, but it has no database failover and is not a production HA database.
Choose a topology
The default SWARM_RUNTIME_MODE=split runs separate web, scheduler, worker, and
status-projector services. Choose it when you need independent scaling and
failure isolation. It requires an explicit image tag or digest built with
split-runtime support; the deployment script rejects latest in this mode.
SWARM_RUNTIME_MODE=integrated runs the web process and background work in one
service. It is simpler for small installations but cannot scale or isolate
notification lanes independently.
In split mode, the default service layout is:
| Service | Replicas | Update order | Responsibility |
|---|---|---|---|
opsknight-web |
2 | start first | UI, API, authentication, readiness |
opsknight-scheduler |
2 | stop first | lease-fenced scheduled work and SLA evaluation |
opsknight-general-worker |
2 | stop first | general queues and integrations |
opsknight-critical-worker |
2 | stop first | urgent paging and notification work |
opsknight-bulk-worker |
2 | stop first | bulk and maintenance work |
opsknight-status-projector |
2 | stop first | incident and public-status projection |
opsknight-pgbouncer |
2, optional | start first | transaction pooling for web traffic only |
opsknight-db |
1, optional | stop first | bundled, node-pinned PostgreSQL |
Schedulers and workers coordinate through database leases and queue claims, so multiple replicas do not mean every task runs twice. Do not start ad hoc scheduler or worker processes outside the stack.
Prerequisites
Prepare the following before the maintenance window:
- Docker Engine with Swarm mode on every node and registry access for the selected OpsKnight image.
- Three managers for manager-quorum tolerance and enough workers to satisfy the selected replicas after one worker is drained.
- A digest-pinned OpsKnight image, for example
ghcr.io/opsknight-labs/opsknight@sha256:<tested-digest>. - External PostgreSQL for production, with TLS, backups, tested restore, and a connection limit that covers all direct and pooled clients.
- DNS and a TLS-terminating reverse proxy or load balancer. Only port 3000 is published by the application stack; TLS is an infrastructure responsibility.
- Production values for the database URL,
NEXTAUTH_SECRET,ENCRYPTION_KEY,NEXTAUTH_URL, andNEXT_PUBLIC_APP_URL.
Open these ports between Swarm nodes:
| Port | Scope | Purpose |
|---|---|---|
2377/tcp |
nodes to managers | cluster management |
7946/tcp and 7946/udp |
all nodes | discovery and gossip |
4789/udp |
all nodes | overlay-network VXLAN traffic |
3000/tcp |
load balancer to Swarm | OpsKnight routing-mesh ingress |
Keep PostgreSQL 5432 and PgBouncer 6432 on private networks. If you use a cloud firewall, allow the overlay ports by node security group rather than broad internet CIDRs.
1. Build the Swarm and verify quorum
On the first manager:
docker swarm init --advertise-addr <manager-private-ip>
docker swarm join-token manager
docker swarm join-token worker
Run the printed join command on the remaining managers and workers. Then, from
a manager, verify that every intended node is Ready and Active:
docker node ls
Use node names and labels that remain meaningful during an incident. If you intend to use bundled PostgreSQL in a multi-node Swarm, select the host with the durable volume and pin the database there:
docker node update --label-add opsknight.database=true <database-node>
The deploy script labels the current manager automatically only for a
single-node Swarm. In a multi-node cluster it fails until you explicitly label a
database node, unless you deliberately set AUTO_LABEL_DATABASE_NODE=true.
Configuration: public URLs and secrets
Run deployments from a manager in a protected administrative session. Export values through your CI secret store or shell without committing them:
export OPSKNIGHT_IMAGE='ghcr.io/opsknight-labs/opsknight@sha256:<tested-digest>'
export NEXTAUTH_URL='https://opsknight.example.com'
export NEXT_PUBLIC_APP_URL='https://opsknight.example.com'
export NEXTAUTH_SECRET='<high-entropy-random-value>'
export ENCRYPTION_KEY='<production-encryption-key>'
The deploy script turns sensitive values into content-hashed Docker secrets and
mounts them under /run/secrets. Application variables use the corresponding
*_FILE form, so the secret is not stored in the service environment. When a
secret value changes, a new versioned Raft secret is created and affected
services roll to it. Retain an old encryption key until data encrypted with it
has been migrated; rotating that key is not equivalent to rotating a password.
Production deployment is fail closed: placeholder database credentials or
default secrets stop the script. ALLOW_INSECURE_SECRETS=true is only for an
isolated disposable environment.
Set both public URLs to the exact browser origin, including https and any
non-default port. Configure the external proxy to preserve Host, forwarded
scheme, client IP, WebSocket upgrades, and long-lived server-sent event
responses. Disable response buffering for streaming endpoints and make the
proxy idle timeout longer than the application heartbeat interval.
3. Connect PostgreSQL
Recommended: external PostgreSQL
Provide the database endpoint and enable TLS verification:
export EXTERNAL_DB=true
export EXTERNAL_DB_HOST='postgres.production.internal'
export EXTERNAL_DB_PORT='5432'
export EXTERNAL_DB_USER='opsknight'
export EXTERNAL_DB_PASSWORD='<database-password>'
export EXTERNAL_DB_NAME='opsknight'
export EXTERNAL_DB_SSLMODE='verify-full'
The application roles and migration task connect directly to PostgreSQL unless PgBouncer is enabled. Use a database role that can run the shipped migrations, then follow your platform policy if separate migration and runtime roles are required.
For a private certificate authority, supply its PEM file:
export PGBOUNCER_TLS_CA_CERT='/secure/path/database-root-ca.crt'
The orchestrator creates a Swarm secret for the CA and configures direct and
pooled connections to use /etc/ssl/certs/custom-ca.crt. A missing file or a
certificate whose hostname does not match the database endpoint must be fixed;
do not downgrade production to sslmode=disable.
Limited use: bundled PostgreSQL
Leave EXTERNAL_DB unset and provide strong POSTGRES_USER,
POSTGRES_PASSWORD, and POSTGRES_DB values. The database runs on the node
labeled opsknight.database=true and stores data in a local volume. A failed
task can restart on that same healthy node. Loss of that node or its disk needs
a restore; Swarm does not replicate the database volume.
4. Decide whether to use PgBouncer
Enable PgBouncer when web replica connection demand would otherwise consume too many PostgreSQL backends:
export PGBOUNCER_ENABLED=true
The supplied pool uses transaction mode and defaults to two replicas, pool size 10, reserve pool 5, and 1,000 client connections. Only split-runtime web traffic uses it. Schedulers, workers, the status projector, and migrations connect directly because their transaction and lease behavior must not be hidden behind the web pool.
Before changing replicas or pool sizes, calculate the direct connections from every non-web role plus PgBouncer's possible backend connections and retain headroom for migrations, administration, autovacuum, and failover. The deploy script runs the repository capacity validator and stops when its configured budget is unsafe. See choose a deployment and benchmark results before increasing replicas.
Deploy in a controlled sequence
From the repository root on a manager, deploy the default split topology:
./deploy/swarm/scripts/deploy.sh
For an integrated installation:
SWARM_RUNTIME_MODE=integrated ./deploy/swarm/scripts/deploy.sh
You can set SWARM_STACK_NAME when this cluster hosts more than one isolated
installation. The script serializes deployments for the same stack with a
host-local lock, creates an attachable overlay network, and then performs these
steps in order:
- Confirm Docker is in active Swarm mode and the command is running on a manager.
- Validate connection capacity for the selected runtime, replicas, and pool.
- Create the overlay network and versioned Raft secrets.
- Start and await bundled PostgreSQL when selected.
- Create a one-shot migration service and require exit code 0.
- Run
docker stack deploy --pruneso topology changes remove obsolete roles. - Wait for every desired replica to converge.
- Probe
/api/health?mode=readinessand fail if it is not healthy.
Do not bypass the script with a manual docker stack deploy during routine
operation. That skips preflight, migration ordering, secret construction, and
the final health gate.
Database migration boundary
The one-shot container applies the shipped schema migrations and standard online indexes before application rollout. It connects directly to PostgreSQL and is removed when it exits. If it fails, inspect its emitted logs and fix the database or image problem before retrying; the application rollout does not continue.
The optional SLA scheduler index has a separate compatibility boundary. It is
optional in LEGACY and SHADOW scheduler modes and required before switching
to INDEXED. The normal Swarm migration task does not install that optional
index. Before setting SLA_SCHEDULER_MODE=INDEXED, run this command once against
the direct database from a compatible image in a controlled job:
DATABASE_URL="$DIRECT_DATABASE_URL" npm run prisma:indexes:sla-scheduler
Verify index creation before enabling INDEXED on scheduler replicas. See
database migrations for the rollout and
rollback boundary.
Verify the installation
The deploy script calls the health checker automatically. Re-run it after load balancer or network changes:
./deploy/swarm/scripts/health-check.sh
It verifies node state, desired versus running replicas, rejected or failed tasks, and the readiness endpoint. Inspect the same state manually:
docker node ls
docker stack services opsknight
docker stack ps opsknight --no-trunc
docker service logs --since 15m opsknight_opsknight-web
docker service logs --since 15m opsknight_opsknight-critical-worker
docker service logs --since 15m opsknight_opsknight-scheduler
curl --fail --silent 'https://opsknight.example.com/api/health?mode=readiness'
Expected result: every service shows its desired replica count, no current task is rejected or repeatedly restarting, readiness reports healthy, sign-in works at the public URL, and a test incident produces the expected notification.
Also verify external behavior: create an incident, acknowledge it, confirm the timeline updates without a refresh, exercise one configured provider, and resolve the incident. A green container alone does not prove paging delivery.
7. Prove node-failure behavior
Before production approval, drain one application worker at a time:
docker node update --availability drain <worker-node>
docker stack ps opsknight --no-trunc
./deploy/swarm/scripts/health-check.sh
Expected result: tasks move to other eligible nodes, web remains reachable, critical notification work continues, scheduler leases transfer without duplicate visible actions, and status projection catches up. Return the node to service afterward:
docker node update --availability active <worker-node>
If replicas cannot reschedule, add capacity or correct placement constraints before launch. Draining the sole bundled-database node is expected to make that database unavailable; it demonstrates the bundled database's single-node boundary and is not a valid HA configuration.
8. Scale and observe
Change the repository-supported replica variables and redeploy instead of making an undocumented service change that the next stack deploy will replace. Review database capacity every time web, scheduler, worker, projector, or PgBouncer replicas change.
Use these commands during an incident:
docker service ps opsknight_opsknight-web --no-trunc
docker service inspect opsknight_opsknight-web --pretty
docker service logs --follow --since 10m opsknight_opsknight-web
docker events --since 30m --filter type=service
Look for image-pull errors, unsatisfied placement constraints, health-check failures, database connection exhaustion, and repeated task restarts. Use the health center, system logs, and Prometheus metrics for application diagnosis.
9. Upgrade and roll back
Before an upgrade:
- Read the target release notes and migration requirements.
- Take and verify a PostgreSQL backup.
- Record
docker stack services opsknightand the current image digest. - Pull or mirror the new image on eligible nodes.
- Set
OPSKNIGHT_IMAGEto the new immutable digest and rundeploy.sh. - Repeat the functional checks above and watch queue depth and provider errors.
The web service updates start-first; stateful background roles update stop-first to reduce overlapping ownership. If the application release fails and the database schema remains forward compatible, roll service specifications back:
./deploy/swarm/scripts/rollback.sh
The rollback script detects integrated or split mode, requests a Swarm rollback for each application service, waits, and runs the health checker. It does not reverse database migrations. Never assume an old image is safe with a newer schema; follow database migrations and the release-specific recovery instructions.
10. Back up, restore, and remove
Back up PostgreSQL independently of Swarm and test restoring it into an isolated database. For bundled PostgreSQL, include the node-local volume in host recovery planning, but use a logical or database-native backup as the portable recovery artifact. Follow backup and restore before the first upgrade and on a recurring schedule.
Remove application services with:
docker stack rm opsknight
Stack removal does not automatically prove that database volumes, old content-versioned secrets, backups, or the external overlay network are safe to delete. Inventory each resource, confirm retention requirements, and remove it separately only when its data is no longer needed.
Troubleshooting
The deployment says Swarm is inactive or this is not a manager
Run docker info and docker node ls. Initialize or join the Swarm if inactive,
and run stack operations on a manager. Do not promote an arbitrary node without
considering manager quorum.
Split mode rejects the image
Set OPSKNIGHT_IMAGE to a tested split-compatible tag or immutable digest. The
script intentionally rejects missing images and latest because historical
monolithic images do not recognize split-runtime roles.
A service remains at 0/N
Run docker service ps <service> --no-trunc. Common causes are registry
authentication, a digest unavailable for the node architecture, insufficient
CPU or memory, a missing placement label, an unavailable secret, or failed
health checks. Correct the cause and rerun the orchestrator.
Migration fails
Read the migration task logs printed by the script. Confirm direct database DNS, TLS trust, credentials, privileges, free disk space, and compatible PostgreSQL version. Do not deploy the web tier around a failed migration.
Readiness works locally but not through DNS
Check the load balancer target, published port 3000, health-check path, TLS certificate, public URL variables, forwarded headers, and firewall. A redirect loop usually means the external scheme or host is not being preserved.
Notifications pause after scaling
Check critical-worker replica state, database saturation, queue metrics, provider credentials, and provider rate limits. Scaling web replicas does not increase critical-worker throughput. Avoid repeatedly redeploying until the failed lane is identified.
Production acceptance checklist
- Every node is ready, manager quorum is healthy, and one worker can be drained without losing application availability.
- The deployed image is an immutable, tested digest;
latestis not used. - Production secrets are non-default, stored as Swarm secrets, and recoverable from the approved secret manager.
- External PostgreSQL uses verified TLS, has tested backups, and has adequate connection headroom.
- The public HTTPS origin, forwarded headers, streaming behavior, and readiness probe work through the real load balancer.
- All services converge, a test incident pages successfully, and duplicate work is not observed during a node drain.
- Upgrade and application rollback have been rehearsed, with the database migration compatibility boundary understood.
The historical Swarm measurements associated with PR #777 are not certified. Use benchmark results only within the stated test conditions; measure your own database, network, provider, and workload before setting production capacity.
Last updated for v2.0.0
Edit this page on GitHub