Operate2 min read

Complete an OpsKnight Swarm installation

Use the maintained orchestrator to validate, create secrets, migrate, deploy, converge, and verify an OpsKnight stack.

Prerequisites

Complete Swarm prerequisites and topology/database configuration. Run from a manager in a protected administrative session with all required values available.

Prepare deployment values

Export the immutable image, public URLs, stable secrets, selected runtime mode, database, optional PgBouncer, stack name, replicas/resources, and capacity ceilings through the approved secret source. Do not commit values.

Install the stack

For default split mode:

sh
./deploy/swarm/scripts/deploy.sh

For integrated mode:

sh
SWARM_RUNTIME_MODE=integrated ./deploy/swarm/scripts/deploy.sh

The script validates manager/capacity, creates network and content-hashed Swarm secrets, starts bundled database when selected, runs one-shot direct migration, deploys with prune, waits for convergence, and checks readiness. Do not replace it with routine manual docker stack deploy because that skips gates.

Verify the installation

sh
docker stack services <stack-name>
docker stack ps <stack-name> --no-trunc
./deploy/swarm/scripts/health-check.sh

Require desired replicas, no recurring task failures, migration exit zero, exclusive topology ownership, and readiness through the external HTTPS load balancer. Confirm load-balancer host, TLS certificate, NEXTAUTH_URL, and NEXT_PUBLIC_APP_URL are the same public origin. Open public HTTPS /setup, verify that exact Application URL, and complete Initial setup. Sign in through the same hostname, confirm Settings → System → App URL, verify an unrelated host returns 421, then run a synthetic incident through notification, acknowledgement, resolution, and status projection.

Operate it in production

Monitor manager quorum, node state, service convergence, role queues/heartbeats, provider/database/pool signals, load balancer, storage, backup, and secrets/certificate expiry. Drain nodes one at a time and observe capacity/disruption.

Troubleshooting

Deploy script stops at preflight: correct the reported image, secret, capacity, node-label, or database issue; do not bypass validation.

Migration service fails: preserve task logs and keep stack rollout blocked until direct database/TLS/privilege/migration succeeds.

Service does not converge: inspect docker service ps --no-trunc, placement, resources, image pull, secrets, network, and health.

External readiness fails: compare an internal task/Service request with load balancer/TLS/proxy behavior.

Change or remove the installation

Use Upgrade and rollback. Before docker stack rm, take a verified backup and understand local volume retention. Stack removal is not database recovery.

Next steps

Last updated for v2.0.0

Edit this page on GitHub