Troubleshoot OpsKnight on Docker Swarm
Diagnose manager, node, migration, secret, network, convergence, database, PgBouncer, and readiness failures.
Before you begin
Capture node/quorum, stack services/tasks, service inspect, recent task logs, images, networks, secrets metadata, and database/provider health. Never print secret contents.
Manager quorum is unhealthy
Check: manager reachability/state and Raft quorum.
Recovery: follow the established Swarm disaster-recovery procedure before any application deployment.
Verify: managers agree and stack/service operations succeed.
Service does not converge
Check: docker service ps --no-trunc, placement constraints, node capacity/labels, image pull, secrets, overlay network, and health checks.
Recovery: correct the exact scheduler/task failure and let the maintained deploy script complete.
Verify: desired/actual replicas match without repeated replacement.
Migration task fails
Check: task logs, direct database route, TLS/CA, credentials/privileges, image revision, and schema state.
Recovery: keep rollout blocked, correct the cause, and rerun the one-shot migration through maintained tooling.
Verify: exit zero precedes application convergence.
Bundled database will not schedule
Check: opsknight.database=true node label, node availability, storage volume, placement, and resources.
Recovery: restore the intended durable node or execute database recovery; do not schedule the service onto an empty volume unintentionally.
Verify: expected database data/readiness is present before application rollout.
Overlay network or load balancer fails
Check: Swarm TCP/UDP ports, MTU, node private addresses, published port, load-balancer target health, proxy headers/TLS/SSE.
Recovery: correct network or proxy infrastructure without exposing database/private pool ports.
Verify: external readiness, sign-in, realtime update, and signed webhook pass.
Queue grows with healthy tasks
Check: role heartbeat, lane age/throughput, provider throttling, database/pool saturation, and correct runtime mode.
Recovery: remove the bottleneck and scale only with capacity headroom.
Verify: oldest age declines and synthetic incident completes.
Next steps
Last updated for v2.0.0
Edit this page on GitHub