OPERATIONS

The platform has to survive the incident too.

Choose integrated or split runtimes and inspect worker health, delivery operations, metrics and logs.

WHAT IT SOLVES

Scale the responsibility that is constrained.

Choose an integrated process or separate Web, Scheduler, General Worker, Critical Worker, Bulk Worker and Status Projector. A split runtime alone does not provide high availability.

Understand the capability
01Health Center and worker visibility
02Prometheus metrics and structured logs
03Independently scalable runtime roles
OPSKNIGHT / PRODUCT VIEW Northstar Systems · v2.0.0
OpsKnight operations product view
OpsKnight operations product view

HOW IT WORKS

A concrete operational path, not a feature list.

RUNTIME RESPONSIBILITIESSplit only the responsibility that needs independent scale or failure isolation.
01Web
02Scheduler
03Critical
04General
05Bulk
06Status projector
PgBouncerPostgreSQLHealth + Prometheus

EXPLORE A WORKFLOW

  1. 01Web + background work
  2. 02PostgreSQL
  3. 03Health Center
  4. 04Metrics & logs

The integrated runtime is the simplest operating model. Monitor database capacity, provider health and background work together.

OPERATIONAL DEPTH

Operations, beyond the happy path.

The details below are the parts teams need when evaluating how the capability behaves during real response, failure, and handoff.

01

Workers, queues and providers tell different stories.

Inspect worker heartbeat, oldest pending job age and provider state together. A growing critical queue has a different operational priority from delayed bulk work.

Read the operational guide
02

Budget database connections before adding replicas.

Include each role’s replica count and pool size in the connection budget. More workers can exhaust database capacity instead of improving queue latency.

Read the operational guide
03

Use PgBouncer where the topology calls for it.

Follow the documented transaction-pooling and direct-connection requirements. Schema management and runtime traffic must use the connection paths appropriate to their work.

Read the operational guide
04

Health Center for investigation. Metrics for monitoring.

Use the Health Center to inspect recorded diagnostics and operational priorities. Connect documented health and Prometheus metrics to your own monitoring, and interpret component status rather than only HTTP status.

Read the operational guide
05

Make scale and recovery a tested plan.

Scale the constrained role only after checking database and provider capacity. Plan backups, upgrades, ingress and recovery together for production availability.

Read the operational guide

KNOW BEFORE PRODUCTION

Validate the boundary, not just the happy path.

Use a test service and representative provider configuration before treating operations as production incident infrastructure.

  1. 01Budget PostgreSQL connections before adding worker or web replicas.
  2. 02Test backup, restore, and the documented upgrade path before production changes.
  3. 03Run a synthetic incident after topology or recovery changes before declaring the platform healthy.

RUNTIME OPTIONS

Choose how the pieces run.

Docker Compose · Integrated Runtime
v2.0.0 CONTRACT
INGRESS / NETWORK BOUNDARY
External Webhook & Client Ingress (HTTPS :443)TLS termination · reverse proxy / ingress controller
APPLICATION RUNTIME TIER · INTEGRATED1 container
Web HTTP / UI
User interface, REST API, webhook endpoints
Scheduler & Rotations
On-call shifts, escalation timers, cron
Notification Workers
Voice, SMS, push, ChatOps dispatch
Status Projector
Public service health page publishing
DURABLE STATE LAYER
PostgreSQL 16+Durable state, relation store & audit ledgers
DOCUMENTATIONSetup, authorization, limits, and troubleshooting.
Explore operations documentation

Your incidents should belong to you.

Run OpsKnight on infrastructure you control.

v2.0.0 · AGPL-3.0-only · Self-hosted · 28 inbound integrations