ON-CALL OPERATIONS

Make responsibility explicit.

Build schedules, apply overrides and route escalation to the responder responsible for the service.

01

THE PROBLEM

Coverage and handoffs are hard to reason about during a page.

02

A PRACTICAL START

Create rotations, validate current coverage, configure escalation targets and confirm delivery outcomes on a test incident.

PRODUCT SURFACES

Assemble only what the workflow needs.

01

On-call & escalation

Build rotations, schedule overrides and escalation policies around the people responsible for each service.

02

Paging

Route urgent pages through configured providers and inspect attempts, retries and terminal delivery outcomes.

03

Mobile

Install the progressive web app and acknowledge, triage and follow incidents from your phone.

IMPLEMENTATION CONTEXT

Turn the use case into an operating model.

The product workflow is only useful when deployment, integrations, access controls and validation are planned together.

DEPLOYMENT

How to run it

The deployment choice matters less than proving the scheduler, worker health, database durability, and notification-provider reachability during an incident.

INTEGRATIONS

What to connect first

Map services to the schedules and escalation policies that actually own them. Enable only the paging channels responders are expected to use and test each configured provider.

SECURITY & OPERATIONS

What to decide up front

Limit schedule and escalation administration to the roles that need it, and protect notification-provider credentials with the same production secret controls as the rest of the platform.

VALIDATION

What to prove before production

Test current coverage, an override, a triggered incident, each required delivery channel, acknowledgement, retry visibility, and escalation behavior.

IMPLEMENTATION PATH

Validate the whole journey before production.

Start with the deployment model, connect the required providers, configure identity and routing, then exercise a synthetic incident from signal to responder action and recovery.

Your incidents should belong to you.

Run OpsKnight on infrastructure you control.

v2.0.0 · AGPL-3.0-only · Self-hosted · 28 inbound integrations