THE PROBLEM
ON-CALL OPERATIONS
Make responsibility explicit.
Build schedules, apply overrides and route escalation to the responder responsible for the service.
A PRACTICAL START
Create rotations, validate current coverage, configure escalation targets and confirm delivery outcomes on a test incident.
PRODUCT SURFACES
Assemble only what the workflow needs.
On-call & escalation
Build rotations, schedule overrides and escalation policies around the people responsible for each service.
Paging
Route urgent pages through configured providers and inspect attempts, retries and terminal delivery outcomes.
Mobile
Install the progressive web app and acknowledge, triage and follow incidents from your phone.
IMPLEMENTATION CONTEXT
Turn the use case into an operating model.
The product workflow is only useful when deployment, integrations, access controls and validation are planned together.
How to run it
The deployment choice matters less than proving the scheduler, worker health, database durability, and notification-provider reachability during an incident.
What to connect first
Map services to the schedules and escalation policies that actually own them. Enable only the paging channels responders are expected to use and test each configured provider.
What to decide up front
Limit schedule and escalation administration to the roles that need it, and protect notification-provider credentials with the same production secret controls as the rest of the platform.
What to prove before production
Test current coverage, an override, a triggered incident, each required delivery channel, acknowledgement, retry visibility, and escalation behavior.
IMPLEMENTATION PATH
Validate the whole journey before production.
Start with the deployment model, connect the required providers, configure identity and routing, then exercise a synthetic incident from signal to responder action and recovery.
Your incidents should belong to you.
Run OpsKnight on infrastructure you control.
v2.0.0 · AGPL-3.0-only · Self-hosted · 28 inbound integrations