Incident lifecycle action is stuck
Diagnose an acknowledgement, assignment, escalation, or resolution that does not converge.
Refresh the canonical incident, inspect the latest timeline and audit entries, and identify whether the request failed before or after the database mutation. Check authorization, current status, idempotency key, outbox state, and worker health. Do not directly edit lifecycle columns. Preserve the incident identifier, request identifier, actor, attempted transition, and timestamps.
Determine where progress stopped
- Reload the incident and compare its canonical status with the action response.
- Inspect timeline and audit entries for the same actor and request identifier.
- Check the HTTP response:
403indicates authorization,409indicates a stale/conflicting transition, and5xxrequires server and worker inspection. - Inspect queued side effects only after confirming the database transition.
| Canonical status | Audit/timeline | Downstream work | Conclusion |
|---|---|---|---|
| unchanged | no accepted action | none | authorization, validation, or conflict before mutation |
| changed | matching audit entry | pending | mutation succeeded; diagnose the responsible worker lane |
| changed | matching audit entry | complete | client/cache is stale; reload canonical incident |
| changed unexpectedly | actor/request does not match | another responder or automation won the race |
Before retrying, copy the incident ID and reload it in a new request. A 409
normally means the command was based on stale state, not that the database is
stuck. If another responder already acknowledged or resolved it, do not force the
old transition.
For ChatOps actions, compare the button response with the canonical web incident and audit actor. Expired or duplicate interaction callbacks should not be replayed manually. For API clients, reuse an idempotency key only for an exact retry of the same logical command; a different action needs a different key.
If status changed but notifications did not, follow the notification runbook; repeating the lifecycle action can create unnecessary downstream work. If status did not change, verify the requested transition is legal from the current state and repeat once with a new operator action after correcting the cause.
Verify recovery
Confirm one canonical transition, one matching audit entry, and eventual message, status-page, and ChatOps convergence. Escalation evidence should include incident ID, old/new status, request ID, actor, response code, and relevant worker errors.
Last updated for v2.0.0
Edit this page on GitHub