Metrics reference
Metric types, labels, scope, aggregation, and alerting guidance for the OpsKnight Prometheus endpoint.
GET /api/metrics exposes Prometheus text format after Admin-session or Bearer-token authentication. See Scrape OpsKnight with Prometheus for setup and security.
Every emitted family includes # HELP and # TYPE. Labels are deliberately bounded: incident, user, service, team, request, email, phone, IP, and other high-cardinality identifiers are forbidden by the registry.
Scope and aggregation
cluster_snapshot: a database-derived current value. When every web replica scrapes the same database, do not sum replicas; use one scrape target ormax without(instance).instance: state held by one process. Preserveinstance; sum only when a cluster total is meaningful.counter: monotonic process counter. Userate()orincrease()and normally sum across instances by the documented bounded labels.histogram: use_bucket,_sum, and_count; aggregate buckets bylebeforehistogram_quantile.
Core platform and queue metrics
| Metric | Type | Labels | Scope | Meaning and use |
|---|---|---|---|---|
opsknight_build_info |
gauge | version |
instance | Value 1 for the running build; alert when unexpected versions coexist after a rollout. |
opsknight_active_incidents |
gauge | — | cluster snapshot | Current active incident count; do not sum replicas. |
opsknight_active_users |
gauge | — | cluster snapshot | Current active account count; capacity context, not login activity. |
opsknight_job_queue |
gauge | status |
cluster snapshot | Legacy durable-job count grouped by bounded status. |
opsknight_jobs_pending |
gauge | type |
cluster snapshot | Pending durable work by job type. Alert on sustained growth. |
opsknight_jobs_processing |
gauge | type |
cluster snapshot | Work currently processing by type. Compare with worker availability. |
opsknight_jobs_oldest_pending_age_seconds |
gauge | type |
cluster snapshot | Age of the oldest pending job. This is usually a stronger paging signal than depth alone. |
opsknight_notifications_undelivered |
gauge | — | cluster snapshot | Pending plus retryable failed notification records. |
opsknight_notifications_oldest_undelivered_age_seconds |
gauge | — | cluster snapshot | Age of the oldest undelivered notification. Alert against the delivery objective. |
opsknight_escalations_overdue |
gauge | — | cluster snapshot | Incidents whose next escalation execution is overdue. |
opsknight_escalation_max_lag_seconds |
gauge | — | cluster snapshot | Largest current escalation delay. Treat sustained non-zero lag as paging-path risk. |
opsknight_rollup_freshness_age_seconds |
gauge | — | cluster snapshot | Age of the newest daily analytics rollup update. |
HTTP and collection metrics
| Metric | Type | Labels | Scope | Meaning and use |
|---|---|---|---|---|
opsknight_http_requests_total |
counter | method, route, status_class |
counter | Completed requests on normalized routes. Use a rate grouped by status class. |
opsknight_http_request_duration_seconds |
histogram | method, route |
counter | Request latency. Aggregate buckets across instances before calculating quantiles. |
opsknight_http_requests_in_flight |
gauge | route |
instance | Requests currently executing on one instance. |
opsknight_metrics_collection_errors |
gauge | — | instance | Number of collectors that failed in the current cached snapshot; alert when greater than zero. |
opsknight_metrics_cache_hits_total |
counter | — | counter | Metrics snapshot cache hits. |
opsknight_metrics_cache_misses_total |
counter | — | counter | Metrics snapshot cache misses. |
opsknight_metrics_cache_age_seconds |
gauge | — | instance | Age of the process-local snapshot; normal cache is 10 seconds and degraded cache is 60 seconds. |
Integration and provider metrics
| Metric | Type | Labels | Scope | Meaning and use |
|---|---|---|---|---|
opsknight_external_operations |
gauge | status |
cluster snapshot | Durable outbound/external work grouped by status. |
opsknight_chatops_intents |
gauge | status |
cluster snapshot | Durable ChatOps intents grouped by status. |
opsknight_inbound_deliveries |
gauge | status |
cluster snapshot | Durable inbound-provider deliveries grouped by status. |
opsknight_provider_cooldown |
gauge | provider |
cluster snapshot | 1 when a provider key class has active distributed cooldowns. |
opsknight_integration_reconciliations_total |
counter | kind |
counter | Expired or ambiguous integration work reclaimed. |
opsknight_provider_rate_limits_total |
counter | provider |
counter | Provider rate-limit responses observed by ChatOps/collaboration paths. |
opsknight_provider_permission_failures_total |
counter | provider |
counter | Provider permission failures. |
opsknight_notification_provider_429_total |
counter | provider |
counter | Notification delivery attempts rejected with provider throttling. |
opsknight_notification_queue_depth |
gauge | traffic_class, provider, channel |
cluster snapshot | Ready notification work by bounded delivery dimensions. |
opsknight_notification_oldest_age_seconds |
gauge | traffic_class |
cluster snapshot | Oldest ready notification by traffic class. |
opsknight_notification_throughput_per_second |
counter | provider, traffic_class |
counter | Provider acceptances; despite the name, apply rate() to the counter. |
opsknight_notification_effective_rate |
gauge | provider, channel |
instance | Current adaptive delivery rate on an instance. |
opsknight_notification_admission_deferred_total |
counter | reason, traffic_class |
counter | Work deferred by provider admission control. |
ChatOps, war-room, and meeting families
ChatOps counters use bounded provider, verb, and/or result labels:
opsknight_chatops_invokes_totalopsknight_chatops_identity_resolution_totalopsknight_chatops_authorization_denied_totalopsknight_chatops_duplicate_totalopsknight_chatops_refresh_totalopsknight_chatops_action_latency_seconds(histogram byprovider)
War-room gauges are opsknight_war_room_health, opsknight_war_room_state, opsknight_war_room_projection_lag_seconds, opsknight_war_room_participant_drift, and opsknight_external_cleanup_pending. Counters are opsknight_war_room_participant_sync_total, opsknight_war_room_reconciliation_total, opsknight_war_room_ambiguous_card_abandon_total, and opsknight_war_room_projection_total.
Meeting gauges are opsknight_meeting_state, opsknight_meeting_health, and opsknight_meeting_cleanup_pending. Provision, close, retry, and reconciliation use opsknight_meeting_provision_total, opsknight_meeting_close_total, opsknight_meeting_retry_total, and opsknight_meeting_reconciliation_total; provision/close duration families are histograms.
Realtime, status-page, and fanout families
- Realtime instance state:
opsknight_realtime_subscribers{stream},opsknight_realtime_observed_generation,opsknight_realtime_change_age_seconds, and counteropsknight_realtime_clock_errors_total. - Snapshot health:
opsknight_status_page_snapshot_dirty,opsknight_status_page_snapshot_oldest_age_seconds,opsknight_status_page_snapshot_bytes, andopsknight_status_snapshot_revision_lag. - Publication health:
opsknight_status_page_publication_failed,opsknight_status_page_fail_closed, plus publication duration/attempt counters labelled by bounded change class and outcome. - Fanout:
opsknight_status_page_fanout_total,opsknight_status_fanout_campaign_total,opsknight_status_fanout_materialized, andopsknight_status_fanout_failed. - Serving:
opsknight_status_serving_store_latency_seconds,opsknight_status_serving_store_errors_total,opsknight_status_page_stale_serves_total, andopsknight_status_page_revocations_total.
SLA, response-policy, and compliance families
SLA scheduler state is exposed through opsknight_sla_scheduler_mode{mode} (one active series has value 1), opsknight_sla_scheduler_null_hints, opsknight_sla_scheduler_due, opsknight_sla_scheduler_shadow_mismatch_total, opsknight_sla_scheduler_config_read_failures_total, opsknight_sla_hint_repairs_total, and histogram opsknight_sla_transition_lag_seconds{kind}.
Response-policy/classification counters include opsknight_incident_classification_total, opsknight_response_policy_conflicts_total, opsknight_response_policy_resolution_errors_total, opsknight_incident_resolution_unknown_total, opsknight_sla_projection_invalid_total, opsknight_engagement_deferred_total, and opsknight_response_policy_restore_total.
Compliance exposes monitor/evaluation/drift counters, opsknight_compliance_drift_open{kind}, histogram opsknight_compliance_drift_projection_lag_seconds, and opsknight_compliance_monitor_last_success_unixtime. Alert on an old last-success timestamp together with failed monitor runs; a zero open-drift gauge alone does not prove the monitor is running.
Example PromQL
# HTTP 5xx rate by route
sum by (route) (rate(opsknight_http_requests_total{status_class="5xx"}[5m]))
# Oldest critical notification work
max(opsknight_notification_oldest_age_seconds{traffic_class="critical"})
# p95 request latency across replicas
histogram_quantile(0.95,
sum by (le, route) (rate(opsknight_http_request_duration_seconds_bucket[5m])))
# Any failed collector
max(opsknight_metrics_collection_errors) > 0
# Cluster escalation lag without double-counting replicas
max without(instance) (opsknight_escalation_max_lag_seconds)
Some runtime series appear only after the associated code path has executed. Absence is therefore not always zero. Use opsknight_metrics_collection_errors, cache age, build information, and scrape health when interpreting a missing family.
Last updated for v2.0.0
Edit this page on GitHub