OpsKnightDocumentation
⌘K
Start here
Concepts
Guides
Integrations
Operate OpsKnight
  • Deployment and capacity
  • Choose and deploy an OpsKnight topology
  • Data operations
  • Reliability operations
    • Health checks and metrics
    • Operate with system logs
    • Scale OpsKnight
    • Scrape OpsKnight with Prometheus
    • Use the Health Center
  • Security operations
  • Upgrade operations
Reference
Troubleshooting
Develop OpsKnight
HomeChangelogGitHub
Home/Docs
⌘K
v2.0.0
DocsOperateReliability
Operate1 min read

Reliability operations

Monitor health, metrics, capacity, and system logs in production.

Use health checks and metrics for probes, Prometheus for scraping, Health Center for application diagnostics, system logs for investigation, and scaling for capacity changes. Alert on user-impacting symptoms and queue latency, not only process availability.

PreviousMaintain data and retentionNextHealth checks and metrics

Last updated for v2.0.0

Edit this page on GitHub