Operate1 min read
Reliability operations
Monitor health, metrics, capacity, and system logs in production.
Use health checks and metrics for probes, Prometheus for scraping, Health Center for application diagnostics, system logs for investigation, and scaling for capacity changes. Alert on user-impacting symptoms and queue latency, not only process availability.
Last updated for v2.0.0
Edit this page on GitHub