Operate2 min read

Build a capacity and connection budget

Model alert, notification, user, realtime, status-fanout, and PostgreSQL demand without inventing limits.

Size from workload shape and measured saturation. Until a matching benchmark is CERTIFIED, do not convert a measured peak into a production promise.

Prerequisites

Collect a normal peak and incident-storm sample. Record alert rate, deduplication ratio, notifications per incident, concurrent users/SSE streams, status fanout, provider quotas, replica counts, concurrency, pools, and the PostgreSQL ceiling.

Configure the workload model

Record peak and sustained alert ingestion, incident deduplication ratio, notifications per incident, escalation rate, concurrent users, SSE streams, status subscribers/fanout, and provider rate limits. Bursty fanout can constrain workers even when average alert ingestion is modest.

Budget PostgreSQL connections

Without PgBouncer, calculate the sum of each role's replicas multiplied by its pool size, plus migrations, administration, monitoring, and safety headroom. Validate the result with deploy/scripts/validate-runtime-capacity.cjs and keep it below PostgreSQL max_connections.

With supported Split + PgBouncer, web traffic uses the transaction pool while scheduler, workers, projector, and migrations retain direct connections. The web contribution becomes PgBouncer replicas × (default pool + reserve pool), but the direct-role pools must still be counted.

DIRECT_DATABASE_URL exists so migrations and roles requiring direct PostgreSQL session behavior bypass transaction pooling. Do not point migration ownership at the pooled URL.

Run a representative load

Use a non-production environment with the same topology, resource requests, PostgreSQL class, and provider emulators. Increase one load dimension at a time and stop on correctness failure, sustained queue-age growth, connection exhaustion, or provider admission failure.

Validation after deployment

Run a representative synthetic load and observe the signals in Scaling signals. Increase one constrained resource at a time and repeat the same workload before changing the documented budget.

Production considerations

Reserve connections for migrations, administration, monitoring, and failover. Treat retry storms and notification fanout separately from average alert rate, and review the budget after topology, pool, concurrency, or provider changes.

Troubleshooting

Calculated demand exceeds the database ceiling: reduce pools/concurrency, add transaction pooling where supported, or increase the database ceiling only after database validation.

Average load passes but storms fail: size from oldest queue age and correctness during the storm, not daily average throughput.

More workers reduce throughput: inspect locks, connections, provider throttling, and job-claim contention before adding capacity.

Last updated for v2.0.0

Edit this page on GitHub