Choose a deployment topology
Select Integrated Compose, Split Compose, Swarm, Helm, or Kustomize from operational requirements.
Do not select Split mode from an undocumented requests-per-second threshold. Select it when independent scaling, failure isolation, or notification-lane protection is an operational requirement.
Deployment Planner
Use the closest profile as a conversation starter, then test your actual alert and delivery mix. The profiles come directly from the current load-fixture source; they are workload shapes, not certified capacity limits and not proof that a topology will support every organization of that size.
| Profile | Users | Services | Integrations/service | SSE sessions | Status subscribers | Starting architecture to evaluate |
|---|---|---|---|---|---|---|
| Small | 40 | 12 | 2 | 100 | 1,000 | Integrated Compose when one host and no HA are acceptable |
| Medium | 120 | 32 | 4 | 500 | Split Compose on one host, or Helm Split when Kubernetes/HA is required | |
| Large | 400 | 80 | 5 | 2,500 | Helm/Kustomize Split with PgBouncer and external PostgreSQL | |
| Storm | 1,000 | 200 | 5 | 5,000 | Multi-replica Split, PgBouncer, external HA PostgreSQL, and workload-specific certification |
Small-profile starting point
Integrated Compose has the lowest operating cost: one runtime ownership model, one host, and straightforward proxy and backup operations. Move to Split before bulk work, provider delays, or independent worker scaling must be isolated from critical paging. Forty users is not a Compose limit.
Medium-profile starting point
Use Split Compose when a single host remains acceptable but notification lanes need separate ownership. Use Helm Split when your organization already operates Kubernetes or requires multi-node scheduling, disruption controls, and independent replicas. Validate the PostgreSQL connection budget in either case.
Large-profile starting point
Start evaluation with Split roles, PgBouncer transaction pooling, and an external PostgreSQL service. Independent Web/worker scaling, connection control, failure isolation, and disruption budgets usually matter more at this shape. This is an operational recommendation, not a claim that 400 users require Kubernetes.
Storm or mission-critical starting point
Use multi-replica Split roles, PgBouncer, and external HA PostgreSQL, then run the certification suite with your alert bursts, responder concurrency, provider quotas, status fanout, retention, CPU/RAM, replicas, and database pools. Do not derive a production RPS promise from the historical benchmark.
Inputs that can change the answer
User count alone is insufficient. Record all of these before choosing:
- normal and burst alert-ingestion rate, deduplication shape, and service count;
- notifications generated per incident and each provider's account/destination limits;
- simultaneous responders, API clients, and SSE/realtime sessions;
- public status subscribers, update frequency, and delivery fanout;
- recovery-time, multi-node availability, and maintenance requirements;
- Web and worker replica counts, per-role pools, and PostgreSQL
max_connections; - operator ownership for Docker/Swarm/Kubernetes, backups, monitoring, and upgrades.
If a single constraint is uncertain, choose the simpler topology that still meets availability requirements, measure it, and keep a documented migration trigger. If critical and bulk work compete, database connections approach the budget, or a single process/host is an unacceptable failure domain, evaluate Split or multi-node operation before increasing raw concurrency.
| Topology | Best fit | Multi-node HA support | Independent scaling | PgBouncer support | Capacity status |
|---|---|---|---|---|---|
| Integrated Compose | Simplest single-host operation | No | No | No | Not certified |
| Split Compose | Single-host workload isolation | No | Yes | Optional | Not certified |
| Split + PgBouncer | Isolation plus web connection pooling | No | Yes | Enabled | Not certified |
| Swarm HA | Docker-native multi-node operation | Yes | Yes | Optional in Split | Not certified |
| Kubernetes Helm | Packaged, schema-validated Kubernetes | Supported when configured | Yes | Optional in Split | Not certified |
| Kubernetes Kustomize | GitOps and environment overlays | Supported when configured | Yes | Optional in Split | Not certified |
Decision path
- One host and minimal operational overhead: use Integrated Compose.
- Independent scheduler and queue-lane scaling: use Split runtime.
- Many web replicas or roles threaten the PostgreSQL connection budget: add PgBouncer in supported Split mode after calculating the budget.
- Docker-native multi-node HA: use Docker Swarm.
- Existing Kubernetes platform: choose Helm for packaged values or Kustomize for overlay ownership.
Why Split exists
Integrated mode runs the web application and background work together. Split mode separates web, scheduler, general, critical, bulk, and status-projector roles. This allows a bulk backlog or provider slowdown to be isolated from critical paging and lets each constraint scale independently.
Validate the choice
- Calculate the database connection budget.
- Run readiness, a real alert-to-acknowledgement journey, and provider tests.
- Exercise the closest current fixture profile, then your own burst/fanout mix.
- Watch latency, errors, oldest queued age, lane starvation, provider throttling, SSE stability, CPU/memory, and PostgreSQL connections.
- Record the exact image, product/harness revision, fixture dimensions, topology, replicas, pools, host resources, passing level, and first failing level. Promote only the exact tested configuration.
See Benchmark results for historical observations and Certification methodology for the proof contract.
Last updated for v2.0.0
Edit this page on GitHub