Database or PgBouncer connections fail
Diagnose URLs, TLS, authentication, pool mode, and connection budgets.
Test name resolution and TCP reachability from the failing runtime, then validate URI encoding, database name, credentials, TLS mode, CA material, and PgBouncer user list. Compare pool mode and direct migration URL requirements. Sum configured connections across replicas and workers; saturation can resemble an authentication failure. Never print full database URLs in shared logs.
Diagnostic sequence
- Identify whether the failure affects the web, scheduler, migration, or worker role.
- From that container or pod, resolve the database hostname and test TCP port 5432.
- Run the readiness endpoint and inspect the first database error in runtime logs.
- Compare
DATABASE_URLwithDIRECT_DATABASE_URL; migrations must not use a transaction-pooling endpoint when a direct endpoint is required. - Check active connections, reserved capacity, and the sum of per-replica limits.
Run tests from the failing runtime namespace or container; a successful test from an operator laptop does not prove the application path works:
getent hosts <database-host>
nc -vz <database-host> 5432
pg_isready -h <database-host> -p 5432 -d <database-name> -U <database-user>
Do not put the password on the command line. If the runtime image does not carry these clients, use an approved diagnostic container with the same DNS, NetworkPolicy, service account, and CA mount.
Classify the first error before changing configuration:
| Error family | Inspect | Repair boundary |
|---|---|---|
| name not known | service name, search domain, external DNS | DNS or hostname |
| connection refused | target service/endpoints, database listener | service or database availability |
| timeout | egress policy, firewall, route, exhausted accept queue | network or capacity |
| password authentication failed | username, database, secret revision, PgBouncer user list | credentials |
| certificate verify failed | server name, CA bundle, certificate chain and expiry | TLS trust |
| too many connections / pool timeout | active sessions and aggregate replica budget | connection budget |
| prepared statement / transaction state error through PgBouncer | pool mode and Prisma URL parameters | pooling compatibility |
When PgBouncer is deployed, test both paths independently. DATABASE_URL is the
pooled application path; DIRECT_DATABASE_URL is the direct administrative and
migration path. A healthy pooled query does not prove that migrations can obtain
the session semantics or locks they require.
An authentication error points to credentials, database name, or PgBouncer user
configuration. A timeout points to DNS, routing, NetworkPolicy, TLS, or exhausted
connection capacity. A certificate error points to CA material or sslmode, not
application credentials.
Recover and verify
Correct the secret or connection budget, restart only affected roles, and confirm readiness plus a successful migration-health check. Preserve redacted URLs, runtime role, PostgreSQL/PgBouncer logs, and connection counts for escalation.
Observe at least two readiness intervals and one real read/write workflow. For a pool-capacity fix, also confirm waiting clients return to baseline instead of merely increasing the database ceiling until the next surge.
Last updated for v2.0.0
Edit this page on GitHub