Skip to content

PostgreSQL disappears

PostgreSQL is essential to creating a durable mapping. Its failure should make that operation unavailable, rather than encourage the application to manufacture success from memory. The resilience hypothesis is therefore controlled degradation and recovery: storage-dependent routes return explicit availability failures, the application process stays responsive, and the database becomes ready again inside the bounded experiment window.

This is a different contract from Redis fallback. The application has no alternative durable store and no write queue that can turn a PostgreSQL outage into a successful creation.

The gate starts from a ready deployed candidate, a ready selected target, and a run-scoped paid load source. The load client checks application /livez and /ready, verifies that an unsigned creation produces an x402 challenge, and checks enough signer slots and cached funds for its configured virtual users. It emits the gate’s readiness marker only after a paid 201 and successful redirect.

That marker proves useful pre-fault work in the current run. It does not establish that every later request is paid, every request succeeds, or database outage leaves creation possible. The actual fault and subsequent behavior still need their own measurements.

The PostgreSQL PodChaos definition requests one primary pod failure for 60s, scoped to the staging namespace and the staging Helm instance. It does not test data deletion, persistent-volume corruption, restore, failover to another primary, or a long database outage.

Follow the routes when storage stops answering

Section titled “Follow the routes when storage stops answering”

POST /shorten first asks the database whether the destination exists. If that lookup fails, the route returns 503 with database unavailable. A new paid request does not proceed to settlement when this initial storage failure is already known.

Storage can also fail later. If a paid request passes the lookup and settles before its insertion fails, it still returns 503, but the payment side effect may already exist. This post-settlement gap is a consistency limit; it is not repaired by a process restart or by the database fault score. The paid request narrative explains its boundaries.

GET /{code} tries Redis first. A cached destination can still return 302 without querying PostgreSQL. A missing or unavailable cache needs a database lookup and returns 503 if storage is unavailable. Thus “database failure makes every redirect fail” would be inaccurate. The readiness contract is broader: /ready always pings PostgreSQL and becomes 503 regardless of whether one popular cached code remains resolvable.

The database wrapper closes a pool after handled PostgreSQL, operating-system, or timeout errors. A subsequent readiness ping can reconnect and reinitialize the idempotent schema. /livez remains process-only. This is the application recovery mechanism; it does not require dependency failure to become a liveness failure.

Read the checks according to what they measure

Section titled “Read the checks according to what they measure”

The PostgreSQL scorecard contains five behavioral checks plus release identity validation. It requires non-probe request increase greater than fifty, application restart increase below 0.5, a target readiness minimum of zero, a final target readiness sample of one, and /shorten HTTP 500 increase below 0.5.

The clean-degradation check is deliberately specific to status 500. It does not forbid all 5xx: expected storage 503 responses are allowed. Nor does that check independently prove every returned 503 contains the exact intended body. Dependency error tests exercise the route’s 503 mapping directly, while the live scorer checks the narrower measured contract.

The traffic counter excludes /livez, /ready, and /metrics but includes responses regardless of success status. It establishes application-facing work beyond probes, rather than successful URL throughput during essential storage loss. A database fault can legitimately produce many counted unavailable requests.

Restart and 500 expressions include or vector(0), so absence of those raw series may produce a valid zero. Target readiness does not use that fallback. The scorer’s collection validity checks and final-sample rules remain necessary to avoid calling missing recovery evidence a pass. The metrics reference gives exact query shapes and units.

Recovery belongs to the same candidate and run

Section titled “Recovery belongs to the same candidate and run”

The retained 1 October PostgreSQL card passes its bounded checks and names the immutable candidate and chaos-gate-fkh9p run. The historical run index also records enclosing workflow and cleanup evidence. This example is historical staging-testnet evidence; it is separate from the later release verification and from production-like smoke.

Dependency readiness returning to one satisfies the scored recovery criterion. Application readiness reconnecting, useful requests resuming, and successful cleanup strengthen the operational interpretation, but each must be tied to its actual observations. A passing per-experiment card cannot substitute for the enclosing gate Job’s result or cleanup.

Inspect application storage code, recovery tests, and PostgreSQL scoring tests. Continue with Redis failure to compare an acceleration dependency, or health and recovery to understand why liveness and readiness remain distinct.

Maintained by Satyam Agnihotri · DevOps & Cloud Engineer