Skip to content

Health, startup, and recovery

Restarting a live process does not repair a failed database. It can make a dependency outage harder to understand by discarding application state and creating more connection work. The application therefore separates process liveness from dependency readiness, and uses readiness checks as bounded opportunities to reconnect storage.

The three useful questions are: can the process answer, can it currently serve its defined workload, and has a newly started process passed its initial dependency conditions? Those questions have distinct consequences in Kubernetes.

Endpoint Implemented check Meaning of success
App GET / Returns configured service metadata without dependency checks. The root route responds; storage may still be unavailable.
App GET /livez Returns {"status":"alive"} without connecting or querying dependencies. The process is responsive.
App GET /ready, before first success Pings PostgreSQL, then Redis. Both initial readiness dependencies responded.
App GET /ready, after first success Pings PostgreSQL only. Essential storage responded; Redis fallback remains permitted.
Signer GET /livez Returns process-only state. The signer process responds.
Signer GET /health Reads bootstrapped wallet/service state. At least one wallet completed bootstrap; it does not freshly query RPC.

PostgreSQL failure makes app readiness return 503 with database not ready. An initial Redis failure produces 503 with cache not ready. The has_been_ready flag changes only after both initial pings succeed and lives in one application process. It resets on replacement or restart, so a cold process must satisfy Redis again.

Startup contains failures without crashing immediately

Section titled “Startup contains failures without crashing immediately”

During the app lifespan, the database makes a bounded connection and schema initialization attempt. If that attempt raises DatabaseUnavailable, startup continues. Redis connection failure is also tolerated by its wrapper, which leaves the client unavailable. The process can consequently answer /livez while /ready records that it should not yet receive regular service traffic.

The database wrapper reconnects during ping() when its pool is absent. Failed database queries caused by PostgreSQL, operating-system, or timeout errors close the pool so a later readiness attempt can establish a fresh one. Connection establishment and schema execution use the configured dependency timeout; the pool’s command timeout bounds database commands. This is recovery through fresh dependency work, rather than restarting the application to reconstruct its process.

Redis connection and ping are bounded as well, with short socket timeouts. Once a client exists, the route uses it for reads and writes; the wrapper translates cache exceptions rather than propagating them to the user. Once the app has passed readiness, /ready no longer actively assesses Redis recovery. Later cache hits and target readiness are therefore more useful recovery observations than the Redis dependency gauge alone.

Kubernetes and Docker use different contracts

Section titled “Kubernetes and Docker use different contracts”

The deployment template uses /ready for startup and readiness, and /livez for liveness. Current chart defaults set a startup probe every two seconds with a failure threshold of thirty; readiness runs every five seconds with a threshold of six; liveness runs every ten seconds with a threshold of three. Probe timeout is two seconds. These are configured budgets, not a guaranteed end-to-end recovery time: scheduling, network behavior, and dependencies also matter.

If the startup probe continues to fail beyond its budget, Kubernetes can restart that starting container even though the eventual liveness endpoint is process-only. During established service, dependency failure makes readiness unsuccessful without intentionally making liveness fail. Avoid turning “dependency failures do not drive liveness restarts” into a claim that no lifecycle condition can ever cause a restart.

The app Dockerfile and default Compose stack use / for their app health checks. Compose also waits for healthy PostgreSQL and Redis before starting the app. Docker health success therefore has a narrower meaning than Kubernetes readiness. The signer Docker health check calls /health, so it reflects its separate bootstrap readiness contract.

Interpret gauges and recovery tests carefully

Section titled “Interpret gauges and recovery tests carefully”

url_shortener_dependency_up reports the result of the most recent readiness check for each dependency. PostgreSQL is refreshed whenever /ready runs. Redis is set during initial readiness checks and then can retain a historical 1 after Redis fails. It is not a continuous all-dependency health monitor.

Health tests exercise the initial Redis requirement and later tolerance; recovery tests exercise a failed connection followed by successful reconnection and readiness recovery while liveness remains successful. These source-level tests do not measure real Kubernetes recovery timing. The PostgreSQL and Redis workflows connect those mechanisms to bounded live observation rules.

Maintained by Satyam Agnihotri · DevOps & Cloud Engineer