Redis disappears
A cache can make an application faster while making a dependency failure harder to interpret. A redirect that succeeds during Redis failure may have fallen back correctly, or it may simply have arrived after recovery. A growing cache-miss counter may reflect an outage, or ordinary first-time reads. Resilience Gate asks for several observations because no single one proves its Redis hypothesis.
The hypothesis is bounded: once the application has achieved initial readiness, a selected staging Redis pod can become unavailable for 60s while the application resolves mappings through PostgreSQL, stays within a redirect latency limit, avoids application restarts, and observes Redis recovery.
Enlarge diagram: Redis outage removes the fast cache path while redirects use PostgreSQL and later cache fills resume · Version-controlled diagram source
Start from an established application
Section titled “Start from an established application”Readiness matters before the fault. /ready always checks PostgreSQL and checks Redis until the process’s first successful readiness. After that success, has_been_ready permits steady-state service without Redis. This allows the cache fault to exercise degradation while the application remains usable.
A cold application process behaves differently. If it starts during Redis failure, it cannot pass that initial cache check even if PostgreSQL is working. The startup probe also calls /ready, so initial unavailability can exhaust its startup budget. The resilience claim is therefore about the established application under this scenario, not an assertion that every new replica starts without Redis. The health explanation covers the lifecycle trade-off.
The orchestration checks ready targets and requires useful paid traffic from its own load Job before injection. Its safe marker follows a successful paid creation and redirect. That precondition establishes that the traffic source exercised the path before the experiment; it does not guarantee every later iteration succeeds throughout the fault.
Remove the cache path and follow one redirect
Section titled “Remove the cache path and follow one redirect”The Redis PodChaos source requests pod-failure, mode: one, and duration: 60s. Its selector requires the staging namespace, the staging Helm instance, and Redis’s master component. It targets neither every Redis pod nor the production-like namespace. The environments still share a lab cluster, so these selectors do not imply complete infrastructure isolation.
The redirect route attempts GET url:<code>. A truthy cached destination increments url_shortener_cache_hits_total and returns 302. If no usable value is present, or the cache wrapper raises CacheUnavailable, the route increments url_shortener_cache_misses_total and queries PostgreSQL.
After a matching database row is found, a SETEX attempts to populate the cache. Cache write errors are ignored for the response; the authoritative lookup has already succeeded. The route returns 302 with Location. An unknown mapping returns 404, and a failed required database lookup returns 503.
The fallback is deliberately simple. Redis socket timeouts are short, and cache operations translate failures into one bounded application exception. PostgreSQL bears the additional work. This saves availability for the exercised request path while sacrificing cache speed and increasing load on essential storage.
A cache miss proves fallback, not an outage
Section titled “A cache miss proves fallback, not an outage”Normal first redirects are expected misses because URL creation does not populate Redis. Expiry, an absent code, and read failures also take the miss path. The counter’s help text explicitly says “Redis misses or unavailable cache fallbacks.” No separate miss label identifies the cause.
This matters especially for paid load: the client creates a new destination for each iteration and then resolves its new code. Even healthy Redis therefore produces many first-read misses. The scorer names its miss check redis-fallback-observed, and correctly establishes only that the fallback path was exercised.
Proof of the fault must come from the target’s readiness series and applied workflow state in the same run. The cache-miss increment is an additional behavioral observation. The app’s Redis dependency gauge is also insufficient: after initial readiness, /ready stops checking Redis, so that gauge may retain an earlier 1 through an outage.
Read the seven-check scorecard as one argument
Section titled “Read the seven-check scorecard as one argument”The scorer adds release identity validation to six Redis behavior checks:
| Check | Required observation | Interpretation |
|---|---|---|
| Meaningful traffic | Non-probe HTTP increase > 50 across fault plus settle duration. |
Requests reached the application; the count alone does not establish every request succeeded. |
| Fallback path | Cache-miss increase > 0 over that duration. |
Database fallback was exercised; it does not establish why Redis missed. |
| Redirect latency | Maximum observed one-minute histogram p95 < 1.2 seconds. |
The degraded resolution path stays below the chosen measured latency bound. |
| Application restarts | Maximum container restart increase < 0.5. |
No application restart is observed by the configured query. |
| Target outage | Minimum Redis pod readiness == 0 in the range. |
The selected dependency was observed unavailable. |
| Target recovery | Final Redis readiness sample == 1. |
The dependency recovered before the observation ended. |
| Release identity | Immutable candidate plus bounded run identity validate. | The result belongs to a specific candidate and run. |
The latency rule queries the /{code} handler and aggregates histogram buckets across matching application series before calculating the p95. The scorer takes the maximum of the resulting time samples. It is not the single worst request, an end-to-end payment duration, or a promise that every redirect completes under 1.2 seconds.
The restart expression deliberately includes or vector(0). An absent raw restart series can therefore produce a valid zero result. Required readiness and fallback series do not receive the same fallback. Collection validity and the exact query expression both matter; see metrics and scoring.
A publicly inspectable historical example
Section titled “A publicly inspectable historical example”The retained 1 October staging record contains a passing Redis scorecard. Its observation starts at 17:49:27 UTC, uses an actual injection anchor at 17:49:47, and ends at 17:51:47, combining 60s of fault and 60s settle time after injection.
That card records approximately 293.714 non-probe requests, 98.286 fallback increments, a maximum observed p95 of 0.32 seconds, and zero observed application restarts. Readiness reaches zero and its final sample reaches one. Counter increases are fractional because Prometheus extrapolates increase; they are not literal fractional HTTP requests. The scorecard binds the full image digest, revision, and run identifier.
This is historical staging evidence for the named candidate. It is not the later v1.0.0 release record, a production-like Redis test, or permission to transfer the pass to any new image. The historical cases and recorded release keep those campaigns separate.
Recovery and the remaining capacity question
Section titled “Recovery and the remaining capacity question”Target readiness returning to one establishes the scorer’s recovery condition. Subsequent cache fills and hits support renewed acceleration, but they are not separate required scored checks. Cleanup and the enclosing Job must also succeed before the gate result is usable; a passing Redis card alone does not approve promotion.
Fallback demonstrates a chosen degradation path, not unlimited database capacity. A larger workload, longer cache outage, simultaneous PostgreSQL failure, or a wave of cold replicas could behave differently. The repository’s cache tests establish deterministic fallback and cache population; Redis scoring tests establish rule behavior. Neither supplies a new live capacity measurement.
Maintained by Satyam Agnihotri · DevOps & Cloud Engineer