Skip to content

Bounded experiments

A chaos experiment is useful when it can distinguish the system’s intended behavior from a plausible failure. Removing a dependency while no useful requests run proves very little. Removing several dependencies at once can make a failure impossible to attribute. Leaving a fault behind turns a test into an uncontrolled operating condition.

Resilience Gate addresses those problems with one bounded, serial staging workflow, run-specific object tracking, and explicit observation requirements.

Each current dependency has a different role. PostgreSQL stores durable URL mappings, so its absence should degrade requests cleanly while the application process remains alive and recovery occurs. Redis accelerates redirects, so losing it should exercise database fallback within the checked latency bound. The signer supplies client payment signatures; its absence should be observed and recover without collateral application restarts or the checked replay symptom.

Those are selected hypotheses, not a catalogue of every failure the platform can survive. Pod failure does not reproduce data corruption, every network partition, a region outage, prolonged resource exhaustion, or simultaneous faults. The failure model and known limitations define the project’s testnet scope.

Establish useful work before disrupting it

Section titled “Establish useful work before disrupting it”

The runner checks that the target Service has at least one ready EndpointSlice address. It then creates a Job from the suspended staging load-generator CronJob. The Job’s fixed marker is emitted only after one paid /shorten request returns 201 and its redirect returns 302. The runner reads that marker from the exact Job’s logs and confirms the Job remains active before creating the Chaos Mesh workflow.

This protects against the most direct vacuous experiment: an absent target or an unfunded/misconfigured load path that never completes a paid request. It proves a bounded startup success, not every later iteration. During the workflow, the runner polls the Job’s active/failed/succeeded status. That catches a terminated generator but does not continuously assert payment success. PostgreSQL and Redis have scored non-probe request checks; the signer rules do not include an equivalent traffic-count check.

The load-generator contract tests and orchestrator tests cover these boundaries with local controls. Retained records include a deliberately degraded pipeline candidate that could not establish paid traffic and stopped before injection; see the verification report.

The workflow selects one matching dependency Pod in url-shortener-staging for each 60-second PodChaos node. It starts with 90 seconds of baseline, waits 120 seconds after PostgreSQL, then 60 seconds after Redis and signer. Declared child windows sum to 510 seconds; the serial parent deadline is 660 seconds to allow controller scheduling headroom.

The runner refuses to start while another labeled gate workflow remains. An atomic Lease in the Kargo project namespace serializes cooperating gate runners. A live holder blocks a competing run; an expired Lease can be replaced only using its observed resource version. A fixed duration outlasts the Job budget rather than being renewed by a heartbeat loop.

Namespace and selector scoping reduce the intended fault boundary. They are not physical isolation: dev, staging, prod-like workloads, and controllers share the cluster. Kubernetes RBAC grants the runner namespace-scoped capabilities for workflow, load, observation, and cleanup operations. The script narrows its own object use; the RBAC permissions are not object-by-run constraints for every action.

A submitted workflow only proves an API accepted a specification. The runner needs a successful Apply event from each exact PodChaos resource. The scorer then needs readiness samples showing the selected dependency unavailable and available again at the scored window’s end. Application, dependency, and controller observations complement one another: none silently replaces the others.

If a timestamp is missing, the runner marks that experiment failed and skips its scorer invocation. If scoring can run, all defined checks are evaluated and failures retained. Prometheus is queried after the workflow observation phase; the runner does not make a Prometheus preflight query before injection. Invalid telemetry can therefore produce a failed result after a fault was genuinely applied.

The exit trap deletes the named workflow, its workflow-labeled PodChaos children, the named load Job, and the Lease it acquired. It uses a shared 120-second cleanup deadline and verifies absence or an empty child list. It does not sweep every gate workflow or every load Job to make the cluster look tidy. An older unknown workflow blocks the next run for inspection.

A cleanup failure forces the process to return failure even if all scorecards passed and a PASS log line was emitted earlier. The implementation’s single-object absence helper treats a failed kubectl get as absence; it does not distinguish a NotFound response from every read error. Therefore independent post-run readback remains valuable, and the website does not claim stronger cleanup verification than the source implements. Cleanup tests cover exact deletion, fault-event retention, and deadline budgets.

The chosen experiment earns a narrow conclusion: this candidate met these checks, under this bounded fault, in this testnet window, with these recorded cleanup observations. Follow inside the gate for every transition and measurements to interpret the observations without overstating them.

Maintained by Satyam Agnihotri · DevOps & Cloud Engineer