Skip to content

Gate and observability design

The gate/observability subsystem asks whether an exact staging candidate meets selected dependency-failure assertions under bounded paid traffic. Its implemented design combines a Kubernetes Job, a serial Chaos Mesh workflow, strict returned-telemetry validation, portable scorecards, and cleanup. This is a source-derived platform design; website rendering and deployment are separate systems.

The gate covers PostgreSQL, Redis, and signer Pod failure in staging. It does not claim arbitrary fault combinations, long outages, data-loss recovery, node/region disaster tolerance, continuous paid success through signer failure, or a public production SLO.

Component Responsibility Source
Kargo staging verification Supplies service URL and candidate identity; invokes health and gate analyses Stage
Job AnalysisTemplate Pins verifier runtime, identity arguments, execution budget, security context, result volume AnalysisTemplate
Orchestrator Validates settings, locks, checks target/load, runs workflow, captures fault events, invokes scores, cleans objects orchestrate.sh
Chaos Mesh Applies serial selected Pod failures and exposes conditions/events workflow.yaml
k6 load client Establishes paid create/redirect startup proof and runs bounded traffic loadgen.js
Prometheus Returns application and Kubernetes metric evidence ServiceMonitor
Pure scorer and HTTP adapter Validates responses/windows, evaluates checks, records JSON verdicts score_experiment.py
Grafana/Loki/Alloy Diagnostic visualization, annotation, and log exploration observability chart

Grafana is an observation interface, not the decision authority. Optional annotation failure cannot change a verdict. The scorer has no Kubernetes or cloud client; its only network adapter reads Prometheus and tests substitute synthetic results.

Gate state and outcome diagram

Enlarge diagram: Gate state and outcome diagram · Version-controlled diagram source

The orchestration verdict precedes exit cleanup. The final Job outcome includes that cleanup path.

State Required transition Failure consequence
Validate Safe revision/digest/run ID and sufficient lifecycle budget Stops before Lease/load/fault creation
Lock Atomic creation or version-checked expired takeover Competing/incomplete lock blocks
Inspect No older labeled workflow; ready target endpoint Stops before paid load
Establish load Exact Job active, paid-success marker observed, active status reconfirmed Cleans created load/Lease; no workflow
Observe workflow Successful Apply events captured, load stays active, workflow finishes in time Orchestration fails; available scoring diagnostics still attempted
Score Timestamp and all returned-evidence/threshold checks pass for each experiment Failed experiment retained; gate fails
Annotate/report Optional diagnostic operation and main verdict log Annotation is non-fatal
Cleanup Exact workflow/fault children/load/Lease deletions and absence checks within bound Forces final nonzero exit

The fixed 1470-second Lease duration covers the 1320-second Job budget and buffer; it is not renewed by the script. The Job has no retry and a 120-second termination grace. The RBAC permits namespace-scoped control/observation, while object-by-run narrowing is script behavior rather than a universal RBAC constraint.

The serial child windows are 90 seconds baseline, 60 seconds PostgreSQL fault, 120 seconds PostgreSQL recovery, 60 seconds Redis fault, 60 seconds Redis recovery, 60 seconds signer fault, and 60 seconds signer recovery. They sum to 510 seconds, with a 660-second parent deadline.

The orchestrator preserves the earliest successful Apply event on each exact typed PodChaos child while nodes still exist. Scheduling time is not application proof. Range scores begin 20 seconds before that event and end 120 seconds after it; counter lookbacks span the final 120 seconds. Every scorer uses 60 seconds of settle time, including PostgreSQL despite its longer workflow pause.

Prometheus reads occur after workflow observation, with up to six 15-second requests per scorer and a 105-second scorer process bound. There is no pre-injection Prometheus preflight. The distinction leaves telemetry defects detectable during scoring but does not prevent every experiment from starting when telemetry is unavailable.

The HTTP adapter accepts only a successful response of the expected vector or matrix kind with finite well-formed samples. Each returned range series needs three samples, boundary coverage within 60 seconds, and no gap over 75 seconds. An instant result must have exactly one sample within 60 seconds of evaluation. Error conversion creates failed checks rather than guessing values.

Expressions can intentionally produce zeros: restart/error/replay rules use or vector(0). Missing left-hand series in those expressions need not yield missing returned evidence. Target readiness, traffic, cache-miss, and latency queries do not have that fallback. The metrics reference is the exact rule catalogue.

Each resilience-gate.scorecard/v1 record carries source/Freight revision, application digest, generated workflow run identity, injection/window, all rendered expressions, counts, observed values, thresholds, and reasons. Required identity validation is appended by CLI use. The /results volume is ephemeral; durable evidence needs the collector or emitted sanitized scorecard log.

No ready endpoint is an invalid experiment target, not resilient behavior. No paid marker is unavailable useful work, not a fault pass. Missing Apply evidence is unproven injection. Bad telemetry is invalid evidence. A measured threshold failure is a behavior failure for that rule. A failed cleanup is an unfinished experiment outcome even if behavior passed.

The current single-object absence helper treats any failed kubectl get as absent. It does not separately prove NotFound under every API/permission error. The design therefore records that implementation limit and relies on the actual cleanup calls/readback evidence available rather than claiming a stronger guarantee.

Signer isolation also has a narrow meaning. The client needs the signer before making paid application requests; failed signing returns early. The gate checks generator activity and selected application symptoms, but the signer score lacks an independent traffic threshold. It does not establish uninterrupted paid throughput.

Orchestrator, cleanup, no-data, query, and scorecard tests exercise offline boundaries. Dependency-specific tests cover positive controls and rejected behavior.

The public evidence index separates Kargo-managed historical passes from direct boundary tests and private fresh bundles. A scorer-only unreachable-endpoint test does not demonstrate a full shared-Prometheus outage; a one-second timeout test does not establish default-duration timeout behavior; uncollected live Lease concurrency must remain uncollected. The recorded release report supplies bounded live observations without transferring them to future candidates.

For the reader-facing sequence see inside the gate. For how the result affects eligibility see score to promotion.

Maintained by Satyam Agnihotri · DevOps & Cloud Engineer