Gate and observability design
The gate/observability subsystem asks whether an exact staging candidate meets selected dependency-failure assertions under bounded paid traffic. Its implemented design combines a Kubernetes Job, a serial Chaos Mesh workflow, strict returned-telemetry validation, portable scorecards, and cleanup. This is a source-derived platform design; website rendering and deployment are separate systems.
Scope and responsibilities
Section titled “Scope and responsibilities”The gate covers PostgreSQL, Redis, and signer Pod failure in staging. It does not claim arbitrary fault combinations, long outages, data-loss recovery, node/region disaster tolerance, continuous paid success through signer failure, or a public production SLO.
| Component | Responsibility | Source |
|---|---|---|
| Kargo staging verification | Supplies service URL and candidate identity; invokes health and gate analyses | Stage |
| Job AnalysisTemplate | Pins verifier runtime, identity arguments, execution budget, security context, result volume | AnalysisTemplate |
| Orchestrator | Validates settings, locks, checks target/load, runs workflow, captures fault events, invokes scores, cleans objects | orchestrate.sh |
| Chaos Mesh | Applies serial selected Pod failures and exposes conditions/events | workflow.yaml |
| k6 load client | Establishes paid create/redirect startup proof and runs bounded traffic | loadgen.js |
| Prometheus | Returns application and Kubernetes metric evidence | ServiceMonitor |
| Pure scorer and HTTP adapter | Validates responses/windows, evaluates checks, records JSON verdicts | score_experiment.py |
| Grafana/Loki/Alloy | Diagnostic visualization, annotation, and log exploration | observability chart |
Grafana is an observation interface, not the decision authority. Optional annotation failure cannot change a verdict. The scorer has no Kubernetes or cloud client; its only network adapter reads Prometheus and tests substitute synthetic results.
State transitions and outcomes
Section titled “State transitions and outcomes”Enlarge diagram: Gate state and outcome diagram · Version-controlled diagram source
The orchestration verdict precedes exit cleanup. The final Job outcome includes that cleanup path.
| State | Required transition | Failure consequence |
|---|---|---|
| Validate | Safe revision/digest/run ID and sufficient lifecycle budget | Stops before Lease/load/fault creation |
| Lock | Atomic creation or version-checked expired takeover | Competing/incomplete lock blocks |
| Inspect | No older labeled workflow; ready target endpoint | Stops before paid load |
| Establish load | Exact Job active, paid-success marker observed, active status reconfirmed | Cleans created load/Lease; no workflow |
| Observe workflow | Successful Apply events captured, load stays active, workflow finishes in time | Orchestration fails; available scoring diagnostics still attempted |
| Score | Timestamp and all returned-evidence/threshold checks pass for each experiment | Failed experiment retained; gate fails |
| Annotate/report | Optional diagnostic operation and main verdict log | Annotation is non-fatal |
| Cleanup | Exact workflow/fault children/load/Lease deletions and absence checks within bound | Forces final nonzero exit |
The fixed 1470-second Lease duration covers the 1320-second Job budget and buffer; it is not renewed by the script. The Job has no retry and a 120-second termination grace. The RBAC permits namespace-scoped control/observation, while object-by-run narrowing is script behavior rather than a universal RBAC constraint.
Fault and measurement timing
Section titled “Fault and measurement timing”The serial child windows are 90 seconds baseline, 60 seconds PostgreSQL fault, 120 seconds PostgreSQL recovery, 60 seconds Redis fault, 60 seconds Redis recovery, 60 seconds signer fault, and 60 seconds signer recovery. They sum to 510 seconds, with a 660-second parent deadline.
The orchestrator preserves the earliest successful Apply event on each exact typed PodChaos child while nodes still exist. Scheduling time is not application proof. Range scores begin 20 seconds before that event and end 120 seconds after it; counter lookbacks span the final 120 seconds. Every scorer uses 60 seconds of settle time, including PostgreSQL despite its longer workflow pause.
Prometheus reads occur after workflow observation, with up to six 15-second requests per scorer and a 105-second scorer process bound. There is no pre-injection Prometheus preflight. The distinction leaves telemetry defects detectable during scoring but does not prevent every experiment from starting when telemetry is unavailable.
Evidence model and validation
Section titled “Evidence model and validation”The HTTP adapter accepts only a successful response of the expected vector or matrix kind with finite well-formed samples. Each returned range series needs three samples, boundary coverage within 60 seconds, and no gap over 75 seconds. An instant result must have exactly one sample within 60 seconds of evaluation. Error conversion creates failed checks rather than guessing values.
Expressions can intentionally produce zeros: restart/error/replay rules use or vector(0). Missing left-hand series in those expressions need not yield missing returned evidence. Target readiness, traffic, cache-miss, and latency queries do not have that fallback. The metrics reference is the exact rule catalogue.
Each resilience-gate.scorecard/v1 record carries source/Freight revision, application digest, generated workflow run identity, injection/window, all rendered expressions, counts, observed values, thresholds, and reasons. Required identity validation is appended by CLI use. The /results volume is ephemeral; durable evidence needs the collector or emitted sanitized scorecard log.
Distinct failure meanings
Section titled “Distinct failure meanings”No ready endpoint is an invalid experiment target, not resilient behavior. No paid marker is unavailable useful work, not a fault pass. Missing Apply evidence is unproven injection. Bad telemetry is invalid evidence. A measured threshold failure is a behavior failure for that rule. A failed cleanup is an unfinished experiment outcome even if behavior passed.
The current single-object absence helper treats any failed kubectl get as absent. It does not separately prove NotFound under every API/permission error. The design therefore records that implementation limit and relies on the actual cleanup calls/readback evidence available rather than claiming a stronger guarantee.
Signer isolation also has a narrow meaning. The client needs the signer before making paid application requests; failed signing returns early. The gate checks generator activity and selected application symptoms, but the signer score lacks an independent traffic threshold. It does not establish uninterrupted paid throughput.
Validation and retained observations
Section titled “Validation and retained observations”Orchestrator, cleanup, no-data, query, and scorecard tests exercise offline boundaries. Dependency-specific tests cover positive controls and rejected behavior.
The public evidence index separates Kargo-managed historical passes from direct boundary tests and private fresh bundles. A scorer-only unreachable-endpoint test does not demonstrate a full shared-Prometheus outage; a one-second timeout test does not establish default-duration timeout behavior; uncollected live Lease concurrency must remain uncollected. The recorded release report supplies bounded live observations without transferring them to future candidates.
For the reader-facing sequence see inside the gate. For how the result affects eligibility see score to promotion.
Maintained by Satyam Agnihotri · DevOps & Cloud Engineer