A green result can be insufficient
A green dashboard is an observation surface. A passing scorecard is a set of measured assertions. A successful gate Job says the orchestrated run finished with a successful process outcome. A successful staging verification supplies normal downstream eligibility. These statements connect, but each has its own boundary.
The release decision becomes credible when those boundaries agree for the same candidate and run. It becomes misleading when the reader jumps from one green layer directly to “promoted.”
First, validate the measurements
Section titled “First, validate the measurements”The scorer treats Prometheus response failures as failed checks. It validates response status, vector or matrix kind, sample structure, numeric finiteness, ordering, freshness, and coverage before aggregation. An HTTP 200 response is not enough. An empty returned series is not implicitly zero.
That rule has query-level exceptions. Several error, replay, and restart expressions end in or vector(0), deliberately allowing absent left-hand series to become a returned zero. This can be appropriate for a counter that appears only when a symptom occurs, but it prevents the parser from distinguishing an absent counter from a real observed zero. Request totals, cache misses, latency, and target readiness require their own returned evidence. Measurement interpretation explains this distinction.
For each experiment, the scorer asks whether the required set of observations jointly supports the hypothesis. PostgreSQL and Redis require more than 50 non-probe requests over the fault-plus-settle lookback. Redis also requires a positive cache-miss increase and maximum range p95 below 1.2 seconds. Dependency readiness must show an unavailable sample and end at one. Other checked symptoms must stay below their explicit bounds. The reference lists exact expressions and units.
The signer set checks signer outage/recovery, application restarts, /shorten 5xx, and payment replays. It has no separate meaningful-traffic check. A signer failure can stop the client’s signing step before it calls /shorten. Consequently, quiet application errors support limited isolation assertions; they do not prove uninterrupted successful payment traffic.
Then, inspect the scorecard’s candidate and window
Section titled “Then, inspect the scorecard’s candidate and window”The CLI requires source/Freight revision, application digest, and bounded run identifier. It appends a release-identity check; an invalid identity makes the card fail. It evaluates all defined metric checks rather than short-circuiting on the first problem, so the JSON retains useful diagnoses.
The resilience-gate.scorecard/v1 record includes experiment, verdict, generated time, release fields, injection/start/end times, fault/settle durations, and every check’s expression, query kind, observed value, comparison, unit, sample count, and reason or evidence error. A failed invocation can have no valid window. Scorecard tests and the schema check this contract.
Do not substitute generated_at for the observation window. Scoring occurs after workflow completion; the returned samples describe the earlier fault interval. Do not substitute the card’s source/Freight revision for the output branch revision. Evidence metadata retains that output separately.
A passing card can coexist with a failing gate
Section titled “A passing card can coexist with a failing gate”The runner joins results that the scorer cannot know. A scorecard has no Kubernetes client and cannot determine whether the overall workflow failed, the generator terminated, the fault actually exposed an Apply timestamp for every dependency, or cleanup completed.
| Observation | Narrow conclusion | Additional layer required |
|---|---|---|
| One scorecard passes | That experiment’s defined returned evidence and identity checks passed | Other experiments and orchestration outcome |
| All three scorecards pass | All scored experiment assertions passed | Workflow/load outcome and cleanup |
GATE VERDICT: PASS log appears |
Main orchestration reached a pass before its exit trap | Final process and Kubernetes Job outcome |
| Gate Job succeeds | The Job-backed chaos metric completed successfully | Staging service-health and AnalysisRun status |
| Staging verification succeeds | Normal upstream verification requirement is satisfied | Explicit downstream promotion and reconciliation |
| Prod readiness/liveness succeeds | Prod-like target passed those post-deploy checks | Historical staging link and scope-correct evidence claim |
Cleanup illustrates the difference. The runner can emit passing cards and a PASS line, then fail to delete a run-scoped object or verify cleanup before the deadline. Its exit trap changes the final exit code to failure. An accepted load/workflow request is even earlier: it proves an API accepted work, not that useful traffic or faults occurred.
The single-object cleanup absence helper also treats a failed read as absence without classifying its reason. Do not amplify the implementation into a guarantee that every API outage is independently detected. Recorded readback and the cleanup tests establish the specific evidence available.
The enclosing analysis decides staging verification
Section titled “The enclosing analysis decides staging verification”The staging Stage references service-health and chaos-gate. The AnalysisTemplate configures the chaos verdict as one Job-backed measurement with zero allowed failures and no Job retry. Grafana annotation remains optional and non-fatal. A dashboard annotation is therefore not the decision source.
The AnalysisRun’s status must be inspected alongside the Job and scorecards. A screenshot of an AnalysisRun can illustrate that state at a captured moment, but its candidate, source, and window still need the linked records. A screenshot capture time may be later than the graph’s displayed observation window.
Once staging verification succeeds, Freight gains normal eligibility for prod. The Project policy still sets autoPromotionEnabled: false for prod. No automatic production-like promotion follows from the score. A manual Freight approval is a distinct override and cannot be cited as proof of staging verification.
Keep the recorded pass narrower than a service guarantee
Section titled “Keep the recorded pass narrower than a service guarantee”The public historical gate retains scorecard output and cleanup observations for its named execution. The later 3 October release report summarizes fresh staging and prod-like results, with detailed bundles held in a private external archive. Public summary and public raw scorecards are different evidence availability states.
These records establish selected bounded behavior for exact identities in an owned testnet lab. They do not define a general availability SLO, prove traffic continuity through every signer fault, or establish every runtime signature was independently admitted. Earlier failed candidates remain valuable negative evidence and must keep their own dates and identities.
When reviewing a release, ask in order: did the query return valid evidence; did the bounded rules pass; did the candidate binding match; did orchestration and cleanup finish; did the enclosing analysis succeed; and was the next promotion explicitly requested and reconciled? That order makes “green” an inspectable result instead of an unqualified conclusion.
Maintained by Satyam Agnihotri · DevOps & Cloud Engineer