Evidence and scorecard formats
Structured records make a narrow claim traceable. Each format answers a different question: what collection belongs to a run, what one chaos experiment measured, what a dev baseline measured, or what a screenshot shows. Shape validity is necessary for tooling but is not execution proof or publication approval.
Run metadata: the collection index
Section titled “Run metadata: the collection index”schemas/run-metadata.schema.json declares draft 2020-12, version 1.0, fourteen required top-level fields, and no additional properties.
| Field | Meaning and constraint |
|---|---|
schema_version |
Exactly 1.0. |
run_id |
Fresh collection ID; 1–81 letters/digits/dot/underscore/dash, starting alphanumeric. |
scenario |
baseline, chaos-gate, regression-blocked, recovery, or prod-smoke. |
status |
Operator-supplied pass, fail, blocked, or unavailable. |
collected_at |
Date-time for collection, separate from experiment and screenshot time. |
repository_revision |
Source/Freight revision supplied to the gate; 7–64 lowercase hex characters or unavailable. |
chart_revision |
Observed rendered environment branch revision; same lexical constraint. This field does not mean chart-source commit. |
release_image_digest |
Application image’s sha256: plus 64 lowercase hex characters, or unavailable. |
gate_runner_image_digest |
Gate runtime content identity, or unavailable. |
signer_image_digest |
Signer runtime content identity, or unavailable. |
loadgen_image_digest |
Load runtime content identity, or unavailable. |
namespace |
One of the three fixed workload namespaces. |
cluster_context |
Nonempty single-line public context identity, at most 200 characters. No credentials. |
evidence |
Nonempty unique files paths and nonempty unique redactions categories. |
Scenario/namespace alignment is enforced: baseline maps to url-shortener-dev; chaos-gate, regression-blocked, and recovery map to url-shortener-staging; prod-smoke maps to url-shortener-prod. Paths reject traversal segments. Redaction categories are credentials, secret-data, wallet-keys, authorization-headers, or none.
The collector imposes additional pass requirements: actual source/render/app identity, exact AnalysisRun, and, for chaos, all supporting image digests plus gate Job/Pod and three scorecards. Those operational guards are stronger than mere schema shape.
The evidence.files array may index a full privately retained bundle while Git contains a reviewed subset. The public recovery metadata, for example, indexes Kubernetes extracts that are deliberately not tracked alongside its public log/scorecards. Do not turn the array into download links without checking actual availability.
Chaos scorecard: one experiment decision
Section titled “Chaos scorecard: one experiment decision”schemas/scorecard.schema.json requires schema_version, experiment, verdict, generated_at, release, window, and checks, with no additional properties. Its version is exactly resilience-gate.scorecard/v1; experiments are only postgres-pod-failure, redis-pod-failure, and signer-pod-failure.
release contains revision, image_digest, and run_id. Null values can describe malformed/incomplete invocation output; they do not establish passing identity. A normal bounded window records UTC inject_at, start, and end, integer positive duration_seconds, and nonnegative settle_seconds. Invalid invocation can produce a null window with failed evidence checks.
Every check includes:
| Fields | Interpretation |
|---|---|
id, name |
Stable rule identity and readable question. |
verdict |
pass or fail; all required checks must pass for the experiment verdict. |
observed, operator, threshold, unit |
Compared measurement, exact comparison (<, <=, >, >=, ==), bound, and units. Null observation is available for metadata checks or failed evidence. |
expression, query_kind |
Exact PromQL or metadata expression; instant/range query. |
sample_count |
Number of usable observations assessed, not a universal guarantee of coverage. |
reason, evidence_error |
Human interpretation and machine-readable evidence defect, nullable when absent. |
Read the actual rule aggregation as well as the JSON. An outage rule and recovery rule may use the same readiness expression while taking different parts or summaries of the window. A successful identity check has no telemetry sample and can have null observed; null is not automatically an evidence failure for every kind of check. Conversely, a stale or non-finite telemetry value must not become zero.
The public Redis scorecard is a real example. Scorer source, scorecard tests, and no-data tests define evaluation beyond field shape. The metrics reference names thresholds, units, and aggregation caveats.
Baseline and load formats are separate
Section titled “Baseline and load formats are separate”scripts/score-baseline.py emits resilience-gate.baseline-scorecard/v1, with scenario: baseline, release, verdict, checks, and explicit started_at/ended_at/duration_seconds. Its implemented validator lives in the scorer; the chaos schema deliberately does not accept it. Its normal score window is 60–150 seconds, with final-one-minute latency assessed at the end rather than across an idle-start range.
run-loadgen.sh writes resilience-gate.baseline-run/v1 after copying a baseline summary and verifying exact cleanup. It distinguishes source namespace/CronJob/Job/Pod from dev target, includes requested and observed timestamps, duration/deadline, and cleanup information. The k6 program emits resilience-gate.loadgen-summary/v1 with traffic mode and safe metric summaries. It deliberately omits options, environment, request bodies, headers, and tags, and emits a log sentinel so a completed Pod’s summary can be recovered without copying from a stopped container. These records establish load execution context; neither is a chaos scorecard or Kargo verdict.
Screenshot manifest
Section titled “Screenshot manifest”docs/screenshots/manifest.json is version 1 and currently describes 29 canonical PNG views plus five companion details. Each capture binds id, relative file, sha256, UTC capture time, surface, evidence role, scenario, visible proof, sensitive-content review, and caption review, with applicable source/Freight/digest/render/run identities and observed window.
Capture time and observed window remain separate. Optional identities are null when a static view does not show a run. Hashes bind bytes; they do not certify a gate. The public gallery uses those reviewed captions. Collection and publication responsibilities remain in publish evidence.
Maintained by Satyam Agnihotri · DevOps & Cloud Engineer