Skip to content

Evidence and scorecard formats

Structured records make a narrow claim traceable. Each format answers a different question: what collection belongs to a run, what one chaos experiment measured, what a dev baseline measured, or what a screenshot shows. Shape validity is necessary for tooling but is not execution proof or publication approval.

schemas/run-metadata.schema.json declares draft 2020-12, version 1.0, fourteen required top-level fields, and no additional properties.

Field Meaning and constraint
schema_version Exactly 1.0.
run_id Fresh collection ID; 1–81 letters/digits/dot/underscore/dash, starting alphanumeric.
scenario baseline, chaos-gate, regression-blocked, recovery, or prod-smoke.
status Operator-supplied pass, fail, blocked, or unavailable.
collected_at Date-time for collection, separate from experiment and screenshot time.
repository_revision Source/Freight revision supplied to the gate; 7–64 lowercase hex characters or unavailable.
chart_revision Observed rendered environment branch revision; same lexical constraint. This field does not mean chart-source commit.
release_image_digest Application image’s sha256: plus 64 lowercase hex characters, or unavailable.
gate_runner_image_digest Gate runtime content identity, or unavailable.
signer_image_digest Signer runtime content identity, or unavailable.
loadgen_image_digest Load runtime content identity, or unavailable.
namespace One of the three fixed workload namespaces.
cluster_context Nonempty single-line public context identity, at most 200 characters. No credentials.
evidence Nonempty unique files paths and nonempty unique redactions categories.

Scenario/namespace alignment is enforced: baseline maps to url-shortener-dev; chaos-gate, regression-blocked, and recovery map to url-shortener-staging; prod-smoke maps to url-shortener-prod. Paths reject traversal segments. Redaction categories are credentials, secret-data, wallet-keys, authorization-headers, or none.

The collector imposes additional pass requirements: actual source/render/app identity, exact AnalysisRun, and, for chaos, all supporting image digests plus gate Job/Pod and three scorecards. Those operational guards are stronger than mere schema shape.

The evidence.files array may index a full privately retained bundle while Git contains a reviewed subset. The public recovery metadata, for example, indexes Kubernetes extracts that are deliberately not tracked alongside its public log/scorecards. Do not turn the array into download links without checking actual availability.

schemas/scorecard.schema.json requires schema_version, experiment, verdict, generated_at, release, window, and checks, with no additional properties. Its version is exactly resilience-gate.scorecard/v1; experiments are only postgres-pod-failure, redis-pod-failure, and signer-pod-failure.

release contains revision, image_digest, and run_id. Null values can describe malformed/incomplete invocation output; they do not establish passing identity. A normal bounded window records UTC inject_at, start, and end, integer positive duration_seconds, and nonnegative settle_seconds. Invalid invocation can produce a null window with failed evidence checks.

Every check includes:

Fields Interpretation
id, name Stable rule identity and readable question.
verdict pass or fail; all required checks must pass for the experiment verdict.
observed, operator, threshold, unit Compared measurement, exact comparison (<, <=, >, >=, ==), bound, and units. Null observation is available for metadata checks or failed evidence.
expression, query_kind Exact PromQL or metadata expression; instant/range query.
sample_count Number of usable observations assessed, not a universal guarantee of coverage.
reason, evidence_error Human interpretation and machine-readable evidence defect, nullable when absent.

Read the actual rule aggregation as well as the JSON. An outage rule and recovery rule may use the same readiness expression while taking different parts or summaries of the window. A successful identity check has no telemetry sample and can have null observed; null is not automatically an evidence failure for every kind of check. Conversely, a stale or non-finite telemetry value must not become zero.

The public Redis scorecard is a real example. Scorer source, scorecard tests, and no-data tests define evaluation beyond field shape. The metrics reference names thresholds, units, and aggregation caveats.

scripts/score-baseline.py emits resilience-gate.baseline-scorecard/v1, with scenario: baseline, release, verdict, checks, and explicit started_at/ended_at/duration_seconds. Its implemented validator lives in the scorer; the chaos schema deliberately does not accept it. Its normal score window is 60–150 seconds, with final-one-minute latency assessed at the end rather than across an idle-start range.

run-loadgen.sh writes resilience-gate.baseline-run/v1 after copying a baseline summary and verifying exact cleanup. It distinguishes source namespace/CronJob/Job/Pod from dev target, includes requested and observed timestamps, duration/deadline, and cleanup information. The k6 program emits resilience-gate.loadgen-summary/v1 with traffic mode and safe metric summaries. It deliberately omits options, environment, request bodies, headers, and tags, and emits a log sentinel so a completed Pod’s summary can be recovered without copying from a stopped container. These records establish load execution context; neither is a chaos scorecard or Kargo verdict.

docs/screenshots/manifest.json is version 1 and currently describes 29 canonical PNG views plus five companion details. Each capture binds id, relative file, sha256, UTC capture time, surface, evidence role, scenario, visible proof, sensitive-content review, and caption review, with applicable source/Freight/digest/render/run identities and observed window.

Capture time and observed window remain separate. Optional identities are null when a static view does not show a run. Hashes bind bytes; they do not certify a gate. The public gallery uses those reviewed captions. Collection and publication responsibilities remain in publish evidence.

Maintained by Satyam Agnihotri · DevOps & Cloud Engineer