Skip to content

Metrics and scoring rules

The implemented rules live in the chaos scorer and baseline scorer. This reference preserves their exact query semantics. The bounds are project verification rules, not an availability SLO or a public production guarantee.

The ServiceMonitor selects the URL-shortener Service in each reviewed application namespace, scrapes /metrics every 15 seconds with a 10-second timeout, and attaches a namespace label. Application HTTP instrumentation excludes /livez, /ready, and /metrics as configured in application source. Kubernetes readiness and restart signals come from the observability stack’s Kubernetes state metrics.

Metric Meaning relevant to review
http_requests_total Instrumented HTTP request count by route/status/method labels; it does not alone identify a paid success.
http_request_duration_seconds_bucket HTTP duration histogram in seconds; scorer uses route-template labels.
url_shortener_cache_hits_total Redirects resolved through a cached value.
url_shortener_cache_misses_total Ordinary cache misses or unavailable-cache fallbacks; it is not an outage-only counter.
url_shortener_urls_created_total Successfully persisted new short URLs.
url_shortener_payment_outcomes_total{outcome=…} Coarse payment decision outcomes, not raw facilitator messages.
url_shortener_payment_replays_total Settlement identifiers rejected as already used by the application.
url_shortener_dependency_up{dependency=…} Dependency reachability at the most recent health check; not the Kubernetes target-readiness expression used for chaos.
url_shortener_facilitator_calls_total{endpoint=…,outcome=…} Facilitator verify/settle call results.
url_shortener_facilitator_request_duration_seconds Per-facilitator-call duration histogram.
kube_pod_status_ready{condition="true"} Kubernetes Pod readiness state for the named target.
kube_pod_container_status_restarts_total Kubernetes restart count for the application container.

Custom metric definitions are in main.py and payment.py; metric tests and payment-metric tests cover their behavior. Payment dashboard display and gate assertions are different consumers of telemetry.

Let T be the earliest successful Apply event on an experiment’s exact PodChaos child. The runner invokes each scorer with duration 60 seconds; each definition uses settle 60 seconds. Therefore:

Item Current meaning
Range start T − 20 seconds
Range end / instant evaluation time T + 120 seconds
Counter lookback W 120s, fault plus scored settling; no lead included
Range step 15 seconds
Prometheus request timeout 15 seconds
Minimum samples 3 per returned range series
Allowed range start/end distance First sample no later than start + 60s; last sample no earlier than end − 60s
Maximum adjacent range gap 75 seconds
Instant freshness Absolute timestamp distance from evaluation time ≤ 60 seconds

The workflow’s PostgreSQL recovery pause is 120 seconds; its scorer still uses 60 seconds of settling. Redis and signer workflow recovery pauses are 60 seconds. A full workflow screenshot interval is not interchangeable with these query windows.

Instant results must contain exactly one sample after the Prometheus expression’s aggregation. Range aggregations operate over all validated returned samples: min, max, and last below mean the Python scorer aggregation, following the PromQL expression. Equality uses absolute tolerance 1e-9; strict comparisons remain strict.

Empty results, wrong response/result shapes, failed transport, non-finite values, unordered timestamps, stale data, or insufficient coverage fail the affected check. The deliberate query fallback or vector(0) can produce a valid returned zero when its left side is absent; the parser does not fail that synthetic result merely because the underlying counter is absent.

The table’s queries show the current staging namespace and W=120s. A namespace CLI override changes the label value, but exact target names in the PostgreSQL and Redis definitions remain staging names.

Check ID Query kind / scorer aggregation Pass rule and unit
meaningful-traffic Instant / value > 50 requests
app-restarts Instant / value < 0.5 restarts
postgres-outage-observed Range / min == 0 state
postgres-recovered Range / last == 1 state
postgres-clean-degradation Instant / value < 0.5 responses
# meaningful-traffic
sum(increase(http_requests_total{namespace="url-shortener-staging", handler!~"/(livez|ready|metrics)"}[120s]))
# app-restarts
max(increase(kube_pod_container_status_restarts_total{namespace="url-shortener-staging", container="url-shortener"}[120s])) or vector(0)
# postgres-outage-observed and postgres-recovered
min(kube_pod_status_ready{namespace="url-shortener-staging", condition="true", pod="url-shortener-staging-postgresql-0"})
# postgres-clean-degradation
sum(increase(http_requests_total{namespace="url-shortener-staging", handler="/shorten", status="500"}[120s])) or vector(0)

Clean degradation checks the specific /shorten status 500 symptom. It does not require all requests to succeed and does not ban every 5xx status; expected database unavailability can produce 503. Traffic totals are non-probe requests, not a dedicated count of paid creations. PostgreSQL scorer tests cover landed fault and recovery requirements.

Check ID Query kind / scorer aggregation Pass rule and unit
meaningful-traffic Instant / value > 50 requests
redis-fallback-observed Instant / value > 0 fallbacks
redirect-latency Range / max < 1.2 seconds
app-restarts Instant / value < 0.5 restarts
redis-outage-observed Range / min == 0 state
redis-recovered Range / last == 1 state
# meaningful-traffic
sum(increase(http_requests_total{namespace="url-shortener-staging", handler!~"/(livez|ready|metrics)"}[120s]))
# redis-fallback-observed
sum(increase(url_shortener_cache_misses_total{namespace="url-shortener-staging"}[120s]))
# redirect-latency: maximum of returned one-minute p95 values
histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket{namespace="url-shortener-staging", handler="/{code}"}[1m])) by (le))
# app-restarts
max(increase(kube_pod_container_status_restarts_total{namespace="url-shortener-staging", container="url-shortener"}[120s])) or vector(0)
# redis-outage-observed and redis-recovered
min(kube_pod_status_ready{namespace="url-shortener-staging", condition="true", pod="url-shortener-staging-redis-master-0"})

The maximum range p95 is not the p95 of all requests pooled over the 140-second range. Each returned point uses a one-minute rate. The positive cache-miss count must be interpreted alongside observed Redis outage; ordinary misses also increment it. Redis scorer tests check slow fallback, missing fallback evidence, route-label preservation, and exact target selection.

Check ID Query kind / scorer aggregation Pass rule and unit
signer-outage-observed Range / min == 0 state
signer-recovered Range / last == 1 state
app-restarts Instant / value < 0.5 restarts
app-5xx Instant / value < 0.5 responses
payment-replays Instant / value < 0.5 replays
# signer-outage-observed and signer-recovered
min(kube_pod_status_ready{namespace="url-shortener-staging", condition="true", pod=~"radius-signer-[a-z0-9]+-[a-z0-9]+"})
# app-restarts
max(increase(kube_pod_container_status_restarts_total{namespace="url-shortener-staging", container="url-shortener"}[120s])) or vector(0)
# app-5xx
sum(increase(http_requests_total{namespace="url-shortener-staging", handler="/shorten", status=~"5.."}[120s])) or vector(0)
# payment-replays
sum(increase(url_shortener_payment_replays_total{namespace="url-shortener-staging"}[120s])) or vector(0)

There is no signer meaningful-traffic check. The load client returns early when it cannot obtain a signature, so low application errors cannot establish continuous paid useful work. Signer scorer tests check dependency recovery and collateral application symptoms.

The baseline scorer uses explicit start/end times rather than a manufactured injection time. Its observation duration must be an integer from 60 to 150 seconds. W below is that exact duration. Instant queries evaluate at the supplied end; range queries span the supplied window and share the coverage validation above.

Check ID Query / evaluation Pass rule
root-get-traffic sum(increase(http_requests_total{namespace="N", handler="/", method="GET"}[Ws])), instant/value > 20 requests
application-5xx sum(increase(http_requests_total{namespace="N", status=~"5.."}[Ws])) or vector(0), instant/value < 0.5 responses
root-get-p95-latency histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket{namespace="N", handler="/", method="GET"}[1m])) by (le)), instant/value < 1.0 seconds
postgres-ready min(kube_pod_status_ready{namespace="N", condition="true", pod=~"url-shortener-[a-z0-9-]+-postgresql-0"}), range/min == 1 state
redis-ready min(kube_pod_status_ready{namespace="N", condition="true", pod=~"url-shortener-[a-z0-9-]+-redis-master-0"}), range/min == 1 state

Baseline p95 is one fresh instant value covering the final minute; unlike Redis chaos latency, it is not a maximum across range points. The baseline produces resilience-gate.baseline-scorecard/v1, distinct from the three chaos experiment cards. Baseline tests cover fresh evidence, invalid windows/identities, and sanitized failure output.

The chaos CLI appends required release-identity validation. It needs a source/Freight revision, lowercase sha256 digest, and a bounded safe run ID. Its card verdict passes only when every included check passes. Query failures retain observed: null, an evidence error, and a reason; they are not represented as a guessed zero.

The scorer does not decide cleanup or Kargo promotion. Those belong to gate orchestration and score-to-promotion interpretation. Read query tests, no-data tests, and scorecard tests when updating this reference against future source.

Maintained by Satyam Agnihotri · DevOps & Cloud Engineer